Methods and systems for distinguishing somatic genomic sequences from germline genomic sequences

By processing, amplifying, capturing, and sequencing nucleic acid molecules in samples, combined with surrogate gene sequences and statistical models, the problem of distinguishing between somatic cell and germ cell gene sequences has been solved, supporting the implementation of precision medicine and improving the targeting of treatments.

JP2026016702APending Publication Date: 2026-02-03FOUNDATION MEDICINE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025184719
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-06-19
Filing Date
2025-10-31
Publication Date
2026-02-03

Smart Images

  • Figure 2026016702000014
    Figure 2026016702000014
  • Figure 2026016702000015
    Figure 2026016702000015
  • Figure 2026016702000016
    Figure 2026016702000016
Patent Text Reader

Abstract

To provide methods and systems for distinguishing somatic genomic sequences from germline genomic sequences.SOLUTION: Described herein are methods for distinguishing between somatic and germline variants, and devices for implementing such methods. In certain implementations of the method, the method may include identifying a genomic sequence of interest in a patient sample at a genomic locus, identifying one or more proxy genomic sequences for the sequence of interest, comparing an observed frequency of the sequence of interest to a centrality measure of the observed frequencies of the one or more proxy genomic sequences, and characterizing the genomic sequence of interest as either germline or somatic based on the comparison.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 63 / 035,572, filed June 5, 2020, and U.S. Provisional Patent Application No. 63 / 041,437, filed June 19, 2020, both of which are incorporated by reference in their entireties.

[0002] Technical Field The present disclosure relates to systems and methods for distinguishing somatic genomic sequences from germline genomic sequences. [Background technology]

[0003] background Germline genomic sequence refers to the sequence that an organism inherits from its parents. In particular, if one or both of an organism's parents have particular genomic mutations (or if the organism experiences particular mutations in its very early development), those mutations may be in the organism's germline and are passed on to the organism's offspring (if any).

[0004] In contrast, somatic genomic sequences are sequences that are not passed from parent to child. For example, organisms can develop genomic mutations due to external factors (e.g., pollution, radiation, diet, smoking, etc.), and those genomic mutations are restricted to specific tissues, body fluids, or other anatomical materials. In some cases, these mutations result in undesirable medical conditions, including, but not limited to, cancer.

[0005] Precision medicine is the field of treating patients with therapies that target their individual characteristics or conditions. For many patients (including cancer patients), this may involve determining genomic information about both the patient's "normal" genomic state and the genomic state of the patient's "abnormal" tissues, fluids, or other anatomical material. This information may be derived from a sample from the patient, such as a tumor biopsy, a blood draw, or any other type of sample that has both normal and abnormal tissues, fluids, or other anatomical material.

[0006] These samples can be assayed to determine (at least in part) the genomic sequence of the material contained therein, however, it is sometimes difficult to identify whether a particular genomic sequence is derived from a patient's normal or abnormal anatomical material, i.e., it is sometimes difficult to determine whether a particular genomic sequence is germline or somatic.

[0007] Understanding whether genetic variants observed in the DNA of cancer patients are of germline or somatic origin is crucial in both clinical practice and cancer research. Somatic / germline differentiation can be achieved, for example, by sequencing matched tumor and normal tissues from the same patient. Variants present in the tumor but absent from the normal tissue are classified as somatic, whereas variants present in both are classified as germline. However, such a dual-sample approach is limited by cost and specimen availability. Typically, matched normal specimens are not available in clinical practice. For example, in the case of tissue biopsies, a single specimen containing both the tumor and its adjacent normal tissue is collected. Therefore, there is a need to develop methods that can reliably classify detected variants as somatic or germline in origin. Summary of the Invention [Means for solving the problem]

[0008] overview Described herein are methods, devices, and computer-readable media for distinguishing somatic genomic sequences from germline genomic sequences.

[0009] Provided herein is a method for identifying a subject's genomic sequence as germline or somatic, the method comprising: providing a plurality of nucleic acid molecules obtained from a sample from the subject, the plurality of nucleic acid molecules comprising a mixture of tumor and non-tumor nucleic acid molecules; optionally ligating one or more adapters to one or more nucleic acids from the plurality of nucleic acid molecules; amplifying nucleic acid molecules from the plurality of nucleic acid molecules; capturing nucleic acid molecules from the amplified nucleic acid molecules, wherein the captured nucleic acid molecules are captured from the amplified nucleic acid molecules by hybridization to one or more bait molecules; and sequencing the captured nucleic acid molecules with a sequencer to identify one or more bait molecules. obtaining a plurality of sequence reads corresponding to at least one genomic locus; selecting, by one or more processors, a genomic sequence of interest at a genomic locus from the one or more genomic loci; selecting, by the one or more processors, one or more proxy genomic sequences for the genomic sequence of interest and determining, by the one or more processors, an allele frequency distance using summary statistics or distributions indicative of observed allele frequencies of the genomic sequence of interest and the one or more proxy genomic sequences; and identifying, by the one or more processors, the genomic sequence of interest as germline or somatic using the allele frequency distances.

[0010] In some embodiments, the subject is a cancer patient. In some embodiments, the sample comprises a tissue biopsy sample, a liquid biopsy sample, a circulating tumor cell (CTC) sample, a cell-free DNA (cfDNA) sample, or a normal control. In some embodiments, the sample is a liquid biopsy sample and comprises blood, plasma, cerebrospinal fluid, sputum, stool, urine, or saliva. In some embodiments, the tumor nucleic acid molecule is derived from a tumor portion of a heterogeneous tissue biopsy sample, and the non-tumor nucleic acid molecule is derived from a normal portion of the heterogeneous tissue biopsy sample. In some embodiments, the tumor nucleic acid molecule is derived from a circulating tumor DNA (ctDNA) fraction of a cell-free DNA sample, and the non-tumor nucleic acid molecule is derived from a non-tumor fraction of the cell-free DNA sample. In some embodiments, the one or more adapters comprise amplification primers or sequencing adapters. In some embodiments, the one or more bait molecules comprise one or more nucleic acid molecules, each nucleic acid molecule comprising a region complementary to a region of a captured nucleic acid molecule. In some embodiments, amplifying the nucleic acid molecule comprises performing a polymerase chain reaction (PCR) or an isothermal amplification technique. In some embodiments, the sequencing involves the use of next-generation sequencing (NGS) technology. In some embodiments, the sequencing involves a next-generation sequencer. In some embodiments, one or more proxy genome sequences are located within a defined segment of the genome sequence of interest, and the selected genome sequence of interest is located within the same defined segment. In some embodiments, the genome sequence of interest is segmented into multiple segments based on the uniformity of copy numbers within each segment. In some embodiments, the summary statistic is the mean allele frequency or median allele frequency. In some embodiments, the allele frequency distance is determined using a distribution showing the observed allele frequencies of the genome sequence of interest and the observed frequencies of the multiple proxy genome sequences, and the genome sequence of interest is identified as germline or somatic based on the probability that the observed allele frequency of the genome sequence of interest fits or does not fit within the distribution.

[0011] In some embodiments, a method for identifying a genomic sequence of interest as germline or somatic includes: selecting, by one or more processors, a genomic sequence of interest at a genomic locus from within a patient genomic sequence obtained for a patient sample comprising a mixture of tumor and non-tumor nucleic acid molecules; selecting, by one or more processors, one or more proxy genomic sequences for the genomic sequence of interest; determining, by one or more processors, an allele frequency distance using summary statistics or distributions indicating the observed allele frequencies of the genomic sequence of interest and the observed allele frequencies of the one or more proxy genomic sequences; and identifying (e.g., classifying), by one or more processors, the genomic sequence of interest as germline or somatic using the allele frequency distance.

[0012] In some embodiments of the method, the summary statistic is a mean allele frequency or a median allele frequency. In some embodiments, an allele frequency distance is determined using a distribution showing the observed allele frequencies of the genomic sequence of interest and the observed frequencies of the multiple proxy genomic sequences, and the genomic sequence of interest is identified as germline or somatic based on the probability that the observed allele frequency of the genomic sequence of interest fits or does not fit within the distribution.

[0013] In some embodiments, the tumor nucleic acid molecules and the non-tumor nucleic acid molecules comprise DNA molecules. In some embodiments, the tumor nucleic acid molecules and the non-tumor nucleic acid molecules comprise RNA molecules.

[0014] In some embodiments, the method further comprises sequencing tumor and non-tumor nucleic acid molecules from the patient sample to determine a patient genomic sequence. In some embodiments, the patient genomic sequence is obtained or determined using next-generation sequencing technology. In some embodiments, the sequencer is a next-generation sequencer.

[0015] In some embodiments of the method, one or more proxy genome sequences are located within a defined segment of the patient genome sequence, and the selected genome sequence of interest is located within the same defined segment. In some embodiments, the patient genome sequence is segmented into multiple segments based on the uniformity of copy numbers within each segment. In some embodiments, the method comprises segmenting the patient genome sequence into multiple segments.

[0016] In some embodiments of the method, the patient genome sequence is determined using targeted sequencing. In some embodiments, the targeted sequencing comprises targeted sequencing of one or more genes or portions thereof associated with cancer. In some embodiments, the targeted sequencing comprises targeted sequencing of one or more exon regions.

[0017] In some embodiments, the method includes identifying, by one or more processors, a genomic sequence of interest in a patient sample at a genomic locus; identifying, by the one or more processors, one or more proxy genomic sequences for the sequence of interest; comparing, by the one or more processors, an observed frequency of the sequence of interest to a centrality measure of the observed frequencies of the one or more proxy genomic sequences; and identifying (e.g., classifying or characterizing) the genomic sequence of interest as either germline or somatic based on the comparison.

[0018] In some embodiments of the method, one or more proxy genome sequences comprise a single nucleotide polymorphism (SNP).

[0019] In some embodiments of the method, the one or more proxy genome sequences comprise an allele.

[0020] In some embodiments, the method further includes identifying, by the one or more processors, a segment of the patient's genome that includes the genomic locus. In some embodiments, identifying the segment, by the one or more processors, includes performing a segmentation procedure on a contiguous portion of the patient's genome. In some embodiments, the portion of the patient's genome is large enough to identify three distinct segments. In some embodiments, the proxy is identified by the one or more processors to be located on the same segment as the genomic locus. In some embodiments, the segmentation procedure identifies the segments according to whether a genomic parameter is equal throughout each individual segment. In some embodiments, the genomic parameter is copy number.

[0021] In some embodiments of any of the above methods for identifying a genomic sequence of interest as germline or somatic, the step of identifying the genomic sequence of interest as germline or somatic by one or more processors comprises: inputting the allele frequency distance into a trained statistical model; and outputting from the trained statistical model a value indicative of the likelihood that the genomic sequence of interest is germline or a value indicative of the likelihood that the genomic sequence of interest is somatic. In some embodiments, the allele frequency distance is adjusted to correct for contamination levels in patient samples, low sequencing read depth, noisy estimates of allele frequencies, low numbers of segmental germline single nucleotide polymorphisms (SNPs), or high variability in segmental germline SNP allele frequencies. In some embodiments, the trained statistical model comprises a function relating the allele frequency distance to a value indicative of the likelihood that the genomic sequence of interest is germline or a value indicative of the likelihood that the genomic sequence of interest is somatic.

[0022] In some embodiments, the trained statistical model is a logistic regression model. In some embodiments, the trained statistical model is trained using tumor samples with known germline sequences. In some embodiments, the trained statistical model is trained using data for tumor samples with known germline sequences and known somatic sequences. In some embodiments, the method further comprises training the statistical model using data for tumor samples with known germline sequences. In some embodiments, the method further comprises training the statistical model using data for tumor samples with known germline sequences and known somatic sequences.

[0023] In some embodiments, the trained statistical model is trained using data on variant allele frequencies that excludes variants located in genomic regions known to have allele frequencies that deviate from expected values. In some embodiments, the method further comprises training the statistical model using data on variant allele frequencies that excludes variants located in genomic regions known to have allele frequencies that deviate from expected values.

[0024] In some embodiments, the trained statistical model is trained using data that incorporates prior knowledge of the likelihood of a variant being a germline, a somatic variant, or a clonal hematopoietic with undetermined potential (CHIP) variant based on historical data or a database. In some embodiments, the method further comprises training the statistical model using data that incorporates prior knowledge of the likelihood of a variant being a germline, a somatic variant, or a clonal hematopoietic with undetermined potential (CHIP) variant based on historical data or a database.

[0025] In some embodiments, the trained statistical model is trained using data describing the noise level for a given variant call and its genomic context. In some embodiments, the method further comprises training the statistical model using data describing the noise level for a given variant call and its genomic context.

[0026] In some embodiments, one or more proxy genome sequences comprise single nucleotide polymorphisms (SNPs). In some embodiments, one or more proxy genome sequences comprise alleles. In some embodiments of the method, the genome sequence of interest comprises a genomic variant.

[0027] In some embodiments of the method, the method further includes generating, by the one or more processors, a report indicating the genomic sequence of interest as either germline or somatic. In some embodiments, the method includes transmitting the report, for example, to a healthcare provider. In some embodiments, the report is transmitted over a computer network or a peer-to-peer connection.

[0028] In some embodiments of any of the above methods, the patient sample is derived from a tissue biopsy comprising tumor tissue and non-tumor tissue. In some embodiments, the tissue biopsy is a solid tissue biopsy or a liquid biopsy. In some embodiments, the tissue biopsy is a liquid biopsy comprising blood, plasma, cerebrospinal fluid, sputum, stool, urine, or saliva. In some embodiments, the patient sample comprises cell-free DNA (cdDNA) obtained from the subject. In some embodiments, the patient sample comprises circulating tumor DNA (ctDNA) obtained from the subject.

[0029] Further described herein are methods for treating cancer in a patient, the methods comprising: identifying, by one or more processors, one or more genomic sequences of interest as somatic using any of the methods described above; selecting a cancer treatment modality based on the one or more identified somatic sequences; and treating the cancer with the selected cancer treatment modality. In some embodiments, the one or more identified somatic sequences are associated with the success of cancer treatment with the selected treatment modality. In some embodiments, the method comprises determining, by one or more processors, a microsatellite instability status of the cancer using the one or more identified somatic sequences; and selecting a cancer treatment modality based on the microsatellite instability status of the cancer. In some embodiments, the method comprises determining, by one or more processors, a tumor mutational burden for the cancer using the one or more identified somatic sequences; and selecting a cancer treatment modality based on the tumor mutational burden being above a predetermined tumor mutational burden threshold. In some embodiments, the cancer treatment modality comprises administering an effective amount of one or more anti-cancer agents to the patient if the tumor mutational burden is above the predetermined threshold. In some embodiments, the one or more anti-cancer agents comprise a cancer immunotherapeutic agent. In some embodiments, the cancer immunotherapeutic agent is an immune checkpoint inhibitor.

[0030] Also described herein are methods for monitoring the progression or recurrence of cancer in a patient, the methods comprising: identifying, by one or more processors, one or more genomic sequences of interest as somatic using any of the methods described above; and detecting, by the one or more processors, the presence or absence of the one or more genomic sequences of interest identified as somatic in a second patient sample obtained from the patient after the cancer has been treated. In some embodiments, the method comprises obtaining a second patient sample from the patient. In some embodiments, the method comprises treating the patient for cancer after the first patient sample has been obtained from the patient and before the second patient sample has been obtained from the patient. In some embodiments, the second patient sample comprises cell-free DNA. In some embodiments, detecting the presence or absence of the one or more genomic sequences of interest identified as somatic in the second patient sample comprises sequencing nucleic acid molecules in the second patient sample.

[0031] Further described herein is a method for selecting neoantigens for a cancer vaccine personalized to a subject with cancer, the method comprising: identifying, by one or more processors, one or more genomic sequences of interest as somatic using any of the methods described above, wherein the one or more genomic sequences of interest identified as somatic are located within an exon region of a gene; and selecting, by the one or more processors, from the one or more genomic sequences of interest identified as somatic, genomic sequences encoding neoantigens suitable as a cancer vaccine for the subject. In some embodiments, the method further comprises producing a vaccine comprising the neoantigen.

[0032] Also described herein are non-transitory computer-readable storage media storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of an electronic device, cause the electronic device to: select a genomic sequence of interest at a genomic locus from a patient genomic sequence obtained for a patient sample containing a mixture of tumor and non-tumor nucleic acid molecules; select one or more proxy genomic sequences for the genomic sequence of interest; determine an allele frequency distance using a summary statistic or distribution indicating the observed allele frequency of the genomic sequence of interest and the observed allele frequency of the one or more proxy genomic sequences; and identify the genomic sequence of interest as germline or somatic using the allele frequency distance. In some embodiments, the summary statistic is the mean allele frequency or median allele frequency. In some embodiments, the allele frequency distance is determined using a distribution indicating the observed allele frequency of the genomic sequence of interest and the observed frequencies of a plurality of proxy genomic sequences, and the genomic sequence of interest is identified as germline or somatic based on the probability that the observed allele frequency of the genomic sequence of interest fits or does not fit within the distribution. In some embodiments, the tumor nucleic acid molecules and the non-tumor nucleic acid molecules comprise DNA molecules. In some embodiments, the tumor nucleic acid molecules and the non-tumor nucleic acid molecules comprise RNA molecules.

[0033] In some embodiments of the non-transitory computer-readable storage medium, the one or more proxy genome sequences are located within a defined segment of the patient genome sequence, and the selected genome sequence of interest is located within the same defined segment. In some embodiments, the patient genome sequence is segmented into multiple segments based on copy number uniformity within each segment.

[0034] In some embodiments of the non-transitory computer-readable storage medium, the one or more programs further include instructions that, when executed by one or more processors of the electronic device, cause the electronic device to segment the patient genomic sequence into a plurality of segments.

[0035] In some embodiments of the non-transitory computer-readable storage medium, the patient genomic sequence is determined using targeted sequencing. In some embodiments, the patient genomic sequence is determined using next-generation sequencing. In some embodiments, the targeted sequencing comprises targeted sequencing of one or more genes or portions thereof associated with cancer. In some embodiments, the targeted sequencing comprises targeted sequencing of one or more exon regions.

[0036] In some embodiments, a non-transitory computer-readable storage medium stores one or more programs that include instructions that, when executed by one or more processors of an electronic device, cause the electronic device to identify a genomic sequence of interest in a patient sample at a genomic locus, identify one or more proxy genomic sequences for the sequence of interest, identify an observed frequency of the sequence of interest relative to a centrality measure of the observed frequencies of the one or more proxy genomic sequences, and characterize the genomic sequence of interest as either germline or somatic based on this comparison.

[0037] In some embodiments of the non-transitory computer-readable storage medium, the one or more programs further include instructions that, when executed by one or more processors of the electronic device, cause the electronic device to generate a report indicating the genomic sequence of interest as either germline or somatic. In some embodiments, the electronic device comprises a display, and the one or more programs further include instructions that, when executed by the one or more processors of the electronic device, cause the electronic device to display the report.

[0038] In some embodiments of the non-transitory computer-readable storage medium, the one or more proxy genome sequences comprise single nucleotide polymorphisms (SNPs).

[0039] In some embodiments of the non-transitory computer-readable storage medium, the one or more proxy genome sequences comprise alleles.

[0040] In some embodiments of the non-transitory computer-readable storage medium, the one or more programs further comprise instructions that, when executed by one or more processors of the electronic device, cause the electronic device to identify a segment of the patient's genome that includes the genomic locus. In some embodiments, identifying the segment comprises performing a segmentation procedure on a contiguous portion of the patient's genome. In some embodiments, the portion of the patient's genome is large enough to identify three distinct segments. In some embodiments, one or more proxy genome sequences are identified as being located on the same segment as the genomic locus. In some embodiments, the segmentation procedure identifies the segment according to whether a genomic parameter is equal across each individual segment. In some embodiments, the genomic parameter is copy number.

[0041] In some embodiments of the non-transitory computer-readable storage medium, the genomic sequence of interest comprises a genomic variant.

[0042] In some embodiments of the non-transitory computer-readable storage medium, one or more programs further comprise instructions that, when executed by one or more processors of the electronic device, cause the electronic device to receive sequencing data associated with a patient genome sequence. In some embodiments, one or more programs further comprise instructions that, when executed by one or more processors of the electronic device, cause the electronic device to assemble a patient genome sequence using the sequencing data. In some embodiments, one or more programs further comprise instructions that, when executed by one or more processors of the electronic device, cause the sequencer to sequence nucleic acid molecules from the patient sample, thereby obtaining sequencing data.

[0043] In some embodiments of the non-transitory computer-readable storage medium, the one or more programs further comprise instructions that, when executed by one or more processors of the electronic device, cause the electronic device to generate a report indicating the genomic sequence of interest as either germline or somatic. In some embodiments, the one or more programs further comprise instructions that, when executed by one or more processors of the electronic device, cause the electronic device to transmit the report using a computer network.

[0044] In some embodiments of the non-transitory computer-readable storage medium, the electronic device comprises a display, and the one or more programs further include instructions that, when executed by one or more processors of the electronic device, cause the electronic device to display the report.

[0045] In some embodiments of the non-transitory computer-readable storage medium, the one or more proxy genome sequences comprise single nucleotide polymorphisms (SNPs).

[0046] In some embodiments of the non-transitory computer-readable storage medium, the one or more proxy genome sequences comprise alleles.

[0047] In some embodiments of the non-transitory computer-readable storage medium, the genomic sequence of interest comprises a genomic variant.

[0048] Also disclosed herein is an electronic device comprising one or more processors and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including: instructions for selecting a genomic sequence of interest at a genomic locus from a patient genomic sequence obtained for a patient sample containing a mixture of tumor and non-tumor nucleic acid molecules; instructions for selecting one or more proxy genomic sequences for the genomic sequence of interest; instructions for determining an allele frequency distance using summary statistics or distributions indicative of the observed allele frequencies of the genomic sequence of interest and the one or more proxy genomic sequences; and instructions for identifying the genomic sequence of interest as germline or somatic using the allele frequency distance. In some embodiments, the summary statistic is a mean allele frequency or a median allele frequency. In some embodiments, the allele frequency distance is determined using a distribution showing the observed allele frequencies of the genomic sequence of interest and the observed frequencies of multiple proxy genomic sequences, and the genomic sequence of interest is identified as germline or somatic based on the probability that the observed allele frequency of the genomic sequence of interest fits or does not fit within the distribution. In some embodiments, the tumor nucleic acid molecule and the non-tumor nucleic acid molecule comprise DNA molecules. In some embodiments, the tumor nucleic acid molecule and the non-tumor nucleic acid molecule comprise RNA molecules. In some embodiments, the patient genomic sequence is determined using next-generation sequencing.

[0049] In some embodiments of the electronic device, the one or more proxy genome sequences are located within a defined segment of the patient genome sequence, and the selected genome sequence of interest is located within the same defined segment. In some embodiments, the patient genome sequence is segmented into multiple segments based on copy number uniformity within each segment. In some embodiments, the one or more programs further comprise instructions for segmenting the patient genome sequence into multiple segments.

[0050] In some embodiments of the electronic device, the patient genome sequence is determined using targeted sequencing. In some embodiments, the targeted sequencing comprises targeted sequencing of one or more genes or portions thereof associated with cancer. In some embodiments, the targeted sequencing comprises targeted sequencing of one or more exon regions.

[0051] In some embodiments, an electronic device comprises one or more processors and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions to identify a genomic sequence of interest in a patient sample at a genomic locus, instructions to identify one or more proxy genomic sequences for the sequence of interest, instructions to compare an observed frequency of the genomic sequence of interest to a centrality measure of the observed frequency of the one or more proxy genomic sequences, and instructions to identify the genomic sequence of interest as germline or somatic based on the comparison.

[0052] In some embodiments of the electronic device, one or more proxy genome sequences comprise a single nucleotide polymorphism (SNP).

[0053] In some embodiments of the electronic device, the one or more proxy genome sequences comprise alleles.

[0054] In some embodiments of the electronic device, the one or more programs further comprise instructions for identifying a segment of the patient's genome in which the genomic locus is included. In some embodiments, identifying the segment comprises performing a segmentation procedure on a contiguous portion of the patient's genome. In some embodiments, the portion of the patient's genome is large enough to identify three distinct segments. In some embodiments, the proxy is identified as being located on the same segment as the genomic locus. In some embodiments, the segmentation procedure identifies the segments according to whether a genomic parameter is equal across each individual segment. In some embodiments, the genomic parameter is copy number.

[0055] In some embodiments of the electronic device, the genomic sequence of interest comprises a genomic variant.

[0056] In some embodiments of the electronic device, the one or more programs further comprise instructions for receiving sequencing data associated with a patient genome sequence. In some embodiments, the one or more programs further comprise instructions for assembling the patient genome sequence using the sequencing data. In some embodiments, the one or more programs further comprise instructions for causing a sequencer to sequence nucleic acid molecules from the patient sample, thereby obtaining sequencing data.

[0057] In some embodiments of the electronic device, one or more proxy genome sequences comprise a single nucleotide polymorphism (SNP).

[0058] In some embodiments of the electronic device, the one or more proxy genome sequences comprise alleles.

[0059] In some embodiments of the electronic device, the genomic sequence of interest comprises a genomic variant.

[0060] In some embodiments of the electronic device, the one or more programs further include instructions for generating a report indicating the genomic sequence of interest as either germline or somatic. In some embodiments, the one or more programs further include instructions for transmitting the report over a computer network or a peer-to-peer connection. In some embodiments, the device further comprises a display, and the one or more programs further include instructions for displaying the report.

[0061] In some embodiments of the electronic device, the patient sample is derived from a tissue biopsy including tumor tissue and non-tumor tissue. In some embodiments, the tissue biopsy is a solid tissue biopsy or a liquid biopsy. In some cases, the tissue sample is a liquid biopsy including blood, plasma, cerebrospinal fluid, sputum, stool, urine, or saliva. In some embodiments, the patient sample includes cell-free DNA (cfDNA) obtained from the subject. In some embodiments, the patient sample includes circulating tumor DNA (ctDNA) obtained from the subject.

[0062] Also described herein are systems that include any of the electronic devices described herein and a sequencer configured to sequence nucleic acid molecules from a patient sample. In some embodiments, the sequencer is a next-generation sequencer.

[0063] Disclosed herein are methods for identifying a genomic sequence of interest as germline or somatic, the methods including: identifying, by one or more processors, the genomic sequence of interest in a patient sample at a genomic locus; identifying, by one or more processors, a proxy genomic sequence for the genomic sequence of interest; comparing, by one or more processors, the observed allele fraction of the genomic sequence of interest to the observed allele fraction of the proxy genomic sequence; and identifying, by the one or more processors, the genomic sequence of interest as germline or somatic based on the comparison. In some embodiments, the proxy genomic sequence has the same copy number as the genomic sequence of interest. In some embodiments, identifying, by the one or more processors, the genomic sequence of interest as germline or somatic includes inputting the allele frequency distance into a trained statistical model; and outputting, from the trained statistical model, a value indicative of the likelihood that the genomic sequence of interest is germline or a value indicative of the likelihood that the genomic sequence of interest is somatic. In some embodiments, the allele fraction of the genomic sequence and the allele fraction of the proxy genomic sequence are determined using next-generation sequencing technology. In some embodiments, the allelic fraction of the genomic sequence and the allelic fraction of the proxy genomic sequence are determined using microarray technology. In some embodiments, the patient sample comprises a solid tissue biopsy or a liquid biopsy. In some embodiments, the patient sample is a liquid biopsy comprising blood, plasma, cerebrospinal fluid, sputum, stool, urine, or saliva. In some embodiments, the patient sample comprises cell-free DNA (cfDNA) obtained from the subject. In some embodiments, the patient sample comprises circulating tumor DNA (ctDNA) obtained from the subject. In some embodiments, the patient is a cancer patient. In certain embodiments, for example, the following items are provided: (Item 1) 1. A method for identifying a genomic sequence of interest as germline or somatic, said method comprising: providing a plurality of nucleic acid molecules obtained from a sample from a subject, wherein the plurality of nucleic acid molecules comprises a mixture of tumor nucleic acid molecules and non-tumor nucleic acid molecules; Optionally, ligating one or more adaptors to one or more nucleic acids from said plurality of nucleic acid molecules; amplifying a nucleic acid molecule from said plurality of nucleic acid molecules; capturing nucleic acid molecules from the amplified nucleic acid molecules, wherein the captured nucleic acid molecules are captured from the amplified nucleic acid molecules by hybridization to one or more bait molecules; sequencing the captured nucleic acid molecules with a sequencer to obtain a plurality of sequence reads corresponding to one or more genomic loci; selecting, by one or more processors, a genomic sequence of interest at a genomic locus from said one or more genomic loci; selecting, by the one or more processors, one or more proxy genome sequences for the genome sequence of interest; determining, by the one or more processors, allele frequency distances using summary statistics or distributions indicative of the observed allele frequencies of the genome sequence of interest and the observed allele frequencies of the one or more proxy genome sequences; identifying, by the one or more processors, the genomic sequence of interest as germline or somatic using the allele frequency distance; A method comprising: (Item 2) Item 10. The method of item 1, wherein the subject is a cancer patient. (Item 3) 3. The method of claim 1 or claim 2, wherein the sample comprises a tissue biopsy sample, a liquid biopsy sample, a circulating tumor cell (CTC) sample, a cell-free DNA (cfDNA) sample, or a normal control. (Item 4) 4. The method of claim 3, wherein the sample is a liquid biopsy sample and comprises blood, plasma, cerebrospinal fluid, sputum, stool, urine, or saliva. (Item 5) 4. The method of any one of items 1 to 3, wherein the tumor nucleic acid molecule is derived from a tumor portion of a heterogeneous tissue biopsy sample, and the non-tumor nucleic acid molecule is derived from a normal portion of the heterogeneous tissue biopsy sample. (Item 6) 4. The method of any one of items 1 to 3, wherein the tumor nucleic acid molecule is derived from a circulating tumor DNA (ctDNA) fraction of a cell-free DNA sample, and the non-tumor nucleic acid molecule is derived from a non-tumor fraction of the cell-free DNA sample. (Item 7) 7. The method of any one of items 1 to 6, wherein the one or more adapters comprise an amplification primer or a sequencing adapter. (Item 8) 8. The method of any one of items 1 to 7, wherein the one or more bait molecules comprise one or more nucleic acid molecules, each nucleic acid molecule comprising a region complementary to a region of a captured nucleic acid molecule. (Item 9) 9. The method of any one of items 1 to 8, wherein amplifying the nucleic acid molecule comprises performing a polymerase chain reaction (PCR) or an isothermal amplification technique. (Item 10) 10. The method of any one of items 1 to 9, wherein the sequencing comprises the use of next-generation sequencing (NGS) technology. (Item 11) 11. The method of any one of items 1 to 10, wherein the sequencer comprises a next-generation sequencer. (Item 12) 12. The method according to any one of items 1 to 11, wherein the one or more proxy genome sequences are located within a defined segment of the subject genome sequence, and the selected genome sequence of interest is located within the same defined segment. (Item 13) 13. The method of claim 12, wherein the subject's genomic sequence is segmented into a plurality of segments based on the uniformity of copy number within each segment. (Item 14) 14. The method of any one of items 1 to 13, wherein the summary statistic is a mean allele frequency or a median allele frequency. (Item 15) 15. The method of any one of items 1 to 14, wherein the allele frequency distance is determined using a distribution showing the observed allele frequency of the genomic sequence of interest and the observed frequencies of a plurality of proxy genome sequences, and the genomic sequence of interest is identified as germline or somatic based on the probability that the observed allele frequency of the genomic sequence of interest fits or does not fit within the distribution. (Item 16) 1. A method for identifying a genomic sequence of interest as germline or somatic, said method comprising: selecting, by one or more processors, a genomic sequence of interest at a genomic locus from within a patient genomic sequence obtained for a patient sample comprising a mixture of tumor and non-tumor nucleic acid molecules; selecting, by the one or more processors, one or more proxy genome sequences for the genome sequence of interest; determining, by the one or more processors, allele frequency distances using summary statistics or distributions indicative of the observed allele frequencies of the genome sequence of interest and the observed allele frequencies of the one or more proxy genome sequences; identifying, by the one or more processors, the genomic sequence of interest as germline or somatic using the allele frequency distance; A method comprising: (Item 17) 17. The method of claim 16, comprising sequencing the tumor nucleic acid molecules and the non-tumor nucleic acid molecules from the patient sample using a sequencer to determine the patient genome sequence. (Item 18) 18. The method of claim 17, wherein the patient genomic sequence is obtained using next-generation sequencing technology. (Item 19) Item 18. The method of item 17, wherein the sequencer is a next-generation sequencer. (Item 20) 20. The method of any one of Items 16 to 19, wherein the one or more proxy genome sequences are located within a defined segment of the patient genome sequence, and the selected genome sequence of interest is located within the same defined segment. (Item 21) 21. The method of claim 20, wherein the patient genome sequence is segmented into a plurality of segments based on copy number uniformity within each segment. (Item 22) 22. The method of item 20 or 21, comprising segmenting the patient genome sequence into a plurality of segments. (Item 23) 23. The method of any one of items 16 to 22, wherein the summary statistic is a mean allele frequency or a median allele frequency. (Item 24) 24. The method of any one of Items 16 to 23, wherein the allele frequency distance is determined using a distribution showing the observed allele frequency of the genome sequence of interest and the observed frequencies of a plurality of proxy genome sequences, and the genome sequence of interest is identified as germline or somatic based on the probability that the observed allele frequency of the genome sequence of interest fits or does not fit within the distribution. (Item 25) 25. The method of any one of items 16 to 24, wherein the tumor nucleic acid molecule and the non-tumor nucleic acid molecule comprise DNA molecules. (Item 26) 26. The method of any one of items 16 to 25, wherein the tumor nucleic acid molecule and the non-tumor nucleic acid molecule comprise RNA molecules. (Item 27) 27. The method of any one of items 16 to 26, wherein the patient genomic sequence is determined using targeted sequencing. (Item 28) 28. The method of claim 27, wherein the targeted sequencing comprises targeted sequencing of one or more genes or portions thereof associated with cancer. (Item 29) 29. The method of claim 27 or 28, wherein the targeted sequencing comprises targeted sequencing of one or more exon regions. (Item 30) 1. A method for identifying a genomic sequence of interest as germline or somatic, said method comprising: identifying, by one or more processors, genomic sequences of interest in the patient sample at genomic loci; identifying, by the one or more processors, one or more proxy genome sequences for the sequence of interest; comparing, by the one or more processors, the observed frequency of the genome sequence of interest to a centrality measure of the observed frequencies of the one or more proxy genome sequences; identifying, by the one or more processors, the genomic sequence of interest as germline or somatic based on the comparison; and A method comprising: (Item 31) 31. The method of claim 30, further comprising identifying, by the one or more processors, a segment of the patient's genome that includes the genomic locus. (Item 32) 32. The method of claim 31, wherein identifying the segments by the one or more processors comprises performing a segmentation procedure on a contiguous portion of the patient's genome. (Item 33) 33. The method of claim 32, wherein the portion of the patient's genome is large enough to identify three distinct segments. (Item 34) 32. The method of claim 31, wherein the proxy is identified by the one or more processors to be located within the same segment as the genomic locus. (Item 35) 33. The method of claim 32, wherein the segmentation procedure identifies segments by the one or more processors according to whether genomic parameters are equal throughout each individual segment. (Item 36) 36. The method of item 35, wherein the genomic parameter is a copy number. (Item 37) identifying, by the one or more processors, the genomic sequence of interest as germline or somatic; inputting the allele frequency distances into a trained statistical model; outputting from the trained statistical model a value indicating the likelihood that the genomic sequence of interest is germline or a value indicating the likelihood that the genomic sequence of interest is somatic; 37. The method according to any one of items 16 to 36, comprising: (Item 38) 38. The method of any one of items 16 to 37, wherein the allele frequency distance is adjusted to correct for contamination levels in the patient sample, low sequencing read depth, noisy estimates of allele frequencies, low segment germline single nucleotide polymorphism (SNP) numbers, or high variability in segment germline SNP allele frequencies. (Item 39) Item 39. The method of Item 37 or Item 38, wherein the trained statistical model comprises a function relating the allele frequency distance to the value indicative of the likelihood that the genomic sequence of interest is germline or the value indicative of the likelihood that the genomic sequence of interest is somatic. (Item 40) 40. The method of any one of items 37 to 39, wherein the trained statistical model is a logistic regression model. (Item 41) 41. The method of any one of items 37 to 40, further comprising training the statistical model using data for tumor samples with known germline sequences. (Item 42) 42. The method of any one of items 37 to 41, further comprising training the statistical model using data for tumor samples with known germline sequences and known somatic sequences. (Item 43) 41. The method of any one of items 37 to 40, wherein the trained statistical model is trained using data for tumor samples with known germline sequences. (Item 44) 44. The method of claim 43, wherein the trained statistical model is trained using data for tumor samples with known germline sequences and known somatic sequences. (Item 45) 45. The method of any one of items 37 to 44, further comprising training the statistical model with data on variant allele frequencies that excludes variants located in genomic regions known to have allele frequencies that deviate from expected values. (Item 46) 45. The method of any one of items 37 to 44, wherein the trained statistical model is trained using data on variant allele frequencies that excludes variants located in genomic regions known to have allele frequencies that deviate from expected values. (Item 47) 47. The method of any one of items 37 to 46, further comprising training the statistical model using data that incorporates prior knowledge of the likelihood of the variant to be a germline, somatic variant, or clonal hematopoietic with undetermined potential (CHIP) variant based on historical data or a database. (Item 48) 47. The method of any one of items 37 to 46, wherein the trained statistical model is trained using data that incorporates prior knowledge of the likelihood of the variant being a germline, somatic variant, or clonal hematopoietic with undetermined potential (CHIP) variant based on historical data or a database. (Item 49) 49. The method of any one of items 37 to 48, further comprising training the statistical model with data that accounts for the noise level for a given variant call and its genomic context. (Item 50) 49. The method of any one of paragraphs 37 to 48, wherein the trained statistical model is trained using data that describes the noise level for a given variant call and its genomic context. (Item 51) 51. The method of any one of items 16 to 50, wherein the one or more proxy genome sequences comprise a single nucleotide polymorphism (SNP). (Item 52) 52. The method of any one of items 16 to 51, wherein the one or more proxy genome sequences comprise alleles. (Item 53) 53. The method of any one of Items 16 to 52, wherein the genomic sequence of interest comprises a genomic variant. (Item 54) 54. The method of any one of items 16 to 53, further comprising generating, by the one or more processors, a report indicating the genomic sequence of interest as germline or somatic. (Item 55) 55. The method of claim 54, further comprising sending the report to a healthcare provider. (Item 56) 56. The method of claim 54 or 55, wherein the report is transmitted over a computer network or a peer-to-peer connection. (Item 57) 57. The method of any one of items 16 to 56, wherein the patient sample is derived from a tissue biopsy comprising tumor tissue and non-tumor tissue. (Item 58) 58. The method of item 57, wherein the tissue biopsy is a solid tissue biopsy or a liquid biopsy. (Item 59) 59. The method of item 58, wherein the tissue biopsy is a liquid biopsy comprising blood, plasma, cerebrospinal fluid, sputum, stool, urine, or saliva. (Item 60) 60. The method of any one of items 16 to 59, wherein the patient sample comprises cell-free DNA (cfDNA) obtained from the subject. (Item 61) 61. The method of any one of items 16 to 60, wherein the patient sample comprises circulating tumor DNA (ctDNA) obtained from the subject. (Item 62) 1. A method of treating cancer in a patient, comprising: Identifying one or more genomic sequences of interest as somatic cells by the one or more processors using the method according to any one of items 16 to 61; selecting a cancer treatment modality based on the one or more identified somatic sequences; treating said cancer with said selected cancer treatment modality; and A method comprising: (Item 63) 63. The method of claim 62, wherein the one or more identified somatic sequences are involved in the success of cancer treatment with the selected therapeutic modality. (Item 64) determining, by the one or more processors, the microsatellite instability status of the cancer using the one or more identified somatic sequences; selecting the cancer treatment modality based on the microsatellite instability status of the cancer; Item 63. The method according to Item 62, comprising: (Item 65) determining, by the one or more processors, a tumor mutational burden for the cancer using the one or more identified somatic sequences; selecting the cancer treatment modality based on the tumor mutation burden being above a predetermined tumor mutation burden threshold; Item 63. The method according to Item 62, comprising: (Item 66) 66. The method of claim 64 or 65, wherein the cancer treatment modality comprises administering to the patient an effective amount of one or more anti-cancer agents if the tumor mutational burden is above a predetermined threshold. (Item 67) Item 67. The method of item 66, wherein the one or more anticancer agents comprise a cancer immunotherapeutic agent. (Item 68) 68. The method of item 67, wherein the cancer immunotherapeutic agent is an immune checkpoint inhibitor. (Item 69) 1. A method for monitoring the progression or recurrence of cancer in a patient, comprising: Identifying, by the one or more processors, one or more genomic sequences of interest as somatic using the method of any one of items 16 to 67, wherein the patient sample is obtained from a patient with cancer; and detecting, by the one or more processors, the presence or absence of the one or more genomic sequences of interest identified as somatic in a second patient sample obtained from the patient after the cancer has been treated; A method comprising: (Item 70) 70. The method of claim 69, comprising obtaining the second patient sample from the patient. (Item 71) 71. The method of claim 69 or claim 70, comprising treating the cancer in the patient after the first patient sample is obtained from the patient and before the second patient sample is obtained from the patient. (Item 72) 72. The method of any one of items 69 to 71, wherein the second patient sample comprises cell-free DNA. (Item 73) 73. The method of any one of items 69-72, wherein detecting the presence or absence of the one or more genomic sequences of interest identified as somatic in the second patient sample comprises sequencing nucleic acid molecules in the second patient sample. (Item 74) 1. A method for selecting neoantigens for a personalized cancer vaccine for a subject with cancer, comprising: Identifying, by the one or more processors, one or more genomic sequences of interest as somatic cells using the method of any one of items 16 to 67, wherein the one or more genomic sequences of interest identified as somatic cells are located within an exon region of a gene; selecting, by the one or more processors, from the one or more genomic sequences of interest identified as somatic, genomic sequences encoding neo-antigens suitable as a cancer vaccine for the subject; A method comprising: (Item 75) 75. The method of claim 74, further comprising producing a vaccine comprising the neoantigen. (Item 76) A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of an electronic device, cause the electronic device to: selecting a genomic sequence of interest at a genomic locus from within a patient genomic sequence obtained for a patient sample containing a mixture of tumor and non-tumor nucleic acid molecules; selecting one or more proxy genome sequences for the genome sequence of interest; determining an allele frequency distance using summary statistics or distributions indicative of the observed allele frequencies of the genome sequence of interest and the observed allele frequencies of the one or more proxy genome sequences; and The allele frequency distance is used to identify the genomic sequence of interest as germline or somatic. A non-transitory computer-readable storage medium. (Item 77) 77. The non-transitory computer-readable storage medium of Item 76, wherein the one or more proxy genome sequences are located within a defined segment of the patient genome sequence, and the selected genome sequence of interest is located within the same defined segment. (Item 78) 80. The non-transitory computer-readable storage medium of Item 77, wherein the patient genomic sequence is segmented into a plurality of segments based on copy number uniformity within each segment. (Item 79) 80. The non-transitory computer-readable storage medium of claim 77 or claim 78, wherein the one or more programs further comprise instructions that, when executed by the one or more processors of the electronic device, cause the electronic device to segment the patient genome sequence into a plurality of segments. (Item 80) 80. The non-transitory computer-readable storage medium of any one of items 76 to 79, wherein the summary statistic is a mean allele frequency or a median allele frequency. (Item 81) 81. The non-transitory computer-readable storage medium of any one of Items 76 to 80, wherein the allele frequency distance is determined using a distribution indicating the observed allele frequency of the genomic sequence of interest and the observed frequencies of a plurality of proxy genome sequences, and the genomic sequence of interest is identified as germline or somatic based on the probability that the observed allele frequency of the genomic sequence of interest fits or does not fit within the distribution. (Item 82) 82. The non-transitory computer-readable storage medium of any one of items 76 to 81, wherein the tumor nucleic acid molecule and the non-tumor nucleic acid molecule comprise DNA molecules. (Item 83) 83. The non-transitory computer-readable storage medium of any one of items 76 to 82, wherein the tumor nucleic acid molecule and the non-tumor nucleic acid molecule comprise RNA molecules. (Item 84) 84. The non-transitory computer-readable storage medium of any one of items 76 to 83, wherein the patient genomic sequence is determined using targeted sequencing. (Item 85) 85. The non-transitory computer-readable storage medium of any one of items 76 to 84, wherein the patient genomic sequence is determined using next generation sequencing. (Item 86) 86. The non-transitory computer-readable storage medium of claim 84 or claim 85, wherein the targeted sequencing comprises targeted sequencing of one or more genes or portions thereof associated with cancer. (Item 87) 87. The non-transitory computer-readable storage medium of any one of items 84 to 86, wherein the targeted sequencing comprises targeted sequencing of one or more exon regions. (Item 88) A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of an electronic device, cause the electronic device to: Identifying a genomic sequence of interest in a patient sample at a genomic locus; identifying one or more proxy genome sequences for said sequence of interest; Identifying the observed frequency of the sequence of interest relative to a centrality measure of the observed frequency of the one or more proxy genome sequences; and A non-transitory computer-readable storage medium that characterizes the genomic sequence of interest as either germline or somatic based on the comparison. (Item 89) Item 89. The non-transitory computer-readable storage medium of Item 88, wherein the one or more programs further comprise instructions that, when executed by the one or more processors of the electronic device, cause the electronic device to identify a segment of the patient's genome in which the genomic locus is included. (Item 90) Item 89. The non-transitory computer-readable storage medium of Item 88, wherein identifying the segments comprises performing a segmentation procedure on a contiguous portion of the patient's genome. (Item 91) 91. The non-transitory computer-readable storage medium of Item 90, wherein the portion of the patient's genome is large enough to identify three distinct segments. (Item 92) 92. The non-transitory computer-readable storage medium of any one of Items 88 to 91, wherein the one or more proxy genome sequences are identified as being located on the same segment as the genomic locus. (Item 93) 93. The non-transitory computer-readable storage medium of any one of items 90 to 92, wherein the segmentation procedure identifies segments according to whether genomic parameters are equal throughout each individual segment. (Item 94) Item 94. The non-transitory computer-readable storage medium of Item 93, wherein the genomic parameter is a copy number. (Item 95) 95. The non-transitory computer-readable storage medium of any one of items 76 to 94, wherein the one or more programs further comprise instructions that, when executed by one or more processors of the electronic device, cause the electronic device to receive sequencing data associated with the patient genomic sequence. (Item 96) 96. The non-transitory computer-readable storage medium of Item 95, wherein the one or more programs further comprise instructions that, when executed by one or more processors of the electronic device, cause the electronic device to assemble the patient genome sequence using the sequencing data. (Item 97) 97. The non-transitory computer-readable storage medium of claim 95 or claim 96, wherein the one or more programs further comprise instructions that, when executed by one or more processors of the electronic device, operate a sequencer to sequence nucleic acid molecules from the patient sample, thereby obtaining the sequencing data. (Item 98) 98. The non-transitory computer-readable storage medium of any one of items 76 to 97, wherein the one or more programs further comprise instructions that, when executed by the one or more processors of the electronic device, cause the electronic device to generate a report indicating the genomic sequence of interest as either germline or somatic. (Item 99) 99. The non-transitory computer-readable storage medium of any one of items 76 to 98, wherein the one or more programs further include instructions that, when executed by the one or more processors of the electronic device, cause the electronic device to transmit the report using a computer network. (Item 100) 99. The non-transitory computer-readable storage medium of any one of items 76 to 99, wherein the electronic device comprises a display, and the one or more programs further include instructions that, when executed by the one or more processors of the electronic device, cause the electronic device to display the report. (Item 101) 101. The non-transitory computer-readable storage medium of any one of Items 76 to 100, wherein the one or more proxy genome sequences comprise single nucleotide polymorphisms (SNPs). (Item 102) 102. The non-transitory computer-readable storage medium of any one of Items 76 to 101, wherein the one or more proxy genome sequences comprise alleles. (Item 103) 103. The non-transitory computer-readable storage medium of any one of items 76 to 102, wherein the genomic sequence of interest comprises a genomic variant. (Item 104) 1. An electronic device comprising: one or more processors; a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs comprising: instructions for selecting a genomic sequence of interest at a genomic locus from within a patient genomic sequence obtained for a patient sample comprising a mixture of tumor and non-tumor nucleic acid molecules; instructions for selecting one or more proxy genome sequences for the genome sequence of interest; instructions for determining an allele frequency distance using summary statistics or distributions indicative of the observed allele frequencies of the genome sequence of interest and the observed allele frequencies of the one or more proxy genome sequences; and a memory storing one or more programs including instructions for using the allele frequency distance to identify the genomic sequence of interest as germline or somatic; An electronic device comprising: (Item 105) Item 105. The electronic device of item 104, wherein the one or more proxy genome sequences are located within a defined segment of the patient genome sequence, and the selected genome sequence of interest is located within the same defined segment. (Item 106) Item 106. The electronic device of Item 105, wherein the patient genome sequence is segmented into a plurality of segments based on copy number uniformity within each segment. (Item 107) 107. The electronic device of any one of items 104 to 106, wherein the one or more programs further comprise instructions for segmenting the patient genome sequence into a plurality of segments. (Item 108) 108. The electronic device of any one of items 104 to 107, wherein the summary statistic is a mean allele frequency or a median allele frequency. (Item 109) 109. The electronic device of any one of Items 104 to 108, wherein the allele frequency distance is determined using a distribution indicating the observed allele frequency of the genomic sequence of interest and the observed frequencies of a plurality of proxy genome sequences, and the genomic sequence of interest is identified as germline or somatic based on the probability that the observed allele frequency of the genomic sequence of interest fits or does not fit within the distribution. (Item 110) 110. The electronic device according to any one of items 104 to 109, wherein the tumor nucleic acid molecule and the non-tumor nucleic acid molecule comprise DNA molecules. (Item 111) 111. The electronic device according to any one of items 104 to 110, wherein the tumor nucleic acid molecule and the non-tumor nucleic acid molecule comprise RNA molecules. (Item 112) 112. The electronic device of any one of items 104 to 111, wherein the patient genome sequence is determined using next generation sequencing. (Item 113) 113. The electronic device of any one of items 104 to 112, wherein the patient genome sequence is determined using targeted sequencing. (Item 114) Item 114. The electronic device of item 113, wherein the targeted sequencing comprises targeted sequencing of one or more genes or portions thereof associated with cancer. (Item 115) 115. The electronic device of claim 113 or 114, wherein the targeted sequencing comprises targeted sequencing of one or more exon regions. (Item 116) 1. An electronic device comprising: one or more processors; a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs comprising: instructions for identifying a genomic sequence of interest in a patient sample at a genomic locus; instructions for identifying one or more proxy genome sequences for said sequence of interest; instructions for comparing the observed frequency of the genome sequence of interest to a centrality measure of the observed frequency of the one or more proxy genome sequences; and a memory storing one or more programs including instructions for identifying the genomic sequence of interest as germline or somatic based on the comparison; An electronic device comprising: (Item 117) Item 117. The electronic device of item 116, wherein the one or more programs further comprise instructions for identifying a segment of the patient's genome in which the genomic locus is included. (Item 118) Item 118. The electronic device of Item 117, wherein identifying the segments comprises performing a segmentation procedure on a contiguous portion of the patient's genome. (Item 119) Item 119. The electronic device of item 118, wherein the portion of the patient's genome is large enough to identify three distinct segments. (Item 120) 120. The electronic device of any one of Items 117 to 119, wherein the one or more proxy genome sequences are identified as being located within the same segment as the genomic locus. (Item 121) 121. The electronic device of any one of items 118 to 120, wherein the segmentation procedure identifies segments according to whether genomic parameters are equal throughout each individual segment. (Item 122) Item 122. The electronic device of item 121, wherein the genomic parameter is a copy number. (Item 123) 123. The electronic device of any one of items 104 to 122, wherein the one or more programs further comprise instructions for receiving sequencing data associated with the patient genome sequence. (Item 124) Item 124. The electronic device of item 123, wherein the one or more programs further comprise instructions for assembling the patient genome sequence using the sequencing data. (Item 125) 125. The electronic device of claim 123 or 124, wherein the one or more programs further comprise instructions for causing a sequencer to sequence nucleic acid molecules from the patient sample, thereby obtaining the sequencing data. (Item 126) 126. The electronic device according to any one of items 104 to 125, wherein the one or more proxy genome sequences comprise a single nucleotide polymorphism (SNP). (Item 127) 127. The electronic device of any one of items 104 to 126, wherein the one or more proxy genome sequences comprise alleles. (Item 128) 128. The electronic device according to any one of Items 104 to 127, wherein the genomic sequence of interest comprises a genomic variant. (Item 129) 129. The electronic device of any one of items 104 to 128, wherein the one or more programs further comprise instructions for generating a report indicating the genomic sequence of interest as either germline or somatic. (Item 130) Item 130. The electronic device of item 129, wherein the one or more programs further include instructions for transmitting the report over a computer network or a peer-to-peer connection. (Item 131) Item 131. The electronic device of item 129 or 130, wherein the device further comprises a display and the one or more programs further comprise instructions for displaying the report. (Item 132) Item 104 to 106, wherein the patient sample is derived from a tissue biopsy containing tumor tissue and non-tumor tissue. 131. An electronic device according to any one of claims 131. (Item 133) Item 133. The electronic device of item 132, wherein the tissue biopsy is a solid tissue biopsy or a liquid biopsy. (Item 134) Item 134. The electronic device of item 133, wherein the tissue biopsy is a liquid biopsy comprising blood, plasma, cerebrospinal fluid, sputum, stool, urine, or saliva. (Item 135) 135. The electronic device of any one of items 104 to 134, wherein the patient sample comprises cell-free DNA (cfDNA) obtained from the subject. (Item 136) 136. The electronic device of any one of items 104 to 135, wherein the patient sample comprises circulating tumor DNA (ctDNA) obtained from the subject. (Item 137) 137. A system comprising the electronic device of any one of items 104 to 136 and a sequencer configured to sequence nucleic acid molecules derived from the patient sample. (Item 138) Item 138. The system of item 137, wherein the sequencer is a next-generation sequencer. (Item 139) 1. A method for identifying a genomic sequence of interest as germline or somatic, said method comprising: identifying, by one or more processors, genomic sequences of interest in the patient sample at genomic loci; identifying, by the one or more processors, proxy genome sequences for the genome sequence of interest; comparing, by the one or more processors, the observed allele fraction of the genome sequence of interest to the observed allele fraction of the proxy genome sequence; identifying, by the one or more processors, the genomic sequence of interest as germline or somatic based on the comparison; and A method comprising: (Item 140) 140. The method of claim 139, wherein the proxy genome sequence has the same copy number as the genome sequence of interest. (Item 141) identifying, by the one or more processors, the genomic sequence of interest as germline or somatic; inputting the allele frequency distances into a trained statistical model; outputting from the trained statistical model a value indicating the likelihood that the genomic sequence of interest is germline or a value indicating the likelihood that the genomic sequence of interest is somatic; The method according to item 139 or 140, comprising: (Item 142) 142. The method of any one of items 139 to 141, wherein the allelic fraction of the genome sequence and the allelic fraction of the proxy genome sequence are determined using next generation sequencing technology. (Item 143) 143. The method of claim 142, wherein the allelic fraction of the genomic sequence and the allelic fraction of the proxy genomic sequence are determined using microarray technology. (Item 144) 144. The method of any one of items 139 to 143, wherein the patient sample comprises a solid tissue biopsy or a liquid biopsy. (Item 145) 145. The method of claim 144, wherein the patient sample is a liquid biopsy comprising blood, plasma, cerebrospinal fluid, sputum, stool, urine, or saliva. (Item 146) 146. The method of any one of items 139 to 145, wherein the patient sample comprises cell-free DNA (cfDNA) obtained from the subject. (Item 147) 147. The method of any one of items 139 to 146, wherein the patient sample comprises circulating tumor DNA (ctDNA) obtained from the subject. (Item 148) 148. The method according to any one of items 139 to 147, wherein the patient is a cancer patient.

[0064] Incorporation by Reference All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference in their entirety to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference in its entirety. In the event of a conflict between a term in this specification and a term in an incorporated reference, the term in this specification shall control. [Brief explanation of the drawings]

[0065] [Figure 1] Figure 1 is a schematic diagram of a section of a patient's genome.

[0066] [Figure 2] FIG. 2 is a flow chart of the process for distinguishing between germline and somatic genomic sequences.

[0067] [Figure 3] FIG. 3 is a schematic diagram of genome segmentation.

[0068] [Figure 4] FIG. 4 illustrates an exemplary system including an electronic device that can be used to perform the methods described herein.

[0069] [Figure 5A] FIG. 5A shows an exemplary process for determining the difference in variant allele fraction expected for somatic and germline variants given the same tumor fraction, ploidy, and copy number.

[0070] [Figure 5B] Figure 5B shows an exemplary method for determining allele frequency distance from expected germline allele frequencies (AFDIS) and an exemplary density distribution of AFDIS, from which an empirical cumulative distribution function (ECDF) can be constructed.

[0071] [Figure 5C] FIG. 5C shows an exemplary plot of AFDIS plotted against the calculated purity of a tumor sample.

[0072] [Figure 5D] FIG. 5D shows a non-limiting example of an ROC curve for classification of somatic and germline variants in a tumor sample according to the methods disclosed herein.

[0073] [Figure 5E] FIG. 5E shows a non-limiting example of a probability plot of an exemplary logistic regression model that may be used in some embodiments.

[0074] [Figure 5F] FIG. 5F shows a plot of the somatic probabilities of different variants determined using an exemplary logistic regression model.

[0075] [Figure 5G] FIG. 5G shows the improvement of the claimed method over the conventional SGZ method.

[0076] [Figure 5H] FIG. 5H shows a non-limiting example of a sensitivity plot of training and testing data used to train and test a logistic regression model according to the exemplary methods disclosed herein.

[0077] [Figure 5I] FIG. 5I shows a non-limiting example of a positive predictive value (PPV) plot for training and test data used to train and test a logistic regression model according to the exemplary methods disclosed herein.

[0078] [Figure 5J] FIG. 5J shows a non-limiting example of data for classification of variants in the BRCA1 and BRCA2 genes using an exemplary embodiment of the described method.

[0079] [Figure 5K] FIG. 5K shows a non-limiting example of data for classification of variants in the STH11 gene using an exemplary embodiment of the described method.

[0080] [Figure 6A] FIG. 6A shows a non-limiting example of a plot of variant allele frequency (AF) versus segment minor allele frequency (MAF) for known germline variants in a tumor sample.

[0081] [Figure 6B] FIG. 6B shows non-limiting examples of density versus variant AF plots corresponding to segment MAF values ​​of 0.1, 0.2, and 0.3, respectively, derived from the data plotted in FIG. 6A. DETAILED DESCRIPTION OF THE INVENTION

[0082] Detailed Description The present specification describes a method, a device, and a computer-readable medium for distinguishing somatic genomic sequences from germline genomic sequences.The genomic sequence of interest in a patient sample can be identified at a genomic locus.Then, one or more proxy genomic sequences can be identified for the sequence of interest.The observed frequency of the sequence of interest can be compared with the centrality measure of the observed frequency of one or more proxy genomic sequences, and based on this comparison, the genomic sequence of interest can be characterized as either a germline sequence or a somatic sequence.

[0083] Several methods have been developed in the past to determine the somatic / germline status of variants in single-sample settings, including matching against public germline databases such as dbSNP or using surrogates constructed from a large number of normal individuals in place of matched normals. See, for example, Hiltemann et al., "Discriminating somatic and germline mutations in tumor DNA samples without matching normals," Genome Res., Vol. 25, No. 9, pp. 1382-1390 (2015). However, such methods are ineffective when dealing with rare germline variants restricted to families or small populations. So-called "basic methods" also exist that consider variants with allele frequencies (or allele fractions) approaching 50% or 100% to be germline and classify those that do not meet this criterion as somatic. Jones et al., "Personalized genomic analyses for cancer mutation discovery and interpretation," Sci. Transl. Med., Vol. 7, No. 283, pp. 283ra53 (2015). This basic method fails to account for the fact that aneuploidy can significantly shift the allele frequency of a germline variant away from the expected 50% or 100%. The terms "allele frequency" and "allelic fraction" are used interchangeably herein and refer to the proportion of sequence reads corresponding to a particular allele relative to the total number of sequence reads for a genomic locus.

[0084] The SGZ (somatic germline zygosity) algorithm, released in early 2018, attempted to provide a solution to the single-sample somatic / germline classification problem by considering tumor content, tumor ploidy, and local copy number. SGZ was demonstrated to significantly outperform "basic methods" in somatic / germline calling accuracy in a validation dataset (Sun et al., "A computational approach to distinguish somatic vs. germline origin of genomic alterations from deep sequencing of cancer specimens without a matched normal," PLoS Comput Biol., Vol. 14, No. 2, e1005965 (2018), which is incorporated herein by reference in its entirety). The application of the SGZ algorithm to FMI's massively parallel sequencing (MPS)-based diagnostic products enabled effective somatic / germline status determination for short variants (substitutions and indels), making it an essential tool for applications such as tumor mutation burden (TMB) estimation.

[0085] The method described herein for somatic / germline classification represents a further improvement over the SGZ approach. The new approach is based on the same fundamental principle: in a tumor / normal mixture, somatic and germline variants often have different expected allele frequencies determined by tumor fraction, tumor ploidy, and local copy number. However, in contrast to SGZ, which estimates expected germline allele frequencies through computational modeling of tumor fraction, tumor ploidy, and local copy number, the new method disclosed herein directly infers expected germline allele frequencies from known germline SNPs located on the same copy number segment as the variant in question. Therefore, using the method described herein, it is not necessary to determine or model copy number or tumor purity to obtain accurate calls for somatic and germline variants.

[0086] In some embodiments, a trained model, such as a logistic regression model, is used to predict the probability that a variant is somatic based on the difference between the observed variant allele frequency and the estimated expected germline variant allele frequency. In some embodiments, the model is trained using data from matched tumor / normal pairs and validated on an independent dataset. In some embodiments, the model is trained using data from tumor samples with known germline (and, optionally, known somatic) sequences. In some embodiments, the model is trained using data from mixed tumor / normal samples with known germline (and, optionally, known somatic) sequences. This validation shows that the new classifier outperforms SGZ in sensitivity and positive predictive value (PPV) for somatic variant classification.

[0087] The determined genomic sequence can be a somatic variant sequence or a germline sequence. Publicly accessible databases of known germline sequences exist (e.g., dbSNP ( 1671675357232_0 (See, for example, gnomAD (available at gnomAD.broadinstitute.org)), and a match between a known germline sequence and a sequence determined by sequencing nucleic acids in a sample obtained from a subject indicates that the sequence associated with the sample is likely a germline sequence. However, a lack of a match to a known germline sequence does not demonstrate that the sequence is a somatic variant sequence, as it may be a previously unknown (or unrecorded) germline sequence for the subject. The methods described herein allow for the classification of a sequence as a germline sequence or a somatic variant sequence.

[0088] Methods for calling somatic or germline sequences The methods described herein allow for the identification of a genomic sequence of interest as a germline sequence or a somatic sequence. In some embodiments, the somatic sequence is associated with a patient's cancer. For example, a patient sample may contain a mixture of tumor nucleic acid molecules (i.e., nucleic acid molecules derived from a tumor, either directly (such as in the case of a tumor biopsy) or indirectly (such as in the case of a liquid biopsy or body fluid sample containing circulating tumor DNA (ctDNA) and cell-free DNA (cfDNA))) and non-tumor nucleic acid molecules (i.e., nucleic acid molecules derived from a non-tumorous, preferably healthy, tissue, cell, liquid biopsy sample, or body fluid sample). The method may include selecting a genomic sequence of interest from within the patient genomic sequence (i.e., the genomic sequence obtained for the patient, which may be the entire genome or a portion thereof (e.g., an exome or target region within the entire genome)) and selecting one or more proxy genomic sequences for the genomic sequence of interest. The patient genomic sequence may include one or more alleles at any given locus (e.g., the somatic and / or germline sequence at any given locus).

[0089] Nucleic acid molecules from a sample (e.g., a mixed tumor / normal tissue sample or a cell-free DNA (cfDNA) sample containing a mixture of ctDNA and non-tumor cfDNA) can be sequenced to determine the patient's genome sequence. A genomic sequence of interest can be identified or selected at a genomic locus from the patient's genome sequence. The selected genomic sequence is a test sequence characterized as germline or somatic. In some embodiments, the genomic sequence of interest differs from a reference sequence. In some embodiments, the genomic sequence of interest differs from a sequence in a selected germline sequence database.

[0090] FIG. 1 is a schematic diagram of a sample genomic region. Region 100 may include an organism's entire genome or may include only a fraction of the entire genome. While region 100 is shown as a continuous line in FIG. 1 , in general, region 100 may include several components physically separated on an organism's chromosome(s). In some implementations, the sample from which region 100 is determined may include normal patient tissue, a fluid containing normal cells or cell-free DNA, or other anatomical material. In some implementations, the sample may include abnormal (e.g., cancerous or genetically mutated) tissue, a fluid containing abnormal cells or circulating tumor DNA, or other anatomical material. In some implementations, the sample may include a combination of normal and abnormal tissue, a bodily fluid, or other anatomical material.

[0091] Genomic region 100 shown in Figure 1 may correspond to a single strand or fragment of DNA, or a strand or fragment of RNA. Although not shown in Figure 1, region 100 includes a sequence of various bases (i.e., cytosine ("C"), guanine ("G"), adenine ("A"), thymine ("T"), or uracil ("U")). The particular sequence of bases can often determine important characteristics of an anatomical material or patient, such as whether the patient has cancer and, if so, which treatments may be effective or ineffective for the treatment.

[0092] The techniques described below involve characterizing a sequence of interest 102 within a genomic region 100 as either germline or somatic. The characterization is aided by the use of a reference sequence 104. The reference sequence 104 is an exemplary genomic sequence representing a "normal" (e.g., non-cancerous) patient. In some implementations, the reference sequence 104 may include a sequence determined by the Human Genome Project, such as hg19.

[0093] The reference sequence 104 contains known polymorphic regions 106a, 106b. The polymorphic regions 106a, 106b are regions (containing any number of bases, from a single base to hundreds or more) where variation in a particular organism's genomic sequence is expected across a population of organisms without corresponding adverse effects. For example, humans have polymorphic regions that correspond to various hair colors, eye colors, or other individualized characteristics. The genomic region 100 corresponding to an actual patient sample has specific base values ​​108a, 108b at positions in the region 100 that correspond to the polymorphic regions 106a, 106b in the reference sequence 104. In other words, the polymorphic regions 106a, 106b in the reference sequence 104 are positions at which a person's specific characteristics (e.g., hair color) are determined. The base values ​​108a, 108b are the individualized determination of those characteristics (e.g., red hair) that describe a particular patient.

[0094] Optionally, the polymorphic regions 106a, 106b include one or more single nucleotide polymorphisms (or "SNPs"). Optionally, the polymorphic regions can include an entire allele or a portion thereof.

[0095] 2 is a flowchart of a process for distinguishing between germline and somatic genomic sequences. Process 200 begins with identifying (i.e., selecting or sorting) a genomic region of interest (step 202). In some implementations, step 202 includes identifying a region of interest (i.e., a sequence of interest) 102 within a larger genomic region 100.

[0096] Determining a genomic sequence (e.g., genomic region 100) from a physical sample can be accomplished in a variety of ways. One such method is described in U.S. Patent No. 9,340,830, and another is described in U.S. Patent Publication No. 2017 / 0356053, both of which are incorporated herein by reference in their entireties. More generally, there is a category of machines operable to determine the genetic sequence of an input sample called genome sequencers. In some cases, the disclosed methods and systems can be implemented using any of a variety of next-generation sequencing (NGS) technologies and sequencers, including circular array sequencers and single-molecule sequencers configured for massively parallel sequencing. Additionally, there are various known subregions of the genomes of humans and other organisms known to be associated with various medical conditions.

[0097] The techniques described herein are not dependent on the use of a particular sequencing platform or a particular sequencing technology; any of these machines and associated technologies can be used in step 202. In some cases, the disclosed methods can be implemented using alternative nucleic acid sequence analysis techniques, such as microarrays, and fluorescence in situ hybridization (FISH).

[0098] In some implementations, the region of interest (i.e., sequence) 102 is identified as corresponding to a known genetic locus within the reference genome 104. In some implementations, the region of interest 102 corresponds to a mutation to the reference sequence 104 (i.e., a subsection of the genomic region 100 other than a polymorphic region having a genetic sequence that differs from the genetic sequence of the corresponding portion of the reference sequence 104). In some implementations, the sequence of interest corresponds to a gene associated with a medical condition possessed by a patient. In some implementations, the region of interest 102 is an oncogene or a portion thereof.

[0099] In step 204, one or more proxy genome sequences for the genome sequence are identified (step 204). The one or more selected proxy genome sequences can be known germline sequences (e.g., based on matching known germline sequences from a database of known germline sequences, or by sequencing healthy tissue, cells, or cell-free DNA from the subject or another healthy individual). Referring to FIG. 1 , one characteristic of proxy 110 is a sequence at a locus that (a) is known to encode germline genetic information and (b) is known to have the same copy number as sequence of interest 102 (e.g., by being confirmed to be physically close to or located within the same copy number segment as sequence of interest 102). An alternative characterization requires that proxy 110 be known to encode somatic genetic information. For convenience, this document assumes that proxy 110 encodes germline information unless otherwise specified, but one of skill in the art will understand the equivalence of the two approaches.

[0100] The germline status of a particular surrogate sequence candidate can be determined from the research literature, publicly available databases (e.g., dbSNP ( 1671675357232_1 Somatic variants may be known from the GnomAD (available at gnomAD.broadinstitute.org) or gnomAD (available at gnomAD.broadinstitute.org) or may be discovered by other ab initio means. On the other hand, somatic variants can be identified from matched tumor / normal samples, i.e., samples from the same patient containing both tumor and non-tumor ("normal") DNA. In particular, variants found in tumor DNA but not in the corresponding normal DNA are necessarily somatic. Known somatic variants can also be discovered by other ab initio means.

[0101] Referring to FIG. 3 , in some implementations, step 204 is performed by employing a segmentation process. In such a process, a portion 100 of a patient's genome is divided into segments (depicted by dashed vertical lines in FIG. 3 ) based on genetic parameters. Segments are defined such that the parameter values ​​within a particular segment are all equal (i.e., within a desired range or within a desired threshold). For example, a segment may be a contiguous sequence with approximately the same sequencing depth or copy number (i.e., within a desired range or within a desired threshold). In some implementations, the genetic parameters used to segment the input include copy number, or frequency of an allelic or suballelic segment of interest, etc. One or more proxy sequences may be located within the same segment as the genome sequence of interest, thus greatly increasing the likelihood that the one or more proxy genome sequences and the genome sequence of interest have the same copy number.

[0102] A variety of segmentation procedures are known in the art. For example, iSeg (described in Girimurugan et al., "iSeg: an Efficient Algorithm for Segmentation of Genomic and Epigenomic Data," BMC Bioinformatics, Vol. 19:131 (2018), which is incorporated herein in its entirety), CBS (described in Olshen et al., "Circular Binary Segmentation for the Analysis of Array-Based DNA Copy Number Data," Biostatistics, October 2004; Vol. 5(4):557-72, which is incorporated herein in its entirety), SLM Suite (described in Orlandini et al., "SLMSuite: A Suite of Algorithms for Segmenting Genomic Profiles," BMC Bioinformatics, Vol. 18:321 (2017), which is incorporated herein in its entirety), Pelt (described in Killick et al., "Optimal detection of changepoints with a linear computational "The cost" in Journal of the American Statistical Association, Vol. 107:500 (2012) are four of many such algorithms. In some embodiments, the patient genomic sequence is segmented into multiple segments based on copy number uniformity within each segment.

[0103] Referring again to FIG. 2 , in some implementations, only proxies 110 that are on the same segment as the region of interest 102 are identified. In some implementations, the proxies 110 include all known germline SNPs that are on the same segment as the region of interest 102. In some implementations, the proxies 110 include all known germline alleles that are on the same segment as the region of interest 102. In some implementations, for example, when it is difficult to correctly segment the genomic sequence into segments corresponding to different copy numbers, only proxies 110 that are no more than a predetermined number of bases away from the region of interest 102 are identified. For example, in some cases, the maximum number of bases separating the region of interest from the proxy sequence can range from about 10 bases to about 1,000 bases. In some cases, the maximum number of bases separating the region of interest from the proxy sequence can be about 10 bases, 20 bases, 30 bases, 40 bases, 50 bases, 60 bases, 70 bases, 80 bases, 90 bases, 100 bases, 200 bases, 300 bases, 400 bases, 500 bases, 600 bases, 700 bases, 800 bases, 900 bases, or 1,000 bases. In some cases, the maximum number of bases separating the region of interest from the proxy sequence can have any value within the range of values ​​described in this paragraph.

[0104] In step 206, the frequency of the proxies 110 is identified. In step 208, the allele frequencies (allele fractions) of sequences from the region of interest (i.e., the genomic sequence of interest) 102 are identified. Here, "frequency" refers to a normalized statistical frequency, e.g., the number of occurrences of a sequence or proxy in a sample divided by the total number of occurrences of any sequence at the same genomic locus. In some implementations, several frequency measurements may be performed. The allele frequencies of the genomic sequence of interest and one or more proxy genomic sequences can be determined by sequencing nucleic acid molecules in a sample from the subject. In some cases, the allele frequencies may be determined using other methodologies, such as microarrays or fluorescent in situ hybridization (FISH) techniques. When using several proxies, outlier proxy frequencies may be discarded, and the remaining frequencies may be combined into a single statistical centrality measure (e.g., a summary statistic such as the mean, median, or mode, or a distribution (e.g., a probability distribution) of the allele frequencies of the proxy sequences), so that step 210 involves a single numerical comparison. For example, in some embodiments, the centrality measure (summary statistic) is the mean allele frequency for one or more proxy sequences. In some embodiments, the centrality measure (summary statistic) is the median allele frequency for one or more proxy sequences. When a single proxy genome sequence is used, the centrality measure of the observed frequency of the proxy genome sequence is the frequency of that proxy sequence. The centrality measure, in some embodiments, can be the distribution of the observed allele frequencies for the proxy sequences.

[0105] In decision 210, the proxy frequency or frequencies (e.g., a centrality measure of the observed frequencies of one or more proxy sequences) are compared to the frequency or frequencies of the region of interest to determine whether they are equal. Throughout this specification and application, the term "equal" includes "equal to within a desired range" or "equal to within a desired threshold," which can be routinely determined based on the desired selectivity and specificity of process 200. The range or threshold can be set, for example, using a statistical threshold or statistical test selected by one of skill in the art. If, instead of combining proxy frequencies as described above, several proxies 110 are used and individual comparisons are made, then decision 210 will be "yes" if a certain percentage of the comparisons (e.g., greater than 50%, greater than 55%, greater than 60%, greater than 65%, greater than 70%, greater than 75%, greater than 80%, greater than 85%, greater than 90%, or greater than 95%) are equal.

[0106] If the proxy frequency is equal to the frequency of the sequence of interest, the sequence of interest is classified as germline (step 212). Otherwise, the sequence of interest is classified as somatic (step 214). Alternatively, if the proxy 110 is selected to be known to encode somatic information (instead of germline), then equal frequencies are interpreted as the sequence of interest being somatic, and unequal frequencies are interpreted as the sequence of interest being germline.

[0107] In some implementations, the comparison in decision 210 may also be used to eliminate potentially erroneous classifications. In particular, the frequency of true somatic variants is necessarily lower than that of true germline variants because both tumor and non-tumor DNA contribute to the frequency count of germline variants, and only tumor DNA contributes to the frequency count of somatic variants. Thus, in some implementations, if the frequency of the sequence of interest exceeds the proxy frequency, then the sequence of interest is classified as germline.

[0108] By way of example, in some embodiments, comparing the observed frequency of a genomic sequence of interest to a centrality measure of the observed frequencies of one or more proxy genomic sequences may include determining the "allele frequency distance" (AFDIS) of the genomic sequence of interest from the expected allele frequency. The expected allele frequency when the genomic sequence of interest is a germline sequence is determined based on the frequency of one or more proxy sequences (or a summary statistic indicating the observed frequency of one or more proxy sequences) that are assumed to be germline based on the selection of one or more proxy sequences. AFDIS, in some embodiments, may be expressed numerically according to: AFDIS=AF 生殖系列 -AF バリアント In the formula, AF 生殖系列 is the allele frequency expected if the genomic sequence of interest were germline, as determined based on the observed allele frequencies of one or more proxy sequences, and AF バリアント is the observed allele frequency of the genomic sequence of interest.

[0109] In some embodiments, the allele frequency distance can be determined using the distribution of observed frequencies of proxy genome sequences. The distribution can be used to determine the probability that a genome sequence of interest is germline or somatic. In some embodiments, the allele frequency distance is the probability that the observed frequency of a genome sequence of interest fits (or does not fit) the distribution of observed frequencies of multiple proxy sequences. For example, if the allele frequency of a genome sequence of interest falls within the distribution, the genome sequence of interest can be identified as a germline sequence. If the allele frequency of a genome sequence of interest does not fit within the distribution, the genome sequence of interest can be identified as somatic. One skilled in the art can select a statistical test or a predetermined threshold to determine whether the allele frequency of a genome sequence of interest fits within the distribution.

[0110] In some embodiments, the allele frequency distance can be used to classify a genomic sequence of interest. For example, in some embodiments, if the allele frequency distance is above a selected threshold, the genomic sequence of interest is classified as somatic. In some embodiments, if the allele frequency distance is below a selected threshold, the genomic sequence of interest is classified as germline. The threshold can be set based on a desired accuracy or specificity tolerance.

[0111] In some embodiments, the classification of a genomic sequence of interest as germline or somatic may involve the use of a statistical model. The statistical model may, for example, receive allele frequency distances for a given genomic sequence of interest and output a classification of the genomic sequence of interest as somatic (or potentially somatic) or germline (or potentially germline). The classification may be based on the probability that the genomic sequence of interest is somatic or germline. In some implementations, a genomic sequence of interest may be classified as ambiguous, for example, if the probability that the sequence is somatic or germline is not sufficiently high. The probability threshold for making a call may be based on the desired specificity and / or accuracy of the call. For example, in some embodiments, if the probability that the genomic sequence of interest is somatic is greater than any one of 0.8, 0.85, 0.9, 0.95, 0.96, 0.97, 0.98, or 0.99 (or any selected value therebetween), the genomic sequence of interest is classified as somatic, and if the probability that the genomic sequence of interest is somatic is less than any one of 0.2, 0.15, 0.1, 0.05, 0.04, 0.03, 0.02, or 0.01 (or any selected value therebetween), the genomic sequence of interest is classified as germline. Genomic sequences of interest that are not classified as somatic or germline based on the statistical model may be labeled as ambiguous.

[0112] In some embodiments, the statistical model is trained using data from one or more matched tumor / normal sample pairs. The normal sample in the matched tumor / normal sample pair can be sequenced to establish a ground truth for the germline sequence, and the tumor sample can be sequenced to establish a ground truth for the somatic variant sequence (i.e., a non-germline sequence according to the matched normal sample). Using sequencing data from tumor samples that may contain a mixture of normal and tumor nucleic acid molecules, the probability of somatic variants (p somatic ) is equal to 1) or germline (p somatic The allele frequency distance for a selected genomic sequence of interest can be determined, labeled as ∑ i = 0 (i = 1, ii = 1, iii = 0). A function relating the allele frequency distance to the probability of being somatic can then be generated using the training data.

[0113] Other methods of training statistical models may also be used, for example, in some embodiments, the model is trained using only data for germline sequences or only data for somatic sequences.

[0114] In some implementations, the comparison of step 210 may be performed indirectly by a statistical model. For example, if the median allele frequency of a set of proxies is used as the central measure in step 206, then a logistic regression model can be constructed to describe the difference in allele frequency of the sequence of interest from the median of the median allele frequencies of the proxies. In some implementations, this logistic regression model satisfies the above statement:

number

[0115] The rationale behind this characterization is that each proxy is physically close to the sequence of interest in the patient's genome. Therefore, the proxy and the sequence of interest are likely to have experienced the same or similar genomic dynamics or mutations, such as duplication events or deletions. Rather than attempting to model the specific dynamics of the sequence of interest to correlate observed frequencies with germline / somatic status, this approach replaces such models with direct empirical measurements. This approach offers advantages insofar as prior art models have historically been somewhat insensitive or inaccurate.

[0116] The methods described herein can further include generating a report indicating one or more genomic sequences of interest as germline or somatic. The generated report can be transmitted (e.g., using a computer network) to a patient, a healthcare provider, or another party. This report is particularly useful for evaluating cancer treatment therapies, determining treatment, monitoring cancer progression or recurrence, designing personalized cancer vaccines, and other beneficial uses.

[0117] Electronic Devices and Systems FIG. 4 illustrates an example of a system according to one embodiment. Device 400 may be a host computer connected to a network. Device 400 may be a client computer or a server. As shown in FIG. 4, device 400 may be any suitable type of microprocessor-based device, such as a personal computer, a workstation, a server, or a handheld computing device (portable electronic device, e.g., a phone or tablet). The device may include, for example, one or more of a processor 410, an input device 420, an output device 430, memory storage 440, and / or a communication device 460. Input device 420 and output device 430 may be either connectable to or integrated with the computer. In some embodiments, the device is configured to operate a sequencer 470 capable of sequencing nucleic acid molecules in a patient sample to obtain sequencing data.

[0118] Input device 420 may be any suitable device that provides input, such as a touchscreen, a keyboard or keypad, a mouse, or a voice recognition device. Output device 430 may be any suitable device that provides output, such as a display, a touchscreen, a tactile device, or a speaker.

[0119] The memory storage 440 may be any suitable device providing storage, such as electrical, magnetic, or optical memory, including RAM, a cache, a hard drive, or a removable storage disk. The communication device 460 may include any suitable device capable of sending and receiving signals over a network, such as a network interface chip or device. The components of the computer may be connected in any suitable manner, such as by a physical bus or wirelessly.

[0120] Software such as SGZ module 450 and other sequence analysis and variant calling program modules can be stored in memory storage 440 and executed by processor(s) 410, and can include, for example, code for an AFDIS-based logistic regression model and other programming that performs the functions of the present disclosure (e.g., as implemented in a device such as those described above).

[0121] Software such as SGZ module 450, as well as other sequence analysis and variant calling program modules, can also be stored in and / or transmitted to any non-transitory computer-readable storage medium for use by or in connection with an instruction execution system, apparatus, or device (such as those described above) that can fetch and execute instructions associated with the software from the instruction execution system, apparatus, or device. In the context of the present disclosure, a computer-readable storage medium can be any medium, such as storage 440, that can contain or store programming for use by or in connection with an instruction execution system, apparatus, or device.

[0122] Software such as the SGZ module 450, as well as other sequence analysis and variant calling program modules, can also be propagated within any transmission medium for use by or in connection with an instruction execution system, apparatus, or device (such as those described above) that can fetch instructions associated with the software from and execute the instructions. In the context of this disclosure, a transmission medium can be any medium capable of communicating, propagating, or transmitting transmission programming for use by or in connection with an instruction execution system, apparatus, or device. Transmission-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, or infrared wired or wireless propagation media.

[0123] Device 400 may be connected to a network, which may be any suitable type of interconnected communications system. The network may implement any suitable communications protocol and may be protected by any suitable security protocol. The network may include any suitable configuration of network links capable of implementing the transmission and reception of network signals, such as a wireless network connection (T1 or T3 line), a cable network, DSL, or telephone lines.

[0124] Device 400 can implement any operating system suitable for operating on a network. Software such as SGZ module 450 and other sequence analysis and variant calling program modules can be written in any suitable programming language, such as C, C++, Java, or Python. In various embodiments, application software embodying functionality of the present disclosure can be deployed in different configurations (e.g., in a client / server configuration, or via a web browser as a web-based application or web service).

[0125] Subjects, Samples, and Sequencing The subject sample (e.g., patient sample) used in the methods described herein can contain a mixture of tumor and non-tumor nucleic acid molecules. Tumor nucleic acid molecules can be obtained directly or indirectly from a tumor. For example, tumor nucleic acid molecules can be obtained from a tissue biopsy of a tumor. Tumor biopsies often contain both tumor and non-tumor tissue, thereby providing a mixture of tumor and non-tumor nucleic acid molecules. In some embodiments, tumor and non-tumor nucleic acid molecules are obtained from a bodily fluid or liquid biopsy sample (e.g., blood, plasma, spinal fluid, etc.), which can contain cell-free (or circulating free) DNA, including tumor (e.g., circulating tumor DNA or ctDNA) and non-tumor cell-free nucleic acid molecules.

[0126] Patient samples may be taken, for example, from subjects with cancer, subjects suspected of having cancer, or subjects who have previously undergone treatment for cancer. In certain embodiments, the sample is obtained from a subject with a solid tumor, a hematological cancer, or a metastatic form thereof. In certain embodiments, the sample is obtained from a subject who has cancer or is at risk of having cancer. In certain embodiments, the sample is obtained from a subject who is not undergoing treatment to treat cancer, is undergoing treatment to treat cancer, or has previously undergone treatment to treat cancer, as described herein.

[0127] Various tissues can be the source of the sample used in this method. Genomic or subgenomic nucleic acid (e.g., DNA or RNA) can be isolated from a subject's sample (e.g., a sample containing tumor cells, a blood sample, a blood component sample, a sample containing cell-free DNA (cfDNA), a sample containing circulating tumor DNA (ctDNA), a sample containing circulating tumor cells (CTC), or any normal control (e.g., normal adjacent tissue (NAT)).

[0128] In some embodiments, the sample is obtained from a liquid biopsy. A liquid biopsy patient sample can be derived from, for example, blood, plasma, cerebrospinal fluid, sputum, stool, urine, or saliva.

[0129] In some embodiments, the patient sample is derived from a solid tissue sample, such as a solid tumor biopsy. Solid tumor biopsies often contain a mixture of tumor and non-tumor tissue. In some embodiments, the solid tissue biopsy sample is a fresh sample. In some embodiments, the solid tissue biopsy sample is a frozen sample or a previously frozen sample. In some embodiments, the solid tissue biopsy sample is a fresh sample. In some embodiments, the solid tissue biopsy sample is an archived sample (e.g., a chemically preserved sample). In certain embodiments, the sample is a formalin-fixed, paraffin-embedded (FFPE) sample.

[0130] In some embodiments, the tumor purity of a patient sample (i.e., the portion of the sample that is tumor nucleic acid molecules compared to total nucleic acid molecules) for any of the sample types disclosed herein is about 1% or more, about 5% or more, about 10% or more, about 15% or more, about 20% or more, about 25% or more, about 30% or more, about 40% or more, about 50% or more, about 60% or more, about 70% or more, or about 80% or more. In some embodiments, the tumor purity of a patient sample is about 99% or less, about 95% or less, about 90% or less, about 85% or less, about 80% or less, about 75% or less, about 70% or less, about 60% or less, about 50% or less, about 40% or less, about 30% or less, about 25% or less, or about 20% or less.

[0131] In one embodiment, the method further includes obtaining a sample, e.g., a patient sample described herein. The sample can be obtained directly or indirectly. In one embodiment, the sample is obtained, e.g., by isolation or purification from a sample containing cfDNA. In one embodiment, the sample is obtained, e.g., by isolation or purification from a sample containing ctDNA. In one embodiment, the sample is obtained, e.g., by isolation or purification, from a sample containing both malignant and non-malignant cells (e.g., tumor-infiltrating lymphocytes). In one embodiment, the sample is obtained, e.g., by isolation or purification from a sample containing CTCs. In some embodiments, the sample is obtained by solid tissue biopsy.

[0132] Sequencing libraries can be prepared from patient samples using known methods. Nucleic acid molecules can be purified or isolated from patient samples. In some embodiments, the isolated nucleic acids are fragmented or sheared using known methods. For example, nucleic acid molecules can be fragmented by physical shearing (e.g., sonication), enzymatic cleavage, chemical cleavage, and other methods known to those skilled in the art. Nucleic acids can be linked to adapter sequences for sequencing. In some cases, the adapters can include amplification primers and / or sequencing adapters. In some cases, nucleic acid molecules purified or isolated from patient samples or sequencing libraries prepared therefrom can be amplified using, for example, polymerase chain reaction (PCR) or isothermal amplification methods known to those skilled in the art.

[0133] In some embodiments, nucleic acid molecules from a patient sample and used to prepare a sequencing library (or a selected (e.g., captured) subset thereof) are sequenced to generate a patient genome sequence. Sequencing methods are well known in the art and can be performed using multiplex (e.g., next-generation) or single-molecule sequencing. The patient genome sequence determined by sequencing need not be the patient's entire genome. For example, in some embodiments, targeted sequencing methods (e.g., the use of specific probe (or bait) molecules for hybridization-based capture) are used to sequence a portion of the patient's genome (i.e., less than the entire genome). See, e.g., U.S. Pat. No. 9,340,830 B2. Targeted sequencing can be used to target, for example, one or more exon regions, one or more intron regions, one or more intragenic regions, one or more 3'-UTRs (untranslated regions), and / or one or more 5'-UTRs.

[0134] In some embodiments, targeted sequencing may be used to sequence one or more genes or portions of one or more genes associated with cancer. Exemplary genes associated with cancer that can be sequenced using targeted sequencing include ABL2, AKT2, AKT3, ARAF, ARFRP1, ARID1A, ATM, ATR, AURKA, AURKB, BCL2, BCL2A1, BCL2L1, BCL2L2, BCL6, BRCA1, BRCA2, CARD11, CBL, CCND1, CCND2, CCND3, CCNE1, CDH1, CDH2, CDH20, CDH5, CDK4, CDK6, CDK8, CDKN2B, CDKN2C, CHEK1, CHEK2, CRKL, CRLF2, DNMT3A, DOT1L, EPHA3, EPHA5, EPHA6, EPHA7, EPHB1, EPHB4, EPHB6, ERBB3, ERBB4, ERG, ETV1, ETV4, ETV5, ETV6, EWSR1, EZ H2, FANCA, FBXW7, FGFR4, FLT1, FLT4, FOXP4, GATA1, GNA11, GNAQ, GNAS, GPR124, GUCY1A2, HOXA3, HSP90AA1, IDH1, IDH2, IGF1R, IGF2R, IKBKE, IKZF1, INHBA, IRS2, JAK1, JAK3, JUN, KDR, LRP1B, LTK, MAP2K1, MAP2K2, MAP2K4, MCL1, MDM2, MDM4, MEN1, MITF, MLH1, MPL, MRE11A, MSH2, MSH6 , MTOR, MUTYH, MYCL1, MYCN, NF2, NKX2-1, NTRK1, NTRK3, PAK3, PAX5, PDGFRB, PIK3R1, PKHD1, PLCG1, PRKDC, PTCH1, PTPN11, PTPRD, RAF1, RARA , RICTOR, RPTOR, RUNX1, SMAD2, SMAD3, SMAD4, SMARCA4, SMARCB1, SMO, SOX10, SOX2, SRC, STK11, TBX22, TET2, TGFBR2, TMPRSS2, TOP1, TSC1, T SC2, USP9X, VHL, WT1, ABL1, AKT1, ALK, APC, AR, BRAF, CDKN2A, CEBPA, CTNNB1, EGFR, ERBB2, ESR1, FGFR1, FGFR2, FGFR3, FLT3, HRAS, JAK2, KIT,These include, but are not limited to, KRAS, MET, MLL, MYC, NF1, NOTCH1, NPM1, NRAS, PDGFRA, PIK3CA, PTEN, RB1, RET, and TP53.

[0135] In certain embodiments, the sample is obtained from a subject with cancer. Exemplary cancers include, but are not limited to, B-cell cancers, such as multiple myeloma, melanoma, breast cancer, lung cancer (such as non-small cell lung cancer or NSCLC), bronchial cancer, colorectal cancer, prostate cancer, pancreatic cancer, gastric cancer, ovarian cancer, bladder cancer, brain or central nervous system cancer, peripheral nervous system cancer, esophageal cancer, cervical cancer, uterine or endometrial cancer, oral or pharyngeal cancer, liver cancer, kidney cancer, testicular cancer, biliary tract cancer, small intestine or adnexal cancer. , salivary gland cancer, thyroid cancer, adrenal adenocarcinoma, osteosarcoma, chondrosarcoma, cancer of blood tissue, adenocarcinoma, inflammatory myofibroblastic tumor, gastrointestinal stromal tumor (GIST), colon cancer, multiple myeloma (MM), myelodysplastic syndrome (MDS), myeloproliferative disorder (MPD), acute lymphocytic leukemia (ALL), acute myeloid leukemia (AML), chronic myeloid leukemia (CML), chronic lymphocytic leukemia (CLL), polycythemia vera, Hodgkin's lymphoma, non-Hodgkin's lymphoma NHL, soft tissue sarcoma, fibrosarcoma, myxosarcoma, liposarcoma, osteogenic sarcoma, chordoma, angiosarcoma, endothelial sarcoma, synovioma, mesothelioma, Ewing's tumor, leiomyosarcoma, rhabdomyosarcoma, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinoma, medullary carcinoma, bronchogenic carcinoma, renal cell carcinoma, hepatocellular carcinoma, bile duct carcinoma, choriocarcinoma, seminoma, embryonal carcinoma, Wilms' tumor, bladder cancer, epithelial carcinoma, glioma, astrocytoma These include cell tumors, medulloblastomas, craniopharyngiomas, ependymoma, pinealoma, hemangioblastoma, acoustic neuroma, oligodendroglioma, meningiomas, neuroblastomas, retinoblastomas, cell lymphomas, mantle cell lymphomas, hepatocellular carcinomas, thyroid cancer, gastric cancer, head and neck cancer, small cell carcinoma, essential thrombocythemia, agnogenic myeloid metaplasia, hypereosinophilic syndrome, systemic mastocytosis, the familiar hypereosinophilia, chronic eosinophilic leukemia, neuroendocrine carcinomas, and cancerous tumors.

[0136] In one embodiment, the cancer is a hematological malignancy (or pre-malignancy). As used herein, hematological malignancy refers to a tumor of the hematopoietic or lymphoid tissues, e.g., a tumor affecting the blood, bone marrow, or lymph nodes. Exemplary hematological malignancies include leukemia (e.g., acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), chronic lymphocytic leukemia (CLL), chronic myeloid leukemia (CML), hairy cell leukemia, acute monocytic leukemia (AMoL), chronic myelomonocytic leukemia (CMML), juvenile myelomonocytic leukemia (JMML), or large granular lymphocytic leukemia), lymphoma (e.g., AIDS-related lymphoma, cutaneous T-cell lymphoma, Hodgkin's lymphoma (e.g., classical Hodgkin's lymphoma or nodular lymphocyte-predominant Hodgkin's lymphoma), mycosis fungoides, non-Hodgkin's lymphoma, and leukemia (e.g., leukemia ... lymphomas (e.g., B-cell non-Hodgkin's lymphoma (e.g., Burkitt's lymphoma, small lymphocytic lymphoma (CLL / SLL), diffuse large B-cell lymphoma, follicular lymphoma, immunoblastic large cell lymphoma, precursor B-lymphoblastic lymphoma, or mantle cell lymphoma) or T-cell non-Hodgkin's lymphoma (mycosis fungoides, anaplastic large cell lymphoma, or precursor T-lymphoblastic lymphoma)), primary central nervous system. As used herein, premalignant tumors refer to tissue that is not yet malignant but is preparing to become malignant.

[0137] In some embodiments, the sample is obtained, e.g., collected, from a subject, e.g., a patient, having a condition or disease, e.g., a hyperproliferative disease (e.g., as described herein) or a non-cancer indication. In some embodiments, the disease is a hyperproliferative disease. In some embodiments, the hyperproliferative disease is a cancer, e.g., a solid tumor or a hematological cancer. In some embodiments, the cancer is a solid tumor. In some embodiments, the cancer is a hematological cancer, e.g., leukemia or lymphoma.

[0138] In some embodiments, the subject has cancer. In some embodiments, the subject has been or is being treated for cancer. In some embodiments, the subject needs to be monitored for progression or regression of cancer, e.g., after being treated with a cancer therapy. In some embodiments, the subject needs to be monitored for recurrence of cancer. In some embodiments, the subject is at risk of having cancer. In some embodiments, the subject has not been treated with a cancer therapy. In some embodiments, the subject has a genetic predisposition to cancer (e.g., having a mutation that increases the baseline risk for developing cancer). In some embodiments, the subject has been exposed to an environment (e.g., radiation or chemicals) that increases the risk of developing cancer. In some embodiments, the subject needs to be monitored for the development of cancer.

[0139] In some embodiments, the patient has been previously treated with targeted therapy, for example, one or more targeted therapies.In some embodiments, for the patient who has been previously treated with targeted therapy, a post-targeted therapy sample, for example, specimen, is obtained, for example, collected.In some embodiments, the post-targeted therapy sample is obtained, for example, collected, after the completion of targeted therapy.

[0140] In some embodiments, the patient has not been previously treated with a targeted therapy. In some embodiments, for patients not previously treated with a targeted therapy, the sample comprises a resection, e.g., the original resection, or a recurrence, e.g., disease recurrence after treatment, e.g., a non-targeted therapy. In some embodiments, the sample is, or is a portion of, a primary tumor or a metastasis, e.g., a metastasis biopsy. In some embodiments, the sample is obtained from a tumor, e.g., a site with the highest percentage of tumor cells, e.g., the tumor site, compared to an adjacent site, e.g., an adjacent site with tumor cells. In some embodiments, the sample is obtained from a site with the largest tumor focus, e.g., the tumor site, compared to an adjacent site, e.g., an adjacent site with tumor cells.

[0141] In some embodiments, the subject is a human.

[0142] Cancer Treatment The genomic profile of a cancer can often affect the likelihood of success of various cancer treatment modalities. For example, a given anticancer drug may be more likely to be successful in treating a particular cancer with one genomic profile than in treating a particular cancer with a different genomic profile. The methods described herein can be used to characterize the genomic profile of a cancer by distinguishing somatic sequences that may be attributable to the cancer from germline sequences.

[0143] For example, a method for treating cancer in a patient may include identifying (e.g., classifying) one or more genomic sequences of interest as somatic using the methods described herein, and selecting a cancer treatment modality based on the one or more identified somatic sequences. An effective amount of the selected cancer treatment modality can then be used to treat the cancer. This allows for personalized cancer treatment for a patient based on somatic sequences specific to that patient's cancer. In contrast, if treatment selection were based on germline variants rather than somatic variants, there is a risk that the selected treatment modality may be ineffective for the patient's cancer.

[0144] Exemplary cancer treatment modalities may include, for example, selected chemotherapeutic agents, selected immuno-oncology agents (such as immune checkpoint inhibitors), resective surgery, radiation therapy, targeted therapy, gene expression modulators, angiogenesis inhibitors, and hormone therapy, among others.

[0145] A cancer treatment can be selected, for example, based on an association between one or more identified somatic sequences and successful cancer treatment using a selected therapeutic modality. Exemplary associations between cancer types, somatic sequences, and therapeutic modalities are listed in Table 1. [Table 1-1] [Table 1-2]

[0146] The microsatellite instability (MSI) status of cancer can be useful for selecting cancer treatment modalities. Microsatellite instability can result from defects in the DNA mismatch repair (MMR) pathway in cancer cells, resulting in an abnormally high frequency of genetic mutations. See Kim et al., "The Landscape of Microsatellite Instability in Colorectal and Endometrial Cancer Genomes," Cell, Vol. 155, No. 4, pp. 858-868 (2013). MSI status is generally characterized based on the MSI signature as high (MSI-H), low (MSI-L), or stable (MSS) (or alternatively, MSI-H or not MSI-H; or MSI-H or MSI-indeterminate). MSI-H status has been detected in multiple types of solid tumors and can be an indicator of the success of cancer treatment with specific cancer treatment modalities. See Cortes-Ciriano et al., "A molecular portrait of microsatellite instability across multiple cancers," Nature Communications, 8, 15180 (2017). Mutations in microsatellites (i.e., MSI events) can be detected by distinguishing somatic from germline sequences using the methods described herein.

[0147] The success of certain cancer treatment modalities is related to the MSI-H status of the cancer. For example, PD-1 inhibitors (i.e., pembrolizumab) have been shown to be particularly effective in treating MSI-H solid tumors (e.g., unresectable or metastatic solid tumors). In some embodiments, a cancer determined to have MSI-H status is treated with an effective amount of an immuno-oncology agent. In some embodiments, a cancer determined to have MSI-H status is treated with an effective amount of an immune checkpoint inhibitor. In some embodiments, the immune checkpoint inhibitor is AMP-224, AMP-514, atezolizumab, AUNP12, avelumab, BGB-A317, BMS-986189, CA-170, canrelizumab, cemiplimab, CK-301, dostarimab, durvalumab, ipilimumab, INCMGA00012, KN035, nivolumab, pembrolizumab, sintilimab, spartalizumab, tislelizumab, or toripalimab. In some embodiments, the cancer determined to have MSI-H status is treated with an effective amount of a PD-1 inhibitor, a PD-L1 inhibitor, or a CTLA-4 inhibitor. In some embodiments, the cancer determined to have MSI-H status is treated with an effective amount of pembrolizumab.

[0148] In some embodiments, the methods of treating cancer include identifying (e.g., classifying) one or more genomic sequences of interest as somatic using the methods described herein; determining the microsatellite instability status of the cancer using the identified somatic sequences; and selecting a cancer treatment modality based on the microsatellite instability status of the cancer. The cancer can then be treated with an effective amount of the selected cancer treatment modality. In some embodiments, the cancer is colorectal cancer, endometrial cancer, biliary tract cancer, bladder cancer, breast cancer, esophageal cancer, gastric cancer, gastroesophageal junction cancer, pancreatic cancer, prostate cancer, renal cell carcinoma, retroperitoneal adenocarcinoma, sarcoma, small cell lung cancer, small intestine cancer, or thyroid cancer.

[0149] In some embodiments, the tumor mutational burden (TMB) of a cancer is determined using one or more somatic sequences identified using the methods described herein to select a treatment modality. TMB is a cancer genomic biomarker that quantifies the frequency of somatic mutations in a patient's tumor. High TMB correlates with higher neoantigen expression, which helps the immune system recognize tumors. It has been detected across multiple tumor types and is associated with improved response rates and longer progression-free survival in patients undergoing immunotherapy. See Goodman et al., "Tumor Mutational Burden as an Independent Predictor of Response to Immunotherapy in Diverse Cancers," Mol. Cancer Ther., Vol. 16, No. 11, pp. 2598-2608 (2017).

[0150] Tumor mutational burden can be determined for a cancer by identifying somatic sequences associated with the cancer using the methods described herein.

[0151] TMB can provide a quantitative value that allows for the selection of cancer treatment modalities based on tumor mutational burden above or below a predetermined tumor mutational burden threshold. In some embodiments, the predetermined threshold is about 5 mutations / Mb, about 10 mutations / Mb, about 15 mutations / Mb, about 20 mutations / Mb, about 25 mutations / Mb, about 30 mutations / Mb, about 40 mutations / Mb, about 50 mutations / Mb, or more, or any number therebetween (e.g., the predetermined threshold can be between 5 mutations / Mb and about 50 mutations / Mb). By way of example, certain immuno-oncology agents have been found to be particularly effective when used to treat tumors with a high tumor mutational burden. See, e.g., Fabrizio et al., "Beyond microsatellite testing: assessment of tumor mutational burden identifies subsets of colorectal cancers who may respond to immune checkpoint inhibition," J. Gastrointestinal Oncology, Vol. 9, No. 4, pp. 610-617 (2018).

[0152] In some embodiments, cancers determined to have a TMB above a predetermined threshold are treated with an effective amount of an immuno-oncology agent. In some embodiments, cancers determined to have a TMB above a predetermined threshold are treated with an effective amount of an immune checkpoint inhibitor. In some embodiments, the immune checkpoint inhibitor is AMP-224, AMP-514, atezolizumab, AUNP12, avelumab, BGB-A317, BMS-986189, CA-170, canrelizumab, cemiplimab, CK-301, dostarimab, durvalumab, ipilimumab, INCMGA00012, KN035, nivolumab, pembrolizumab, sintilimab, spartalizumab, tislelizumab, or toripalimab. In some embodiments, cancers determined to have a TMB above a predetermined threshold are treated with an effective amount of a PD-1 inhibitor, a PD-L1 inhibitor, or a CTLA-4 inhibitor. In some embodiments, cancers determined to have a TMB above a predetermined threshold are treated with an effective amount of pembrolizumab. In some embodiments, cancers determined to have a TMB above a predetermined threshold are treated with an effective amount of pembrolizumab, the predetermined threshold being about 10 mutations / Mb.

[0153] In some embodiments, a method for treating cancer includes identifying one or more genomic sequences of interest as somatic using the methods described herein; determining the tumor mutation burden for the cancer using the one or more identified somatic sequences; and selecting a cancer treatment modality based on the tumor mutation burden exceeding a predetermined tumor mutation burden threshold. The cancer can then be treated with an effective amount of the selected cancer treatment modality. In some embodiments, the cancer is colorectal cancer, endometrial cancer, biliary tract cancer, bladder cancer, breast cancer, esophageal cancer, gastric cancer, gastroesophageal junction cancer, pancreatic cancer, prostate cancer, renal cell carcinoma, retroperitoneal adenocarcinoma, sarcoma, small cell lung cancer, small intestine cancer, or thyroid cancer.

[0154] Monitoring cancer progression Monitoring cancer progression and / or detecting minimal residual disease is useful for evaluating cancer treatment plans and / or monitoring patients for cancer recurrence. Cancer patients can receive cancer treatment until the cancer is undetectable. Nevertheless, patients may remain susceptible to recurrence. Patients can be monitored for cancer recurrence by detecting nucleic acid molecules derived from recurrent tumors (e.g., ctDNA molecules). In other embodiments, cancer patients can be treated for the disease, and the progression of the cancer (e.g., an increase or decrease in the amount of cancer) can be monitored by quantifying the amount of tumor nucleic acid molecules detected in the patient (e.g., ctDNA levels).

[0155] Identification of somatic sequences can be particularly useful in monitoring cancer progression or detecting minimal residual cancer disease. Somatic sequences provide a genomic signature of the cancer and can be used to distinguish tumor nucleic acid molecules from non-tumor nucleic acid molecules.

[0156] Patient samples can be obtained and analyzed at two or more time points to monitor cancer progression or cancer recurrence. A first sample is analyzed to identify one or more somatic sequences according to the methods described herein. The first sample can be obtained before, during, or after cancer treatment, although the patient generally has some detectable cancer.

[0157] A second sample may be obtained at a later time point after the patient has been treated for cancer and analyzed to determine whether one or more of the identified somatic sequences are present in the sample. The presence of a somatic sequence indicates that the patient still has cancer or that the cancer has recurred. The inability to detect a somatic sequence does not conclusively prove that the patient does not have cancer, but indicates that the level of cancer may be low.

[0158] The second patient sample may be the same type of sample as the first patient sample type, or may be a different sample type. In some embodiments, the second patient sample is obtained from a liquid biopsy. For example, the liquid biopsy patient sample may be blood, plasma, cerebrospinal fluid, sputum, stool, urine, or saliva. In some embodiments, the patient sample is obtained from a solid tissue sample, such as a solid tumor biopsy. In some embodiments, the solid tissue biopsy sample is a fresh sample. In some embodiments, the solid tissue biopsy sample is a frozen sample or a previously frozen sample. In some embodiments, the solid tissue biopsy sample is a fresh sample. In some embodiments, the solid tissue biopsy sample is an archived sample (e.g., a chemically preserved sample). In certain embodiments, the sample is a formalin-fixed, paraffin-embedded (FFPE) sample.

[0159] Somatic sequences can be detected in DNA or RNA (or both) from the second sample. The presence or absence of somatic sequences in the second sample can be detected by sequencing, quantitative PCR (qPCR), reverse transcription PCR (RT-PCR), fluorescence in situ hybridization (FISH), or any other suitable method for specific detection of one or more somatic sequences. In certain embodiments, nucleic acid molecules are isolated from the second sample. In some embodiments, nucleic acid molecules are directly detected from the second sample.

[0160] In some embodiments, the presence of one or more somatic sequences is identified in the second sample, and the patient may be treated for cancer using the same or a different treatment modality with which the cancer was previously treated.

[0161] In some embodiments, a method for monitoring the progression or recurrence of cancer in a patient includes identifying one or more genomic sequences of interest as somatic using the methods described herein, wherein the patient sample is obtained from a patient with cancer; obtaining a second patient sample from the patient after the cancer has been treated; and detecting the presence or absence of the one or more genomic sequences of interest identified as somatic in the second patient sample. For example, the one or more genomic sequences of interest can be identified as somatic by selecting a genomic sequence of interest at a genomic locus from a patient genomic sequence obtained for a patient sample containing a mixture of tumor and non-tumor nucleic acid molecules; selecting one or more proxy genomic sequences for the genomic sequence of interest; determining an allele frequency distance using summary statistics indicating the observed allele frequency of the genomic sequence of interest and the observed frequency of the one or more proxy genomic sequences; and using the allele frequency distance to identify the genomic sequence of interest as germline or somatic. In some embodiments, the method includes treating cancer in the patient after a first patient sample is obtained from the patient and before a second patient sample is obtained from the patient, in some embodiments, the method includes treating cancer in the patient if the presence of one or more genomic sequences of interest identified as somatic is detected in the second patient sample.

[0162] Neoantigen selection and cancer vaccine production Somatic sequences detected in the exon regions of various genes may be suitable as neoantigens, for example, in the development of personalized cancer vaccines. Peptides can be generated based on the nucleic acid sequences encoded by somatic variant sequences, which can stimulate the immune system to kill cancer cells. See, for example, Richters et al., "Best practices for bioinformatics characterization of neoantigens for clinical utility," Genome Medicine, Vol. 11, p. 56 (2019).

[0163] In some embodiments, a method for selecting neoantigens for a cancer vaccine personalized to a subject with cancer includes identifying one or more genomic sequences of interest as somatic using the methods described herein, wherein the one or more genomic sequences of interest identified as somatic are located within exon regions of genes, and selecting from the one or more genomic sequences of interest identified as somatic, genomic sequences encoding neoantigens suitable for a cancer vaccine for the subject. For example, the one or more genomic sequences of interest can be identified as somatic by: selecting a genomic sequence of interest at a genomic locus from a patient genomic sequence obtained for a patient sample containing a mixture of tumor and non-tumor nucleic acid molecules; selecting one or more proxy genomic sequences for the genomic sequence of interest; determining an allele frequency distance using summary statistics indicating the observed allele frequency of the genomic sequence of interest and the observed frequency of the one or more proxy genomic sequences; and identifying the genomic sequence of interest as germline or somatic using the allele frequency distance.

[0164] In some embodiments, the method further comprises producing a vaccine comprising the neoantigen. [Example]

[0165] Example 1 - Discrimination between somatic and germline variants based on allele frequency distance (AFDIS) The following examples are provided to illustrate exemplary embodiments of the invention described herein and are not intended to limit the scope of the invention.

[0166] Previously described SGZ algorithms (e.g., Sun et al. (2018), ibid.) can be used to determine the difference in expected variant allele frequencies for somatic and germline variants (e.g., mutations substituting C for T), where the tumor fraction of the sample, the allele number of the variant, and the copy number of the genomic locus have been determined, as shown in Figure 5A. The expected variant allele frequencies (VAFs) for somatic and germline variants can be determined as follows:

number

[0167] This example provides an alternative approach to the previously described SGZ algorithm that does not require modeling of tumor purity, variant allele number, or copy number values. The allele frequency distance from the expected germline allele frequency (AFDIS) is determined as follows: AFDIS=AF 生殖系列 -AF バリアント AF 生殖系列is the allele frequency of the sequence, assuming it is the definitive germline sequence, as defined by the allele frequency of the corresponding proxy sequence. AF バリアント where is the observed allele frequency of a given sequence being characterized. To understand the allele frequency distance distribution of germline variants, we segmented genomic sequences from 3,802 tumor samples based on copy number uniformity using the Circular Binary Segmentation algorithm described by Olshen et al., "Circular Binary Segmentation for the Analysis of Array-Based DNA Copy Number Data," Biostatistics, Vol. 5, No. 4, pp. 557-572 (October 2004). Approximately 2.1 million known germline variants (identified in the dbSNP and / or gnomAD databases) were selected from the 3,802 samples, and the allele frequency (based on sequencing) of each germline variant was compared to the median allele frequency of proxy sequences within the same segment to determine the allele frequency distance for each germline variant. The probability densities of the approximately 2.1 million germline variants from the 3,802 samples are shown in Figure 5B, and the selected values ​​are listed in Table 2. An empirical cumulative distribution function (ECDF) can be constructed from this germline AFDIS distribution data and used to assess the probability that a given AFDIS is derived from a germline variant. [Table 2]

[0168] A threshold of 0.1 AFDIS, corresponding to a cumulative distribution of 0.993 based on the ECDF described above, was empirically determined to be able to effectively separate somatic from germline variants. As shown in Table 2, AFDIS thresholds ranging from approximately 0.05 to 0.1 all provided good discrimination between somatic and germline variants. Nevertheless, as described below, a trained statistical model was constructed to understand the probability that any given sequence is germline or somatic.

[0169] Allele frequency distances were then determined for 92 genotype-matched high-purity / low-purity tumor samples with known germline, somatic, and tumor purity. Low-purity samples were generally considered approximations of normal samples, allowing for reliable determination of the somatic versus germline status of variants within them, and so low-purity samples were used to establish ground truth for the somatic / germline status of selected sequences. Figure 5C shows the variant AFDIS for germline and somatic sequences from 92 tumor samples plotted against the calculated purity of the samples. Gray circles represent ground truth somatic sequences, and black circles represent ground truth germline sequences.

[0170] Example 2 - Logistic regression of somatic / germline status based on AFDIS Using available data from 21 matched tumor / normal pairs (lung squamous cell carcinoma (n = 5), ovarian serous carcinoma (n = 4), lung adenocarcinoma (n = 3), breast invasive ductal carcinoma (n = 2), anal carcinoma (n = 1), bladder urothelial carcinoma (n = 1), CRC (n = 1), renal clear cell carcinoma (n = 1), ovarian high-grade serous carcinoma (n = 1), cutaneous sarcoma (n = 1), and endometrial adenocarcinoma (n = 1), a logistic regression model was generated. The matched tumor / normal pairs allowed for reliable determination of somatic and germline sequences. Figure 5D shows the receiver operating characteristic (ROC) curve of this approach, i.e., the classification model in discriminating between somatic and germline variants. Figure 1 shows a graphical plot of the true positive (TP) and false positive (FP) performance of the model. The "leave-one-out cross validation" (LOOCV) results for the model showed an accuracy of 0.97 (95% confidence interval = [0.95, 0.99]) and a Cohen's (unweighted) kappa statistic of 0.93. The model was trained using matched tumor / normal paired data to output the probability that a given sequence is a somatic sequence. For known germline sequences in the training data, the probability that the sequence is somatic is 0. For known somatic sequences in the training data, the probability that the sequence is somatic is 1. A logistic regression model was trained using the training dataset according to the following function:

number

[0171] The AFDIS data calculated as above for variants in a total of 188 tumor samples in three different test sets was input into the trained model to determine the probability that each selected sequence is somatic or germline. Based on the somatic variant probability, the variant sequence was labeled as somatic (if it is above the somatic probability threshold), germline (if it is below the germline probability threshold), or ambiguous (i.e., between the somatic probability threshold and the germline probability threshold). See Figure 5F.

[0172] As shown in Figure 5G, classification results using the AFDIS classifier on a set of 93 tumor samples with matched normal samples used to validate the traditional SGZ method demonstrate improvement over the traditional SGZ method. The genomic sequences of the 93 tumor samples were obtained using a hybrid capture bait set different from that used in the training dataset, demonstrating that the AFDIS classifier is robust and applicable to genomic data collected by various methods. Non-limiting examples of the method's performance at various levels (# true positives, # false positives, and positive predictive value) are outlined in Table 3. [Table 3]

[0173] Non-limiting example data regarding sample-level sensitivity performance of the method is shown in Figure 5H, and non-limiting example data regarding positive predictive value (PPV) performance is shown in Figure 5I. The "violin plots" shown in Figures 5H and 5I show that the shape of the plot indicates the probability density of the values ​​on the vertical axis. Nested box plots within the violin plots show the median, first and third quartiles, minimum, maximum, and outliers for the parameter plotted on the vertical axis. In this example PPV plot, the majority of samples have a PPV of 100%, and therefore the median, maximum, and first and third quartile indices are compressed.

[0174] A non-limiting example of data for classification of variants in the BRCA1 and BRCA2 genes is shown in Figure 5J. A non-limiting example of data for classification of variants in the STK11 gene is shown in Figure 5K. As expected, BRCA1 and BRCA2 mutations were found to be enriched in germline-origin variants in breast cancer compared to other cancer types (p=0.025 chi-squared test), and STK11 mutations were found to be enriched in somatic-origin variants in lung cancer compared to other cancer types (p=0.0026 chi-squared test).

[0175] Example 3 - Logistic regression of somatic / germline status based on AFDIS The disclosed method for distinguishing between somatic variants and germline variants is based on comparing the allele frequency (AF) of the variant in question with the allele frequency of known variants adjacent to its genomic location. In some cases, as described above, known germline variants in a germline database (e.g., a public database) can be used for comparison. If the AF of the variant in question is very similar or very different from the AF of the known germline variant located nearby, it will be concluded that the variant in question is very likely or unlikely to be germline, respectively.

[0176] Generally, the AF of a given variant is mainly determined by its copy number and the tumor fraction of sample.Tumor fraction is a constant for a particular sample, and therefore the AF of a given variant in a given sample is largely determined by its copy number.This means that AF can be compared with the AF of the germline variant with the same copy number to infer the somatic / germline status of variant.Two non-limiting examples of implementing such comparison are described below and in Example 4.

[0177] In one implementation, an "allele frequency distance" (AFDIS) is calculated, which represents the distance between the AF of the variant in question and the median AF of germline variants located on the same copy number segment (e.g., located in the same physically contiguous part of the genomic segment, or located in non-contiguous parts of the genomic segment, as long as the segment is present in the same copy number as the variant in question). First, the AFDIS was calculated as follows: AFDIS=|MAF バリアント -MAF セグメント | where MAF = minor allele frequency, i.e., the absolute distance between the minor allele frequency of both the variant of interest and the median minor allele frequency for the segment germline variants was calculated. A logistic regression model was then trained with a training dataset consisting of known somatic and germline variants to capture the relationship between "somatic probability" and AFDIS. The model then computed the AFDIS using directional distances, i.e., AFDIS = AFDIS = AF セグメント -AF バリアント where AF セグメントis the median allele frequency for segmental germline variants. In this formula, the sign of AFDIS accounts for somatic variants that have a lower allele frequency compared to germline variants of the same copy number when normal tissues, cells, or cfDNA are mixed in the sample. This is because sequencing reads derived from normal parts of the sample or normal cells in the blood carry germline variants rather than somatic variants. A logistic regression model is trained to recognize that a negative AFDIS is associated with a lower probability that the variant is somatic. The use of directional AFDIS calculations improved the model's performance for distinguishing between somatic and germline variants.

[0178] The AFDIS-based approach has the advantage of being computationally simple and easy to calculate, and therefore can be easily modified to incorporate other considerations in a given implementation. Specifically, because AFDIS is a single predictor variable in a logistic regression model, the AFDIS value can be easily adjusted to modify results to account for other potential technical issues. For example, to account for increased uncertainty introduced by mild contamination of nucleic acid samples, adjustments can be applied to the AFDIS value depending on the contamination level, moving the AFDIS value into a range corresponding to more accurate classification of somatic / germline variants by the model. Similar adjustments can be made to account for additional uncertainty introduced by factors such as low read depth, noisy AF estimates, low segmental germline SNP counts, and high variability in segmental germline SNP AFs. The extent and manner of implementing these adjustments can be designed and tuned using a training dataset containing known somatic and germline variants.

[0179] Example 4 - Germline exclusion based on probability distribution of germline allele frequencies In this particular implementation, a large dataset of known germline variants is constructed, each with its own AF and corresponding segmental MAF, which is the median MAF of other known germline variants located in the same copy number segment. Figure 6A shows a plot of variant AF versus segmental MAF. For an unknown variant to be classified, its AF and corresponding segmental MAF are determined. To classify an unknown variant, data is obtained from a known germline dataset containing a subset of known germline variants with segmental MAFs similar to that of the unknown variant (e.g., one of the three density versus mutation AF plots shown in Figure 6B, corresponding to variant allele frequency distributions at segmental MAFs around 0.1, 0.2, and 0.3, respectively, as shown in Figure 6A). This data can be used to establish the distribution of germline AF values ​​for a given segmental MAF (i.e., a given copy number, since the segmental MAF is essentially determined by the copy number of the segment). The AF of the unknown variant is compared to this germline AF distribution to estimate the probability that the unknown variant is a germline variant. For example, an unknown variant with an AF of either 0.1 or 0.9 and a segmental MAF of 0.1 is likely to be a germline variant, while an unknown variant with an AF of 0.4 and a segmental MAF of 0.1 is likely to be a somatic variant.

[0180] Example 5 - Performance Verification The disclosed methods provide exemplary techniques for selecting somatic variants from baseline tissue or liquid biopsy samples for plasma surveillance. To further improve performance for this specific purpose, several additional measures were devised, including (i) selecting well-behaved variants (e.g., by excluding variants located in genomic regions where allele frequencies are known or predicted to deviate from expected values, such as regions with repetitive sequences or regions that share homology with other regions of the genome) for building logistic regression models; (ii) incorporating prior knowledge of the likelihood of a variant to be a germline, somatic, or clonal hematopoietic with undetermined potential (CHIP) variant based on historical data and public databases; and (iii) taking into account the noise level of the variant call and its genomic context. These measures were found to improve the performance of somatic variant classification.

[0181] The ability of the disclosed AFDIS-based logistic regression model to distinguish somatic variants from germline variants in a sample was validated, for example, using data from matched tumor / normal pairs. Non-limiting examples of the initial training and test datasets used to develop the logistic regression model, as well as the performance metrics (# false positives (FP), sensitivity, and positive predictive value (PPV)) obtained for various levels and sample-level performance, are outlined in Tables 4 and 5, respectively. [Table 4] [Table 5]

[0182] The dataset used in the variant calling pipeline validation study contained data from 86 matched tissue / peripheral blood mononuclear cell (PBMC) sample pairs. The various-level and sample-level performance metrics are summarized in Tables 6 and 7, respectively. [Table 6] [Table 7]

[0183] The dataset used in an additional variant calling pipeline validation study included data from 746 matched tissue / peripheral blood mononuclear cell (PBMC) sample pairs. Various-level and sample-level performance metrics are summarized in Tables 8 and 9, respectively. [Table 8] [Table 9]

[0184] It will be understood that the above-described methods and systems are presented by way of example, and not by way of limitation. Numerous variations, additions, omissions, and other modifications will be apparent to those skilled in the art. Additionally, the order or presentation of method steps in the above description and drawings is not intended to require this order of performing the recited steps unless a particular order is explicitly required or is clear from the context.

[0185] The method steps of the present invention described herein are intended to include any suitable manner of having one or more other parties or entities perform the steps, unless a different meaning is expressly provided or clear from the context. In some embodiments, such parties or entities need not be under the direction or control of the other parties or entities, and need not be located in any particular jurisdiction. Thus, for example, a statement or recitation of "adding a first number to a second number" includes having one or more parties or entities add the two numbers together. For example, if person X enters into an arm's length transaction with person Y to add two numbers, and person Y actually adds the two numbers, then both person X and person Y perform the steps indicated by person Y actually adding the numbers and by person X having person Y add the numbers. Furthermore, if person X is located within the United States and person Y is located outside the United States, the method is performed in the United States by person X's involvement in causing the steps to be performed.

[0186] The terminology used in the description of the various embodiments set forth herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used in the description of the various described embodiments and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. The term "and / or," as used herein, will also be understood to refer to and encompass any and all possible combinations of one or more of the associated listed items. It will be further understood that, as used herein, the terms "includes," "including," "comprises," and / or "comprising" specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0187] The disclosures of all publications, patents, and patent applications referenced herein are each incorporated herein by reference in their entirety. To the extent that a reference incorporated by reference conflicts with the present disclosure, the present disclosure shall control.

[0188] While particular embodiments of the present invention have been shown and described, it will be apparent to those skilled in the art that various changes and modifications in form and detail can be made therein without departing from the spirit and scope of the invention as defined by the following claims, which are intended to embrace all such changes and modifications as may be within their scope and should be interpreted in the broadest sense permitted by law.

Claims

[Claim 1] The invention described in the specification.