Methods of improved somatic mutation detection

The use of a germline mutational signature addresses the contamination issue in somatic mutation detection, improving the accuracy and sensitivity of somatic mutation identification by filtering out germline variants.

WO2025221921A1PCT designated stage Publication Date: 2025-10-23MYRIAD GENETICS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/025011
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-18
Filing Date
2025-04-16
Publication Date
2025-10-23

AI Technical Summary

Technical Problem

Existing methods struggle to accurately detect somatic mutations and signatures in cancer genomes due to the challenge of distinguishing them from germline mutations, leading to potential contamination and inaccurate analysis.

Method used

Utilizing a novel germline mutational signature, referred to as 'Signature. gDNA', which decomposes sequencing reads to identify and remove germline contamination, allowing for the detection of somatic mutations and signatures with improved accuracy.

Benefits of technology

The method effectively filters out germline mutations, enhancing the sensitivity and accuracy of somatic mutation detection, reducing noise from germline variants and enabling precise identification of somatic signatures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025025011_23102025_PF_FP_ABST
    Figure US2025025011_23102025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are germline mutational signatures that can improve somatic calling and detection of somatic mutational signatures.
Need to check novelty before this filing date? Find Prior Art

Description

METHODS OF IMPROVED SOMATIC MUTATION DETECTIONCROSS-REFERENCE TO RELATED APPLICATION

[0001] This application claims the benefit of priority under 35 U.S.C. § 119(e) of U.S. Provisional Patent Application No. 63 / 636,066, filed April 18, 2024, the entire contents of which is incorporated herein by reference in its entirety.TECHNICAL FIELD

[0002] This disclosure relates to methods of improving the detection of somatic mutation signatures in a sample. More particularly, the disclosure relates to methods of identifying possible germline mutations in a sample by identifying mutations from sequenced DNA derived from a subject’s biological sample and utilizing a set of germline mutational signatures to identify mutations or variants that are likely to be a germline mutations or variants and not somatic mutations or variants.BACKGROUND

[0003] The following description of the background of the present technology is provided simply as an aid in understanding the present technology and is not admitted to describe or constitute prior art to the present technology.

[0004] Somatic mutations accumulate in the genomes of all cells of the human body throughout an individual’s lifetime. These mutations arise from different endogenous and exogenous mutational processes, with each process generating a characteristic pattern of mutations, known as a mutational signature. Somatic mutations found in cancer genomes may be the consequence of the intrinsic slight infidelity of the DNA replication machinery, exogenous or endogenous mutagen exposures, enzymatic modification of DNA, or defective DNA repair. Moreover, various signatures may be informative regarding patient prognosis, applicable treatments, likeliness of metastasis, and disease progression. Accordingly, there is a need to improve the capacity to detect somatic mutations and somatic mutation signatures with high fidelity and without inadvertent inclusion of germline variants.SUMMARY

[0005] The present disclosure provides methods for improved somatic called and improved methods for detecting somatic mutations and somatic mutation signatures. In particular, the present disclosure provides a novel germline mutational signature that can be used to betterfilter out germline mutations when it is desirable to analyze or detect only the somatic mutations in a given sample.

[0006] In one aspect, the present disclosure provides methods of detecting a somatic mutation signature in a sample, comprising: sequencing DNA from a sample obtained from a subject, thereby obtaining sequencing reads for the subject; detecting a plurality of mutations in the sequencing reads; decomposing the plurality of mutations with a reference set of somatic mutation signatures and a germline mutation signature comprising a frequency of a classification set unique to a germline genome comprising (i) nucleotide changes, (ii) at least one unchanged position that is 5’ of each nucleotide change in the set, and (iii) at least one unchanged position that is 3’ of each nucleotide change in the set, thereby removing germline contamination from the sequencing reads and detecting at least one somatic mutation signature for the subject.

[0007] In some embodiments, the classification set comprises 96 triplicates, wherein (i) the nucleotide changes comprise six types of substitutions, (ii) the at least one unchanged position that is 5’ of each nucleotide change in the set comprises one of four types of bases 5’ to the substitution, and (iii) the at least one unchanged position that is 3’ of each nucleotide change in the set comprises one of four types of bases 3’ to the substitution.

[0008] In some embodiments, the frequency of the classification set is the frequency provided in FIG. 1.

[0009] In some embodiments, (i) the nucleotide changes comprise six types of substitutions, (ii) the at least one unchanged position that is 5’ of each nucleotide change in the set comprises at least two, at least three, or at least four bases 5’ to the substitution, and (iii) the at least one unchanged position that is 3’ of each nucleotide change in the set comprises at least two, at least three, or at least four bases 3’ to the substitution.

[0010] In some embodiments, the methods may further comprise removing germline mutations from the plurality of mutations in the sequence reads by (a) comparing the sequences reads for the subject to a set of germline mutations, wherein the set of germline mutations are (i) determined by sequencing a germline sample from the subject, (ii) compiled from a database of germline mutations, or (iii) a combination of mutations found in sequences from a germline sample from the subject and mutations from a database of germline mutations; (b) determining the frequency of the sequence reads comprising mutations and comparing the frequency to a database of germline mutations; or (c)determining a tumor mutation burden (TMB) score based on the number of somatic variants having a somatic mutation significance score above a threshold.

[0011] In some embodiments, the reference set of somatic mutation signatures comprises Signature 1 A, Signature IB, Signature 2, Signature 3, Signature 4, Signature 5, Signature 6, Signature 7, Signature 8, Signature 9, Signature 10, Signature 11, Signature 12, Signature 13, Signature 14, Signature 15, Signature 16, Signature 17, Signature 18, Signature 19, Signature 20, and Signature 21 (see Table 1, below).

[0012] In some embodiments, the subject has cancer, previously had cancer, or is at risk of developing cancer.

[0013] In some embodiments, the sample is a tumor sample or a fluid sample. In some embodiments, the cancer is selected from adrenal cancer, anal cancer, bile duct cancer, bladder cancer, bone cancer, a brain / CNS tumor, breast cancer, Castleman disease, cervical cancer, colon or rectum cancer, endometria cancer, esophagus cancer, a Ewing tumor, eye cancer, gallbladder cancer, a gastrointestinal carcinoid tumor, a gastrointestinal stromal tumor (GIST), gestational trophoblastic disease, Hodgkin disease, Kaposi sarcoma, kidney cancer, laryngeal and hypopharyngeal cancer, leukemia, liver cancer, lung cancer, lymphoma, malignant mesothelioma, multiple myeloma, myelodysplastic syndrome, nasal cavity or paranasal sinus cancer, nasopharyngeal cancer, neuroblastoma, oral cavity or oropharyngeal cancer, osteosarcoma, ovarian cancer, pancreatic cancer, penile cancer, a pituitary tumor, prostate cancer, retinoblastoma, rhabdomyosarcoma, salivary gland cancer, skin cancer, small intestine cancer, stomach cancer, testicular cancer, thymus cancer, thyroid cancer, uterine sarcoma, vaginal cancer, vulvar cancer, Waldenstrom macroglobulinemia, and Wilms tumor. In some embodiments, the fluid sample is a blood sample or a plasma sample. In some embodiments, the subject has, previously had, or is at risk of developing leukemia or lymphoma

[0014] In some embodiments, sequencing DNA from the sample comprises whole genome sequencing, whole exome sequencing, or targeted gene sequencing.

[0015] In some embodiments, the at least one somatic mutation signature is selected from Signature 1 A, Signature IB, Signature 2, Signature 3, Signature 4, Signature 5, Signature 6, Signature 7, Signature 8, Signature 9, Signature 10, Signature 11, Signature 12, Signature 13, Signature 14, Signature 15, Signature 16, Signature 17, Signature 18, Signature 19, Signature 20, and Signature 21, or any combination thereof (see Table 1, below).

[0016] In some embodiments, decomposing the plurality of mutations comprises normalizing the at least one somatic mutation signature.

[0017] In another aspect, the present disclosure provides methods of improved somatic mutation calling, comprising: obtaining a tumor sample from a subject; sequencing DNA from the tumor sample, thereby obtaining sequencing reads; and calling a somatic mutation signation for the tumor sample by: detecting somatic mutations in the sequencing reads, and removing germline contamination from the sequencing reads by decomposing the sequencing reads with a germline mutation signature; and detecting at least one somatic mutation signature; wherein the somatic mutation signature has a reduced negative effect from germline contamination relative a somatic mutation call that does not account for the germline mutational signature.

[0018] In some embodiments, the germline mutation signature comprises a frequency of a classification set that is unique to a germline genome, the classification set comprising (i) nucleotide changes, (ii) at least one unchanged position that is 5’ of each nucleotide change in the set, and (iii) at least one unchanged position that is 3’ of each nucleotide change in the set. In some embodiments, the classification set comprises 96 triplicates, wherein (i) the nucleotide changes comprise six types of substitutions, (ii) the at least one unchanged position that is 5’ of each nucleotide change in the set comprises one of four types of bases 5’ to the substitution, and (iii) the at least one unchanged position that is 3’ of each nucleotide change in the set comprises one of four types of bases 3’ to the substitution. In some embodiments, the frequency of the classification set is the frequency provided in FIG. 1. In some embodiments, (i) the nucleotide changes comprise six types of substitutions, (ii) the at least one unchanged position that is 5’ of each nucleotide change in the set comprises at least two, at least three, or at least four bases 5’ to the substitution, and (iii) the at least one unchanged position that is 3’ of each nucleotide change in the set comprises at least two, at least three, or at least four bases 3’ to the substitution.

[0019] In some embodiments, removing germline contamination further comprises: (a) comparing the sequences reads for the subject to a set of germline mutations, wherein the set of germline mutations are (i) determined by sequencing a germline sample from the subject, (ii) compiled from a database of germline mutations, or (iii) a combination of mutations found in sequences from a germline sample from the subject and mutations from a database of germline mutations; (b) determining the frequency of the sequence reads comprisingmutations and comparing the frequency to a database of germline mutations; or (c) determining a tumor mutation burden (TMB) score based on the number of somatic variants having a somatic mutation significance score above a threshold.

[0020] In some embodiments, the subject has or previously had a cancer selected from adrenal cancer, anal cancer, bile duct cancer, bladder cancer, bone cancer, a brain / CNS tumor, breast cancer, Castleman disease, cervical cancer, colon or rectum cancer, endometria cancer, esophagus cancer, a Ewing tumor, eye cancer, gallbladder cancer, a gastrointestinal carcinoid tumor, a gastrointestinal stromal tumor (GIST), gestational trophoblastic disease, Hodgkin disease, Kaposi sarcoma, kidney cancer, laryngeal and hypopharyngeal cancer, leukemia, liver cancer, lung cancer, lymphoma, malignant mesothelioma, multiple myeloma, myelodysplastic syndrome, nasal cavity or paranasal sinus cancer, nasopharyngeal cancer, neuroblastoma, oral cavity or oropharyngeal cancer, osteosarcoma, ovarian cancer, pancreatic cancer, penile cancer, a pituitary tumor, prostate cancer, retinoblastoma, rhabdomyosarcoma, salivary gland cancer, skin cancer, small intestine cancer, stomach cancer, testicular cancer, thymus cancer, thyroid cancer, uterine sarcoma, vaginal cancer, vulvar cancer, Waldenstrom macroglobulinemia, Wilms tumor, leukemia, and lymphoma.

[0021] In some embodiments, the sample is a tumor sample. In some embodiments, the sample is a fluid sample. In some embodiments, the fluid sample is a blood sample or a plasma sample.

[0022] In some embodiments, sequencing DNA from the tumor sample comprises whole genome sequencing, whole exome sequencing, or targeted gene sequencing.

[0023] In some embodiments, the at least one somatic mutation signature is selected from Signature 1 A, Signature IB, Signature 2, Signature 3, Signature 4, Signature 5, Signature 6, Signature 7, Signature 8, Signature 9, Signature 10, Signature 11, Signature 12, Signature 13, Signature 14, Signature 15, Signature 16, Signature 17, Signature 18, Signature 19, Signature 20, and Signature 21, or any combination thereof (see Table 1, below).

[0024] In another aspect, the present disclosure provides computer implemented methods for improved detection of a somatic mutation signature, comprising: retrieving, by a processor, data associated with a tumor sample, wherein the data comprises DNA sequencing reads from the tumor sample; generating, by a computer device, a set of mutations in the sequencing reads; and decomposing, by a computer processor, the set of mutations with a reference set of somatic mutation signatures and a germline mutation signature, thereby removing germlinecontamination from the sequencing reads; and detecting at least one somatic mutation signature.

[0025] In some embodiments, the reference set of somatic mutation signatures comprises Signature 1 A, Signature IB, Signature 2, Signature 3, Signature 4, Signature 5, Signature 6, Signature 7, Signature 8, Signature 9, Signature 10, Signature 11, Signature 12, Signature 13, Signature 14, Signature 15, Signature 16, Signature 17, Signature 18, Signature 19, Signature 20, and Signature 21 (see Table 1, below).

[0026] In some embodiments, the germline mutation signature comprises a frequency of a classification set that is unique to a germline genome, the classification set comprising (i) nucleotide changes, (ii) at least one unchanged position that is 5’ of each nucleotide change in the set, and (iii) at least one unchanged position that is 3’ of each nucleotide change in the set. In some embodiments, the classification set comprises 96 triplicates, wherein (i) the nucleotide changes comprise six types of substitutions, (ii) the at least one unchanged position that is 5’ of each nucleotide change in the set comprises one of four types of bases 5’ to the substitution, and (iii) the at least one unchanged position that is 3’ of each nucleotide change in the set comprises one of four types of bases 3’ to the substitution. In some embodiments, the frequency of the classification set is the frequency provided in FIG. 1. In some embodiments, (i) the nucleotide changes comprise six types of substitutions, (ii) the at least one unchanged position that is 5’ of each nucleotide change in the set comprises at least two, at least three, or at least four bases 5’ to the substitution, and (iii) the at least one unchanged position that is 3’ of each nucleotide change in the set comprises at least two, at least three, or at least four bases 3’ to the substitution.

[0027] In some embodiments, decomposing the set of mutations comprises normalizing the at least one somatic mutation signature.

[0028] In another aspect, the present disclosure provides method of improved detection of tumor mutation burden (TMB), comprising: obtaining a tumor sample from a subject; sequencing DNA from the tumor sample, thereby obtaining sequencing reads; and detecting TMB by: quantifying somatic mutations in the sequencing reads, and removing germline contamination from the sequencing reads by decomposing the sequencing reads with a germline mutation signature.

[0029] In some embodiments, the germline mutation signature comprises a frequency of a classification set that is unique to a germline genome, the classification set comprising (i)nucleotide changes, (ii) at least one unchanged position that is 5’ of each nucleotide change in the set, and (iii) at least one unchanged position that is 3’ of each nucleotide change in the set. In some embodiments, the classification set comprises 96 triplicates, wherein (i) the nucleotide changes comprise six types of substitutions, (ii) the at least one unchanged position that is 5’ of each nucleotide change in the set comprises one of four types of bases 5’ to the substitution, and (iii) the at least one unchanged position that is 3’ of each nucleotide change in the set comprises one of four types of bases 3’ to the substitution. In some embodiments, the frequency of the classification set is the frequency provided in FIG. 1.

[0030] In some embodiments, the subject has or previously had a cancer selected from adrenal cancer, anal cancer, bile duct cancer, bladder cancer, bone cancer, a brain / CNS tumor, breast cancer, Castleman disease, cervical cancer, colon or rectum cancer, endometria cancer, esophagus cancer, a Ewing tumor, eye cancer, gallbladder cancer, a gastrointestinal carcinoid tumor, a gastrointestinal stromal tumor (GIST), gestational trophoblastic disease, Hodgkin disease, Kaposi sarcoma, kidney cancer, laryngeal and hypopharyngeal cancer, leukemia, liver cancer, lung cancer, lymphoma, malignant mesothelioma, multiple myeloma, myelodysplastic syndrome, nasal cavity or paranasal sinus cancer, nasopharyngeal cancer, neuroblastoma, oral cavity or oropharyngeal cancer, osteosarcoma, ovarian cancer, pancreatic cancer, penile cancer, a pituitary tumor, prostate cancer, retinoblastoma, rhabdomyosarcoma, salivary gland cancer, skin cancer, small intestine cancer, stomach cancer, testicular cancer, thymus cancer, thyroid cancer, uterine sarcoma, vaginal cancer, vulvar cancer, Waldenstrom macroglobulinemia, Wilms tumor, leukemia, and lymphoma.

[0031] In some embodiments, the sample is a tumor sample. In some embodiments, the sample is a fluid sample. In some embodiments, the fluid sample is a blood sample or a plasma sample.

[0032] In some embodiments, sequencing DNA from the tumor sample comprises whole genome sequencing, whole exome sequencing, or targeted gene sequencing.BRIEF DESCRIPTION OF THE DRAWINGS

[0033] FIG. 1 shows a histogram with the triplicate prevalence for the disclosed germline signature, which is also referred to as “Signature. gDNA” herein. This graph shows the prevalence (proportion of counts for each triplicate / divided by sum of all triplicates) for each of the 96 triplicate types for a set of known germline samples.

[0034] FIG. 2 shows a pie graph with signature proportions for a whole exome sequencing sample referred to as “TCGA-A2-A0T0,” which without germline contamination decomposes entirely into Signature.3 (HRD pathway deficiency).

[0035] FIG. 3 shows three pie graphs with signature proportions for TCGA-A2-A0T0 after it was informatically spiked (i.e., in silico) with 50% germline SNP contamination. The left graph shows the results of a detection model with deconvolution without taking into account the disclosed germline signature (i.e., “Signature. gDNA”), the middle graph shows the model with the disclosed germline signature (i.e., “Signature. gDNA”) included, and the right graphs shows the results of an adjusted detection model that was deconvoluted with the disclosed germline signature (i.e., “Signature. gDNA”).

[0036] FIG. 4 shows three pie graphs with signature proportions for TCGA-A2-A0T0 after it was spiked with 100% germline SNP contamination. The left graph shows the results of a detection model with deconvolution without taking into account the disclosed germline signature (i.e., “Signature. gDNA”), the middle graph shows the model with the disclosed germline signature (i.e., “Signature. gDNA”) included, and the right graphs shows the results of an adjusted detection model that was deconvoluted with the disclosed germline signature (i.e., “Signature. gDNA”).

[0037] FIG. 5 shows three pie graphs with signature proportions for TCGA-A2-A0T0 after it was spiked with 300% germline SNP contamination. The left graph shows the results of a detection model with deconvolution without taking into account the disclosed germline signature (i.e., “Signature. gDNA”), the middle graph shows the model with the disclosed germline signature (i.e., “Signature. gDNA”) included, and the right graphs shows the results of an adjusted detection model that was deconvoluted with the disclosed germline signature (i.e., “Signature. gDNA”).

[0038] FIG. 6 shows three pie graphs with signature proportions for TCGA-A2-A0T0 after it was spiked with 500% germline SNP contamination. The left graph shows the results of a detection model with deconvolution without taking into account the disclosed germline signature (i.e., “Signature. gDNA”), the middle graph shows the model with the disclosed germline signature (i.e., “Signature. gDNA”) included, and the right graphs shows the results of an adjusted detection model that was deconvoluted with the disclosed germline signature (i.e., “Signature. gDNA”).

[0039] FIG. 7 shows three pie graphs with signature proportions for TCGA-A2-A0T0 after it was spiked with 800% germline SNP contamination. The left graph shows the results of a detection model with deconvolution without taking into account the disclosed germline signature (i.e., “Signature. gDNA”), the middle graph shows the model with the disclosed germline signature (i.e., “Signature. gDNA”) included, and the right graphs shows the results of an adjusted detection model that was deconvoluted with the disclosed germline signature (i.e., “Signature. gDNA”).

[0040] FIG. 8 shows three pie graphs with signature proportions for TCGA-A2-A0T0 after it was spiked with 900% germline SNP contamination. The left graph shows the results of a detection model with deconvolution without taking into account the disclosed germline signature (i.e., “Signature. gDNA”), the middle graph shows the model with the disclosed germline signature (i.e., “Signature. gDNA”) included, and the right graphs shows the results of an adjusted detection model that was deconvoluted with the disclosed germline signature (i.e., “Signature. gDNA”).

[0041] FIG. 9 shows three pie graphs with signature proportions for TCGA-A2-A0T0 after it was spiked with 1500% germline SNP contamination. The left graph shows the results of a detection model with deconvolution without taking into account the disclosed germline signature (i.e., “Signature. gDNA”), the middle graph shows the model with the disclosed germline signature (i.e., “Signature. gDNA”) included, and the right graphs shows the results of an adjusted detection model that was deconvoluted with the disclosed germline signature (i.e., “Signature. gDNA”).

[0042] FIG. 10 shows a pie graph with signature proportions for a hypothetical whole genome sequencing sample referred to as “TCGA-A8-A07R,” which without germline contamination decomposes into three significant signatures, Signature.3, Signature.4, and Signature.9.

[0043] FIG. 11 shows three pie graphs with signature proportions for TCGA-A8-A07R after it was spiked with 10% germline SNP contamination. The left graph shows the results of a detection model without deconvolution without taking into account the disclosed germline signature (i.e., “Signature. gDNA”), the middle graph shows the model with the disclosed germline signature (i.e., “Signature. gDNA”) included, and the right graphs shows the results of an adjusted detection model that was deconvoluted with the disclosed germline signature (i.e., “Signature. gDNA”).

[0044] FIG. 12 shows three pie graphs with signature proportions for TCGA-A8-A07R after it was spiked with 50% germline SNP contamination. The left graph shows the results of a detection model without deconvolution without taking into account the disclosed germline signature (i.e., “Signature. gDNA”), the middle graph shows the model with the disclosed germline signature (i.e., “Signature. gDNA”) included, and the right graphs shows the results of an adjusted detection model that was deconvoluted with the disclosed germline signature (i.e., “Signature. gDNA”).

[0045] FIG. 13 shows three pie graphs with signature proportions for TCGA-A8-A07R after it was spiked with 100% germline SNP contamination. The left graph shows the results of a detection model without deconvolution without taking into account the disclosed germline signature (i.e., “Signature. gDNA”), the middle graph shows the model with the disclosed germline signature (i.e., “Signature. gDNA”) included, and the right graphs shows the results of an adjusted detection model that was deconvoluted with the disclosed germline signature (i.e., “Signature. gDNA”).

[0046] FIG. 14 shows three pie graphs with signature proportions for TCGA-A8-A07R after it was spiked with 200% germline SNP contamination. The left graph shows the results of a detection model without deconvolution without taking into account the disclosed germline signature (i.e., “Signature. gDNA”), the middle graph shows the model with the disclosed germline signature (i.e., “Signature. gDNA”) included, and the right graphs shows the results of an adjusted detection model that was deconvoluted with the disclosed germline signature (i.e., “Signature. gDNA”).

[0047] FIG. 15 shows three pie graphs with signature proportions for TCGA-A8-A07R after it was spiked with 300% germline SNP contamination. The left graph shows the results of a detection model without deconvolution without taking into account the disclosed germline signature (i.e., “Signature. gDNA”), the middle graph shows the model with the disclosed germline signature (i.e., “Signature. gDNA”) included, and the right graphs shows the results of an adjusted detection model that was deconvoluted with the disclosed germline signature (i.e., “Signature. gDNA”).

[0048] FIG. 16 shows three pie graphs with signature proportions for TCGA-A8-A07R after it was spiked with 400% germline SNP contamination. The left graph shows the results of a detection model without deconvolution without taking into account the disclosed germline signature (i.e., “Signature. gDNA”), the middle graph shows the model with the disclosedgermline signature (i.e., “ Signature. gDN A”) included, and the right graphs shows the results of an adjusted detection model that was deconvoluted with the disclosed germline signature (i.e., “Signature. gDNA”).

[0049] FIG. 17 shows three pie graphs with signature proportions for TCGA-A8-A07R after it was spiked with 700% germline SNP contamination. The left graph shows the results of a detection model without deconvolution without taking into account the disclosed germline signature (i.e., “Signature. gDNA”), the middle graph shows the model with the disclosed germline signature (i.e., “Signature. gDN A”) included, and the right graphs shows the results of an adjusted detection model that was deconvoluted with the disclosed germline signature (i.e., “Signature. gDNA”).

[0050] FIG. 18 shows three pie graphs with signature proportions for TCGA-A8-A07R after it was spiked with 1000% germline SNP contamination. The left graph shows the results of a detection model without deconvolution without taking into account the disclosed germline signature (i.e., “Signature. gDNA”), the middle graph shows the model with the disclosed germline signature (i.e., “Signature. gDN A”) included, and the right graphs shows the results of an adjusted detection model that was deconvoluted with the disclosed germline signature (i.e., “Signature. gDNA”).

[0051] FIG. 19 shows three pie graphs with signature proportions for TCGA-A8-A07R after it was spiked with 1500% germline SNP contamination. The left graph shows the results of a detection model without deconvolution without taking into account the disclosed germline signature (i.e., “Signature. gDNA”), the middle graph shows the model with the disclosed germline signature (i.e., “Signature. gDN A”) included, and the right graphs shows the results of an adjusted detection model that was deconvoluted with the disclosed germline signature (i.e., “Signature. gDNA”).

[0052] FIG. 20 shows graphs graphs that illustrate the effects of Signature. gSNP on detection of tumor mutation burden (TMB) for the simulations with TCGA-A2-A0T0 and TCGA-A8-A07R. The gray dots (lower curve) in each of the graphs represent the TMB as calculated when results of an adjusted detection model that was deconvoluted with the disclosed germline signature (i.e., “Signature. gDNA”), whereas the black dots (higher curve) in each graph represent the TMB as calculated when results with deconvolution without taking into account the disclosed germline signature (i.e., “Signature. gDNA”). As can be seenfrom these graphs, without the disclosed germline signature (i.e., “ Signature. gDNA”), TMB was dramatically overestimated.

[0053] FIG. 21 shows the somatic mutation signatures established in Alexandrov et al., Signatures of mutational processes in human cancer, NATURE, 2013, 500:415-421.DETAILED DESCRIPTION

[0054] Described herein are germline mutational signatures and methods of utilizing the germline mutational signatures to identify true somatic variants or more accurately detect a somatic mutation signature in a sample. The disclosed germline mutational signature (also referred to herein as “Signature. gDNA”) can be applied to a set of putative somatic mutations or signatures and decomposing the set of putative somatic mutations / signatures to reduce noise from germline variants. The disclosed signatures and methods allow for more sensitive detect of somatic signatures in a sample, and allow for the detection and correct calling of somatic signatures that would otherwise be lost in the noise from germline variants that are not properly removed from the putative somatic mutation sets.

[0055] It is to be appreciated that certain aspects, modes, embodiments, variations and features of the present methods are described below in various levels of detail in order to provide a substantial understanding of the present technology. It is to be understood that the present disclosure is not limited to particular uses, methods, reagents, compounds, compositions or biological systems, which can, of course, vary. It is also to be understood that the terminology used herein for the purpose of describing particular embodiments only and is not intended to be limiting.I. Definitions

[0056] Unless defined otherwise, all technical and scientific terms used herein generally have the same meaning as commonly understood by one of ordinary skill in the art to which this technology belongs. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” include plural referents unless the context clearly dictates otherwise. For example, reference to a “a cell” includes a combination of two or more cells, and the like. Generally, the nomenclature used herein and the laboratory procedures in cell culture, molecular genetics, organic chemistry, analytical chemistry and nucleic acid chemistry and hybridization described below are those well-known and commonly employed in the art.

[0057] As used herein, the term “about” in reference to a number is generally taken to include numbers that fall within a range of 1%, 5%, or 10% in either direction (greater than or less than) of the number unless otherwise stated or otherwise evident from the context (except where such number would be less than 0% or exceed 100% of a possible value).

[0058] The term “sample” herein refers to any substance containing or presumed to contain nucleic acid. The sample can be a biological sample obtained from a subject. The nucleic acids can be RNA, DNA, e.g., genomic DNA, mitochondrial DNA, viral DNA, synthetic DNA, or cDNA reverse transcribed from RNA. The nucleic acids in a nucleic acid sample generally serve as templates for extension of a hybridized primer. In some embodiments, the biological sample is a biological fluid sample. The fluid sample can be whole blood, plasma, serum, ascites, cerebrospinal fluid, sweat, urine, tears, saliva, buccal sample, cavity rinse, or organ rinse. The fluid sample can be an essentially cell-free liquid sample (e.g., plasma, serum, sweat, urine, tears, etc.). In other embodiments, the biological sample is a solid biological sample, e.g., feces or tissue biopsy, e.g., a tumor biopsy. A sample can also comprise in vitro cell culture constituents (including but not limited to conditioned medium resulting from the growth of cells in cell culture medium, recombinant cells and cell components). In some embodiments, the sample is a biological sample that is a mixture of nucleic acids from multiple sources, i.e., there is more than one contributor to a biological sample, e.g., two or more individuals. A “sample” may include, but is not limited to, blood, plasma, saliva, urine, semen, amniotic fluid, oocytes, skin, hair, feces, cheek swabs, or pap smear lysate from an individual.

[0059] The term “mutation” herein refers to a change introduced into a reference sequence, including, but not limited to, substitutions, insertions, deletions (including truncations) relative to the reference sequence. Mutations can involve large sections of DNA (e.g., copy number variation). Mutations can involve whole chromosomes (e.g., aneuploidy). Mutations can involve small sections of DNA. Examples of mutations involving small sections of DNA include, e.g., point mutations or single nucleotide polymorphisms (SNPs), multiple nucleotide polymorphisms, insertions (e.g., insertion of one or more nucleotides at a locus but less than the entire locus), multiple nucleotide changes, deletions (e.g., deletion of one or more nucleotides at a locus), and inversions (e.g., reversal of a sequence of one or more nucleotides). The consequences of a mutation include, but are not limited to, the creation of a new character, property, function, phenotype or trait not found in the protein encoded by thereference sequence. In some embodiments, the reference sequence is a parental sequence. In some embodiments, the reference sequence is a reference human genome, e.g., hl9. In some embodiments, the reference sequence is derived from a non-cancer (or non-tumor) sequence. In some embodiments, the mutation is inherited. In some embodiments, the mutation is spontaneous or de novo.

[0060] “ Germline mutations” are mutations that occurred in gametes that gave rise to a resulting embryo and individual. The mutation will be present in all of the cells of the person’s body and not only in a specific subset of cells, such as cancer cells.

[0061] “ Somatic mutations” are the result of mutations in a single cell or a few cells in the individual and are not present in every cell of the individual’s body. Somatic mutations are known to accumulate in a cancer cell’s genome and may drive progression of the disease.

[0062] As used herein, a quantity related to the frequency of somatic variants can be defined as “tumor mutation burden” (TMB). TMB can be calculated as a count of somatic variants in a cancer sample normalized to the total number of genomic positions assayed in determining the count of somatic variants. TMB can be expressed as a number of mutations per megabase ofDNA.

[0063] A reference value for TMB may be a TMB level in a population of subjects having cancer who have been treated with an anticancer agent. In some embodiments, the population may comprise a group of subjects who have been treated with a particular anticancer agent and a different group of subjects that have been treated with a different anticancer agent. A reference value for TMB may be a TMB level in population of subjects having cancer who do not respond to treatment with an anticancer agent. A reference value with respect to a tumor mutation burden may represent the average TMB level in a plurality of training patients, for example cancer patients, with similar outcomes whose clinical and follow-up data are available and sufficient to define and categorize the patients by disease outcome, for example recurrence or prognosis.

[0064] A “gene” refers to a DNA segment that is involved in producing a polypeptide and includes regions preceding and following the coding regions as well as intervening sequences (introns) between individual coding segments (exons).

[0065] The term “Next Generation Sequencing (NGS)” herein refers to sequencing methods that allow for massively parallel sequencing of clonally amplified and of single nucleic acidmolecules during which a plurality, e.g., millions, of nucleic acid fragments from a single sample or from multiple different samples are sequenced in unison. Non-limiting examples of NGS include sequencing-by-synthesis, sequencing-by-ligation, real-time sequencing, and nanopore sequencing.II. Somatic Variants and Signatures

[0066] Cancers are known to be caused by somatic mutations in a patient’s genome, which may be a result of, but not limited to, infidelity during DNA replication, exposure to mutagens and carcinogens, and defects in DNA repair mechanisms. Different mutational processes may generate different combinations of mutations, which have been termed signatures. However, most work has been done looking at mutational signatures for a small number of frequently mutated cancer genes.

[0067] To generate the mutational signatures, individual mutations from a plurality of cancer samples are compiled. However, development of these signatures is dependent upon the presence of an individual or limited number of mutational processes, as the presence of multiple mutational processes results in noisy composite signatures that may not be predictive of disease presence, progression, prognosis, or responsiveness.

[0068] Alexandrov et al., Signatures of mutational processes in human cancer, NATURE, 2013, 500:415-421 describes a method to determine single base somatic variants, as well as immediately adjacent bases, thereby forming a triplicate of three bases, in order to generate a signature that is informative about the various mutational processes occurring within the sample. These signatures can be informative of, for example, the types of stresses or insults that lead to the development of a cancer or the types of repair mechanisms that are not functioning appropriately in a cancer. This can be aid in selecting treatments that are most likely to be effective for the cancer.

[0069] In particular, Alexandrov et al. utilized whole genome or exome sequencing to identify thousands of somatic mutations beyond simply those at known cancer genes, which were analyzed to identify more than one mutational signature when more than one mutational process is present. The investigators developed an algorithm to extract mutational signatures from somatic mutations and applied it to numerous cancer whole-genome or whole-exome sequences. The algorithm revealed a number of novel and known signatures along with the contribution of each signature to each cancer sample and the estimated timing of the activityof the signatures. The prevalence of somatic mutations is highly variable between and within cancer classes. The mutational signatures used base substitutions and additionally included information on the sequence context of each mutation. As there are six classes of base substitution (C>A, OG, OT, T>A, T>C, T>G; all substitutions are referred to by the pyrimidine of the mutated Watson-Crick base pair) and information on the bases immediately 5’ and 3’ to each mutated base are incorporated, there are 96 possible mutations in this classification. The classification is useful for determining and distinguishing mutational signatures that cause the same substitutions but in different sequence contexts.

[0070] Application of this approach to 30 different cancer types revealed 22 distinct somatic mutational signatures (see Table 1, below, and FIG. 21). The signatures were characterized by prominence of only one or two of the 96 possible substitution mutations, suggesting that there is high specificity of mutation types and sequence context for some cancer types. Other cancer types exhibit nearly equal representation of the 96 possible substitution mutations. Most individual cancer genomes sequenced included more than one mutational signature, and individual cancer genomes included different combinations of signatures and different patterns of contribution.Table 1

[0071] The gold standard for generating a list of somatic variants for a given tumor sample involves subtracting out germline variants. This can be done by running both a tumor sample and a non-tumor sample (e.g., a blood sample or buffy coat sample), where any overlapping variants shared between the samples are considered germline and the remaining variants within the tumor sample are considered somatic. However, this process requires that a nontumor sample is available, and it requires significant additional sequencing costs, as a high level of sequencing coverage is needed to provide assurance that somatic variants are being appropriately selected. In the absence of this method, using database-based exclusion (e.g., using dbSNP or GNOME-DB) and allele frequency exclusion can remove many germline variants, but some germline variant contamination should be expected.

[0072] Identifying somatic mutations is dependent upon the proper identification and filtering out of germline mutations. Existing methods of identifying putative germline mutations (such as those used in Alexander et al.) are costly, slow, and may result in the missorting of mutations based on, for instance, sequencing errors. Similarly, reliance on database-based exclusion (i.e., subtracting out know / common germline mutations found in a database) and allele frequency exclusion are prone to error and do not provide consistent, high-quality results. In other words, in order to be truly useful and informative of patient needs, somatic signatures should contain a pure somatic variant list, such that one can generate accurate signature weights. Thus, germline variant contamination is a serious concern because it canimpact both weights and presence of observed signatures from the affected sample. The present disclosure provides novel methods of identifying germline mutations that allow for the faster, more accurate, and more cost-effective identification of somatic mutational signatures in an individual patient’s genome.

[0073] The disclosed methods may be used for assessing somatic mutations in any cancer patient. For example, the cancer patient may have a cancer selected from, but not limited to, adrenal cancer, anal cancer, bile duct cancer, bladder cancer, bone cancer, a brain / CNS tumor, breast cancer, Castleman disease, cervical cancer, colon or rectum cancer, endometria cancer, esophagus cancer, a Ewing tumor, eye cancer, gallbladder cancer, a gastrointestinal carcinoid tumor, a gastrointestinal stromal tumor (GIST), gestational trophoblastic disease, Hodgkin disease, Kaposi sarcoma, kidney cancer, laryngeal and hypopharyngeal cancer, leukemia, liver cancer, lung cancer, lymphoma, malignant mesothelioma, multiple myeloma, myelodysplastic syndrome, nasal cavity or paranasal sinus cancer, nasopharyngeal cancer, neuroblastoma, oral cavity or oropharyngeal cancer, osteosarcoma, ovarian cancer, pancreatic cancer, penile cancer, a pituitary tumor, prostate cancer, retinoblastoma, rhabdomyosarcoma, salivary gland cancer, skin cancer, small intestine cancer, stomach cancer, testicular cancer, thymus cancer, thyroid cancer, uterine sarcoma, vaginal cancer, vulvar cancer, Waldenstrom macroglobulinemia, and Wilms tumor. The cancer may be a blood-borne or hematological cancer, such as leukemia or lymphoma.III. Germline Mutational Signatures

[0074] A major issue presented in prior methods of detecting somatic variants or somatic variant signatures is effective elimination of germline mutational signatures when identifying somatic variants in a sample. This is because germline mutations vastly outnumber somatic mutations. Germline mutations occur with a frequency of 1 : 1,000, whereas somatic mutations commonly occur at a frequency of around 1 : 1,000,000. When sequencing DNA from a tumor sample to detect somatic mutations, it is difficult to ensure that only somatic mutations are being counted because even if 99.9% of germline mutations are identified and removed, the remaining germline mutations may still outnumber the true somatic mutations in the sample.

[0075] Described herein are germline mutational signatures that can be used in improving the identification of true somatic mutations or variants, and improve calling of somatic signatures. In order to counteract the effects of germline variant contamination, the present disclosure provides a signature based on only germline variants, which may be referred toherein as “Signature. gSNP” and a histogram of which is provided in FIG. 1. Signature. gSNP was created by aggregating all variants from germline / blood samples and detecting the prevalence of each triplicate from these variants is the disclosed germline signature. After sequencing a sample and during signature decomposition, germline variants will be attributed to the disclosed germline signature (i.e., Signature. gSNP), while somatic variants will decompose into the various somatic signatures that they most closely represent. Thus, the disclose germline signature improves the ability of a sequencer to accurate detect and identify somatic mutation signatures.

[0076] Briefly, a mutational signature, such as the disclosed germline signature, can be understood as the proportion or frequency of a set number of nucleotide changes and the invariable / unchanged positions flanking the site of the change. For example, if a change is a substitution, and only one flanking base on each side is considered, then combinatorically it the equation defining the set number of nucleotide changes can be exemplified as:4 types of bases before the substitution (e.g., A, T, G, or C) times 6 types of substitutions (e.g., OA, OG, OT, T>A, T>C, T>G) times 4 types of bases after the substitution (e.g., A, T, G, or C), which is 96 possible changes.

[0077] The disclosed germline mutational signature may be decomposed or deconvoluted from a subject whole genome sequencing reads to increase the signal of actual somatic mutational signatures. This decomposition may permit for the identification of somatic mutational signatures that would be masked by the signal of germline contamination, permitting more accurate analysis of the subject’s mutational signatures, which can be used to determine the mutational processes underlying the cancer, provide prognostic information to the subject, and / or identify possible treatment candidates.IV. Methods of Using Germline Mutational Signatures

[0078] The present disclosure provides methods for utilizing a germline mutation signature (e.g., Signature. gSNP, which is visualized in FIG. 1) to improve detection of somatic signature detection and somatic calling in general. As explained above, germline variants vastly outnumber somatic variants within the genome, and therefore germline contamination is a significant problem for detection of somatic signatures and other somatic alterations. The disclosed germline signatures can be used to deconvolute otherwise weak signals of somatic mutations signatures present in a given sample.

[0079] In particular, the present disclosure provides methods of detecting a somatic mutation signature in a sample, comprising: sequencing DNA from a sample obtained from a subject, thereby obtaining sequencing reads for the subject; detecting a plurality of mutations in the sequencing reads; decomposing the plurality of mutations with a reference set of somatic mutation signatures and a germline mutation signature comprising a frequency of a classification set unique to a germline genome comprising (i) nucleotide changes, (ii) at least one unchanged position that is 5’ of each nucleotide change in the set, and (iii) at least one unchanged position that is 3’ of each nucleotide change in the set, thereby removing germline contamination from the sequencing reads and detecting at least one somatic mutation signature for the subject.

[0080] Additionally, the present disclosure provides methods of improved somatic mutation calling, comprising: obtaining a tumor sample from a subject; sequencing DNA from the tumor sample, thereby obtaining sequencing reads; and calling a somatic mutation signation for the tumor sample by: detecting somatic mutations in the sequencing reads, and removing germline contamination from the sequencing reads by decomposing the sequencing read with a germline mutation signature; and detecting at least one somatic mutation signature; wherein the somatic mutation signature has a reduced negative effect from germline contamination relative a somatic mutation call that does not account for the germline mutational signature.

[0081] Additionally, the present disclosure provides computer implemented methods for improved detection of a somatic mutation signature, comprising: retrieving, by a processor, data associated with a tumor sample, wherein the data comprises DNA sequencing reads from the tumor sample; generating, by a computer device, a set of mutations in the sequencing reads; and decomposing, by a computer processor, the set of mutations with a reference set of somatic mutation signatures and a germline mutation signature, thereby removing germline contamination from the sequencing reads; and detecting at least one somatic mutation signature.

[0082] The “classification set” may comprise 96 triplicates, as described above. When analyzing a particular substitution site, there is a possibility of 4 types of bases before the substitution (e.g., A, T, G, or C), 6 types of substitutions (e.g., OA, OG, OT, T>A, T>C, T>G), and 4 types of bases after the substitution (e.g., A, T, G, or C), which is 96 possible changes. Accordingly, when the 96 triplicate arrange of the classification set can be understood as embodiments in which (i) the nucleotide changes comprise six types ofsubstitutions, (ii) the at least one unchanged position that is 5’ of each nucleotide change in the set comprises one of four types of bases 5’ to the substitution, and (iii) the at least one unchanged position that is 3’ of each nucleotide change in the set comprises one of four types of bases 3’ to the substitution. The germline signature visualized in FIG. 1 utilizes the 96 triplicate approach.

[0083] However, it is also possible to include more than one base 5’ or 3 ’of a given substitution of mutation site. For example, the classification could include 2, 3, 4, 5, 6, 7, 8, 9, or 10 based upstream and / or downstream (i.e., 5’ or 3’) of the given substitution or mutation site. In such embodiments, a classification set may be used in which (i) the nucleotide changes comprise six types of substitutions, (ii) the at least one unchanged position that is 5’ of each nucleotide change in the set comprises at least two, at least three, or at least four bases 5’ to the substitution, and (iii) the at least one unchanged position that is 3’ of each nucleotide change in the set comprises at least two, at least three, or at least four bases 3’ to the substitution.

[0084] Additionally, while the disclosed germline signature improves somatic calling and detection of somatic mutation signatures on its own, the disclosed methods of removing germline contamination can be combined with prior methods of removing germline contamination. This may occur either before or after decomposition with the disclosed germline signatures. For example, any of the disclosed methods may comprise removing germline mutations from the plurality of mutations in the sequence reads by comparing the sequences reads for the subject to a set of germline mutations, wherein the set of germline mutations are (i) determined by sequencing a germline sample from the subject, (ii) compiled from a database of germline mutations, or (iii) a combination of mutations found in sequences from a germline sample from the subject and mutations from a database of germline mutations. Useful databases for this purpose include, but are not limited to dbSNP and GNOME-DB. Generally, if these type of additional contamination steps are included in the disclosed methods, they will occur prior to decomposition with the disclosed germline signatures. Additionally or alternatively, any of the disclosed methods may comprise removing germline mutations from the plurality of mutations in the sequence reads by determining the frequency of the sequence reads comprising mutations and comparing the frequency to a database of germline mutations. Additionally or alternatively, any of the disclosed methods may comprise removing germline mutations from the plurality ofmutations in the sequence reads by determining a tumor mutation burden (TMB) score based on the number of somatic variants having a somatic mutation significance score above a threshold.

[0085] In general, the “reference set” of somatic mutation signatures will comprise the somatic signatures disclosed in Table 1, which include Signature 1A, Signature IB, Signature 2, Signature 3, Signature 4, Signature 5, Signature 6, Signature 7, Signature 8, Signature 9, Signature 10, Signature 11, Signature 12, Signature 13, Signature 14, Signature 15, Signature 16, Signature 17, Signature 18, Signature 19, Signature 20, and Signature 21. However, it should be understood that additional or alternative somatic signatures can be utilized in the disclosed methods.

[0086] A further aspect of the present disclosure is to provide methods for accurately quantifying, calculating, or estimating tumor mutation burden (TMB), but deconvoluting sequencing data (e.g., whole genome or whole exome sequencing reads) with the disclosed germline signature. A TMB value can distinguish between subjects who have different responsiveness to treatment with an anticancer agent; distinguish subjects who have increased overall survival, or progression-free survival after treatment with an anticancer agent from subjects who do not have increased survival; identify subjects of a population who benefit from or respond to a therapeutic treatment; or any combination thereof.

[0087] TMB values can be calculated using sequencing data obtained from a single sample from a subject (see, e.g., WO 2020 / 102661) and subtracting out germline mutations by deconvoluting the sequencing data with the disclosed germline signature. The sequencing data can be obtained by various methods known in the art including microelectrophoretic methods, sequencing by hybridization, real-time observation of single molecules, and cyclic- array sequencing. TMB values can be calculated using fragmentation sequencing data obtained from a single sample from a subject using an algorithm such as the one disclosed in WO 2020 / 102661 or other methods of calculating TMB. A SNP region of about one read length can be used to detect a variant near a SNP position. The read length can be sufficient to cover both the SNP position and the variant position. A set of SNP regions can provide the sequencing data needed to detect somatic variants and quantify a value of TMB for a sample. As used herein, a variant may be “near” a SNP position when the variant is within about one sequencing read length of the SNP position. A SNP region may be ±1 read length about aSNP position. Examples of human SNP position sets known in the art include SNP Array 6.0 (Affymetrix).

[0088] The sequencing data from a set of SNP regions can be plotted to show the number of variant positions (y axis) versus the Allele Ratio (x axis). The area under the curve can be an estimate of the presence of somatic variants. Using this arrangement of the sequencing data, by integrating the area under the curve a value for the total number of variants that are identified as somatic variants can be obtained. The value for the total number of variants that are identified as somatic variants can be a measure of TMB. Thus, a measure of TMB can be obtained as the area under a curve from an Allele Ratio of about 15% up to an Allele Ratio of about 85%, or up to an Allele Ratio of about 65%, where the curve plots the number of variant positions (y axis) in a set of SNP regions against the Allele Ratio (x axis) of the variants.

[0089] In some embodiments, only sequence reads having a length spanning both variant and SNP positions may be included in the assembly of a count matrix. In general, the read should cover the SNP and the position to be counted. A set of SNP positions can be used to obtain the sequencing data. The allele frequency of the SNP can be compared with the variant to determine whether the variant was germline or somatic. A SNP region of about one read length can be used to detect a variant near a SNP position. The read length can be sufficient to cover both the SNP position and the variant position. A set of SNP regions can provide the sequencing data needed to detect somatic variants and quantify a value of TMB for a sample.

[0090] A “good prognosis value” can be generated from a plurality of training cancer patients characterized as having “good outcome,” for example those who have not had cancer recurrence for a period of time, such as five years, or ten years, or more after initial treatment, or who have not had progression in their cancer five years, or ten years, or more after initial diagnosis. In contrast, a “poor prognosis value” can be generated from a plurality of training cancer patients defined as having “poor outcome,” for example those who have had cancer recurrence within five years, or ten years, or more after initial treatment, or who have had progression in their cancer within five years, or ten years, or more after initial diagnosis. Thus, a good prognosis value may represent an average level of TMB in patients having a “good outcome,” whereas a poor prognosis value may represent an average level of TMB in patients having a “poor outcome.” When a value of TMB is increased, a subject may have a poor prognosis.

[0091] The present disclosure can provide methods for treating a cancer patient or providing guidance for selecting the treatment of a patient. In this method, evaluation of TMB and one or more recurrence-associated clinical parameters may be determined. Active treatment may be recommended, initiated or continued if a sample from the patient has an elevated TMB and the patient has one or more recurrence-associated clinical parameters. Active surveillance may be recommended, or initiated, or continued if the patient has neither an elevated TMB, nor a recurrence-associated clinical parameter. In certain embodiments, TMB, or TMB and one or more clinical parameters may indicate that active treatment is recommended, or that a particular active treatment is recommended, or that aggressive treatment is recommended. In general, adjuvant therapy (e.g., chemotherapy, radiotherapy, HIFU, hormonal therapy, etc. after prostatectomy or radiotherapy) may be recommended for aggressive disease.

[0092] For the purposes of the disclosed methods, generally the subject has cancer, previously had cancer, or is at risk of developing cancer. The cancer can include, but is not limited to, any cancer selected from adrenal cancer, anal cancer, bile duct cancer, bladder cancer, bone cancer, a brain / CNS tumor, breast cancer, Castleman disease, cervical cancer, colon or rectum cancer, endometria cancer, esophagus cancer, a Ewing tumor, eye cancer, gallbladder cancer, a gastrointestinal carcinoid tumor, a gastrointestinal stromal tumor (GIST), gestational trophoblastic disease, Hodgkin disease, Kaposi sarcoma, kidney cancer, laryngeal and hypopharyngeal cancer, leukemia, liver cancer, lung cancer, lymphoma, malignant mesothelioma, multiple myeloma, myelodysplastic syndrome, nasal cavity or paranasal sinus cancer, nasopharyngeal cancer, neuroblastoma, oral cavity or oropharyngeal cancer, osteosarcoma, ovarian cancer, pancreatic cancer, penile cancer, a pituitary tumor, prostate cancer, retinoblastoma, rhabdomyosarcoma, salivary gland cancer, skin cancer, small intestine cancer, stomach cancer, testicular cancer, thymus cancer, thyroid cancer, uterine sarcoma, vaginal cancer, vulvar cancer, Waldenstrom macroglobulinemia, Wilms tumor, leukemia, and lymphoma.

[0093] The sample utilized for sequencing may be a tumor sample or a fluid sample. The fluid sample is a blood sample or a plasma sample. In general, a tumor sample, such as a biopsy, may be preferred for solid tumors, whereas a fluid sample (e.g., blood or plasma) may be preferred for hematological cancers like leukemia or lymphoma.

[0094] As described herein, the disclosed methods can improve detection of any somatic mutation signature. For example, the disclosed methods can be used to detect and / orimproved detection of a somatic mutation signature is selected from Signature 1 A, Signature IB, Signature 2, Signature 3, Signature 4, Signature 5, Signature 6, Signature 7, Signature 8, Signature 9, Signature 10, Signature 11, Signature 12, Signature 13, Signature 14, Signature 15, Signature 16, Signature 17, Signature 18, Signature 19, Signature 20, and Signature 21, or any combination thereof. Further, decomposing the mutations comprise normalizing at least one of the foregoing somatic mutation signatures.

[0095] It should be understood that, for the purposes of the disclosed methods, the type of sequencing performed on the sample is not particularly limited. Somatic mutation signatures may be discernable from whole genome sequencing, whole exome sequencing, or targeted sequencing (i.e., sequencing of particular, predetermined genes or locations within the genome). Similarly, somatic calling, in general, can be performed using whole genome sequencing, whole exome sequencing, or targeted sequencing. As such, any of the disclosed methods may comprise performing whole genome sequencing, whole exome sequencing, or targeted gene sequencing.

[0096] Sequencing may be performed by, for example, Next Generation Sequencing (NGS). Deep sequencing may allow for more sensitive detection, and so the depth of the sequencing may be at least 50X, at least 100X, at least 150X, at least 200X, at least 250X, at least 300X, at least 350X, at least 400X, at least 450X, at least 500X, at least 550X, at least 600X, at least 650X, at least 700X, at least 750X, at least 800X, at least 850X, at least 900X, at least 950X, or at least 1000X. In other words, the depth of the sequencing may be about 50X, about 100X, about 150X, about 200X, about 250X, about 300X, about 350X, about 400X, about 450X, about 500X, about 550X, about 600X, about 650X, about 700X, about 750X, about 800X, about 850X, about 900X, about 950X, or about 1000X.

[0097] An individual’s cancer may be characterized by the combination of somatic mutations that the cancer cells have accumulated over time and may be correlated with, but not limited to, disease progression, disease characterization, and treatment outcomes or efficacy. Thus, detection of the somatic mutations that are specific to the cancer cells allows for more accurate test results based on the genetic profile of the cancer cells specifically. Thus, being able to accurately detect somatic mutation signatures using the disclosed germline mutation signature can improve determination of cancer risk, improve determination of cancer status, and improve determination of responsiveness to chemotherapy or other forms of therapy, among other applications.EXAMPLES

[0098] The present technology is further illustrated by the following Examples, which should not be construed as limiting in any way.Example 1

[0099] This example shows how signature models perform with and without the inclusion of the disclosed germline signature, as germline contamination is increased. Signature weights were produced for two samples, TCGA-A2-A0T0 and TCGA-A8-A07R, as germline contamination goes from zero to 1500% of the somatic variant count.

[0100] TCGA-A2-A0T0:

[0101] This sample without germline contamination decomposed entirely into Signature.3 (HRD pathway deficiency), shown in FIG. 2.

[0102] At 50% germline SNP contamination, the model without the disclosed germline signature produced a new Signature.1 A, while the model with the disclosed germline signature attributed the germline contamination correctly (see FIG. 3).

[0103] At 100% germline SNP contamination, a spurious signature was generated within the disclosed germline signature model (see FIG. 4). This resulted in the apparent detection of a new Signature.20, which was a false signal that was attributable to the added contamination.

[0104] This pattern, of an additional spurious signature generation, continued from 100%- 800% germline contamination (see FIGs 5-7). Indeed, as the contamination levels rose, additional false signatures appeared even when the disclosed germline signature model was used to deconvolute the data.

[0105] Starting at the 900% germline contamination level, the model without the disclosed germline signature was no longer able to detect the presence of Signature.3. In contrast, the model with the disclosed germline signature was able to detect Signature.3 all the way up to 1500% contamination (see FIGs 8-9).

[0106] Thus, the inclusion of the disclosed germline signature allowed the model to continue to call the correct signature (Signature.3) for at higher weight for a given contamination level. The correct call of Signature.3 would have been undetectable at physiological levels of germline contamination if the model was not deconvoluted with the disclosed germline signature.

[0107] TCGA-A8-A07R:

[0108] TCGA-A8-A07R was a sample with a more complicated starting signature, as the decomposition results in three significant weights, Signature.3, Signature.4, and Signature.9 (FIG. 10).

[0109] Immediately upon the addition of germline contamination, Signature.9 was lost from all models and spurious secondary signatures were generated (see FIG. 11).

[0110] This pattern, of additional spurious signature inclusion, continued from 10% to 300% contamination (see FIGs 12-15), and at 100% contamination Signature.4 was lost within the disclosed germline signature model, whereas it was lost at 200% within the model without disclosed germline signature. After 300% contamination, the model without disclosed germline signature lost Signature.3 as well (FIG. 16).[OHl] Between 400-1500%, the model with the disclosed germline signature continued to express Signature.3, while the model without the disclosed germline signature contained no remaining initial signatures (FIGs 17-19).

[0112] Thus, the inclusion of the disclosed germline signature (i.e., “Signature. gSNP”) allows for detection of Signature.3, both presence of and at higher weights, as germline SNP contamination increases. Moreover, the inclusion of the disclosed germline signature (i.e., “Signature. gSNP”) allows for more accurate determination of TMB. As shown in FIG. 20, without deconvulation with the disclosed germline signature (i.e., “Signature. gSNP”), the TMB was consistently determined to be significantly higher than the estimated TMB for each sample. In contrast, once germline mutations were subtracted out by virtue of deconvolution with the disclosed germline signature (i.e., “Signature. gSNP”), the TMB calculations were much closer to the estimated TMB for each sample.EQUIVALENTS

[0113] The present technology is not to be limited in terms of the particular embodiments described in this application, which are intended as single illustrations of individual aspects of the present technology. Many modifications and variations of this present technology can be made without departing from its spirit and scope, as will be apparent to those skilled in the art. Functionally equivalent methods and apparatuses within the scope of the present technology, in addition to those enumerated herein, will be apparent to those skilled in the art from the foregoing descriptions. Such modifications and variations are intended to fall within the scope of the present technology. It is to be understood that this present technology is not limited to particular methods, reagents, compounds, compositions or biological systems, which can, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting.

[0114] All patents, patent applications, provisional applications, and publications referred to or cited herein are incorporated by reference in their entirety, including all figures and tables, to the extent that are not inconsistent with the explicit teachings of this specification.

Claims

CLAIMSWhat is claimed is:

1. A method of detecting a somatic mutation signature in a sample, comprising: sequencing DNA from a sample obtained from a subject, thereby obtaining sequencing reads for the subject; detecting a plurality of mutations in the sequencing reads; decomposing the plurality of mutations with a reference set of somatic mutation signatures and a germline mutation signature comprising a frequency of a classification set unique to a germline genome comprising (i) nucleotide changes, (ii) at least one unchanged position that is 5’ of each nucleotide change in the set, and (iii) at least one unchanged position that is 3’ of each nucleotide change in the set, thereby removing germline contamination from the sequencing reads and detecting at least one somatic mutation signature for the subject.

2. The method of claim 1, wherein the classification set comprises 96 triplicates, wherein (i) the nucleotide changes comprise six types of substitutions, (ii) the at least one unchanged position that is 5’ of each nucleotide change in the set comprises one of four types of bases 5’ to the substitution, and (iii) the at least one unchanged position that is 3’ of each nucleotide change in the set comprises one of four types of bases 3’ to the substitution.

3. The method of claim 1 or 2, wherein the frequency of the classification set is the frequency provided in FIG. 1.

4. The method of claim 1, wherein (i) the nucleotide changes comprise six types of substitutions, (ii) the at least one unchanged position that is 5’ of each nucleotide change in the set comprises at least two, at least three, or at least four bases 5’ to the substitution, and (iii) the at least one unchanged position that is 3’ of each nucleotide change in the set comprises at least two, at least three, or at least four bases 3 ’ to the substitution.

5. The method of any one of claims 1-4, further comprising removing germline mutations from the plurality of mutations in the sequence reads by(a) comparing the sequences reads for the subject to a set of germline mutations, wherein the set of germline mutations are (i) determined by sequencing a germline sample from the subject, (ii) compiled from a database of germline mutations, or (iii) a combination of mutations found in sequences from a germline sample from the subject and mutations from a database of germline mutations;(b) determining the frequency of the sequence reads comprising mutations and comparing the frequency to a database of germline mutations; or(c) determining a tumor mutation burden (TMB) score based on the number of somatic variants having a somatic mutation significance score above a threshold.

6. The method of any one of claims 1-5, wherein the reference set of somatic mutation signatures comprises Signature 1 A, Signature IB, Signature 2, Signature 3, Signature 4, Signature 5, Signature 6, Signature 7, Signature 8, Signature 9, Signature 10, Signature 11, Signature 12, Signature 13, Signature 14, Signature 15, Signature 16, Signature 17, Signature 18, Signature 19, Signature 20, and Signature 21.

7. The method of any one of claims 1-6, wherein the subject has cancer, previously had cancer, or is at risk of developing cancer.

8. The method of any one of claims 1-7, wherein the sample is a tumor sample or a fluid sample.

9. The method of claim 8, wherein the cancer is selected from adrenal cancer, anal cancer, bile duct cancer, bladder cancer, bone cancer, a brain / CNS tumor, breast cancer, Castleman disease, cervical cancer, colon or rectum cancer, endometria cancer, esophagus cancer, a Ewing tumor, eye cancer, gallbladder cancer, a gastrointestinal carcinoid tumor, a gastrointestinal stromal tumor (GIST), gestational trophoblastic disease, Hodgkin disease, Kaposi sarcoma, kidney cancer, laryngeal and hypopharyngeal cancer, leukemia, liver cancer, lung cancer, lymphoma, malignant mesothelioma, multiple myeloma, myelodysplastic syndrome, nasal cavity or paranasal sinus cancer, nasopharyngeal cancer, neuroblastoma, oral cavity or oropharyngeal cancer, osteosarcoma, ovarian cancer, pancreatic cancer, penile cancer, a pituitary tumor, prostate cancer, retinoblastoma, rhabdomyosarcoma, salivary gland cancer, skin cancer, small intestine cancer, stomach cancer, testicular cancer, thymuscancer, thyroid cancer, uterine sarcoma, vaginal cancer, vulvar cancer, Waldenstrom macroglobulinemia, and Wilms tumor.

10. The method of claim 8, wherein the fluid sample is a blood sample or a plasma sample.

11. The method of claim 10, wherein the subject has, previously had, or is at risk of developing leukemia or lymphoma.

12. The method of any one of claims 1-11, wherein sequencing DNA from the sample comprises whole genome sequencing, whole exome sequencing, or targeted gene sequencing.

13. The method of any one of claims 1-12, wherein the at least one somatic mutation signature is selected from Signature 1 A, Signature IB, Signature 2, Signature 3, Signature 4, Signature 5, Signature 6, Signature 7, Signature 8, Signature 9, Signature 10, Signature 11, Signature 12, Signature 13, Signature 14, Signature 15, Signature 16, Signature 17, Signature 18, Signature 19, Signature 20, and Signature 21, or any combination thereof.

14. The method of any one of claims 1-13, wherein decomposing the plurality of mutations comprises normalizing the at least one somatic mutation signature.

15. A method of improved somatic mutation calling, comprising: obtaining a tumor sample from a subject; sequencing DNA from the tumor sample, thereby obtaining sequencing reads; and calling a somatic mutation signation for the tumor sample by: detecting somatic mutations in the sequencing reads, and removing germline contamination from the sequencing reads by decomposing the sequencing reads with a germline mutation signature; and detecting at least one somatic mutation signature; wherein the somatic mutation signature has a reduced negative effect from germline contamination relative a somatic mutation call that does not account for the germline mutational signature.

16. The method of claim 15, wherein the germline mutation signature comprises a frequency of a classification set that is unique to a germline genome, the classification set comprising (i) nucleotide changes, (ii) at least one unchanged position that is 5’ of each nucleotide change in the set, and (iii) at least one unchanged position that is 3’ of each nucleotide change in the set.

17. The method of claim 16, wherein the classification set comprises 96 triplicates, wherein (i) the nucleotide changes comprise six types of substitutions, (ii) the at least one unchanged position that is 5’ of each nucleotide change in the set comprises one of four types of bases 5’ to the substitution, and (iii) the at least one unchanged position that is 3’ of each nucleotide change in the set comprises one of four types of bases 3’ to the substitution.

18. The method of claim 16 or 17, wherein the frequency of the classification set is the frequency provided in FIG. 1.

19. The method of claim 16, wherein (i) the nucleotide changes comprise six types of substitutions, (ii) the at least one unchanged position that is 5’ of each nucleotide change in the set comprises at least two, at least three, or at least four bases 5’ to the substitution, and (iii) the at least one unchanged position that is 3’ of each nucleotide change in the set comprises at least two, at least three, or at least four bases 3 ’ to the substitution.

20. The method of any one of claims 15-19, wherein removing germline contamination further comprises:(a) comparing the sequences reads for the subject to a set of germline mutations, wherein the set of germline mutations are (i) determined by sequencing a germline sample from the subject, (ii) compiled from a database of germline mutations, or (iii) a combination of mutations found in sequences from a germline sample from the subject and mutations from a database of germline mutations;(b) determining the frequency of the sequence reads comprising mutations and comparing the frequency to a database of germline mutations; or(c) determining a tumor mutation burden (TMB) score based on the number of somatic variants having a somatic mutation significance score above a threshold.

21. The method of any one of claims 15-20, wherein the subject has or previously had a cancer selected from adrenal cancer, anal cancer, bile duct cancer, bladder cancer, bone cancer, a brain / CNS tumor, breast cancer, Castleman disease, cervical cancer, colon or rectum cancer, endometria cancer, esophagus cancer, a Ewing tumor, eye cancer, gallbladder cancer, a gastrointestinal carcinoid tumor, a gastrointestinal stromal tumor (GIST), gestational trophoblastic disease, Hodgkin disease, Kaposi sarcoma, kidney cancer, laryngeal and hypopharyngeal cancer, leukemia, liver cancer, lung cancer, lymphoma, malignant mesothelioma, multiple myeloma, myelodysplastic syndrome, nasal cavity or paranasal sinus cancer, nasopharyngeal cancer, neuroblastoma, oral cavity or oropharyngeal cancer, osteosarcoma, ovarian cancer, pancreatic cancer, penile cancer, a pituitary tumor, prostate cancer, retinoblastoma, rhabdomyosarcoma, salivary gland cancer, skin cancer, small intestine cancer, stomach cancer, testicular cancer, thymus cancer, thyroid cancer, uterine sarcoma, vaginal cancer, vulvar cancer, Waldenstrom macroglobulinemia, Wilms tumor, leukemia, and lymphoma.

22. The method of any one of claims 15-21, wherein the sample is a tumor sample.

23. The method of any one of claims 15-21, wherein the sample is a fluid sample.

24. The method of claim 23, wherein the fluid sample is a blood sample or a plasma sample.

25. The method of any one of claims 15-24, wherein sequencing DNA from the tumor sample comprises whole genome sequencing, whole exome sequencing, or targeted gene sequencing.

26. The method of any one of claims 15-25, wherein the at least one somatic mutation signature is selected from Signature 1 A, Signature IB, Signature 2, Signature 3, Signature 4, Signature 5, Signature 6, Signature 7, Signature 8, Signature 9, Signature 10, Signature 11, Signature 12, Signature 13, Signature 14, Signature 15, Signature 16, Signature 17, Signature 18, Signature 19, Signature 20, and Signature 21, or any combination thereof.

27. A computer implemented method for improved detection of a somatic mutation signature, comprising:retrieving, by a processor, data associated with a tumor sample, wherein the data comprises DNA sequencing reads from the tumor sample; generating, by a computer device, a set of mutations in the sequencing reads; and decomposing, by a computer processor, the set of mutations with a reference set of somatic mutation signatures and a germline mutation signature, thereby removing germline contamination from the sequencing reads; and detecting at least one somatic mutation signature.

28. The method of claim 27, wherein the reference set of somatic mutation signatures comprises Signature 1 A, Signature IB, Signature 2, Signature 3, Signature 4, Signature 5, Signature 6, Signature 7, Signature 8, Signature 9, Signature 10, Signature 11, Signature 12, Signature 13, Signature 14, Signature 15, Signature 16, Signature 17, Signature 18, Signature 19, Signature 20, and Signature 21.

29. The method of claim 27 or 28, wherein the germline mutation signature comprises a frequency of a classification set that is unique to a germline genome, the classification set comprising (i) nucleotide changes, (ii) at least one unchanged position that is 5’ of each nucleotide change in the set, and (iii) at least one unchanged position that is 3’ of each nucleotide change in the set.

30. The method of claim 29, wherein the classification set comprises 96 triplicates, wherein (i) the nucleotide changes comprise six types of substitutions, (ii) the at least one unchanged position that is 5’ of each nucleotide change in the set comprises one of four types of bases 5’ to the substitution, and (iii) the at least one unchanged position that is 3’ of each nucleotide change in the set comprises one of four types of bases 3’ to the substitution.

31. The method of claim 29 or 30, wherein the frequency of the classification set is the frequency provided in FIG. 1.

32. The method of claim 29, wherein (i) the nucleotide changes comprise six types of substitutions, (ii) the at least one unchanged position that is 5’ of each nucleotide change in the set comprises at least two, at least three, or at least four bases 5’ to the substitution, and (iii) the at least one unchanged position that is 3’ of each nucleotidechange in the set comprises at least two, at least three, or at least four bases 3 ’ to the substitution.

33. The method of any one of claims 27-32, wherein decomposing the set of mutations comprises normalizing the at least one somatic mutation signature.

34. A method of improved detection of tumor mutation burden (TMB), comprising: obtaining a tumor sample from a subject; sequencing DNA from the tumor sample, thereby obtaining sequencing reads; and detecting TMB by: quantifying somatic mutations in the sequencing reads, and removing germline contamination from the sequencing reads by decomposing the sequencing reads with a germline mutation signature.

35. The method of claim 34, wherein the germline mutation signature comprises a frequency of a classification set that is unique to a germline genome, the classification set comprising (i) nucleotide changes, (ii) at least one unchanged position that is 5’ of each nucleotide change in the set, and (iii) at least one unchanged position that is 3’ of each nucleotide change in the set.

36. The method of claim 35, wherein the classification set comprises 96 triplicates, wherein (i) the nucleotide changes comprise six types of substitutions, (ii) the at least one unchanged position that is 5’ of each nucleotide change in the set comprises one of four types of bases 5’ to the substitution, and (iii) the at least one unchanged position that is 3’ of each nucleotide change in the set comprises one of four types of bases 3’ to the substitution.

37. The method of claim 35 or 36, wherein the frequency of the classification set is the frequency provided in FIG. 1.

38. The method of any one of claims 34-37, wherein the subject has or previously had a cancer selected from adrenal cancer, anal cancer, bile duct cancer, bladder cancer, bone cancer, a brain / CNS tumor, breast cancer, Castleman disease, cervical cancer, colon or rectum cancer, endometria cancer, esophagus cancer, a Ewing tumor, eye cancer, gallbladder cancer, a gastrointestinal carcinoid tumor, a gastrointestinal stromal tumor (GIST), gestational trophoblastic disease, Hodgkin disease, Kaposisarcoma, kidney cancer, laryngeal and hypopharyngeal cancer, leukemia, liver cancer, lung cancer, lymphoma, malignant mesothelioma, multiple myeloma, myelodysplastic syndrome, nasal cavity or paranasal sinus cancer, nasopharyngeal cancer, neuroblastoma, oral cavity or oropharyngeal cancer, osteosarcoma, ovarian cancer, pancreatic cancer, penile cancer, a pituitary tumor, prostate cancer, retinoblastoma, rhabdomyosarcoma, salivary gland cancer, skin cancer, small intestine cancer, stomach cancer, testicular cancer, thymus cancer, thyroid cancer, uterine sarcoma, vaginal cancer, vulvar cancer, Waldenstrom macroglobulinemia, Wilms tumor, leukemia, and lymphoma.

39. The method of any one of claims 34-38, wherein the sample is a tumor sample.

40. The method of any one of claims 34-38, wherein the sample is a fluid sample.

41. The method of claim 40, wherein the fluid sample is a blood sample or a plasma sample.

42. The method of any one of claims 34-41, wherein sequencing DNA from the tumor sample comprises whole genome sequencing, whole exome sequencing, or targeted gene sequencing.

Citation Information

Patent Citations

  • Alkaline pouch cell with coated terminals

    WO2020102661A1