Detection of non-cancerous somatic mutations

By analyzing copy number variants in buffy coat DNA samples using NGS, the method addresses the limitations of existing techniques in detecting non-cancerous somatic mutations, providing a comprehensive and cost-effective means to identify age-related health risks in both humans and animals.

JP2025536559APending Publication Date: 2025-11-07ZOETIS SERVICES LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025524586
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-01
Filing Date
2023-10-31
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Current methods fail to effectively detect non-cancerous somatic mutations, particularly clonal hematopoiesis of indeterminate potential (CHIP), which are associated with an increased risk of hematologic malignancies and cardiovascular diseases, especially in aging populations, and are limited by targeted deep sequencing techniques that focus on single-point mutations.

Method used

A method involving the analysis of copy number abnormalities in buffy coat DNA samples using next-generation sequencing (NGS) to identify recurrent copy number variants (CNVs) in white blood cell genomic DNA (WBC gDNA) and the absence of these variants in cell-free DNA (cfDNA), providing an early indication of age-related somatic changes.

Benefits of technology

This approach enhances the detection of non-cancerous somatic mutations, offering a cost-effective and broad-spectrum analysis applicable to both humans and animal models, enabling early identification of health risks such as hematologic malignancies and cardiovascular diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025536559000001_ABST
    Figure 2025536559000001_ABST
Patent Text Reader

Abstract

The methods, systems, and compositions provided herein enable improved methods for detecting non-cancerous somatic mutations in subjects by measuring buffy coat somatic mutations to improve identification of samples with age-related somatic mutations.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to methods for detecting, characterizing, or managing non-cancerous somatic mutations in a subject by analyzing copy number abnormalities in a buffy coat DNA sample. [Background technology]

[0002] background Somatic mutations are a normal consequence of aging, despite the existence of effective DNA repair mechanisms in human cells. While most somatic mutations are silent, some affect genes important for self-renewal and differentiation, leading to the clonal expansion of specific cells with selective proliferation under positive clonal selection pressure. In aging hematopoietic tissues, hematopoietic stem cells (HSCs) lose their self-renewal capacity and become functionally restricted due to pathogenic mutations, leading to clonal hematopoiesis (CH), or HSC expansion and dominance. CH is frequent and can affect more than 10% of individuals over the age of 50 (Park, Curr Stem Cell Rep, 2018; Gondek, Hematol, 2021; Kusne, Leuk Res, 2022).

[0003] The phenomenon of somatic genomic alterations in hematopoietic cell lines, commonly referred to as clonal hematopoiesis of indeterminate potential (CHIP) or age-associated clonal hematopoiesis (ARCH), has been well characterized in humans and refers to somatic changes that occur as a result of clonal expansion of the blood system, particularly white blood cell populations. These alterations tend to be associated with aging, often show consistent signals over time, and are frequently found at recurrent locations within the genome (Jaiswal, Science, 2019; Jaiswal, N Eng J Med, 2014; Genovese, N Engl J Med, 2014; Gao, Nat Commun, 2021; Saiki, Nat Med, 2021). Most human patients with these somatic alterations do not have concurrent cancer. In patients with concurrent cancer (e.g., solid tumors), these alterations are typically not found in the corresponding tumor tissue. In humans, CHIP has been shown to be associated with an increased risk of developing primary or secondary hematologic malignancies and represents a novel biomarker for cancer risk prediction. Two population-based studies, each involving over 10,000 individuals, independently demonstrated that the presence of CHIP alterations is associated with an approximately 10-fold increased relative risk of developing hematologic cancer compared with age-matched controls, with an absolute risk of approximately 0.5% per year (Jaiswal, N Eng J Med, 2014; Genovese, N Engl J Med, 2014). For this reason, major medical centers have established "CHIP clinics" for close clinical monitoring of patients with known CHIP findings (Grisham, MSKCC, 2023; Mayo Clinic, 2023; Cleveland Clinic, 2023; Siteman Cancer Center ARCH Clinic, 2023). Summary of the Invention

[0004] Described herein are methods and compositions for the detection of non-cancerous somatic mutations in a subject. In some embodiments, the methods disclosed herein are capable of identifying such mutations where other methods known in the art have failed to identify these mutations.

[0005] Some embodiments provided herein relate to a method for detecting non-cancerous somatic mutations in a subject. In some embodiments, the method includes obtaining a sample from the subject, isolating genomic DNA from the sample, preparing a DNA sequencing library from the genomic DNA (gDNA), sequencing the gDNA to generate sequencing data, and detecting at least one copy number abnormality in the sequencing data. In some embodiments, the at least one copy number abnormality indicates the non-cancerous somatic mutation. In some embodiments, the subject is a canine subject. In some embodiments, the sample is a whole blood sample. In some embodiments, the sequencing data is aligned to a reference genome. In some embodiments, the gDNA is extracted from white blood cells. In some embodiments, the white blood cells are present in a buffy coat. In some embodiments, the method further includes isolating cell-free DNA (cfDNA). In some embodiments, the cfDNA is extracted from plasma. In some embodiments, the gDNA is matched to the cfDNA. In some embodiments, the absence of the non-cancerous somatic mutation is detected in the cell-free DNA sequencing library.

[0006] Some embodiments provided herein relate to a method for detecting non-cancerous somatic mutations in a subject. In some embodiments, the method includes isolating white blood cell (WBC) genomic DNA (gDNA) from a buffy coat sample from the subject, creating a DNA sequencing library from the WBC gDNA, generating sequencing data by sequencing the DNA sequencing library, aligning the sequencing data to a reference genome, and detecting at least one copy number abnormality. In some embodiments, the at least one copy number abnormality indicates a non-cancerous somatic mutation. In some embodiments, the method further includes isolating cell-free DNA (cfDNA) from a plasma sample from the subject, creating a cfDNA sequencing library from the cfDNA, generating cell-free sequencing data by sequencing the cfDNA sequencing library, aligning the cell-free sequencing data to a reference genome, and detecting the absence of non-cancerous somatic mutations in the cell-free alignment. In some embodiments, the method further includes isolating non-buffy coat DNA from a non-buffy coat sample from the subject, creating a non-buffy coat DNA sequencing library from the non-buffy coat DNA, generating non-buffy coat sequencing data by sequencing the non-buffy coat DNA sequencing library, aligning the non-buffy coat sequencing data to a reference genome, and detecting the absence of non-cancerous somatic mutations in the non-buffy coat alignment. In some embodiments, one or more non-buffy coat somatic mutations are detected within the non-buffy coat alignment. In some embodiments, the non-buffy coat somatic mutations are not matched to the non-cancerous somatic mutations. In some embodiments, the subject is a canine subject. In some embodiments, the buffy coat sample and the plasma sample are matched.

[0007] Some embodiments provided herein relate to a method for measuring age-related somatic changes in a canine subject. In some embodiments, the method comprises measuring copy number variations (CNVs) from white blood cell genomic DNA (WBC gDNA) obtained from a buffy sample from the canine subject, and detecting the absence of CNVs from cell-free DNA (cfDNA) obtained from a matched plasma sample from the canine subject. [Brief explanation of the drawings]

[0008] [Figure 1A] An example of a patient with a WBC gDNA-specific CNV is shown, where the CNV was identified in WBC gDNA (Figure 1A) but not in cfDNA (Figure 1B) in an 11-year-old neutered male mixed-breed dog with no clinical evidence of cancer. [Figure 1B] An example of a patient with a WBC gDNA-specific CNV is shown, where the CNV was identified in WBC gDNA (Figure 1A) but not in cfDNA (Figure 1B) in an 11-year-old neutered male mixed-breed dog with no clinical evidence of cancer. [Figure 2A] Figure 2A shows an exemplary CNV identified in WBC gDNA that is distinct from the CNVs in cfDNA (Figure 2B) and tumor tissue (Figure 2C) in a 10-year-old neutered male mixed-breed dog with hepatocellular carcinoma. [Figure 2B] Figure 2A shows an exemplary CNV identified in WBC gDNA that is distinct from the CNVs in cfDNA (Figure 2B) and tumor tissue (Figure 2C) in a 10-year-old neutered male mixed-breed dog with hepatocellular carcinoma. [Figure 2C] Figure 2A shows an exemplary CNV identified in WBC gDNA that is distinct from the CNVs in cfDNA (Figure 2B) and tumor tissue (Figure 2C) in a 10-year-old neutered male mixed-breed dog with hepatocellular carcinoma. [Figure 3A]Figure 3 shows longitudinal blood samples from a 13-year-old, spayed, female mixed-breed dog with no history of cancer and no suspicion of cancer, demonstrating persistence of chromosome 25 (CFA25) gain / loss and chromosome 36 (CFA36) gain in WBC gDNA, with little or no change in signal amplitude and absence of this finding in cfDNA over a 9-month period. Figure 3A shows WBC gDNA and cfDNA at time 1, Figure 3B shows WBC gDNA and cfDNA at time 2 (3 months), and Figure 3C shows WBC gDNA and cfDNA at time 3 (9 months). [Figure 3B] Figure 3 shows longitudinal blood samples from a 13-year-old, spayed, female mixed-breed dog with no history of cancer and no suspicion of cancer, demonstrating persistence of chromosome 25 (CFA25) gain / loss and chromosome 36 (CFA36) gain in WBC gDNA, with little or no change in signal amplitude and absence of this finding in cfDNA over a 9-month period. Figure 3A shows WBC gDNA and cfDNA at time 1, Figure 3B shows WBC gDNA and cfDNA at time 2 (3 months), and Figure 3C shows WBC gDNA and cfDNA at time 3 (9 months). [Figure 3C] Figure 3 shows longitudinal blood samples from a 13-year-old, spayed, female mixed-breed dog with no history of cancer and no suspicion of cancer, demonstrating persistence of chromosome 25 (CFA25) gain / loss and chromosome 36 (CFA36) gain in WBC gDNA, with little or no change in signal amplitude and absence of this finding in cfDNA over a 9-month period. Figure 3A shows WBC gDNA and cfDNA at time 1, Figure 3B shows WBC gDNA and cfDNA at time 2 (3 months), and Figure 3C shows WBC gDNA and cfDNA at time 3 (9 months). [Figure 4] 1 shows an exemplary plot showing the normalized age distribution of dogs with and without cancer and the absence of WBC gDNA-specific CNVs compared to dogs with and without cancer and the presence of WBC gDNA-specific CNVs. DETAILED DESCRIPTION OF THE INVENTION

[0009] In the following detailed description, reference is made to the accompanying drawings, which form a part of this specification. In the drawings, like symbols typically identify like elements unless otherwise dictated by context. The illustrative embodiments described in the detailed description, drawings, and claims are not meant to be limiting. Other embodiments may be utilized, and other changes may be made, without departing from the spirit or scope of the subject matter presented herein. It will be readily understood that the aspects of the present disclosure, as generally described and illustrated in the figures herein, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations, all of which are expressly contemplated herein. All references cited herein are expressly incorporated by reference in their entirety and for the specific disclosures cited herein.

[0010] Embodiments of the present disclosure relate to methods, systems, and compositions for identifying non-cancerous somatic mutations in a sample from a subject by isolating buffy coat DNA from the buffy coat sample from the subject.

[0011] Embodiments relate to methods, systems, and compositions for screening a subject for the possibility of having a non-cancerous somatic mutation. In some embodiments, a buffy coat sample is screened by isolating a buffy coat sample from a subject, such as a canine, and creating a DNA sequencing library from the buffy coat DNA, sequencing the DNA sequencing library to generate sequencing data, aligning the sequencing data to a reference genome, and detecting at least one copy number abnormality. In some embodiments, the at least one copy number abnormality indicates a non-cancerous somatic mutation.

[0012] The term "buffy coat" has its ordinary meaning as understood in light of this specification and refers to the dull yellow layer resulting from centrifugation of an anticoagulated blood sample. The buffy coat contains most of the blood cells and platelets. Typically, the buffy coat represents less than about 1% of the blood sample, with the plasma accounting for about 55% and the red blood cell layer accounting for about 45%. The buffy coat layer is formed or "sandwiched" between the underlying red blood cell or red blood cell layer and the overlying clear plasma layer. Buffy coats are useful for extracting relatively large amounts of genomic DNA for the methods described herein.

[0013] The term "somatic" as used herein has its ordinary meaning as understood in the context of this specification and refers to a type of genetic variant in a subject that is found in a cancer tumor or cells derived therefrom. These genetic changes are thought to occur during cell division that leads to tumor expansion, but can also occur in the cancer stem cell lineage that leads to tumor initiation. Somatic mutations can cause cancer or other non-cancerous diseases. Non-cancerous somatic mutations can occur during development and can affect cell proliferation or alter cell function. In addition to cancer, other diseases such as immunodeficiency, neurofibromatosis, and hypertension are the product of somatic mutations.

[0014] The term clonal hematopoiesis of undetermined potential (CHIP) was introduced in 2015 to describe individuals who harbored somatic mutations known to be associated with hematologic malignancies but who were healthy and did not exhibit any other symptoms or diagnostic criteria for such hematologic malignancies (Steensma, Blood, 2015). Additional studies have shown that clonally restricted somatic mutations are not limited to individuals with hematologic cancers but can also be detected in healthy individuals with normal blood cell counts (Shlush, Nature, 2014; Busque, Blood 2015). Table 1 lists the top eight genes most frequently observed in CHIP and the conditions in which they are more common. Over 70% of all mutations identified in CHIP occur in two genes with opposing functions—DNMT3A and TET2—and result in a similar clinical outcome: increased all-cause mortality (Jaiswal, N Eng J Med, 2014). [Table 1]

[0015] Numerous large-scale sequencing studies have demonstrated that 10% of people over 65 years of age have detectable CHIP mutations, with the prevalence increasing to up to 20% in people over 90 years of age, depending on the study and sequencing method used (McKerrel, Cell Rep, 2015; Jaiswal, N Engl J Med, 2014; Genovese, N Engl J Med, 2014; Kusne, Leuk Res, 2022). CHIP mutations have also been detected in people with different conditions, including cancer patients after treatment with chemotherapy or CAR T-cell therapy (Miller, Blood Adv, 2021), ANCA-associated autoimmune vasculitis (Arends, Haematologica, 2020), rheumatoid arthritis (Savola, Nat Commun, 2017), ulcerative colitis (Zhang, Exp Hematol, 2019), and those with HIV (Dharan, Nat Med, 2021).

[0016] CHIP-positive individuals have a higher risk of developing hematologic malignancies and higher all-cause mortality (Xie, Nat Med, 2014; Jaiswal, N Eng J Med, 2014; Genovese, N Engl J Med, 2014). CHIP-positive individuals have a 0.5-1% / year progression risk of developing hematologic cancer, but the higher all-cause mortality is mediated by an increased risk of developing cardiovascular disease, myocardial infarction, and stroke (Jaiswal, N Engl J Med, 2017; Gibson, Clin Canc Res, 2018; Min, J Intern Med, 2020; Evans, Ann Rev Path, Mechanisms of Disease, 2020).

[0017] In addition to being a well-documented risk factor for malignancies in humans, CHIP is also associated with an increased risk of cardiovascular disease (HR 1.9; 95% CI, 1.4-2.7) (Jaiswal, N Engl J Med, 2017) and may be a risk factor for cerebrovascular events (e.g., stroke) (Mayerhover, Stroke, 2023; Steensma, Blood, 2020). Although there is a growing literature surrounding CHIP in humans, the population-level prevalence and clinical correlates of CHIP in dogs or other species have not been systematically studied.

[0018] Apart from aging, other selective pressures known to contribute to CH include tobacco use, cancer treatments, particularly radiation and platinum-based chemotherapy, and certain inflammatory and autoimmune diseases, including ulcerative colitis (UC) and rheumatoid arthritis (RA) (Arends, Haematologica, 2020; Savola, Nat Commun, 2017; Zhang, Exp Hematol, 2019; Dharan, Nat Med, 2021; Bekele, Rhem Dis Clin No Am, 2020).

[0019] Although significant advances in genome sequencing technology have improved the detection of CHIP mutations and uncovered new genes associated with CHIP, further research is needed to understand the broad range of conditions associated with CHIP and identify ways to ameliorate its negative impact in aging societies. To achieve this, cost-effective testing that can be applied not only to humans but also to other animal models is crucial but lacking. Most existing techniques used to study CHIP involve targeted deep sequencing focused on single-point mutations, limiting testing to humans.

[0020] The ability to detect somatic mutations such as CHIP has benefited from advances in precision medicine, such as tumor tissue-based molecular testing, and more recently, non-invasive blood tests (liquid biopsies) that rely on the analysis of cellular cfDNA shed into the bloodstream by somatic cells (Cohen, Science, 2018; Gale, PLoS One 2018; Plagnol, PLoS One 2018; Jahangiri, Cancers (Basel), 2019). Following the reported detection of cancer-associated alterations in circulating tumor DNA in the plasma of cancer patients (Chen, Nat Med, 1996; Nawroz, Nat Med, 1996), significant efforts have been made to develop tests to detect the presence of somatic alterations in blood using next-generation sequencing (NGS) (Cohen, Science, 2018; Diehl, Nat Med, 2008; Bettegowda, Sci Transl Med, 2014; Aravanis, Cell, 2017; Liu, Ann Oncol, 2020). Recent studies have shown that CHIP is frequently observed in patients with solid tumors, with approximately 30% harboring CH mutations in their blood (Bekel, Rhem Dis Clin NO Am, 2020). In recent years, the increased use of tumor tissue and plasma-based sequencing has facilitated clinical decision-making and facilitated the unintended discovery of CH mutations through blood-based testing, ultimately increasing the prevalence of CH (Cohen, Science, 2018; Gale, PLoS One, 2018; Plagnol, PLoS One, 2018). Given the significant clinical implications that CH mutations can have on cancer patients, several institutions have established CH clinics to treat patients whose genetic testing identifies CH (Ehrhart, Small An Clin Oncol, 2020; Uzuelli, Clin Chim Acta, 2009).

[0021] Recently, a blood-based liquid biopsy test using next-generation sequencing (NGS) was developed and clinically deployed for cancer detection in dogs. Clinical validation of this test included 1,100 client-owned dogs diagnosed with cancer and probable cancer-free dogs (Flory, PLoS One, 2022). Since becoming commercially available in 2021, this test has been further performed in thousands of dogs. This recent ability to test large numbers of dogs using liquid biopsy provides an unprecedented opportunity to study the genomic profiles of a broad canine patient population. During canine liquid biopsy testing, cell-free DNA (cfDNA) is extracted from plasma, and genomic DNA (gDNA) is extracted from white blood cells (WBCs) present in the buffy coat. The extracted DNA is then subjected to NGS to identify genomic alterations. Cell-free DNA contains DNA removed from various tissues throughout the body, including tumors (if present). If genomic alterations are identified in cfDNA, this indicates a high probability of cancer being present in the body. If a genomic alteration is identified in gDNA, this may indicate the presence of constitutional (germline) abnormalities in the patient, including mosaicism, certain hematologic malignancies (if the corresponding CNV is also identified in cfDNA), or age-related somatic alterations (e.g., CHIP).

[0022] While most human studies have focused on characterizing single nucleotide variants (SNVs) specifically in genes associated with CHIP, recent reports have shown that copy number variants (CNVs) associated with CHIP are also important risk factors for the development of leukemia and cardiovascular disease (Gao, Nat Commun, 2021; Saiki, Nat Med, 2021; Jaiswal, N Engl J Med, 2017). NGS-based liquid biopsies offer an opportunity to study whether similar changes are found in dogs and whether they may be useful as biomarkers for predicting the risk of cancer or other diseases. In some embodiments, the methods described herein use NGS. As used herein, the term "next-generation sequencing" (NGS) generally refers to a technology for massively parallel measurement of the sequences of nucleic acid molecules, such as deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) molecules. NGS has largely replaced Sanger sequencing, which was later developed and considered the first-generation DNA sequencing technology.

[0023] Embodiments of the methods described herein relate to the identification of recurrent CNVs in canine WBC gDNA. In some embodiments, the recurrent CNVs persist with little or no change in amplitude over time and are absent in cfDNA (and tumor tissue, as well as in dogs with concurrent cancer). In some embodiments, the methods described herein provide an early indication of the presence of age-related somatic changes in canines, similar to the previously described phenomenon of CHIP in humans.

[0024] Previous studies related to CHIP in humans have involved the analysis of SNVs, and recent studies have also reported that CNVs associated with CHIP are similarly age-related and associated with a higher risk of developing leukemia in individuals with or without other pre-existing cancers. Additionally, these studies have shown that the coexistence of CNVs and SNVs associated with CHIP is associated with a higher cumulative incidence of leukemia than the exclusive occurrence of either type of genomic alteration. A recent study using blood samples from dogs diagnosed with cancer found that a small percentage (4.3%, 4 / 93) of the genes associated with CHIP in human patients had genomic alterations (SNVs). However, no tumor tissue samples were available to confirm that these alterations were not actually derived from the cancer itself, and no comparison was made between WBC-derived gDNA and plasma-derived cfDNA.

[0025] Embodiments of the present disclosure relate to the analysis of WBC gDNA-specific CNVs in canine subjects. Some embodiments provided herein relate to assessing the prevalence and distribution of WBC gDNA-specific SNVs in large dog populations. In some embodiments, the methods relate to assessing the interaction between genomic alterations associated with multiple types of CHIP and clinical outcomes.

[0026] Embodiments of the present disclosure relate to measuring the population-level frequency of age-associated somatic copy number aberrations in canine WBC gDNA and characterizing the types of alterations observed in these subjects.

[0027] As used herein, "copy number abnormality" or CNA has its ordinary meaning as understood in light of the present specification and refers to a change in the copy number of a particular gene sequence or component within an individual's genome, which can range from a loss of one or more copies of the gene component (deletion) to a gain of many additional copies of the gene component (amplification). One type of CNA is "aneuploidy," which generally refers to an abnormal number of total chromosomes. Typically, aneuploidy can result from a genetic imbalance resulting from cancer or other disease. In some embodiments, aneuploidy results in either three ("trisomy") or only one ("monosomy") chromosome. In some embodiments, measuring aneuploidy can be used in the context of cancer diagnosis, as described above.

[0028] Sequencing of a DNA sequencing library can be performed via any method recognized by those skilled in the art, including, for example, targeted sequencing or genome-wide sequencing. Other non-limiting examples include methods using nanopores, emulsions, and "sequencing-by-binding" cycle sequencing. As used herein, a "nucleic acid library" is an intentionally created collection of nucleic acids that can be prepared synthetically or biosynthetically in a variety of different formats (e.g., libraries of soluble molecules and libraries of oligonucleotides tethered to resin beads, silica chips, or other solid supports). A DNA sequencing library is a sample of genomic DNA fragments purified from a particular biological source representing the entire genome of that biological source. In a DNA sequencing library, genomic DNA fragments can be 3'- and 5'-ligated to primer and adapter sequences for further analysis (e.g., sequencing analysis) of the genomic DNA fragments.

[0029] For example, preparation of a DNA library for sequencing can begin with fragmentation of a DNA sample purified from a specific biological source. Fragmentation defines the molecular entry point for sequencing reads. In the next step, DNA ends can be enzymatically repaired, and adenine (A) can be added to the 3' ends of the DNA fragments. The (A)-tailed DNA fragments can then be amplified as templates to ligate double-stranded, partially complementary adapters to the DNA fragments. The DNA library can then be size-selected and amplified to improve the quality of the sequence reads. The amplification reaction introduces PCR primers specific to the adapter sequences required for flow cell sequencing.

[0030] In some embodiments, non-cancerous somatic mutations are screened for by isolating cell-free DNA from a plasma sample from a subject, creating a cell-free DNA sequencing library from the cell-free DNA, sequencing the cell-free DNA sequencing library to generate cell-free sequencing data, aligning the cell-free sequencing data to a reference genome, and detecting the absence of non-cancerous somatic mutations in the cell-free alignment. The model is derived from the fragment size distribution profile of at least one fragment. As used herein, "fragment distribution" has its ordinary meaning as understood by those skilled in the art and thus refers to the length, sequence, fragmentation, and other distribution characteristics of at least one DNA fragment collected from a cfDNA sample. "Fragment size distribution" is understood as a fragment distribution that focuses on the size of the fragments, including length or fragmentation. As disclosed herein, a model can be created for a subject suspected of having cancer or a tumor, as well as a model for one or more healthy subjects. These models can then be compared to each other to monitor significant differences. Non-limiting examples of models include summary statistics, the number and shape of nucleosome peaks, the proportion of fragments longer or shorter than a certain threshold, the proportion of fragments in a certain interval, approximating the data with a statistical distribution, and discriminative learning methods such as support vector machines or neural networks. Non-limiting examples of detectable differences include peak location (mode), peak height (weight), peak spread (scale), the proportion of fragments longer or shorter than a certain threshold, the amplitude of oscillations, the overall shape of the fragment size distribution, principal component values, and the Kullback-Leibler (KL) divergence between two models. In some embodiments, a statistically significant difference between the fragment size distribution in a subject suspected of having cancer or a tumor and the fragment size distribution in one or more healthy subjects indicates the presence of cancer or a tumor. In some embodiments, a non-statistically significant difference between the fragment size distribution in a subject suspected of having cancer or a tumor and the fragment size distribution in one or more healthy subjects indicates the absence of cancer or a tumor.

[0031] As used herein, the term "sequence alignment" or "sequence mapping" refers to a method of aligning DNA or RNA sequences relative to one another to identify regions of interest, such as regions of similarity or variation. Such alignments may be the result of functional, structural, or evolutionary relationships between sequences. Aligned sequences of nucleotides are typically represented as rows in a matrix. Well-known algorithms for sequence alignment include, for example, the Needleman-Wunsch algorithm, the Smith-Waterman algorithm, the Waterman-Eggert algorithm, or the Burrows-Wheeler transformation. Well-known tools for sequence alignment include, for example, BLAST, BLAT, WMBOSS, Clustal, BWA, and Bowtie.

[0032] In some embodiments, non-cancerous somatic mutations are screened for by isolating non-buffy coat DNA from a non-buffy coat sample from the subject, creating a cell-free DNA sequencing library from the cell-free DNA, generating non-buffy coat sequencing data by sequencing the non-buffy coat DNA sequencing library, aligning the non-buffy coat sequencing data to a reference genome, and detecting the absence of non-cancerous somatic mutations in the cell-free alignment.

[0033] In some embodiments, one or more non-buffy coat somatic mutations are detected within the non-buffy coat alignment. In some embodiments, the non-buffy coat somatic mutations are not matched with non-cancerous somatic mutations.

[0034] There are various methods for determining the fragment size distribution of cfDNA in a subject. In one embodiment, a blood sample is collected from a subject. Circulating free DNA (cfDNA) is obtained from blood. In some embodiments, the blood sample contains circulating tumor DNA (ctDNA). cfDNA is isolated by removing blood cells from the sample, so that only cfDNA remains in the sample. In some embodiments, a random PCR primer set for whole genome sequencing is added to the sample, and fragments are amplified while maintaining the original fragment length in the sample.

[0035] A polymerase is then added to the mixture to extend the primers over the entire length of each fragment. The amplified fragments, in one embodiment, may include sequencing ends formatted for use in a next-generation sequencing (NGS) system to identify the nucleotide sequence within the fragment.

[0036] The methods and compositions provided herein improve the detection, diagnosis, staging, screening, treatment, and management of cancer in subjects, particularly humans, mammals, and other types of subjects. As mentioned above, embodiments include identifying the distribution of cfDNA fragments in circulating biological fluids, such as blood. In one embodiment, the nucleic acid sequence elements are found in circulating tumor DNA in blood. In some embodiments, the nucleic acid sequence elements can be found in cell-free DNA, saliva, or urine.

[0037] As used herein, "detection," with respect to measuring cancer or tumors, includes the use of instruments used to observe and record a signal corresponding to the level or measurement of cancer, or the materials necessary to generate such a signal. In various embodiments, detection includes any suitable method, including amplification, sequencing, arrays, fluorescence, chemiluminescence, surface plasmon resonance, surface acoustic wave, mass spectrometry, infrared spectroscopy, Raman spectroscopy, atomic force microscopy, scanning tunneling microscopy, electrochemical detection methods, nuclear magnetic resonance, quantum dots, etc.

[0038] Some embodiments provided herein relate to a kit. In some embodiments, the kit is for determining cancer in a subject. In some embodiments, the kit includes a whole genome sequencing primer for amplifying cfDNA in a biological sample from a subject, and a polymerase for amplifying the primer.

[0039] It should be understood that the analysis described herein can be part of a larger diagnostic suite used to determine the overall health status of a subject.For example, analysis of the fragment size distribution of cfDNA in a subject can be used simultaneously or sequentially with other methods for cancer detection, diagnosis, stage determination, screening, monitoring, treatment, and management, including additional genetic variance analysis.These procedures can be useful for detecting various cancers, including leukemia, squamous cell carcinoma, feline mammary gland carcinoma, mast cell tumor, bladder cancer, osteosarcoma, hemangiosarcoma, or various other cancers that afflict a subject.

[0040] In some embodiments, the method includes obtaining or having obtained a biological sample from a subject. In some embodiments, the sample is a liquid biopsy sample, such as a blood sample. In some embodiments, the sample comprises cfDNA. In some embodiments, the sample is provided in a volume of less than 10 mL, such as 10 mL, 9 mL, 8 mL, 7 mL, 6 mL, 5 mL, 4 mL, 3 mL, 2 mL, 1 mL, 500 μL, 250 μL, or 100 μL, or an amount within a range defined by any two of the aforementioned values. In some embodiments, the sample comprises DNA in an amount of 10 μg or less, such as 10 μg, 5 μg, 1 μg, 500 ng, 100 ng, 50 ng, 10 ng, 5 ng, 1 ng, 500 pg, 100 pg, 50 pg, 10 pg, 9 pg, 8 pg, 7 pg, 6 pg, 5 pg, 4 pg, 3 pg, 2 pg, or 1 pg, or an amount within a range defined by any two of the aforementioned values. In some embodiments, the method includes purifying DNA from the sample. Purification of DNA can be achieved using DNA purification techniques, including, for example, extraction techniques, precipitation, chromatography, bead-based methods, or commercially available kits for DNA purification. In some embodiments, the method can be used to determine a putative cancer type or cancer tissue of origin based on one or more of the fragment size distribution characteristics.

[0041] definition Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art. All patents, applications, published applications, and other publications referenced herein are incorporated by reference in their entirety unless otherwise specified. In the event that there are multiple definitions for terms herein, the definitions in this section shall prevail unless otherwise specified.

[0042] As used herein, "a" or "an" can mean one or more than one.

[0043] As used herein, the terms "about" or "approximately" have their ordinary meaning as understood by one of ordinary skill in the art, and thus indicate that a value includes the inherent variation of error for the method used to determine the value or the variation that exists among multiple determinations.

[0044] The dimensions and values ​​disclosed herein are not to be understood as being strictly limited to the exact numerical values ​​recited. Instead, unless otherwise specified, each such dimension is intended to mean both the recited value and a functionally equivalent range surrounding that value. For example, a dimension disclosed as "20 mm" is intended to mean "about 20 mm."

[0045] Throughout this specification, unless the context requires otherwise, the words "comprise," "comprises," and "comprising" will be understood to imply the inclusion of the stated step or element or group of steps or elements, but not the exclusion of any other step or element or group of steps or elements. "Consisting of" means including, but not limited to, what follows the word "consisting of." Thus, the word "consisting of" indicates that the recited elements are necessary or mandatory, and that no other elements may be present. "Consisting essentially of" means including all elements listed after the word, and is limited to other elements that do not interfere with or contribute to the activity or action specified in this disclosure for those recited elements. Thus, the phrase "consisting essentially of" indicates that the recited elements are necessary or mandatory, but that other elements are optional and may or may not be present depending on whether they materially affect the activity or action of the recited elements.

[0046] As used herein, the terms "function" and "functional" have their plain and ordinary meaning as understood in light of this specification and refer to biological, enzymatic, or therapeutic function.

[0047] As used herein, the term "yield" of any given substance, compound, or material has its plain and ordinary meaning as understood in light of this specification, and refers to the actual total amount of the substance, compound, or material relative to the expected total amount. For example, the yield of a substance, compound, or material is, is about, is at least, or is 80, 81, 82, 83, 84, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the expected total amount. or at least about 80, 81, 82, 83, 84, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100%, or less than or equal to 80, 81, 82, 83, 84, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100%, including all decimal points therebetween. Yields can be affected by reaction or process efficiency, undesired side effects, decomposition, quality of input substances, compounds, or materials, or loss of desired substances, compounds, or materials during any step of manufacturing.

[0048] As used herein, the term "isolated" has its plain and ordinary meaning as understood in light of this specification and refers to substances and / or entities that (1) have been separated from at least some of the components with which they are associated when originally produced (in nature and / or in an experimental setting) and / or (2) have been produced, prepared, and / or manufactured by the hand of man. Isolated substances and / or entities may be at least about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, about 95%, about 98%, about 99%, substantially 100%, or 100% of the other components with which they are initially associated (or ranges including and / or spanning the aforementioned values). The separation may be 8%, about 99%, substantially 100%, or 100%, at least about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, about 95%, about 98%, about 99%, substantially 100%, or 100%, or less than about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, about 95%, about 98%, about 99%, substantially 100%, or 100%.In some embodiments, the isolated agent is 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, substantially 100%, or 100% pure (or ranges including and / or ranging from the aforementioned values), or is about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, substantially 100%, or 100% pure. 6%, about 97%, about 98%, about 99%, substantially 100%, or 100% pure (or ranges including and / or ranging from the aforementioned values), or at least 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, substantially 100%, or 100% pure (or ranges including and / or ranging from the aforementioned values). range), or is at least about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, substantially 100%, or 100% pure (or ranges including and / or ranging from the foregoing values), or is at least about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, substantially 100%, or 100% pure or less (or ranges inclusive and / or ranging therein), or about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, substantially 100%, or 100% pure or less (or ranges inclusive and / or ranging therein). As used herein, an "isolated" material can be "pure" (e.g., substantially free of other components). As used herein, the term "isolated cell" can refer to a cell that is not contained in a multicellular organism or tissue.

[0049] As used herein, "in vivo" is given its plain and ordinary meaning as understood in light of the present specification and refers to the performance of a method within a living organism, usually an animal, including a human, a mammal, or a plant, or the living cells that make up these organisms, as opposed to a tissue extract or a dead organism.

[0050] As used herein, "ex vivo" is given its plain and ordinary meaning as understood in light of this specification and refers to the practice of methods outside of a living organism with little change in natural conditions.

[0051] As used herein, "in vitro" is given its plain and ordinary meaning as understood in light of the present specification and refers to the performance of methods outside biological conditions, e.g., in a petri dish or test tube.

[0052] As used herein, "nucleic acid," "nucleic acid molecule," or "nucleotide" refers to a polynucleotide or oligonucleotide, such as deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), oligonucleotides, fragments produced by polymerase chain reaction (PCR), and fragments produced by any of ligation, cleavage, endonuclease action, exonuclease action, and synthetic production. Nucleic acid molecules can be composed of monomers that are naturally occurring nucleotides (such as DNA and RNA), or analogs of naturally occurring nucleotides (e.g., enantiomeric forms of naturally occurring nucleotides), or combinations of both. Modified nucleotides can have changes in the sugar moiety and / or the pyrimidine or purine base moiety. Sugar modifications include, for example, replacement of one or more hydroxyl groups with halogens, alkyl groups, amines, and azide groups, or the sugar can be functionalized as an ether or ester. Additionally, the entire sugar moiety can be replaced with sterically and electronically similar structures, such as azasugars and carbocyclic sugar analogs. Examples of modifications in the base moiety include alkylated purines and pyrimidines, acylated purines or pyrimidines, or other well-known heterocyclic substituents. Nucleic acid monomers can be linked by phosphodiester bonds or analogs of such linkages. Phosphodiester bond analogs include phosphorothioates, phosphorodithioates, phosphoroselenoates, phosphorodiselenoates, phosphoroanilothioates, phosphoranilidates, phosphoramidates, and the like. The term "nucleic acid molecule" also includes so-called "peptide nucleic acids," which contain naturally occurring or modified nucleobases attached to a polyamide backbone. Nucleic acids can be either single-stranded or double-stranded.

[0053] As used herein, the terms "peptide," "polypeptide," and "protein" have their plain and ordinary meaning as understood in light of the present specification and refer to macromolecules composed of amino acids linked by peptide bonds. The numerous functions of peptides, polypeptides, and proteins are known in the art and include, but are not limited to, enzymatic, structural, transport, defensive, hormonal, or signal transduction functions. Peptides, polypeptides, and proteins are often, but not always, produced biologically by ribosomal complexes using nucleic acid templates, although chemical synthesis is also available. By manipulating nucleic acid templates, peptide, polypeptide, and protein mutations can be made, such as substitutions, deletions, truncations, additions, duplications, or fusions of more than one peptide, polypeptide, or protein.These fusions of two or more peptides, polypeptides, or proteins can be adjacently linked on the same molecule or can include extra amino acids in between, such as linkers, repeats, epitopes, or tags, or any other sequence, which can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 150, 200, or 300 bases. or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 150, 200, or 300 bases in length 95, 100, 150, 200, or 300 bases in length, or at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 150, 200, or 300 bases in length The length of the polypeptide may be up to 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 150, 200, or 300 bases, or up to about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 150, 200, or 300 bases, or any length within the range defined by any two of the foregoing lengths. As used herein, the term "downstream" with respect to a polypeptide has its plain and ordinary meaning as understood in light of the present specification, and refers to sequences that are after the C-terminus of the preceding sequence.As used herein, the term "upstream" on a polypeptide has its plain and ordinary meaning as understood in light of this specification, and refers to a sequence that precedes the N-terminus of a subsequent sequence.

[0054] The terms "DNA fragment" and "nucleic acid fragment" have their ordinary meaning as understood by those of skill in the art and refer to polynucleotide sequences that can be obtained from a genome at any point along the genome and encompass any sequence of nucleotides.

[0055] The term "fragment size distribution" has its ordinary meaning as understood by those of skill in the art and refers to information regarding one or more of the total number of nucleic acid fragments present in a sample, the size of one or more nucleic acid fragments in a sample, the absolute or relative abundance levels of nucleic acid fragments of a particular size or size range, and the absolute or relative abundance levels of nucleic acid fragments of different sizes present in a sample.

[0056] The term "fragment size" has its ordinary meaning as used herein with respect to nucleic acid molecules, as understood by one of skill in the art, and refers to the number of base pairs of the nucleic acid, indicating the length of the molecule.

[0057] As used herein, the term "gene" has its plain and ordinary meaning as understood in light of the present specification and generally refers to a portion of a nucleic acid that encodes a protein or functional RNA, although the term may optionally encompass regulatory sequences. Those skilled in the art will understand that the term "gene" can include gene regulatory sequences (e.g., promoters, enhancers, etc.) and / or intron sequences. It will be further understood that the definition of a gene includes reference to nucleic acids that do not encode proteins, but rather encode functional RNA molecules such as tRNAs and miRNAs. In some cases, a gene includes regulatory sequences involved in transcription or message production or composition. In other embodiments, a gene includes a transcribed sequence that encodes a protein, polypeptide, or peptide. In accordance with the terminology described herein, an "isolated gene" can include transcribed nucleic acid(s), regulatory sequences, coding sequences, etc., isolated substantially free from other such sequences, e.g., other naturally occurring genes, regulatory sequences, polypeptide- or peptide-encoding sequences, etc. In this regard, the term "gene" will be used for simplicity to refer to a nucleic acid comprising a transcribed nucleotide sequence and its complement. As will be understood by those skilled in the art, the functional term "gene" includes both genomic sequences, RNA or cDNA sequences, or smaller engineered nucleic acid segments, including nucleic acid segments of the non-transcribed portions of a gene, including, but not limited to, the non-transcribed promoter or enhancer regions of a gene. Smaller engineered gene nucleic acid segments may be expressed or adapted to express using nucleic acid engineering techniques, proteins, polypeptides, domains, peptides, fusion proteins, mutants, and / or the like.

[0058] The terms "cancer" and "cancerous" have their ordinary meanings as understood in light of this specification and refer to or describe a physiological condition in an animal that is typically characterized by uncontrolled cell growth. A "tumor" contains one or more cancerous cells. In some embodiments, a tumor is a solid tumor. There are several main types of cancer. Carcinomas are cancers that originate in epithelial cells, such as skin cells or the lining of the intestinal tract. Sarcomas are cancers that originate in mesenchymal cells, such as bone, cartilage, fat, muscle, blood vessels, or other connective or supportive tissues. Leukemia is a cancer that originates in hematopoietic cells, such as the bone marrow, and causes large numbers of abnormal blood cells to be produced and enter the blood. Lymphoma and multiple myeloma are cancers that originate in the lymphatic cells of the lymph nodes. Central nervous system cancers are cancers that originate in the central nervous system and spinal cord.

[0059] As used herein, the phrase "allele" or "allelic variant" has its ordinary meaning as understood in light of the present specification and refers to a variant of a genetic locus or gene. In some embodiments, a particular allele of a genetic locus or gene is associated with a particular phenotype, e.g., an altered risk of developing a disease or condition, likelihood of progressing to a particular stage of the disease or condition, amenability to a particular therapeutic agent, susceptibility to infection, immune function, etc.

[0060] As used herein, the term "amplification" has its ordinary meaning as understood in light of the present specification and refers to any method known in the art for copying a target nucleic acid, thereby increasing the number of copies of a selected nucleic acid sequence. Amplification can be exponential or linear. The target nucleic acid can be either DNA or RNA. Typically, the sequence so amplified forms an "amplicon." Amplification can be achieved in a variety of ways, including, but not limited to, polymerase chain reaction ("PCR"), transcription-based amplification, isothermal amplification, rolling circle amplification, and the like. Amplification can be performed with relatively equal amounts of each primer in a primer pair to generate a double-stranded amplicon. However, asymmetric PCR can be used to amplify primarily or exclusively single-stranded products, as is well known in the art (e.g., Poddar et al., Molec. And Cell. Probes 14:25-32 (2000)). This can be achieved by using each pair of primers by significantly reducing the concentration of one primer in a pair compared to the other primer in that pair (for example, 100-fold difference).The amplification by asymmetric PCR is generally linear.Those skilled in the art will understand that different amplification methods can be used together.

[0061] As used herein, "amplicon" has its ordinary meaning as understood in light of this specification and refers to the nucleic acid sequence being amplified, as well as the nucleic acid polymer resulting from the amplification reaction. Amplicons can be formed artificially, for example, via polymerase chain reaction (PCR) or ligase chain reaction (LCR), or naturally, via gene replication.

[0062] As used herein, the terms "individual," "subject," "host," or "patient" have their ordinary meanings as understood by those of skill in the art, and thus include a human or non-human mammal. The term "mammal" is used in its ordinary biological sense. As such, it specifically includes, but is not limited to, primates, including monkeys (chimpanzees, apes, monkeys), humans, cows, horses, sheep, goats, pigs, rabbits, dogs, cats, rodents, rats, mice, or guinea pigs.

[0063] As used herein, the term "liquid biopsy" has its ordinary meaning as understood in light of this specification and refers to the collection and testing of a sample, the sample being a non-solid biological tissue such as blood.

[0064] As used herein, the term "cfDNA" has its ordinary meaning as understood herein and refers to circulating cell-free DNA, including DNA fragments released into plasma. cfDNA can include circulating tumor deoxyribonucleic acid (ctDNA).

[0065] As used herein, the term "ctDNA" has its ordinary meaning as understood in the context of this specification and refers to circulating tumor DNA, including tumor-derived fragmented DNA in the bloodstream that is not associated with cells.

[0066] Some embodiments provided herein are described in the following enumerated alternative embodiments.

[0067] 1. A method for detecting a non-cancerous somatic mutation in a subject, the method comprising: obtaining a sample from the subject; isolating genomic DNA from the sample; preparing a DNA sequencing library from the genomic DNA (gDNA); sequencing the gDNA to generate sequencing data; and detecting at least one copy number abnormality in the sequencing data, wherein the at least one copy number abnormality is indicative of the non-cancerous somatic mutation.

[0068] 2. The method of alternative embodiment 1, wherein the subject is a canine subject.

[0069] 3. The method of alternative embodiment 1 or 2, wherein the sample is a whole blood sample.

[0070] 4. The method of any one of alternative embodiments 1 to 3, wherein the sequencing data is aligned to a reference genome.

[0071] 5. The method of any one of alternative embodiments 1-4, wherein gDNA is extracted from white blood cells.

[0072] 6. The method of alternative embodiment 5, wherein the leukocytes are present in the buffy coat.

[0073] 7. The method of any one of alternative embodiments 1-6, further comprising isolating cell-free DNA (cfDNA).

[0074] 8. The method of alternative embodiment 7, wherein cfDNA is extracted from plasma.

[0075] 9. The method of alternative embodiment 7 or 8, wherein the gDNA is matched to cfDNA.

[0076] 10. The method of any one of alternative embodiments 7-9, wherein the absence of non-cancerous somatic mutations is detected in a cell-free DNA sequencing library.

[0077] 11. A method for detecting a non-cancerous somatic mutation in a subject, the method comprising: isolating white blood cell (WBC) genomic DNA (gDNA) from a buffy coat sample from the subject; creating a DNA sequencing library from the WBC gDNA; generating sequencing data by sequencing the DNA sequencing library; aligning the sequencing data to a reference genome; and detecting at least one copy number abnormality, wherein the at least one copy number abnormality is indicative of the non-cancerous somatic mutation.

[0078] 12. The method of alternative embodiment 11, further comprising isolating cell-free DNA (cfDNA) from a plasma sample from the subject; creating a cfDNA sequencing library from the cfDNA; generating cell-free sequencing data by sequencing the cfDNA sequencing library; aligning the cell-free sequencing data to a reference genome; and detecting the absence of non-cancerous somatic mutations in the cell-free alignment.

[0079] 13. The method of alternative embodiment 11 or 12, further comprising isolating non-buffy coat DNA from a non-buffy coat sample from the subject; creating a non-buffy coat DNA sequencing library from the non-buffy coat DNA; generating non-buffy coat sequencing data by sequencing the non-buffy coat DNA sequencing library; aligning the non-buffy coat sequencing data to a reference genome; and detecting the absence of non-cancerous somatic mutations in the non-buffy coat alignment.

[0080] 14. The method of alternative embodiment 13, wherein one or more non-buffy coat somatic mutations are detected within the non-buffy coat alignment, and the non-buffy coat somatic mutations are not matched with non-cancerous somatic mutations.

[0081] 15. The method of any one of alternative embodiments 11-14, wherein the subject is a canine subject.

[0082] 16. The method of any one of alternative embodiments 11-15, wherein the buffy coat sample and the plasma sample are matched.

[0083] 17. A method for measuring age-related somatic changes in a canine subject, comprising measuring copy number variations (CNVs) from white blood cell genomic DNA (WBC gDNA) obtained from a buffy coat sample from the canine subject, and detecting the absence of CNVs from cell-free DNA (cfDNA) obtained from a matched plasma sample from the canine subject. [Example]

[0084] Embodiments of the present disclosure are further defined in the following examples. It should be understood that these examples are provided by way of illustration only. From the above discussion and these examples, those skilled in the art can ascertain the essential features described herein and can make various changes and modifications to the embodiments described herein to adapt to various uses and conditions without departing from the spirit and scope thereof. Thus, various modifications of the embodiments described herein, in addition to those shown and described herein, will be apparent to those skilled in the art from the foregoing description. Such modifications are also intended to fall within the scope of the appended claims. The disclosures of each reference set forth herein are incorporated herein in their entirety for the purposes of the present disclosure, which is hereby incorporated by reference.

[0085] Example 1 Isolation of buffy coat DNA from samples Blood samples were processed using a dual centrifugation protocol to separate plasma from the buffy coat (white blood cells). Cell-free DNA was extracted from the plasma using a proprietary bead-based chemistry protocol optimized to maximize cell-free DNA yield. DNA was extracted from the buffy coat using standard DNA extraction methods, such as the QIAamp DNA Mini Blood Kit (Qiagen). Libraries were prepared by incorporating universal adapters and barcodes into the sample DNA via ligation and universal PCR amplification. The amplified libraries were subjected to genome-wide sequencing for CNV analysis. All libraries were sequenced using an Illumina NovaSeq 6000.

[0086] Example 2 Copy number variant analysis Copy number variants were identified in dogs (4.8%) from the CANDiD study and a commercial cohort (N = 2925). Figures 1A and 1B show examples of CHIP CNVs in gDNA that were either distinct from cancer-derived somatic CNVs in cfDNA (Figure 1A) or clearly visible despite no evidence of signal in cfDNA in presumed non-cancer subjects (Figure 1B). Many of these CNVs recurred across patients, consistent with observations of recurrent genes affected by human CHIP mutations.

[0087] In 50 of the 139 affected patients, multiple plasma samples from follow-up time points were available for analysis. Of this group, 37 dogs showed persistent CHIP signals across time points, without any visible increase in signal amplitude suggestive of cancer progression. Matched tumor tissue was available for 11 patients with CHIP, demonstrating a copy number signature. Genomic profiles of tumor tissue and gDNA from these patients corresponded to a tissue confirmation rate of 9.1%; 10 of these cases were unrelated. In contrast, for subjects with cfDNA CNVs, matched tumor tissue was available and demonstrated a copy number signature in 94 subjects, representing a tissue confirmation rate of 92.6% (87 / 94). These results suggest that the coexistence of gDNA-specific somatic alterations was associated with the patients' cancer. Finally, we compared the normalized ages of subjects with CHIP (necessary to account for breed-specific differences in expected lifespan, calculated as age normalized to expected lifespan, which is a linear function of body weight) with the normalized ages of subjects with confirmed cancer who lacked a CHIP signal and with the normalized ages of subjects in a high-risk screening population with no known cancer and no evidence of CHIP. As shown in Figure 4, the ages of patients with CHIP were significantly higher than those of either of the other cohorts (p<0.001 for both comparisons).

[0088] Example 3 A blood-based liquid biopsy test using next-generation sequencing (NGS) has been developed and clinically deployed for cancer detection in dogs. During the canine liquid biopsy test, cell-free DNA (cfDNA) was extracted from plasma, and genomic DNA (gDNA) was extracted from white blood cells (WBCs) present in the buffy coat. The extracted DNA was then subjected to NGS to identify genomic alterations. Cell-free DNA contains DNA removed from various tissues throughout the body, including tumors (if present). Genomic alterations identified in cfDNA indicate a high likelihood of cancer in the body. Genomic alterations identified in gDNA may indicate the presence of constitutional (germline) abnormalities in patients, including mosaicism, certain hematologic malignancies (if the corresponding CNVs are also identified in cfDNA), or age-related somatic alterations (e.g., CHIP).

[0089] While most human studies have focused on characterizing single nucleotide variants (SNVs) specifically in genes associated with CHIP, recent reports have shown that copy number variants (CNVs) associated with CHIP are also important risk factors for the development of leukemia and cardiovascular disease (Gao, Nat Commun, 2021; Saiki, Nat Med, 2021; Jaiswal, N Engl J Med, 2017). NGS-based liquid biopsies offer an opportunity to study whether similar changes are found in dogs and whether they may be useful as biomarkers for predicting risk of cancer or other diseases.

[0090] As detailed in Table 2, a total of 4,870 client-owned dogs were evaluated in this example. Of this total, whole blood samples from 3,595 dogs, with or without clinical suspicion of cancer, were submitted by their veterinarians for commercial liquid biopsy testing at PetDx Laboratories in La Jolla, CA ("clinical cohort"); whole blood samples from 1,275 dogs, with or without a cancer diagnosis, were obtained as part of a larger study collection to support clinical validation of the NGS-based liquid biopsy test ("research cohort") (Flory, PLos One, 2022). All study subjects were enrolled under protocols approved by the Institutional Animal Care and Use Committee (IACUC) or facility-specific ethics (Flory, PLos One, 2022). All methods were performed in accordance with relevant guidelines and regulations and followed the recommendations of the ARRIVE guidelines. All study subjects were client-owned, and written informed consent was obtained from all owners. In a subset of patients diagnosed with cancer from the study cohort, matched tumor tissue samples were also available for analysis. In addition, a subset of patients across both the clinical and study cohorts submitted whole blood samples at multiple time points, allowing for longitudinal monitoring of genomic alterations. [Table 2] *Body weights were available for 3144 dogs in the clinical cohort.

[0091] Whole blood samples were collected from each patient. cfDNA was extracted from the plasma, and gDNA was extracted from WBCs present in the buffy coat. In a subset of patients with available matched tumor tissue, DNA was also extracted from the tissue (Flory, PLoS One, 2022; Kruglyak, Front Vet Sci, 2021). All extracted DNA specimens were subjected to library preparation and NGS as previously described (Flory, PLoS One, 2022). Sequencing data were analyzed using an internally developed bioinformatics pipeline to determine the presence of genomic alterations.

[0092] This example focuses on a large cohort of canine patients in whom specific CNVs were identified in WBC gDNA but not in the matched cfDNA samples. These findings are referred to herein as "WBC gDNA-specific CNVs," and examples are shown in Figures 1A and 1B. A subset of patients with WBC gDNA-specific CNVs also had concurrent CNVs identified in cfDNA (and / or tissue, if available), but these CNVs were distinct from the CNVs identified in WBC gDNA.

[0093] A subset of dogs with WBC gDNA-specific CNVs underwent clinical evaluation to determine the presence of cancer. The components of each evaluation varied but typically included a thorough physical examination of any observed tumors, lesions, and / or enlarged lymph nodes, laboratory tests (CBC, chemistry panel, and urinalysis), imaging (chest radiograph and / or abdominal ultrasound), and tissue sampling (via fine needle aspiration or biopsy).

[0094] If clinical evaluation determined that cancer was present, a definite or presumptive cancer diagnosis was assigned as shown in Table 3. A "definite" diagnosis was one in which cancer was confirmed by tissue-based testing (cytology or histopathology). A "presumptive" diagnosis was based on imaging, direct visualization / examination, or strong suspicion from cytology or histopathology. [Table 3] *The term "cancer evaluation" in the context of this example refers to a thorough workup performed when cancer is suspected (because of liquid biopsy results and / or because of the patient's clinical findings). The components of this workup are described herein. Cancer-free dogs in the study cohort had a clinical history and physical examination at the time of study enrollment.

[0095] To compare the age profiles of adult dogs with and without WBC gDNA-specific CNVs, a "normalized age" was calculated relative to life expectancy, and life expectancy was calculated as follows: Equation: Lifespan = 13.987-0.035 x Weight The mean age of the subjects was 18 years and the mean age of the subjects was 18 years. ...

[0096] The population of 4870 client-owned dogs included 3595 dogs from the clinical cohort and 1275 dogs from the research cohort (Table 3).

[0097] In the clinical cohort, cancer status was unknown at the time of sample submission. 962 dog samples were submitted with clinical suspicion of cancer (liquid biopsy was used as a diagnostic aid), 2399 dog samples were submitted without a documented suspicion of cancer (liquid biopsy was used as a screening test), and 234 dog samples were submitted without documentation of cancer suspicion. The mean age of dogs in the clinical cohort was 9.09 years (SD = 3.06; n = 3595), and the mean weight was 25.80 kg (SD = 13.48; n = 3144).

[0098] In the study cohort, 569 dogs had a definitive diagnosis of cancer and 706 dogs were presumed cancer-free. The mean age of dogs in the study cohort was 7.17 years (SD = 3.68; n = 1275), and the mean weight was 28.28 kg (SD = 12.04; n = 1275). A wide range of purebred and mixed breed dogs was represented in both cohorts.

[0099] WBC gDNA-specific CNVs were identified in 164 dogs (129 from the clinical cohort and 35 from the research cohort) out of 4870 dogs in the total study population (3.4%; 95% CI: 2.9-3.9). For 126 of these dogs (107 from the clinical cohort and 19 from the research cohort), CNVs were identified only in WBC gDNA; no CNVs of any kind were observed in plasma-derived cfDNA.

[0100] Cancer clinical evaluation was performed in 42 of 126 dogs (28 from the clinical cohort and 14 from the research cohort) using WBC gDNA specific CNVs.

[0101] Of the 28 patients in the clinical cohort, 18 provided specimens not suspected of cancer at the time of blood collection (e.g., tests used for cancer screening). Within this group, 78% (14 / 18; 95% CI: 51.9–92.6) had no evidence of cancer after clinical evaluation, while 22% (4 / 18; 95% CI: 7.4–48.1) received a definitive or presumptive cancer diagnosis. Four cancer diagnoses in screened patients included splenic stromal sarcoma (definitively diagnosed via histopathology), a splenic mass (presumptive) suspected of being cancer on imaging but without confirmatory histology, a liver mass (presumptive), and an adrenal tumor (presumptive).

[0102] The remaining 10 of the 28 patients in the clinical cohort submitted samples for testing due to suspicion of cancer (e.g., testing was used as an aid to diagnosis). Within this group, 50% (5 / 10; 95% CI: 20.1–79.9) had no evidence of cancer after clinical evaluation, while 50% (5 / 10; 95% CI: 20.1–79.9) received a definite or presumptive cancer diagnosis. Five cancer diagnoses in the aided-diagnosis patients included a histiocytic sarcoma (definite), a mixed germ cell-sex cord stromal tumor (definite), two bone tumors (both presumptive), and an adrenal tumor (presumptive).

[0103] All 14 dogs whose samples were submitted as part of the research collection had a definitive cancer diagnosis at the time of blood collection: lymphoma / acute lymphocytic leukemia (4 cases), osteosarcoma (3 cases), mast cell tumor (2 cases), hemangiosarcoma, anal gland adenocarcinoma, chronic lymphocytic leukemia, soft tissue sarcoma, and transitional cell carcinoma.

[0104] In addition to WBC gDNA-specific CNVs, 38 dogs (22 from the clinical cohort and 16 from the research cohort) had CNVs identified in plasma-derived cfDNA; 26 of these dogs (11 from the clinical cohort and 15 from the research cohort) underwent clinical cancer evaluation. All of the dogs in the research cohort were among 569 dogs enrolled with a definitive cancer diagnosis. The cancer confirmation rate in these 26 dogs was 100% (26 / 26; 95% CI: 84.0-100), demonstrating that the observation of somatic CNVs in plasma cfDNA is highly predictive of concurrent cancer.

[0105] Tumor tissue samples (from the same collection time point as the blood samples with WBC gDNA-specific findings) were available for testing by NGS from 11 of the 35 dogs diagnosed with cancer in the study cohort. CNVs observed in gDNA were absent in matched tumor tissue in 91% of cases (10 / 11). In one case, the same CNV observed in WBC gDNA was present in matched tissue (from a fine-needle aspirate of the left mandibular lymph node), and this subject received a definitive diagnosis of T-zone lymphoma (stage IIa).

[0106] Tumor tissue samples (from the same collection time point as the blood samples with cfDNA-specific CNV findings) were available for testing by NGS from 105 dogs diagnosed with cancer in the study cohort. In contrast to WBC gDNA-specific CNVs, CNVs observed in cfDNA did not match CNVs found in tumor tissue in the majority of cases (92%; 97 / 105). Figures 2A–2C show matched CNV profiles from one subject with CNVs present in tissue, cfDNA, and gDNA. The WBC gDNA-specific CNV on CFA25 (canine chromosome 25, Figure 2A) is not visible in tumor tissue (Figure 2C), whereas the cfDNA-specific CNVs on CFA26, CFA32, and CFA35 (Figure 2B) are clearly visible in tumor tissue.

[0107] These results suggest that WBC gDNA-specific CNVs, when present, typically coexist with (but are not biologically associated with) the patient's cancer, whereas cfDNA-specific CNVs typically originate from the patient's cancer.

[0108] Longitudinal blood samples were available for a total of 76 patients with WBC gDNA-specific findings. 71% of patients (54 / 76; 95% CI: 60.0–80.0) demonstrated persistence of gDNA CNV over time. Figures 3A–3C show the CNV profile from one of these patients with persistent WBC gDNA-specific CNV at CFA25 and CFA36 at three consecutive time points over a 9-month period. In the remaining 29% of cases (22 / 76; 95% CI: 19.4–40.7), gDNA CNV did not persist across all subsequent time points.

[0109] Several WBC gDNA-specific CNVs recurred across patients; the most common CNV was a partial loss of chromosome 25, typically accompanied by a partial gain on the same chromosome (Table 4; Figures 1A and 2A). This pattern is consistent with human literature reporting CHIP at recurrent locations within the genome (Gao, Nat Commun, 2021; Saiki, Nat Med, 2021; Takahashi, Blood Adv, 2017). [Table 4] *Among the 164 dogs with WBC gDNA-specific CNV(s), findings were not mutually exclusive. Some dogs had multiple concurrent WBC gDNA findings, either within a single chromosome (e.g., partial gain and partial loss of chromosome 25) or affecting multiple chromosomes (e.g., gain on both chromosomes 14 and 25).

[0110] The study cohort was analyzed to assess the normalized age of patients without WBC gDNA-specific CNVs and their cancer status. Among patients without WBC gDNA-specific CNVs, dogs with a cancer diagnosis (definite or presumptive) were significantly older than dogs that submitted samples as part of the clinical cohort for screening and no evidence of cancer (p<0.0001). Additionally, patients with WBC gDNA-specific findings (whether with or without cancer) were significantly older than dogs diagnosed with cancer without WBC gDNA-specific CNVs (p<0.0001) (Figure 4). No significant differences were observed in the age distribution of patients with WBC gDNA-specific CNVs as a function of cancer status (p=0.5864).

[0111] Embodiments of the methods described herein for large-scale NGS-based genomic profiling of dogs with and without cancer surprisingly and unexpectedly revealed evidence of age-associated somatic copy number changes in dogs at the population level.

[0112] As used herein, section headings are for organizational purposes only and should not be construed as limiting the subject matter described in any way. All literature and similar materials cited in this application, including but not limited to patents, patent applications, papers, books, articles, and internet web pages, are expressly incorporated by reference in their entirety for any purpose, including the disclosures specifically referenced herein. If the definitions of terms in the incorporated references appear to differ from the definitions provided in the present teachings, the definitions provided in the present teachings shall prevail. It will be understood that there is an implicit "about" before temperatures, concentrations, times, etc. discussed in the present teachings so that slight and insignificant deviations are within the scope of the present teachings herein.

[0113] While the embodiments described herein have been disclosed in the context of specific embodiments and examples, those skilled in the art will appreciate that the disclosure extends beyond the specifically disclosed embodiments to other alternative embodiments and / or uses of the embodiments described herein, as well as obvious modifications and equivalents thereof. In addition, while several variations of the embodiments have been shown and described in detail, other modifications that are within the scope of the present disclosure will be readily apparent to those skilled in the art based on this disclosure. It is also contemplated that various combinations or subcombinations of specific features and aspects of the embodiments may be made and still fall within the scope of the present disclosure. It should be understood that various features and aspects of the disclosed embodiments can be combined with or substituted for one another to form various modes or embodiments described herein. Accordingly, it is not intended that the scope of the present disclosure should be limited by the specific disclosed embodiments described above.

[0114] It should be understood, however, that this detailed description, while indicating several embodiments, is given by way of example only, since various changes and modifications within the spirit and scope of the disclosure will become apparent to those skilled in the art.

[0115] The terms used in the descriptions presented herein are not intended to be construed in any limiting or restrictive manner. Rather, the terms are utilized simply in conjunction with detailed descriptions of embodiments of systems, methods, and related components. Furthermore, embodiments may include several novel features, only one of which is not solely responsible for its desirable attributes, nor is it considered essential to practicing the embodiments described herein.

Claims

1. 1. A method for detecting a non-cancerous somatic mutation in a subject, comprising: obtaining a sample from a subject; isolating genomic DNA from the sample; preparing a DNA sequencing library from the genomic DNA (gDNA); sequencing the gDNA to generate sequencing data; detecting at least one copy number abnormality in the sequencing data, wherein the at least one copy number abnormality is indicative of a non-cancerous somatic mutation; The method comprising:

2. The method of claim 1 , wherein the subject is a canine subject.

3. The method of claim 1 , wherein the sample is a whole blood sample.

4. 10. The method of claim 1, wherein the sequencing data is aligned to a reference genome.

5. The method of claim 1 , wherein the gDNA is extracted from white blood cells.

6. The method of claim 5, wherein the leukocytes are present in a buffy coat.

7. 10. The method of claim 1, further comprising isolating cell-free DNA (cfDNA).

8. 8. The method of claim 7, wherein the cfDNA is extracted from plasma.

9. 8. The method of claim 7, wherein the gDNA matches the cfDNA.

10. 8. The method of claim 7, wherein the absence of non-cancerous somatic mutations is detected in a cell-free DNA sequencing library.

11. 1. A method for detecting a non-cancerous somatic mutation in a subject, comprising: isolating white blood cell (WBC) genomic DNA (gDNA) from a buffy coat sample from said subject; generating a DNA sequencing library from the WBC gDNA; generating sequencing data by sequencing the DNA sequencing library; aligning the sequencing data to a reference genome; detecting at least one copy number abnormality, wherein said at least one copy number abnormality is indicative of a non-cancerous somatic mutation; The method comprising:

12. isolating cell-free DNA (cfDNA) from a plasma sample from the subject; generating a cfDNA sequencing library from the cfDNA; generating cell-free sequencing data by sequencing the cfDNA sequencing library; aligning the cell-free sequencing data to a reference genome; detecting the absence of said non-cancerous somatic mutation in a cell-free alignment; The method of claim 11 further comprising:

13. isolating non-buffy coat DNA from a non-buffy coat sample from said subject; generating a non-buffy coat DNA sequencing library from said non-buffy coat DNA; generating non-buffy coat sequencing data by sequencing the non-buffy coat DNA sequencing library; aligning the non-buffy coat sequencing data to a reference genome; detecting the absence of said non-cancerous somatic mutation in a non-buffy coat alignment; The method of claim 11 further comprising:

14. 14. The method of claim 13, wherein one or more non-buffy coat somatic mutations are detected within the non-buffy coat alignment, and the non-buffy coat somatic mutations do not match the non-cancerous somatic mutations.

15. 12. The method of claim 11, wherein the subject is a canine subject.

16. The method of claim 11 , wherein the buffy coat sample and the plasma sample are matched.

17. 1. A method for measuring age-related somatic changes in a canine subject, comprising: measuring copy number variations (CNV) from leukocyte genomic DNA (WBC gDNA) obtained from a buffy coat sample from said canine subject; detecting the absence of CNV from cell-free DNA (cfDNA) obtained from a matched plasma sample from said canine subject; The method comprising: