Methods and compositions for detecting cancer using fragmentomics

JP2024519975A5Pending Publication Date: 2025-05-27PETDX INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023572512
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-05-21
Filing Date
2022-05-20
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Current cancer diagnostic methods for pets, particularly dogs and cats, are invasive, expensive, and inconvenient, making early detection of cancers like lymphoma, squamous cell carcinoma, breast cancer, and osteosarcoma difficult.

Method used

Analyzing the fragment size distribution of circulating cell-free DNA (cfDNA) to detect, diagnose, and screen for cancer by isolating cfDNA, sequencing it, and comparing the fragment size distributions with control subjects to determine the presence of cancer.

Benefits of technology

Provides a non-invasive, cost-effective method for early cancer detection in pets by accurately distinguishing between normal and cancerous DNA fragment profiles, improving diagnostic accuracy and reducing the need for invasive biopsies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present invention provides methods and kits for measuring the fragment size distribution of DNA fragments in a subject-derived sample for the purposes of cancer or tumor detection, characterization and / or management.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to methods for detecting, characterizing or managing cancer or tumors in a subject by analyzing the fragment size distribution of DNA fragments in a sample. [Background technology]

[0002] As veterinary medicine continues to improve, the lifespans of pet animals such as dogs and cats continue to increase. However, this increased lifespan has led to a higher incidence of cancer in pets. Some estimates indicate that more than 50% of dogs over the age of 10 will die from cancer-related health issues. Cats also develop various types of cancer. The most common cancers in these animals include lymphoma, squamous cell carcinoma (skin cancer), breast cancer, mast cell tumor, oral tumors, fibrosarcoma (soft tissue cancer), osteosarcoma (bone cancer), respiratory cancer, intestinal adenocarcinoma, and pancreatic / hepatic adenocarcinoma.

[0003] Certain breeds of cats are more susceptible to certain cancers than others. Signs and symptoms vary depending on the type and stage of the cancer. Unfortunately, such cancers are often difficult to detect and diagnose, and an invasive biopsy is usually required to make an accurate diagnosis.

[0004] A similar situation exists in dogs. Certain breeds of dogs are known to be more susceptible to certain types of cancer (Rafalko, BIORXIV, 2022). For example, large dogs are more susceptible to osteosarcoma. In particular, German shepherds, golden retrievers, Labrador retrievers, English pointers, boxers, English setters, Great Danes, poodles and Siberian huskies are more susceptible to hemangiosarcoma (HSA). Hemangiosarcoma tends to develop in larger dog breeds than in smaller ones.

[0005] Current cancer diagnostic methods include imaging, radiolabeling, and biopsy. Liquid biopsy can provide diagnostic information that can only be obtained by invasive biopsy. The primary indication for liquid biopsy is based on the detection of genetic markers such as sex differences, gene polymorphisms, and mutations. Non-invasive prenatal testing is used worldwide to screen fetuses for chromosomal aneuploidies, contributing to a significant decrease in the number of invasive prenatal tests, such as amniocentesis. Liquid biopsies performed on organ transplant patients are used to monitor graft dysfunction. Cancer liquid biopsies are used to select targeted therapies and monitor cancer progression. However, currently available biopsy-based techniques for cancer or tumor detection are relatively expensive and not easily performed. Summary of the Invention [Means for solving the problem]

[0006] Described herein are methods and compositions for measuring the fragment size distribution of DNA obtained from a sample obtained from a subject. In some embodiments, the compositions and methods are used to detect, diagnose, and screen for cancer in a subject.

[0007] Some embodiments provided herein relate to a method of detecting cancer or a tumor in a subject. In some embodiments, the method comprises: isolating a circulating cell-free DNA (cfDNA) sample from a subject; sequencing the cfDNA sample to measure one or more fragment size distributions; comparing the one or more fragment size distributions to a second fragment size distribution obtained from one or more control subjects; and determining the presence or absence of cancer or a tumor based on a comparison of the two fragment size distributions. Includes. In some embodiments, the one or more control subjects include the subject or one or more healthy subjects. In some embodiments, the sequencing of the cfDNA sample is whole genome sequencing or next generation sequencing.

[0008] In some embodiments, the subject is a mammal. In some embodiments, the subject is a dog, a cat, a horse, or a human. In some embodiments, the cfDNA sample is isolated from the subject's blood. In some embodiments, the subject's blood further comprises circulating tumor DNA (ctDNA). In some embodiments, the cancer is a blood cancer. In some embodiments, the cancer is a lymphoma.

[0009] In some embodiments, the method further comprises creating a model of the one or more fragment size distributions. In some embodiments, the model of the one or more fragment size distributions is a statistical model. In some embodiments, the model of the one or more fragment size distributions is obtained from one or more features extracted from the one or more fragment size distributions. In some embodiments, the one or more features include median, mean, area under the curve (AUC), oscillation amplitude, variance, standard deviation, fragment length interval, or a combination thereof.

[0010] In some embodiments, the method further comprises classifying the sample as tumor or normal based on the one or more features. In some embodiments, the model of the second fragment size distribution is a statistical model. In some embodiments, the comparison of the one or more fragment size distributions with the second fragment size distribution is performed by KL divergence. In some embodiments, the one or more fragment size distributions are calculated from at least one of the lengths or sequences of cfDNA fragments in the sample. In some embodiments, the second fragment size distribution is a baseline fragment size distribution.

[0011] In some embodiments, the method further comprises ligating an adaptor to the isolated cfDNA and using a universal primer targeting the adaptor to generate amplified fragments. In some embodiments, the one or more fragment size distributions are measured by measuring the number and distribution of amplified fragment sizes using whole genome sequencing or next generation sequencing. In some embodiments, the comparison of the one or more fragment size distributions with a second fragment size distribution is performed by comparing the number and distribution of amplified fragment sizes with one or more healthy subjects, and the comparison determines whether the number and distribution of amplified fragment sizes of the subject is different from the number and distribution of amplified fragment sizes of the one or more healthy subjects. In some embodiments, the universal primer further comprises a sequence-specific primer. In some embodiments, a statistically significant difference between the one or more fragment size distributions of the subject and the second fragment size distribution of the one or more healthy subjects indicates the presence of cancer or tumor. In some embodiments, a lack of a statistically significant difference between the one or more fragment size distributions of the subject and the second fragment size distribution of the one or more healthy subjects indicates the absence of cancer or tumor.

[0012] Some embodiments provided herein relate to a method for predicting a cancer signal of origin (CSO) in a subject in which a cancer positive signal has been detected. In some embodiments, the method includes: isolating a circulating cell-free DNA (cfDNA) sample from a subject; sequencing the cfDNA sample to determine fragment size distribution and copy number profile; detecting a cancer positive signal from the copy number profile; comparing the fragment size distribution of the copy number gain and / or copy number loss regions with the fragment size distribution of a control copy number region; and Predicting a cancer signal of origin (CSO) based on a difference between the fragment size distributions of the copy number gain and / or copy number loss regions and the control copy number region, or based on the absence of a difference between these fragment size distributions. Includes. In some embodiments, the absence of a difference between the fragment size distribution of the copy number gain region and / or the copy number loss region and the fragment size distribution of the control copy number region is predictive of hematological cancer.

[0013] Some embodiments provided herein relate to a method of detecting cancer or a tumor in a subject. In some embodiments, the method comprises: isolating a circulating cell-free DNA (cfDNA) sample from a subject; sequencing the cfDNA sample to measure one or more fragment size distributions; generating an empirical model of the one or more fragment size distributions; comparing the one or more fragment size distributions to a second fragment size distribution obtained from one or more control subjects; and determining the presence or absence of cancer or a tumor based on a comparison of the two fragment size distributions. Includes. In some embodiments, the one or more control subjects include the subject or one or more healthy subjects. In some embodiments, the one or more empirical models of fragment size distributions are statistical models. In some embodiments, the one or more empirical models of fragment size distributions are derived from one or more features extracted from the one or more fragment size distributions. In some embodiments, the one or more features include a mean, an area under the curve (AUC), an amplitude of oscillation, a standard deviation, a fragment length interval, or a combination thereof.

[0014] In some embodiments, the method further comprises comparing the experimental model obtained from the cfDNA sample with a control model obtained from a control cfDNA sample of an individual known not to have cancer or tumor. In some embodiments, the experimental model is compared with the control model to determine the likelihood of the subject having cancer or tumor. In some embodiments, the one or more features of the experimental model are compared with the one or more features of the control model to determine the likelihood of the subject having cancer or tumor. In some embodiments, the comparison of the one or more fragment size distributions with a second fragment size distribution of at least one healthy subject is performed by KL divergence.

[0015] Some embodiments provided herein relate to a method for measuring a fragment size distribution in a sample. In some embodiments, the method comprises: isolating a DNA sample from a subject; sequencing the DNA sample to determine fragment size distribution; measuring one or more features from the fragment size distribution; and generating an empirical model of said fragment size distribution; Includes. In some embodiments, the subject is a subject suffering from or suspected of suffering from cancer. In some embodiments, the experimental model is a statistical model. In some embodiments, the experimental model is derived from the one or more features. In some embodiments, the one or more features include a mean value, an area under the curve (AUC), an amplitude of oscillation, a standard deviation, a fragment length interval, or a combination thereof.

[0016] In some embodiments, the method further comprises identifying the sample as a tumor sample or a normal sample based on the one or more features. In some embodiments, the fragment size distribution is calculated from at least one of the lengths or sequences of DNA fragments in the sample. In some embodiments, the DNA sample is a cell-free DNA (cfDNA) sample. In some embodiments, the DNA sample is isolated from the subject's blood. In some embodiments, the blood further comprises circulating tumor DNA (ctDNA). In some embodiments, the sequencing comprises whole genome sequencing or next generation sequencing. In some embodiments, the method further comprises ligating an adaptor to the isolated DNA and generating amplified fragments by using a universal primer that targets the adaptor. In some embodiments, the one or more fragment size distributions are measured by measuring the number and distribution of amplified fragment sizes using whole genome sequencing or next generation sequencing. In some embodiments, the universal primer further comprises a sequence-specific primer. [Brief description of the drawings]

[0017] [Figure 1] 1 is a line graph showing a representative profile of the average density of cfDNA having specific fragment lengths in cfDNA samples taken from normal healthy subjects.

[0018] [Figure 2A-2C] Representative line graphs of the fragment size distribution transformed into a negative binomial mixture model (Figure 2A), a Gaussian mixture model (Figure 2B), and a simple mixture model (Figure 2C) are shown. In each graph, the grey line is the sample and the black line is the model fit of the sample. In Figure 2C, the grey line is the sample and the circles indicate the position and height of each identified peak.

[0019] [Figure 3A-3C]Representative dot graphs showing the distribution of modes using a negative binomial mixed model (FIG. 3A), a Gaussian mixed model (FIG. 3B), and a simple mixed model (FIG. 3C) for the inverted data are shown. Normal samples were either from baseline measurements (circles) or from test measurements (herein referred to as "test-normal") samples (triangles). "Mode3" indicates the scaling used in each graph, with larger modes being shown with larger circles or triangles.

[0020] [Figure 4A-4B] Representative dot graphs showing the distribution of weights using negative binomial mixture models (Figure 4A) and Gaussian mixture models (Figure 4B) for the inverted data are shown. Normal samples were either from the baseline measurement (circles) or from the test measurement (herein referred to as "test-normal") samples (triangles). "Weights" are the proportions of each component (each nucleosome peak) in the mixture distribution model. "Weight3" indicates the scaling used in each graph, with larger weights shown as larger circles or triangles.

[0021] [Figure 5A-5B] Representative dot graphs showing the distribution of measures using a negative binomial mixed model (FIG. 5A) or a Gaussian mixed model (FIG. 5B) for the inverted data are shown. Normal samples are either from baseline measurements (circles) or from samples from test measurements (herein referred to as "test-normal") (triangles). The "measures" for the negative binomial mixed model for the inverted data are overdispersed, i.e., small values ​​cause large variance. "Measure 3" indicates the scaling used in each graph, with larger scales shown as larger circles or triangles.

[0022] [Figure 6A-6B]Representative dot plots showing principal component analysis (PCA) using a negative binomial mixed model (Figure 6A) or a Gaussian mixture model (Figure 6B) of the inverted data are shown. Normal samples were either from baseline measurements (circles) or from test measurements (herein referred to as "test-normal") samples (triangles). The extracted features do not distinguish between samples from each test. Nearly all variation is captured in one principal component. "PC3" indicates the scaling used in each graph, with larger principal component values ​​shown as larger circles or triangles.

[0023] [Figure 7] 1 is a dot graph showing a PCA plot of normalized fragment length data comparing PC values ​​for batch 1, batch 2, and batch 3. "PC3" indicates the scaling used in each graph, with larger principal component values ​​indicated by larger circles.

[0024] [Figure 8A-8D] Box plots are shown for the PC values ​​of all samples (FIG. 8A), non-normal samples (FIG. 8B), normal samples (FIG. 8C) and baseline samples (FIG. 8D) for batches 1, 2 and 3. The baseline samples are a subset of the normal samples disclosed herein.

[0025] [Figure 9A-9B] Representative line graphs showing the density profiles of cfDNA with specific fragment lengths in batch 1, batch 2 and batch 3 cfDNA samples obtained from a normal subject (Figure 9A) and a baseline normal subject (Figure 9B).

[0026] [Figure 10A-10B] Representative dot graphs are shown comparing the percentage of peaks using a set of initial statistics for all normal samples from batches 1-3 combined (Figure 10A) and for only the baseline normal samples when normal samples are combined (Figure 10B). "Peak 3" indicates the scaling used in each graph, with larger peaks indicated by larger circles.

[0027] [Figure 11A-11B] Representative dot graphs are shown plotting the vibration values ​​(FIG. 11A) and AUC values ​​(FIG. 11B) for batch 1, batch 2, and batch 3, separated into baseline, non-normal, and normal groups.

[0028] [Figure 12] Box plots showing the age distribution of subjects in batch 1, batch 2 and batch 3 by batch are shown.

[0029] [Figure 13] 1 is a dot graph of the KL divergence values ​​of each sample divided into the baseline group, normal group, and tumor group.

[0030] [Figure 14] 1 is a dot graph of the KL divergence values ​​of each sample of batches 4 to 7 and batch 12, divided into normal and tumor groups.

[0031] [Figure 15] Correlation between features (mean, AUC, oscillations and standard deviation) extracted from Gaussian mixture models. Parameters of these distributions were estimated by Markov chain Monte Carlo. Means, SD and weights are obtained from the mixture distribution models of all samples. AUC of short snippets is the AUC relative to the first mode of each sample. Oscillations were calculated from peaks and valleys identified in baseline samples.

[0032] [Figures 16A-16D] The distributions of accuracy, sensitivity, specificity, positive predictive value (PPV) and F-1 score for each threshold were calculated using probabilistic methods, and the optimized distributions are shown for specificity (Figure 16A), F-1 score (Figure 16B), PPV score (Figure 16C) and sensitivity (Figure 16D), respectively.

[0033] [Figure 17]A profile of the fragment length difference between the normalized counts of the average normal sample and the normalized counts of the average tumor sample is shown.

[0034] [Figure 18] Profiles of mean normalized counts of cfDNA for specific fragment lengths for normal or tumor samples from batches 1–3 are shown.

[0035] [Figure 19] PCA analysis of all samples from batches 1-3 and the 2D density contour of the normal sample are shown. "PC3" indicates the scaling used in this graph, with larger principal component values ​​indicated by larger circles.

[0036] [Figure 20] A dot plot of KL divergence values ​​determined from the mean values ​​of baseline, normal and tumor samples is shown.

[0037] [Figure 21] A dot plot of KL divergence values ​​determined from the mean values ​​of baseline, normal and tumor samples after removing two outlier samples from the baseline mixture distribution model is shown.

[0038] [Figure 22] 1 shows a graph plotting the prior distribution of tumor cell fraction as a function of the tumor cell fraction value.

[0039] [Figure 23A] A graph plotting the estimated tumor cell fraction against the expected tumor cell fraction in samples 201-20885 mixed with healthy cfDNA samples is shown.

[0040] [Figure 23B] Graph showing estimated tumor cell fraction plotted against expected tumor cell fraction in sample 201-00316 mixed with healthy cfDNA samples.

[0041] [Figure 24A] A graph plotting the estimated tumor cell fraction against the expected tumor cell fraction in sample 201-00015 mixed with a healthy cfDNA sample is shown.

[0042] [Figure 24B] A graph plotting the estimated tumor cell fraction against the expected tumor cell fraction in samples 301-30640 mixed with healthy cfDNA samples is shown.

[0043] [Diagram 25] The fragment length distribution of chromosomes with copy number loss, neutral or gain in sample 201-00015 is shown.

[0044] [Figure 26] Adjusted separation values ​​for samples by cancer type are shown, samples with multiple tumor types in a single sample and samples with low tumor cell content despite no separation between fragment length curves are not shown.

[0045] [Figure 27] A plot for threshold selection is shown, where thresholds from 135 to 175 were tested in increments of 1, and the raw values ​​showing the separation between the fragment length curves were plotted against the threshold at which the separation between the fragment length curves was maximal.

[0046] [Figure 28] The effect of data smoothing on threshold selection is shown as the change in threshold for selected samples with and without spline smoothing.

[0047] [Figure 29A]The linear relationship between the segregation values ​​calculated by the equation "decrease" minus "gain" and the equation "no change" minus "gain" (left panel) or between the segregation values ​​calculated by the equation "decrease" minus "gain" and the equation "decrease" minus "no change" (right panel) is shown for selected samples with all three copy number (CN) groups. The cutoff for reads was set to 0.

[0048] [Figure 29B] The correlation between the residuals of the corrected separation values ​​in the linear relationship between the separation values ​​calculated using the "decrease"-"increase" formula and the separation values ​​calculated using the "no change"-"increase" formula and the minimum number of reads examined (M) (left panel) or the correlation between the residuals of the corrected separation values ​​in the linear relationship between the separation values ​​calculated using the "decrease"-"increase" formula and the separation values ​​calculated using the "decrease"-"no change" formula and the minimum number of reads examined (M) (right panel). The cutoff for the reads was set to 0.

[0049] [Figure 29C] The linear relationship between the separation values ​​calculated by the equation "decrease" - "increase" and the separation values ​​calculated by the equation "no change" - "increase" after correction (left panel) or the linear relationship between the separation values ​​calculated by the equation "decrease" - "increase" and the separation values ​​calculated by the equation "decrease" - "no change" (right panel) is shown. The cutoff for reads was 200,000.

[0050] [Diagram 30] The accuracy of adjustment for the "decrease" minus "no change" and "no change" minus "increase" equations is shown plotted as the difference between the adjusted value and the expected value.

[0051] [Diagram 31] The average KL value per chromosome per sample plotted against the number of reads per chromosome is shown.

[0052] [Diagram 32] Changes in KL divergence per chromosome using fragmentomics in a genome-wide approach. Solid horizontal lines indicate potential thresholds.

[0053] [Diagram 33] A graph showing predicted KL values ​​using chromosome-specific hyperbolic curves whose parameters were trained using Model 6 plotted against the true KL values. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0054] In the following detailed description, the present invention will be described with reference to the accompanying drawings, which form a part of this specification. Unless otherwise specified, similar symbols in the drawings generally indicate similar components. The embodiments described in the detailed description, drawings, and claims are for illustrative purposes only and are not intended to limit the present invention in any way. Other embodiments may be adopted and other changes may be made without departing from the spirit or scope of the subject matter described herein. It is readily understood that the aspects of the present disclosure as outlined herein and shown in the drawings can be arranged, substituted, combined, separated, and designed in various configurations, and all such aspects are expressly contemplated herein. All documents cited herein are expressly incorporated in their entirety as part of this specification to present the specific disclosures cited herein.

[0055] Embodiments of the present invention relate to methods, systems and compositions for screening a subject for the likelihood of having a cancer or tumor. In some embodiments, screening for a cancer or tumor includes: isolating a circulating cell-free DNA (cfDNA) sample from a subject (e.g., a dog) suspected of having cancer or a tumor; sequencing the cfDNA fragments in said sample; calculating a size distribution based on the at least one cfDNA fragment; creating a model or summary statistics of the fragment size distribution; comparing the model of fragment size distribution to a second model derived from at least one healthy subject; and determining the presence or absence of cancer or a tumor based on a comparison of the two models; This is carried out by. Sequencing of cfDNA can be performed by any method known to those skilled in the art, such as targeted sequencing or genome-wide sequencing. Other examples include, but are not limited to, nanopore-based methods, emulsion-based methods, and "sequencing by binding" cycle sequencing methods.

[0056] In some embodiments, cancer or tumors are screened by comparing between models. In some embodiments, these models are mixture distribution models. The models are derived from the fragment size distribution profile of at least one fragment. As used herein, the term "fragment distribution" has its usual meaning as understood by those skilled in the art, and thus refers to the length, sequence, fragmentation and other distribution characteristics of at least one DNA fragment obtained from a cfDNA sample. "Fragment size distribution" is interpreted as a fragment distribution focusing on the size of the fragment, including the length or fragmentation of the fragment. As disclosed herein, models can be created for subjects suspected of having cancer or tumors, and models can also be created for one or more healthy subjects. These models can then be compared to each other to see if there are significant differences. Examples of models include, but are not limited to, summary statistics, the number of nucleosome peaks and their shapes, the proportion of fragments longer or shorter than a certain threshold, the proportion of fragments in a certain interval, approximations of data with statistical distributions, and discriminative learning methods such as support vector machines and neural networks. Examples of detectable differences include, but are not limited to, the location of the peak (mode), the height of the peak (weight), the spread of the peak (scale), the percentage of fragments longer or shorter than a certain threshold, the amplitude of oscillations, the overall shape of the fragment size distribution, the principal component values, and the Kullback-Leibler (KL) divergence between the two models. In some embodiments, a statistically significant difference between the fragment size distribution of the subject suspected of having a cancer or tumor and the fragment size distribution of one or more healthy subjects indicates the presence of a cancer or tumor. In some embodiments, a lack of a statistically significant difference between the fragment size distribution of the subject suspected of having a cancer or tumor and the fragment size distribution of one or more healthy subjects indicates the absence of a cancer or tumor.

[0057] There are various methods for measuring the fragment size distribution of cfDNA in a subject. In one embodiment, a blood sample is obtained from the subject. Circulating cell-free DNA (cfDNA) is obtained from the blood. In some embodiments, the blood sample contains circulating tumor DNA (ctDNA). cfDNA is isolated by removing blood cells from the sample so that only cfDNA remains in the sample. In some embodiments, a random PCR primer set for whole genome sequencing is added to the sample to amplify fragments while maintaining the original fragment length in the sample.

[0058] A polymerase is then added to the mixture to extend the full length of each fragment with the primers. The amplified fragments may contain sequence ends, which in one embodiment are formatted for use in a next generation sequencing (NGS) system to identify the nucleotide sequence in the fragment.

[0059] The methods and compositions provided herein can improve the detection, diagnosis, staging, screening, treatment and management of cancer in subjects, particularly humans, mammals and other types of subjects.As mentioned above, embodiments of the present invention include identifying the fragment distribution of cfDNA circulating in bodily fluids, such as blood.In one embodiment, the nucleic acid sequence element is found in circulating tumor DNA in blood.In some embodiments, the nucleic acid sequence element may be found in cell-free DNA in saliva or urine.

[0060] As used herein, "detection" with respect to measuring cancer or tumors includes the use of instruments used to observe and record a signal corresponding to the extent or measurement of cancer, or the substances necessary to generate such a signal. In various embodiments, "detection" includes any suitable method, including amplification, sequencing, arrays, fluorescence, chemiluminescence, surface plasmon resonance, surface acoustic waves, mass spectrometry, infrared spectroscopy, Raman spectroscopy, atomic force microscopy, scanning tunneling microscopy, electrochemical detection, nuclear magnetic resonance, quantum dots, and the like.

[0061] Some embodiments provided herein relate to kits.In some embodiments, the kits are for determining cancer in a subject.In some embodiments, the kits include a whole genome sequencing primer for amplifying the cfDNA in a biological sample obtained from a subject, and a polymerase for amplifying the primers.

[0062] It should be understood that the analysis described herein may be part of a broad set of diagnostic tests used to diagnose the overall health of a subject.For example, the analysis of the fragment size distribution of cfDNA in a subject may be performed simultaneously or sequentially with other methods for cancer detection, diagnosis, stage classification, screening, monitoring, treatment and management (such as additional genetic variance analysis).These techniques may be useful for detecting various cancers, such as leukemia, squamous cell carcinoma, feline breast cancer, mast cell tumor, bladder cancer, osteosarcoma, hemangiosarcoma, or various other cancers that a subject suffers from.

[0063] In some embodiments, the method includes obtaining a biological sample from a subject suspected of having cancer. In some embodiments, the sample is a liquid biopsy sample, for example, a blood sample. In some embodiments, the sample includes cfDNA. In some embodiments, the sample is provided in a volume of less than 10 mL, for example, 10 mL, 9 mL, 8 mL, 7 mL, 6 mL, 5 mL, 4 mL, 3 mL, 2 mL, 1 mL, 500 μL, 250 μL, or 100 μL, or a volume within a range of any two of these values. In some embodiments, the amount of DNA contained in the sample is 10 μg or less, e.g., 10 μg, 5 μg, 1 μg, 500 ng, 100 ng, 50 ng, 10 ng, 5 ng, 1 ng, 500 pg, 100 pg, 50 pg, 10 pg, 9 pg, 8 pg, 7 pg, 6 pg, 5 pg, 4 pg, 3 pg, 2 pg, or 1 pg, or an amount within a range of any two of these values. In some embodiments, the method includes purifying DNA from the sample. Purification of DNA may be performed using DNA purification techniques, such as extraction techniques, precipitation, chromatography, bead-based methods, or commercially available kits for DNA purification. In some embodiments, the method can be used to predict cancer type or cancer tissue origin based on one or more features of fragment size distribution.

[0064] Definition of Terms Unless otherwise defined, all technical and scientific terms used herein have the meanings commonly understood by those skilled in the art. All patents, applications, published applications and other publications cited herein are incorporated herein by reference in their entirety unless otherwise stated. In the event that there are a plurality of definitions for terms herein, the definitions set forth in this section shall prevail unless otherwise stated.

[0065] As used herein, "a" or "an" may mean one or more than one.

[0066] As used herein, the term "about" has its ordinary meaning as understood by one of ordinary skill in the art, and thus indicates that a particular numerical value includes the variation of error inherent in the method employed to determine the numerical value or variation among multiple measurements.

[0067] The dimensions and values ​​disclosed herein are not to be construed as being strictly limited to the numerical values ​​set forth herein. Instead, a dimension set forth herein is intended to mean both the numerical value set forth herein and a functionally equivalent numerical range surrounding that numerical value, unless otherwise specified. For example, a dimension disclosed as being "20 mm" is intended to be "about 20 mm."

[0068] Throughout this specification, unless otherwise stated, the terms "comprise" and "comprising" are meant to include the steps or components or steps or components described herein, but not to exclude other steps or components or steps or components. "Consisting of" means to include only those listed before this term. Thus, the term "consisting of" means that the components listed before this term are necessary or essential, and other components may not be included. "Consisting essentially of" means to include the components listed before this term, and also includes other components that do not interfere with or contribute to the activity or action described in connection with the disclosure of these components. Thus, the term "consisting essentially of" means that the components listed before this term are necessary or essential, but other components are optional and may or may not be included depending on whether they have a substantial effect on the activity or action of the components listed before this term.

[0069] As used herein, the terms "function" and "functionality" have their common and ordinary meaning as understood in the context of this specification and refer to a biological function, an enzymatic function or a therapeutic function.

[0070] As used herein, the term "yield" of a substance, compound or material has its general and ordinary meaning as understood in light of the present specification and refers to the actual total amount of that substance, compound or material relative to the expected total amount. For example, the yield of a substance, compound or material may be 80%, 81%, 82%, 83%, 84%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% of the expected total amount; about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99% or about 100%; or less than 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99% or about 100% of the expected total amount. at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or at least 100%; at least about 80%, at least about 81%, at least about 82%, at least about 83%, or at least about 84%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or at least about 100%; or 80% or less, 81% or less, 82% or less, 83% or less, 84% or less, 85% or less, 90% or less, 91% or less, 92% or less, 93% or less, 94% or less , 95% or less, 96% or less, 97% or less, 98% or less, 99% or less, or 100% or less; about 80% or less, about 81% or less, about 82% or less, about 83% or less, about 84% or less, about 85% or less, about 90% or less, about 91% or less, about 92% or less, about 93% or less, about 94% or less, about 95% or less, about 96% or less, about 97% or less, about 98% or less, about 99% or less, or 100% or less; or any decimal value therebetween. The yield may be affected by the efficiency of the reaction or process; undesired side reactions; decomposition; the quality of the input substances, compounds or materials; or loss of desired substances, compounds or materials in the manufacturing process.

[0071] As used herein, the term "isolated" has its common and ordinary meaning as understood in light of the present specification and refers to (1) a substance and / or object that is separated from at least some components that accompany it when it is first produced (in nature and / or in experimental conditions) and / or (2) a substance and / or object that is manufactured, prepared and / or produced by the hand of man. An isolated substance and / or object may be separated from other components that are first associated with it to about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, about 95%, about 98%, about 99%, substantially 100% or about 100% of the above, at least about 100%, at least about 100%, or less than or equal to about 100% of the above (or to a range including and / or including these values). In some embodiments, the purity of the isolated material is about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, substantially 100% or 100% pure, approximately these purities, at least these purities, at least approximately these purities, less than these purities, or less than these purities (or ranges of purity inclusive of these values ​​and / or ranges of purity bounded by these values). As used herein, an "isolated" material may be "pure" (e.g., substantially free of other components). As used herein, an "isolated cell" may refer to a cell separated from an organism or tissue that is composed of a large number of cells.

[0072] As used herein, "in vivo" has its common and ordinary meaning as understood in the context of this specification and refers to the performance of a particular method within the body of a living organism, usually an animal, a mammal such as a human or a plant, or within the living cells that make up these living organisms, as opposed to a tissue extract or a dead organism.

[0073] As used herein, "ex vivo" has its common and ordinary meaning as understood in the context of this specification and refers to the performance of a particular method outside the body of a living organism with minor modifications to the natural condition.

[0074] As used herein, "in vitro" has its common and ordinary meaning as understood in the context of this specification, and refers to carrying out a particular method outside biological conditions, such as in a petri dish or test tube.

[0075] As used herein, "nucleic acid", "nucleic acid molecule" or "nucleotide" refers to a polynucleotide or oligonucleotide, including, for example, deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), oligonucleotides, fragments obtained by polymerase chain reaction (PCR), and fragments obtained by either ligation, cleavage, endonuclease action, exonuclease action, or synthesis. Nucleic acid molecules may be composed of natural nucleotide monomers (such as DNA or RNA), or analogs of natural nucleotides (e.g., enantiomers of natural nucleotides), or combinations thereof. Modified nucleotides may have modifications in the sugar moiety and / or the pyrimidine or purine base moiety. Modifications in the sugar moiety include, for example, replacement of one or more hydroxyl groups with halogens, alkyl groups, amines, or azide groups, and the sugar moiety may be etherified or esterified. Additionally, the entire sugar moiety may be replaced with conformationally or electronically similar structures, including, for example, azasugars and carbocyclic sugar analogs. Modified base moieties include alkylated purines, alkylated pyrimidines, acylated purines, acylated pyrimidines, and other known heterocyclic substituents. Nucleic acid monomers can be linked by phosphodiester bonds or similar bonds. Linkages similar to phosphodiester bonds include phosphorothioate, phosphorodithioate, phosphoroselenoate, phosphorodiselenoate, phosphoroanilothioate, phosphoranilidate, phosphoroamidate, and the like. "Nucleic acid molecule" also includes so-called "peptide nucleic acids," which contain natural or modified nucleobases attached to a polyamide backbone. Nucleic acids can be single-stranded or double-stranded.

[0076] As used herein, the terms "peptide", "polypeptide" and "protein" have their usual and ordinary meaning as understood in the context of this specification and refer to macromolecules composed of amino acids linked by peptide bonds. Numerous functions of peptides, polypeptides and proteins are known in the art, including, but not limited to, enzymatic, structural, transport, defensive, hormonal or signal transduction functions. Peptides, polypeptides and proteins are often, but not always, produced biologically by ribosomal complexes using nucleic acid templates and can also be produced by chemical synthesis. Mutations of peptides, polypeptides or proteins can be produced by genetic manipulation of nucleic acid templates, including substitutions, deletions, truncations, additions, duplications, or fusions of two or more peptides, polypeptides or proteins. Fusion of two or more peptides, polypeptides or proteins can be made so that they are adjacently linked within one molecule, or they can be linked via additional amino acids, such as linkers, repeats, epitopes or tags, or amino acids that are 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 14, 16, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 104, 109, 108, 109, 109, 108, other sequences that are 2, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 150, 200 or 300 bases long;about 1 base length, about 2 base length, about 3 base length, about 4 base length, about 5 base length, about 6 base length, about 7 base length, about 8 base length, about 9 base length, about 10 base length, about 11 base length, about 12 base length, about 13 base length, about 14 base length, about 15 base length, about 16 base length, about 17 base length, about 18 base length, about 19 base length, about 20 base length, about 25 base length, about 30 base length, about 35 base length, about 40 base length, about 45 base length, about 50 base length, about 55 base length other sequences of at least 1 base, at least 2 bases, at least 3 bases, at least 4 bases, at least 5 bases, at least 6 bases, at least 7 bases, at least 8 bases, at least 9 bases, at least 10 bases, at least 20 bases, at least 30 bases, at least 40 bases, at least 5 bases, at least 6 bases, at least 7 bases, at least 8 bases, at least 9 bases, at least 10 bases, at least 15 bases, at least 16 bases, at least 17 bases, at least 18 bases, at least 19 bases, at least 20 bases, at least 21 bases, at least 22 bases, at least 23 bases, at least 24 bases, at least 25 bases, at least 26 bases, at least 27 bases, at least 28 bases, at least 29 bases, at least 30 bases, at least 31 bases, at least 32 bases, at least 33 bases, at least 34 bases, at least 35 bases, at least 36 bases, at least 37 bases, at least 38 bases, at least 39 bases, at least 40 bases, at least 41 bases, at least 42 bases, at least 43 bases, at least 44 bases, at least 45 bases, at least 46 bases, at least 47 bases, at least 48 bases, at least 49 bases, at least 50 bases, at least 51 bases, at least 52 bases, at least 53 bases, at least 54 bases, at least 55 bases, at least 56 bases, at least 57 bases, at least 58 bases, at least 59 bases, at least 60 bases, at least other sequences at least 9 bases long, at least 10 bases long, at least 11 bases long, at least 12 bases long, at least 13 bases long, at least 14 bases long, at least 15 bases long, at least 16 bases long, at least 17 bases long, at least 18 bases long, at least 19 bases long, at least 20 bases long, at least 25 bases long, at least 30 bases long, at least 35 bases long, at least 40 bases long, at least 45 bases long, at least 50 bases long, at least 55 bases long, at least 60 bases long, at least 65 bases long, at least 70 bases long, at least 75 bases long, at least 80 bases long, at least 85 bases long, at least 90 bases long, at least 95 bases long, at least 100 bases long, at least 150 bases long, at least 200 bases long or at least 300 bases long;at least about 1 base length, at least about 2 base length, at least about 3 base length, at least about 4 base length, at least about 5 base length, at least about 6 base length, at least about 7 base length, at least about 8 base length, at least about 9 base length, at least about 10 base length, at least about 11 base length, at least about 12 base length, at least about 13 base length, at least about 14 base length, at least about 15 base length, at least about 16 base length, at least about 17 base length, at least about 18 base length, at least about 19 base length, at least about 20 base length, at least about 25 base length, at least about 30 base length, at least about 35 base length, at least about 40 base length, at least about 45 base length, at least about 50 base length, at least about 55 base length, at least about 60 base length, at least about 65 base length, at least about 70 base length, at least about 75 base length, at least about 80 base length, at least about 85 base length, at least about 90 base length, other sequences having a length of at least about 95 bases, at least about 100 bases, at least about 150 bases, at least about 200 bases, or at least about 300 bases; 1 base or less, 2 bases or less, 3 bases or less, 4 bases or less, 5 bases or less, 6 bases or less, 7 bases or less, 8 bases or less, 9 bases or less, 10 bases or less, 11 bases or less, 12 bases or less, 13 bases or less, 14 bases or less, 15 bases or less, 16 bases or less, other sequences up to 17 bases in length, 18 bases in length, 19 bases in length, 20 bases in length, 25 bases in length, 30 bases in length, 35 bases in length, 40 bases in length, 45 bases in length, 50 bases in length, 55 bases in length, 60 bases in length, 65 bases in length, 70 bases in length, 75 bases in length, 80 bases in length, 85 bases in length, 90 bases in length, 95 bases in length, 100 bases in length, 150 bases in length, 200 bases in length, or 300 bases in length;about 1 base length or less, about 2 base lengths or less, about 3 base lengths or less, about 4 base lengths or less, about 5 base lengths or less, about 6 base lengths or less, about 7 base lengths or less, about 8 base lengths or less, about 9 base lengths or less, about 10 base lengths or less, about 11 base lengths or less, about 12 base lengths or less, about 13 base lengths or less, about 14 base lengths or less, about 15 base lengths or less, about 16 base lengths or less, about 17 base lengths or less, about 18 base lengths or less, about 19 base lengths or less, about 20 base lengths or less, about 25 base lengths or less, about 30 base lengths or less, about 35 base lengths or less, about 4 0 bases or less, about 45 bases or less, about 50 bases or less, about 55 bases or less, about 60 bases or less, about 65 bases or less, about 70 bases or less, about 75 bases or less, about 80 bases or less, about 85 bases or less, about 90 bases or less, about 95 bases or less, about 100 bases or less, about 150 bases or less, about 200 bases or less, or about 300 bases or less in length; or other sequences having a length within a range of any two of these lengths. As used herein, the term "downstream" on a polypeptide has its general, ordinary meaning as understood in the context of this specification, and refers to a sequence following the C-terminus of a preceding sequence. As used herein, the term "upstream" on a polypeptide has its general, ordinary meaning as understood in the context of this specification, and refers to a sequence preceding the N-terminus of a following sequence.;

[0077] The terms "DNA fragment" and "nucleic acid fragment" have their ordinary meaning as understood by those of skill in the art and refer to a polynucleotide sequence obtained from a genome, may be obtained from any site along the genome, and may contain any nucleotide sequence.

[0078] The term "fragment size distribution" has its general meaning as understood by one of skill in the art and refers to information regarding one or more of the following: the total number of nucleic acid fragments in a sample, the size of one or more nucleic acid fragments in a sample, the absolute or relative abundance of nucleic acid fragments of a particular size or range of sizes, and the absolute or relative abundance of nucleic acid fragments of different sizes in a sample.

[0079] The term "fragment size" has its general meaning as understood by one of skill in the art and is used herein in reference to nucleic acid molecules to refer to the number of base pairs of a nucleic acid and to indicate the length of the nucleic acid molecule.

[0080] The term "gene" has its common and ordinary meaning as understood in the context of this specification and generally refers to a portion of a nucleic acid that codes for a protein or functional RNA, although the term may encompass regulatory sequences. Those skilled in the art will appreciate that the term "gene" may include gene regulatory sequences (e.g., promoters, enhancers, etc.) and / or intron sequences. It will further be appreciated that the definition of "gene" also includes nucleic acids that do not code for proteins but code for functional RNA molecules such as tRNAs and miRNAs. In some cases, genes include regulatory sequences involved in transcription, message production or composition. In another embodiment, genes include transcribed sequences that code for proteins, polypeptides or peptides. In accordance with the definition of the term herein, an "isolated gene" may include transcribed nucleic acids, regulatory sequences, coding sequences, etc. that are substantially isolated from other sequences, such as other naturally occurring genes, other naturally occurring regulatory sequences, other naturally occurring sequences that code for polypeptides or peptides. In this regard, the term "gene" is used simply to refer to a nucleic acid that includes a transcribed nucleotide sequence and its complementary strand. As is understood in the art, the functional term "gene" encompasses genomic sequences, as well as RNA or cDNA sequences, or smaller recombinant nucleic acid segments, including nucleic acid segments of non-transcribed gene portions, such as, but not limited to, non-transcribed promoter or enhancer regions of a gene, which may express or be configured to express proteins, polypeptides, domains, peptides, fusion proteins, mutants, and / or the like, using nucleic acid engineering techniques.

[0081] The terms "cancer" and "cancerous" have their general meaning as understood in the context of this specification and refer to a physiological condition in an animal that is generally characterized by unregulated cell proliferation. A "tumor" contains one or more cancerous cells. In some embodiments, the tumor is a solid tumor. There are several major types of cancer. Carcinomas are cancers that originate from epithelial cells, such as skin cells and the lining cells of the intestinal tract. Sarcomas are cancers that originate from mesenchymal cells, such as bone, cartilage, fat, muscle, blood vessels, and other connective and supporting tissues. Leukemias are cancers that originate from hematopoietic cells, such as the bone marrow, which produce large numbers of abnormal blood cells that enter the blood circulation. Lymphomas and multiple myelomas are cancers that originate from lymphoid cells in the lymph nodes. Central nervous system cancers are cancers that originate from the central nervous system and spinal cord.

[0082] As used herein, the term "allele" or "allelic variant" has its general meaning as understood in the context of this specification and refers to a variant of a genetic locus or gene. In some embodiments, a particular allele of a genetic locus or gene is associated with a particular phenotype, such as an altered risk of developing a disease or condition, likelihood of progressing to a particular disease stage or stage of a particular condition, treatability with a particular therapeutic method, susceptibility to infectious diseases, immune function, etc.

[0083] As used herein, the term "amplification" has its general meaning as understood in light of the present specification and refers to any method known in the art for replicating a target nucleic acid to increase the number of copies of a selected nucleic acid sequence. Amplification may be exponential or linear. The target nucleic acid may be DNA or RNA. Typically, sequences amplified in such a manner form an "amplicon." Amplification may be performed in a variety of ways, including but not limited to polymerase chain reaction ("PCR"), transcription-based amplification, isothermal amplification, rolling circle amplification, and the like. Amplification may be performed using a primer pair consisting of relatively equal amounts of each primer to produce a double-stranded amplicon. However, as is well known in the art, asymmetric PCR may be used to amplify products of one strand preferentially or exclusively (see, e.g., Poddar et al. Molec. And Cell. Probes 14:25-32 (2000)). This can be done by significantly decreasing the concentration of one primer of the primer pair relative to the other primer (e.g., 100-fold difference). Amplification by asymmetric PCR is approximately linear. Those skilled in the art will appreciate that different types of amplification methods may be used in combination.

[0084] As used herein, the term "amplicon" has its general meaning as understood in the context of this specification and refers to a nucleic acid sequence to be amplified and the nucleic acid polymer resulting from the amplification reaction. Amplicons can be formed artificially, such as by polymerase chain reaction (PCR) or ligase chain reaction (LCR), or can be formed naturally by gene duplication.

[0085] As used herein, the terms "individual," "subject," "host," or "patient" have their usual meaning as understood by those skilled in the art, and thus include human or non-human mammals. The term "mammal" is used in the biological sense that the term normally denotes. Thus, "mammals" specifically include, but are not limited to, primates, such as monkeys (chimpanzees, apes, monkeys) and humans; as well as cows, horses, sheep, goats, pigs, rabbits, dogs, cats, rodents, rats, mice, and guinea pigs.

[0086] As used herein, the term "liquid biopsy" has its general meaning as understood in the context of this specification and refers to taking a sample of non-solid biological tissue, such as blood, and testing this sample.

[0087] As used herein, the term "cfDNA" has its general meaning as understood in the context of this specification and refers to circulating cell-free DNA, including DNA fragments released into plasma. cfDNA may also include circulating tumor deoxyribonucleic acid (ctDNA).

[0088] As used herein, the term "cfDNA" has its general meaning as understood in the context of the present specification and refers to circulating tumor DNA, including tumor-derived fragmented DNA that is not bound to cells and circulates in the bloodstream. EXAMPLES

[0089] The following examples further define the embodiments of the present invention. These examples are to be construed as being for illustrative purposes only. Those skilled in the art can understand the essential characteristics of the present invention from the above discussion and these examples, and can make various changes and improvements to the embodiments of the present invention to adapt the present invention to various applications and conditions without departing from the spirit and scope of the present invention. Therefore, various improvements of the embodiments of the present invention other than those shown and described herein will be apparent to those skilled in the art from the above description. Such improvements are also within the scope of the appended claims. The disclosures of each reference cited in this specification are incorporated herein in their entirety by reference for the purpose of citing the disclosures herein.

[0090] Example 1 Extraction of cfDNA from subjects The embodiment of cfDNA isolation described herein was performed by a series of extraction steps. A blood sample was collected from a canine subject into an anticoagulated blood collection tube (BCT) containing a cell-free DNA stabilizing component. Examples of blood collection tubes that can be used include, but are not limited to, Roche cell-free DNA extraction blood collection tubes, Streck blood collection tubes, Biomatrica blood collection tubes, MagMax blood collection tubes, and Norgen blood collection tubes. The anticoagulated blood collection tubes were then centrifuged to separate the plasma fraction and red blood cells. The cell-free plasma layer was collected from the anticoagulated blood collection tubes and either stored or used directly for cell-free DNA (cfDNA) extraction.

[0091] cfDNA was extracted from 2–8 mL of plasma using a commercially available extraction kit (MagMax Cell-Free DNA Isolation Kit) that uses magnetic beads. Other similar extraction methods / kits can also be used to perform this step, including column-based solid-phase methods and precipitation-based methods. cfDNA was eluted and quantified by fluorometry and electrophoresis (TapeStation).

[0092] Whole genome library is made from cfDNA by contacting cfDNA sample with random primers that are configured to amplify whole genome for sequencing.However, those skilled in the art will understand that any suitable method for amplifying sequence can be used, such as next generation sequencing.In one embodiment, when making library, unique molecular identifiers or unique sample-specific barcodes can be incorporated to perform multiplex analysis of samples from different subjects.

[0093] Example 2 Fragmentomics analysis based on fragment size distribution The embodiment of the analysis of cfDNA from subjects was carried out by comparing the size and distribution of cfDNA fragments in dog plasma with that of healthy dog ​​subjects. The library was quantified to measure the total concentration, and the fragment size was analyzed by sequencing the cfDNA fragments and analyzing the length of the cfDNA fragments obtained in the sequencing process.

[0094] The whole genome sequencing of the library was performed by paired-end sequencing at 2x100 cycles using NovaSeq 6000. However, those skilled in the art will understand that various other cycle numbers, such as 2x50 cycles, are also suitable for paired-end sequencing and can be used. After performing the sequencing run, the size of DNA fragments was measured by counting the number of nucleotides of each amplified cfDNA fragment in the library.

[0095] Example 3 Data Analysis to Identify Subjects with Tumors or Cancer as Positive The examples below demonstrate that by analyzing the fragment size distribution of cfDNA fragments, it is possible to determine whether a subject has a tumor or cancer.

[0096] Twelve sample batches containing a mixture of cancer cfDNA, tumor cfDNA, and normal cfDNA were obtained from a population of canine subjects. cfDNA was isolated, sequenced, and analyzed to calculate the fragment size distribution based on the size of each cfDNA fragment. In a series of eight studies, batches ID1-7 and batch ID12 were analyzed. These library batches contained approximately 2-5 million fragments. Figure 1 shows the fragment length distribution in the cfDNA sample batches taken from normal healthy subjects.

[0097] The fragment size distributions of multiple batches were measured directly. Furthermore, the fragment size distributions of multiple batches were used to create and compare mixture distribution models for each batch. As shown in Figure 1, the fragment length distribution is multimodal, with one mode per nucleosome, and oscillations are observed on the short and long sides of each nucleosome peak. Based on the above, the mixture distribution model was automatically selected.

[0098] As used herein, "mixture distribution model" has its usual meaning as understood by those skilled in the art and refers to a probability model that describes subpopulations within a population. This model does not require identification of the subpopulation to which each observation in the observed data set belongs. Each fragment length count is modeled with a probability distribution. In one embodiment, these fragment length distributions are overdispersed Poisson distributions, also known as negative binomial distributions. Because overdispersed Poisson distributions are positively skewed, but the data are negatively skewed, the model is fitted to the inverted data (e.g., length 1 becomes length 1000 and vice versa) and the results are inverted again. An example of this is shown in Figure 2A, where a control sample (grey line) and its model fit (black line) are shown below the four-component negative binomial mixture model. In a second embodiment, the data is modeled with a Gaussian mixture model (Figure 2B). In this model, the fragment size distribution is approximated by a Gaussian distribution. A Gaussian distribution has the advantage that the mode is approximately symmetric. In the Gaussian mixture model, the first peak is well modeled at the expense of the second peak. In a third embodiment, the model essentially consists of a smoothing of the profile and identification of the location of the peak and its maximum height (FIG. 2C).

[0099] Despite the differences observed in Figures 2A-2C, the mode distributions of these mixture distribution models are similar (Figures 3A-3C). Normal samples are either from baseline measurements (circles) or from "outlier" samples (triangles) from threshold measurements (referred to herein as "test"). The labels in Figures 3A-3C indicate the names of the patient samples that were found to be normal. All test samples, named normal clusters, included control samples, but several other sample clusters were also observed. From a fragmentomics perspective, these test samples could indeed be considered normal. The distributions of the weights estimated from the mixture distribution models are shown in Figures 4A-4B, and the distributions of the estimated scale parameters (overdispersion for negative binomial mixture models and standard deviation for Gaussian mixture models) are shown in Figures 5A-5B. For classification by calculation of p-values ​​in multivariate analysis or by machine learning methods, the location of the mode, the scale parameters, and the weights can be considered independently or together. Additional features may include measurements of the amplitude of oscillations, the area under the curve (AUC) of the fragment size profile for short fragments (FIGS. 11A-11B), and other fragment length intervals. The correlation of these features is shown in FIG.

[0100] As disclosed herein, the values ​​of the extracted features are affected by the batch (FIGS. 6A-6B and 13-14). Specifically, the higher the batch number, the more the samples are generally shifted to the upper left corner of the PCA analysis shown in FIG. 7. This trend can also be observed in the box plot analysis of the PC values ​​calculated per batch for PC1-PC3, which captured 99.65% of the data variance (FIGS. 8A-8D). Indeed, this trend leads to higher peaks of amplified nucleosomes as the batch number increases (FIGS. 9A-9B). Also, as disclosed herein, the baseline setting is slightly biased towards older subjects (FIG. 12). With this in mind, it is envisioned that in some embodiments, the correlation between fragment size profile and age may be included in the analysis.

[0101] The initial set of statistics used to characterize and classify samples was for a reference normal sample consisting of one or more combined normal samples. The normal sample profile was smoothed to identify peak locations and the percentage of peaks at these peak locations in the normalized profile was calculated. Additionally, the KL divergence was calculated from the combined normal data. From this statistic, a batch effect was observed (Figure 9A). Next, the absolute difference in the percentage of all peaks between the combined normal samples and the sample was examined. This statistical analysis showed lower values ​​for the normal samples, meaning that the percentage of peaks in the normal samples is more similar to the combined normal samples, i.e., samples containing tumor material have an altered percentage of observable nucleosomal peaks (Figures 10A-10B). Additionally, the KL divergence, which calculates the distance between two probability distributions, compares the probability distribution of fragment sizes between 51 and 1000 bp between the sample and the reference normal sample.

[0102] As disclosed herein, the KL divergence of the normal samples is smaller than that of the reference normal samples, whether the reference normal samples are composed of all normal samples or the reference normal samples are composed of only the baseline samples. However, there is a large overlap between these two distributions. Another statistic is obtained from a Gaussian mixture model fitted to the fragment size distribution from 51 to 1000 bp. This Gaussian mixture model has four components, each derived from one of the observable nucleosome peaks. Markov chain Monte Carlo (MCMC) was used to train each parameter of each sample (four means, four standard deviations, and three mixture weights as the fourth weight obtained by subtracting the sum of the first three weights from one).

[0103] A subset of known normal fragment size distributions was used to form a baseline set from which the reference was calculated. The KL divergence values ​​differed between baseline, normal, and tumor samples, with maximum values ​​in tumor samples and minimum values ​​in baseline samples (Figures 13-14). Four thresholds were considered when comparing the KL values ​​of the experimental samples with those of the normal samples: (1) The maximum value observed in the normal group ("max") (2) The average of the two maximum values ​​observed in the normal group (the "mean"). (3) Three standard deviations ("3sd") of the mean value of KL in the normal group (4) 4 standard deviations ("4sd") of the mean value of KL in the normal group

[0104] The threshold may be selected to optimize a particular criterion. For example, the data from batches 1-3 were used to calculate the accuracy, sensitivity, specificity, positive predictive value (PPV) and F1 score thresholds (Table 1). Table 2 shows the performance metrics for batches 4-7 and batch 12. Various results were obtained by optimizing each of these metrics (Figures 16A-16D). As disclosed herein, optimization of sensitivity and specificity was not preferred as objective variables because a pathological solution was preferred. The conclusion of optimizing PPV by prioritizing specificity seems to be a good compromise. Classifying samples as tumor or normal is a highly discriminative learning task and can usually be explored with alternative methods that do not rely on a baseline set consisting of only normal samples but rely on training data with both normal and tumor labels. An example of such a method is a) Logistic regression (LR) with penalty regularization, e.g., ridge regression, lasso regression, grouped lasso regression, fused lasso regression, etc. b) Support Vector Machine (SVM), c) Neural networks (NNs) with one or more hidden layers These include, but are not limited to:

[0105] These classification methods can utilize normalized counts in a selected range (e.g., 51-1000 bp) as features, or can utilize features extracted from the data, as described above for mixture distribution models. [Table 1] [Table 2]

[0106] Analysis of batches 4-7 and batch 12 identified 47 true positive results (i.e., 47 samples were correctly identified as being tumor or cancer), and 4 true positive results were identified from batches 1-3.

[0107] For classification, it is crucial to distinguish the distribution of data between the two classes. When the average profile for each class was taken and subtracted from the average profile of one class by the average profile of the other class, the normal samples were enriched with fragments of about 150 bp, whereas the tumor samples were enriched with multiple peaks of longer fragments (Figure 17). However, when the average profiles of the two groups were plotted together, the difference in the average profiles was small and barely visible (Figure 18). PCA analysis of all samples shows significant overlap between the normal and tumor groups (Figure 19). This is probably due to the low tumor cell content in the true tumor samples, which makes the fragment size profile look practically normal. The normal samples formed a strong cluster, suggesting that the normal sample profiles are robustly reproducible. On the other hand, the tumor samples are variously different from the normal samples, which complicates the interpretation of the tumor class distribution. Therefore, outlier detection, for example by logistic regression, may be preferable to classification. Due to (i) the observation that all normal samples are similar to each other, (ii) the observation that tumor samples may vary widely, and (iii) the desire to use all data without approximate modeling, another method for outlier detection was implemented, using a distance function, e.g., KL divergence, between the test sample and the mean of the baseline samples (Figure 20). Figure 20 shows the baseline sample with the highest KL divergence value, and all tumor samples are shown to be above this threshold.

[0108] Previously reported fragment size analysis of cfDNA for cancer detection is usually designed based on human data. Surprisingly, by using the technology described herein, it has been found that a characteristic element of the sample (pet sample) is that the number of peaks in the fragment size profile is higher in pet samples than in humans. The method described herein can indirectly benefit from the presence of the aforementioned additional peaks by taking into account the entire fragment size profile. Therefore, the presence of multiple additional peaks is more advantageous than the prior art methods and has not been reported before.

[0109] Removing outlier samples from the baseline increased the number of samples whose KL divergence was significantly different from normal samples, but some false positives were also observed (Figure 21). Testing the dataset with the analytical methodology disclosed herein performed slightly better than the probabilistic method previously described in terms of high specificity and high sensitivity.

[0110] Example 4 Fragmentomics-based estimation of tumor cell content A summary of the methodology for performing fragmentomics-based tumor cell content estimation is provided in the Examples below.

[0111] 1. Probabilistic model:

number

[0112] The above probability model defines three unknown parameters and their prior probabilities, deterministic calculations and likelihood models. The observed data are stored in a matrix Y, with one column for each copy number (CN) (e.g., 1 copy, 2 copies, 3 copies).

[0113] The first unknown is the tumor cell fraction (TC)t, with a prior value that is preferably small. No information about the sample was taken into account in the prior probability. By using a monotonically decreasing curve, we avoided exploring alternatives with each deconvolutional layer by swapping the parameters θ of the normal label and the parameters θ of the tumor label. The prior distribution is shown in Figure 22.

[0114] θN is the pure normal profile and θT is the pure tumor profile. Their prior probabilities are Dirichlet distributed, which guarantees non-negativity and constant sum constraints. The parameters of the prior distribution were obtained from the estimates of the Model 5 profile by deconvolution according to the following formula: These prior probabilities were therefore empirical probabilities based on several aspects of the data.

[0115] 2. Model 5 equation:

number

[0116] Here, Y represents the normalized count data, e.g., Y is the gain profile (CN3) divided by the total number of leads in the gain profile.

number

[0117] Since the above formulas all depend on the unknown tumor cell content (TC) t, the formulas were calculated in 1% increments for each t value ranging from 1 to 99%. The above formulas were solved for any value of t to obtain an estimate of the pure profile. For every value of t, an estimate of the data was obtained (see Q in the model above) and compared to the observed data. Model 5 used a normal distribution to select the t value (and thus the pure profile for the given data) that gave the highest log-likelihood (best fit). The solutions to the above formulas may contain negative numbers. These negative numbers were replaced by 0 before the estimate of the data was obtained. Solutions with more than 20% non-positive entries were ignored. Such entries generally occurred near the extreme TC values ​​in the group of samples with moderate TC.

[0118] We biased the estimates in Model 5 by rescaling the parameter α.

number

[0119] The scaling factor was 6M, where M is the length of the fragmentomics profile (the number of rows in the Y count matrix). The bias was 1.

[0120] The equation at the bottom of the above model is the multinomial likelihood.

[0121] For computational efficiency, instead of analyzing every M position in the fragment length profile (51-260, M=210), the amount of data was halved by considering only every third position (i.e., M=105 in increments of 2 in the range 51-259).

[0122] The above model was run using Stan software, with 12,000 iterations as a warm-up and 3,000 iterations as sampling. The following parameters were initialized and run in parallel with 4 iterations: (1) t was set to 5%; (2) θN was sampled from its prior probability (Dirichlet distribution (αN)); (3) θT was sampled from its prior probability (Dirichlet distribution (αT)).

[0123] The following one control parameter was set: max_treedepth=20

[0124] The model was built using two sets of in silico mixtures and tested using another two sets of in silico mixtures. Each mixture was diluted and analyzed in triplicate to evaluate the robustness of the model in terms of data resampling and convergence to a solution. The model failed to converge for one sample out of 225 (=57×3+54) diluted samples. A total of 57 samples were generated by diluting the samples in 19 steps and analyzing each in triplicate. Three sets of mixtures generated 19×3=57 samples, while one set of mixtures had a tumor cell content (TC) lower than the maximum dilution ratio, resulting in 18×3=54 samples.

[0125] The first mixture was made by mixing samples 201-20885 and 201-00316 with healthy cfDNA. The normal samples were selected to have a fragmentomics profile as similar as possible to the pure normal signal observed in the cancer samples.

[0126] The KL divergence of the pure normal signal between sample 201-20885 and its "matched normal" was about 0.003, and the KL divergence of the pure normal signal between sample 201-00316 and its "matched normal" was 0.03. Sample 201-00316 was difficult to analyze because the KL divergence value was large, breaking down the assumption that there were only two signals (normal and cancer) in the data. Instead, there were two normal signals and a cancer signal in various proportions.

[0127] In both samples, TC was slightly overestimated (Figures 23A and 23B). This effect was more pronounced at low TC in sample 201-20885 and at high TC in sample 201-00316. Expected TC values ​​were derived from the original TC estimates of the undiluted samples.

[0128] Since the model construction was dependent on the mixtures, two other sets of mixtures were made to test the performance (Figure 24A and Figure 24B). Normal samples were obtained that closely matched these cancer samples (KL divergence: ~0.003). In this test, TC was also overestimated, but this overestimation was only seen at small TC values ​​(<10%).

[0129] Example 5 in silico mixture The method for preparing an in silico mixture of tumor and normal samples is summarized in the following examples. In order to obtain a true value when evaluating the performance of estimating the tumor cell content ratio, an in silico mixture of tumor and normal samples was prepared, the mixture ratio of pure profiles was examined, and the mixture ratio was adjusted.

[0130] Tumor-rich cfDNA samples were mixed with healthy cfDNA samples. To create a fragmentomics baseline, the fragment length profile of the healthy cfDNA sample must match the fragment length profile of the normal component of the cancer-containing cfDNA sample.

[0131] Multiple samples were screened to identify samples with clear signals and high tumor cell content. Sample 201-20885 had well resolved copy number specific fragment length profiles consistent with ichorCNA results, with an estimated tumor cell content (TC) of 51% ([45%-56.8%]).

[0132] Sample 201-00316 was confined to copy 1 and copy 3 regions, but potential copy 4 and copy 5 regions were also observed. As expected, the gains in copy 4 and copy 5 appeared to be biased toward short fragments (Figure 25). Based on copies 1, 2, and 3, we estimated the tumor cell content (TC) to be 43.7% ([33.5%-58.1%]). This estimate was lower than the tumor cell content (TC) predicted by ichorCNA.

[0133] Sample 101-10849 was selected as a normal sample representing the normal components of sample 201-20885 (KL=0.003 from the pure profile), and sample 101-00013 was selected as a normal sample representing the normal components of sample 201-00316 (KL=0.036 from the pure profile).

[0134] The tumor cell content of the selected sample 201-20885 was about 51%, and the tumor cell content of another selected sample 201-00316 was about 44%. It should be noted that the mixture of sample 201-00316 was difficult to deconvolute due to its lower tumor cell content, normal sample with slightly different normal signal, and fewer total reads. Various mixtures containing each of these two cancer samples were made in triplicate with the ratios shown in Table 3. [Table 3]

[0135] Example 6 Quantification of separation between copy number specific fragment length curves This example outlines a methodology for quantification of the separation between fragment length curves in sample analysis.

[0136] Fragment length profiles can be calculated and plotted not only genome-wide (as shown in Example 3), but also according to copy number (as shown in Example 4).

[0137] In the calculation of profiles in regions of copy number gain (gain), an increase in the proportion of short fragments is predicted, whereas in the calculation of profiles in regions of copy number loss (loss), an increase in the proportion of long fragments is predicted.

[0138] Due to these differences, there exists a critical fragment length below which an increasing profile is observed above a decreasing profile, and above which a decreasing profile is observed above an increasing profile, with an unchanged profile lying somewhere between the increasing and decreasing profiles.

[0139] The separation between fragment length curves in a single sample is quantified according to the following scheme: In the region below the critical fragment length, the decrease profile is subtracted from the increase profile and the resulting differences are summed to give the quantity A. In the region above the critical fragment length, the increase profile is subtracted from the decrease profile and the resulting differences are summed to give the quantity B. The separation between the fragment length curves is the sum of A and B.

[0140] If either the decrease or increase profile is unavailable, the no change profile is used instead. The separation between the fragment length curves can then be calculated using three formulas: decrease-increase, decrease-no change, and no change-increase. The names in parentheses indicate the names of the profiles used.

[0141] The calculation of the value indicating the separation between the fragment length curves depends on the location of the critical fragment length. Various methods can be considered for dealing with this unknown: a single threshold can be used for all samples, the central interval including the critical fragment length can be ignored, or the threshold can be optimized for each sample (Figure 27).

[0142] Fragment length profiles derived from a small number of reads may be unreliable. For this reason, profiles with read counts below a certain threshold can be eliminated, including but not limited to 100,000 reads, 200,000 reads, 500,000 reads, and 1,000,000 reads.

[0143] Furthermore, the profiles can be smoothed using a spline curve, which does not affect the value of the separation between the fragment length curves, and the threshold for large separation remains stable, except in a few cases where there is practically no separation between the fragment length profiles (Figure 28).

[0144] Because the no-change profile lies between the decrease and increase profiles, the separation values ​​calculated using the equations "decrease" - "no change" or "no change" - "increase" are smaller than the separation values ​​calculated using the equations "decrease" - "increase". This effect was quantified by analyzing samples where all three levels were available (Figure 29A). The resulting linear relationship could be used to obtain a simple linear correction.

[0145] The residuals of this correction did not correlate well with the estimated tumor cell fraction determined from the minimum number of reads required to construct a fragment length profile (Figure 29B), but large residuals were observed for profiles constructed with a small number of reads.

[0146] When the 200,000 reads were subjected to a lead filter, the adjusted separation values ​​calculated using the "decrease"-"no change" equation and the "no change"-"increase" equation closely matched the separation values ​​calculated using the "decrease"-"increase" equation (Figure 29C). This result indicates that the separation values ​​calculated using all three equations can be treated equally and analyzed together (Figure 30).

[0147] Example 7 Fragmentomics in Hematological Cancer Fragmentomics analysis of blood cancer samples is shown in the following examples. Based on the observation that blood cancer samples with clear copy number profiles do not show separation in fragment length curves, other confirmed blood cancer cases are examined to investigate whether the lack of separation in fragment length curves is a feature that indicates blood cancer. If this hypothesis is correct, this method can be used as an alternative method to classify samples as blood cancer.

[0148] They examined 112 confirmed lymphoma samples and classified copy number variations (CNVs) as loss, neutral, or gain.

[0149] Fragment length curves were plotted for each copy number (CN) group. It was expected that the contribution of tumors to each copy number level would vary, but the fragment length profile would not change. This prediction could be explained by the fact that in healthy subjects, the majority of cfDNA is released from leukocytes, and therefore, if a malignant tumor originates from leukocytes, it would show the same nucleosome organization and the same DNA fragmentation as healthy subjects.

[0150] In 90 of the 112 samples, copy number variation (CNV) could be visualized by annotation. In 18 of the 90 cases (20%), separation between the fragmentomics curves was observed, defined as (i) a clear line on the long flank of the major nucleosome peak, and (ii) below about 150 bp, the most tumor-rich profile (e.g., copy number gain of 5) is found above the least tumor-rich profile (copy number 1), and above about 150 bp, the most tumor-rich profile (e.g., copy number gain of 5) is found below the least tumor-rich profile (copy number 1).

[0151] The main analysis was performed by considering all CNV-positive clinical evaluation (CV) samples in which classification of copy number variation (CNV) as copy number loss, no change, or gain was possible. Thus, 245 samples were considered. Segregation scores were calculated using the upper minus lower curve formula for comparing the default loss and gain curves ("full" formula), or by comparison with the "loss" minus "no change" formula, or with the "no change" minus "gain" formula ("partial" formula). Copy number levels calculated from less than 200,000 reads were ignored. 214 samples remained. Segregation values ​​calculated with the "partial" formula were corrected to match those calculated with the "full" formula. The corrected regression model did not include an intercept.

[0152] The resulting regression models were manually reviewed to assign a label to each sample ("separation", "no separation at low tumor cell fraction" or "no separation at high tumor cell fraction"). Separation between the fragment length curves was expected. For short fragments, the increase curve would be located above the no change curve and the no change curve would be located above the decrease curve, with a similar separation expected in the reverse order for longer fragments after passing a change point around 150 bp (Figure 27). Thus, the labels were in good agreement with the separation scores calculated above.

[0153] It was interesting to identify samples that showed no separation between the fragment length curves despite evidence of high tumor cell content in the CNV data. We filtered for only samples with separation between the fragment length curves and samples with high tumor cell content but no separation between the fragment length curves and examined various call thresholds. A threshold of 0.0173 was found to give the best results (sensitivity 97.7%, specificity 98.5%). This analysis was performed blinded to tumor type.

[0154] Further examination of the above two categories revealed two false positive samples at this threshold (samples with separation between fragment length curves but low separation values). Separation was identified in these samples, and samples similar to these samples were identified on re-examination and excluded from the call. For these excluded samples, no cancer signal of origin (CSO) prediction based on separation between fragmentomics curves was performed.

[0155] Before unblinding, we assessed for any apparent batch effect that may affect the performance of the method. A slight increasing trend was observed during the sequencing runs. However, the 95% confidence interval of the regression line included a horizontal line, meaning that we could not reject the null hypothesis of no batch effect. The covariates age and sex led to a similar conclusion.

[0156] In unblinded studies, B-cell lymphomas tended to show no separation between the fragment length curves despite having high tumor cell content, whereas other cancer types such as T-cell lymphomas showed separation between the fragment length curves (Figure 26). After review, the final call was for samples labeled with no separation / high tumor cell content (TC) with an adjusted separation value of <0.01727873, as shown in Table 4. [Table 4]

[0157] The performance of the test set was then calculated, as shown in Table 5. The fragmentomics method requires confirmation of at least two copy number levels (decrease and no change; no change and increase; or decrease and increase) built with at least 200,000 reads. Samples that did not meet these criteria were not given a separation score. The heme prediction currently used in the commercially available OncoK9 test is based on copy number profiles, which have previously been reported to have features associated with hematological cancers (https: / / pubmed.ncbi.nlm.nih.gov / 14562028 / ). The fragmentomics analysis of hematological cancers described in this example had improved sensitivity, as shown in Table 5. [Table 5]

[0158] Example 8 Chromosome-based fragmentomics The use of chromosome-based fragmentomics to detect cancer signals in cfDNA samples is demonstrated in the Examples below.

[0159] One way to improve the sensitivity of cancer detection is to increase the signal from the tumor. The percentage of tumor cells in the blood is usually small. Chromosomal gains increase tumor DNA, and chromosomal losses decrease tumor DNA. Standard fragmentomics analysis considers all reads genome-wide, reducing the signal in gains with the signal in regions with no copy number change, and the "worse" signal from chromosomal losses.

[0160] In this analysis, we investigated chromosome-based fragmentomics in the hope of finding and exploiting differences between individual chromosomes, based on the finding that only about 100,000 fragments are required to obtain a fragment length profile.

[0161] The chromosomes in each sample were compared. For each chromosome pair, the KL divergence between their fragment length distributions was compared. In the presence of tumor and copy number alterations (CNAs), the copy number altered chromosomes consistently showed larger deviations from the copy number unchanged (CNN) chromosomes.

[0162] By understanding what KL divergence values ​​are expected in healthy samples, a call threshold can be established to identify cancer positive samples without the need to compare the test sample to a normal sample.

[0163] To identify the call threshold, we performed pairwise comparisons between chromosomes in euploid normal samples. We examined the mean KL divergence per chromosome in normal samples and found that baseline values ​​were heterogeneous and inversely proportional to chromosome length. Longer chromosomes yield more reads, and more reads yield smoother fragment length profiles and less KL increase due to noise and artifacts.

[0164] Furthermore, chromosome 9 was an outlier: although it is only about the same length as chromosomes 14 and 16, its average KL divergence was much higher than expected, likely due to the high GC content of the sequence.

[0165] To address the artifactual increase in KL divergence observed for certain chromosomes (e.g., chromosome 9 and short chromosomes), the observed KL can be corrected using the chromosome length or the average KL value across samples as a single factor. Alternatively, the relationship between the average KL value and the number of reads can be modeled.

[0166] Focusing on the autosomes, we modeled each KL curve as a hyperbola expressed by the following equation: y=a / (x+b)

[0167] This function was then fitted for each chromosome using a least squares method, resulting in the model shown in Figure 32. In this model, the KL divergence values ​​observed for each chromosome were normalized as a function of the number of reads mapped to each chromosome.

[0168] This correction eliminated the gradient of KL with read number and chromosome-specific artifacts. Although this correction focused on the number of reads, it also focused on each chromosome, implicitly taking into account differences in GC content between chromosomes. If necessary, a more refined correction can be performed that takes into account both the number of reads in a particular region and the GC content of that region.

[0169] In this chromosome-based fragmentomics approach, positive samples are called by identifying chromosomes with high normalized mean KL values: if the normalized mean KL value is above a certain threshold, the sample is called positive for cancer.

[0170] The threshold can be defined in a variety of ways, for example, the threshold can be the number of standard deviations above the mean value of normal samples, a threshold that optimizes the accuracy of the data set, etc.

[0171] First, we considered the average KL value adjusted by chromosome and defined a single threshold using the mean and standard deviation (SD). The results compared to the genome-wide fragmentomics baseline are shown in Table 6. [Table 6]

[0172] Despite the small number of normal samples, the chromosome-based fragmentomics approach outperformed the genome-wide baseline in this dataset.

[0173] The model corrections described above worked on average values ​​per chromosome, not on single values. Thus, small deviating cancer signals may be diluted by the calculation of average values. Instead, pairwise comparisons between chromosomes per sample could be considered to further improve sensitivity.

[0174] To implement this method, a normalization of the KL values ​​can be performed for pairs of chromosomes. As mentioned before, each chromosome is described by a hyperbola with chromosome-specific parameters. The KL divergence by pairwise comparisons is modeled as the sum of two chromosome-specific hyperbolae by the following formula: y=a / (x1+b)+c / (x2+d)

[0175] Chromosome-specific parameters were trained by Markov chain Monte Carlo using pairwise KL divergence values ​​from a set of normal samples as data. The correlation between predicted and observed KL values ​​as a function of the number of reads mapped to a chromosome was 0.9728 (Figure 33).

[0176] Normalization of the divergence values ​​by pairwise comparisons in the test samples allowed us to eliminate chromosome-specific and length-related artifacts.

[0177] A call threshold was established for each pairwise comparison, and may be selected, including but not limited to, the number of standard deviations away from the mean of the control samples, or the control values ​​may be modeled with a probability distribution and a percentile selected from that probability distribution, as shown in Table 7. [Table 7]

[0178] Genome-wide fragmentomics methods compare the genome-wide fragment length profile of a test sample with that of a set of normal samples. Another chromosome-based fragmentomics method applies the same genome-wide methodology to each individual chromosome of a test sample, comparing each individual chromosome with an external standard. The status of the test sample is then determined, for example, by taking into account the most extreme chromosome-level results or the most significant changes compared to the genome-wide method.

[0179] Figure 31 and Table 8 show the change in KL divergence and lines 3, 4 or 5 SD away from the mean of normal samples. Nine samples were called at the 3 SD threshold. [Table 8]

[0180] The section headings used herein are provided for the sole purpose of organizing the present invention and should not be construed as limiting the subject matter described herein. All academic literature and similar materials cited in this application, including but not limited to patents, patent applications, articles, books, papers, and Internet web pages, are expressly incorporated herein in their entirety, including the disclosures specifically cited herein, by reference for any purpose. Where the definition of a term in an incorporated reference differs from the definition of the term in the present teachings, the definition of the term in the present teachings shall be adopted. When a term including the meaning of "about" is placed before the temperature, concentration, time, etc., stated in the present teachings, it is fully understood that very slight deviations are within the scope of the present teachings.

[0181] Although the present invention has been disclosed in specific embodiments and examples, those skilled in the art will understand that the present invention can be extended from the specifically disclosed embodiments to other embodiments and / or uses of the invention, modifications and equivalents thereof that will be apparent to those skilled in the art. Moreover, while several variations of the present invention have been shown and described in detail, other modifications within the scope of the present invention will be readily apparent to those skilled in the art in light of this disclosure. Moreover, it is contemplated that various combinations or subcombinations of specific features and aspects of the embodiments of the present invention are possible and such combinations are within the scope of the present invention. It should be understood that various features and aspects of the embodiments disclosed herein may be combined with or substituted for other features and aspects in order to form various aspects or embodiments of the invention disclosed herein. Thus, the scope of the invention disclosed herein should not be limited by the specific embodiments disclosed herein above.

[0182] However, since various changes and modifications within the spirit and scope of the invention will become apparent to those skilled in the art, the detailed description indicating the preferred embodiment of the invention should be interpreted as being for the purpose of illustrating the invention only.

[0183] The terms used in the description herein are not to be construed as limiting or restrictive. These terms are merely used in conjunction with the detailed description of the embodiments of the systems, methods and related elements of the present invention. Furthermore, the embodiments of the present invention may include several novel features, and no single one of these features contributes to the desirable properties, nor is any one of these features to be construed as essential to the implementation of the invention described herein.

Claims

**Claim 1** A method for detecting cancer or a tumor in a subject, comprising: isolating a cell-free DNA (cfDNA) sample from the subject; sequencing the cfDNA sample to measure one or more fragment size distributions; comparing the one or more fragment size distributions with a second fragment size distribution obtained from one or more control subjects; and determining the presence or absence of cancer or a tumor based on the comparison of the two fragment size distributions. The method, wherein the subject is a non-human subject. **Claim 2** The method according to claim 1, wherein (i) the one or more control subjects include the subject or one or more healthy subjects; or (ii) the sequencing of the cfDNA sample is whole genome sequencing or next generation sequencing. **Claim 3** The method according to claim 1, further comprising creating a model of the one or more fragment size distributions. **Claim 4** The method according to claim 3, wherein (a) the model of the one or more fragment size distributions is a statistical model; or (b) the model of the one or more fragment size distributions is obtained from one or more features extracted from the one or more fragment size distributions. **Claim 5** The method according to claim 4, wherein (a') the one or more features include a median value, an average value, an area under the curve (AUC), an amplitude of vibration, a variance, a standard deviation, an interval of fragment length, or a combination thereof; or (b') the method further comprises classifying a sample as tumor or normal based on the one or more features. **Claim 6** The method according to claim 1, characterized by any one of the following (i) to (vi): (i) the model of the second fragment size distribution is a mixture distribution model; (ii) the comparison of the one or more fragment size distributions and the second fragment size distribution is performed by measuring a distance or similarity; (iii) the one or more fragment size distributions are calculated from at least one of the length or sequence of cfDNA fragments in the sample; (iv) the second fragment size distribution is a baseline fragment size distribution; (v) the subject is a mammal; (vi) the cfDNA sample is isolated from the blood of the subject. **Claim 7**: The method according to claim 1, wherein the comparison of the one or more fragment size distributions and the second fragment size distribution is performed by measuring a distance or similarity, and the measurement of the distance or similarity is the KL divergence. **Claim 8** The method according to claim 1, wherein the subject is a dog, a cat or a horse. **Claim 9**: The method according to claim 1, wherein the cfDNA sample is isolated from the blood of the subject, and the blood of the subject further contains circulating tumor DNA (ctDNA). **Claim 10**: The method according to claim 1, further comprising the step of ligating an adapter to the isolated cfDNA and producing an amplified fragment by using a universal primer targeting the adapter. **Claim 11**: The method according to claim 10, characterized by any one of the following (a) to (c): (a) The measurement of the one or more fragment size distributions is performed by measuring the number and distribution of amplified fragment sizes using whole genome sequencing or next-generation sequencing; (b) The comparison of the one or more fragment size distributions and the second fragment size distribution is performed by comparing the number and distribution of the amplified fragment sizes with the subject or one or more healthy subjects, and from this comparison result, it is determined whether the number and distribution of the amplified fragment sizes of the subject are different from the number and distribution of the amplified fragment sizes of the one or more healthy subjects; (c) The universal primer further includes a sequence-specific primer. **Claim 12**: (i) It is indicated that cancer or a tumor is present because there is a statistically significant difference between the one or more fragment size distributions of the subject and the second fragment size distributions of one or more healthy subjects, or (ii) It is indicated that cancer or a tumor is absent because there is no statistically significant difference between the one or more fragment size distributions of the subject and the second fragment size distributions of one or more healthy subjects. The method according to claim 1. **Claim 13** The method according to claim 1, wherein the cancer is a blood cancer or lymphoma. **Claim 14** A method for predicting the cancer signal origin (CSO) in a subject in whom a cancer positive signal has been detected, comprising: isolating a circulating cell-free DNA (cfDNA) sample from the subject; Sequencing the cfDNA sample to measure the fragment size distribution and the copy number (CN) profile; Detecting a cancer positive signal from the copy number profile; Comparing the fragment size distribution of the copy number increase region and / or the copy number decrease region with the fragment size distribution of the control copy number region, or comparing the fragment size distribution of the copy number increase region with the fragment size distribution of the copy number decrease region; and Predicting the cancer signal origin (CSO) based on the difference between the fragment size distribution of the copy number increase region and / or the copy number decrease region and the fragment size distribution of the control copy number region, or based on the absence of a difference between these fragment size distributions comprising a method, wherein the subject is a non-human subject.

15. The method according to claim 14, wherein it is predicted that the subject has a blood cancer because there is no difference between the fragment size distribution of the copy number increase region and / or the copy number decrease region and the fragment size distribution of the control copy number region.

16. A method for detecting cancer or a tumor in a subject, comprising isolating a circulating cell-free DNA (cfDNA) sample from the subject; sequencing the cfDNA sample to measure one or more fragment size distributions; creating an experimental model of the one or more fragment size distributions; comparing the one or more fragment size distributions with a second fragment size distribution obtained from the subject or one or more control subjects; and determining the presence or absence of cancer or a tumor based on the comparison of the two fragment size distributions comprising a method, wherein the subject is a non-human subject.

17. The method according to claim 16, characterized by any one of the following (i) to (iv): (i) the one or more control subjects include the subject or one or more healthy subjects; (ii) the experimental model of the one or more fragment size distributions is a statistical model; (iii) the experimental model of the one or more fragment size distributions is obtained from one or more feature quantities extracted from the one or more fragment size distributions; (iv) further comprising comparing the experimental model obtained from the cfDNA sample with a control model obtained from a control cfDNA sample of an individual known to have no cancer or tumor.

18. The experimental model of the one or more fragment size distributions is obtained from one or more feature quantities extracted from the one or more fragment size distributions, and the one or more feature quantities include a median value, an average value, an area under the curve (AUC), an amplitude of vibration, a variance, a standard deviation, an interval of fragment length, or a combination thereof. The method according to claim 16.

19. The method further includes a step of comparing the experimental model obtained from the cfDNA sample with a control model obtained from a control cfDNA sample of an individual who has been determined to have no cancer or tumor, and (a) by comparing the experimental model with the control model, it is determined whether the subject has a possibility of having cancer or a tumor, or (b) by comparing one or more feature quantities of the experimental model with one or more feature quantities of the control model, it is determined whether the subject has a possibility of having cancer or a tumor. The method according to claim 16.

20. The method according to claim 16, wherein the comparison between the one or more fragment size distributions and the second fragment size distribution obtained from at least one healthy subject is performed by measuring a distance or similarity.

21. The method according to claim 20, wherein the measurement of the distance or similarity is KL divergence.

22. A method for measuring a fragment size distribution in a sample, comprising: sequencing a DNA sample isolated from a subject who has cancer or is suspected of having cancer to measure a fragment size distribution; measuring one or more feature quantities from the fragment size distribution; and creating an experimental model of the fragment size distribution The method comprising.

23. The method according to claim 22, characterized in that any one of the following (i) to (v): (i) the experimental model is a statistical model; (ii) the experimental model is obtained from the one or more feature quantities; (iii) the one or more feature quantities include a median value, an average value, an area under the curve (AUC), an amplitude of vibration, a variance, a standard deviation, an interval of fragment length, or a combination thereof; (iv) further comprising a step of identifying the sample as a tumor sample or a normal sample based on the one or more feature quantities; (v) the fragment size distribution is calculated from at least one of the length or sequence of DNA fragments in the sample.

24. The method according to claim 22, characterized in that any one of the following (i) to (v): (i)The DNA sample is a cell-free DNA (cfDNA) sample; (ii)The DNA sample is isolated from the blood of the subject; (iii)The DNA sample is isolated from the blood of the subject, and the blood further contains circulating tumor DNA (ctDNA); (iv)The sequencing includes whole genome sequencing or next-generation sequencing; (v)The method further includes a step of producing amplified fragments by ligating an adapter to the isolated DNA and using a universal primer targeting the adapter. **Claim 25** The method further includes a step of producing amplified fragments by ligating an adapter to the isolated DNA and using a universal primer targeting the adapter, wherein (a) the measurement of the one or more fragment size distributions is performed by measuring the number and distribution of amplified fragment sizes using whole genome sequencing or next-generation sequencing, or (b) the universal primer further includes a sequence-specific primer, according to the method of claim 22.