Systems and computer-implemented methods for cancer detection
Computer-implemented methods analyzing DNA sequencing data for chromosomal variations address the limitations of current cancer detection and treatment response monitoring, enhancing sensitivity and specificity for cancer detection and MRD assessment.
Patent Information
- Application Number
- PCT/US2024/061637
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-22
- Filing Date
- 2024-12-23
- Publication Date
- 2025-06-26
AI Technical Summary
Current methods for detecting cancer and characterizing responses to therapies in patients with minimal residual disease (MRD) lack sensitivity and specificity due to limitations in measuring DNA mutations at nonpolymorphic loci in liquid biopsies.
The development of computer-implemented methods and systems that analyze DNA sequencing data from blood samples to identify chromosomal variations, including unique chromosomal instabilities (CINs), to detect solid tumors or neoplasms and assess treatment responses.
These methods enable more sensitive and specific detection of cancer and characterization of treatment responses, improving the accuracy of minimal residual disease monitoring and prognosis.
Smart Images

Figure US2024061637_26062025_PF_FP_ABST
Abstract
Description
SYSTEMS AND COMPUTER-IMPLEMENTED METHODS FOR CANCER DETECTIONCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to and the benefit of U.S. Provisional Application No. 63 / 613,916, entitled "SYSTEMS AND COMPUTER-IMPLEMENTED METHODS FOR CANCER DETECTION", filed on December 22, 2023, the content of which is incorporated by reference herein in its entirety.BACKGROUND
[0002] There are currently many methods of characterizing response to one or more therapies in patients with clinically apparent cancer, or recurrence of disease in treated cancer patients with no clinically apparent cancer, i.e., minimal residual disease (MRD). These techniques perform mutational analysis of nonpolymorphic loci. However, response to therapy or MRD using this prior art of measuring DNA mutations at nonpolymorphic loci in a liquid biopsy lacks sensitivity and specificity. What are needed are automated and computer-implemented methods of detecting cancers and characterizing response to one or more therapies that do not suffer from these limitations.SUMMARY
[0003] Disclosed are computer-implemented methods, systems, and apparatuses for the detection of cancer, characterization of responses to one or more therapies in patients, and assessing the efficacy of a treatment regimen.
[0004] In some implementations, computer-implemented methods for detecting presence of solid tumors or neoplasms in a blood sample are described herein. The computer- implemented method can include: receiving Deoxyribonucleic acid (DNA) sequencing data for a subject's blood sample; identifying a plurality of chromosomal variations in the DNA sequencing data, wherein the plurality of chromosomal variations includes a plurality of different chromosomal instabilities (CINs); and detecting a solid tumor or neoplasm in the subject's blood sample based, at least in part, on analysis of the plurality of chromosomal variations.
[0005] In some implementations, each of the plurality of different CINs is a unique CIN.
[0006] In some implementations, each of the plurality of different CINs is a partial chromosomal instability or whole chromosomal instability.
[0007] In some implementations, the analysis of the plurality of chromosomal variations includes comparing the DNA sequencing data to a healthy control DNA sequencing dataset.
[0008] In some implementations, the analysis of the plurality of chromosomal variations reveals that the plurality of different CINs include at least one of a plurality of candidate CINs for the solid tumor or neoplasm in the subject's blood sample.
[0009] In some implementations, the solid tumor or neoplasm is in one of a plurality of solid tumor primary sites including lung, trachea, breast, uterine, ovarian, skin, bladder, kidney, liver, pancreas, soft tissue, bone, mouth, salivary glands, bowel, rectum, gallbladder, urethra, fallopian tube, vulva, testicle, prostate, penis, thyroid, adrenal, parathyroid, heart, vessels, thymus, nerve, eye, ear, and subcutaneous tissue.
[0010] In some implementations, at least one of the plurality of chromosomal variations involves partial instability of chromosome 1, chromosome 2, chromosome 3, chromosome 4, chromosome 5, chromosome 6, chromosome 7, chromosome 8, chromosome 9, chromosome 10, chromosome 11, chromosome 12, chromosome 13, chromosome 14, chromosome 15, chromosome 16, chromosome 17, chromosome 18, chromosome 19, chromosome 20, chromosome 21, chromosome 22, chromosome X, or chromosome Y.
[0011] In some implementations, at least one of the plurality of chromosomal variations involves whole instability of chromosome 1, chromosome 2, chromosome 3, chromosome 4, chromosome 5, chromosome 6, chromosome 7, chromosome 8, chromosome 9, chromosome 10, chromosome 11, chromosome 12, chromosome 13, chromosome 14, chromosome 15, chromosome 16, chromosome 17, chromosome 18, chromosome 19, chromosome 20, chromosome 21, chromosome 22, chromosome X, or chromosome Y.
[0012] In some implementations, the DNA sequencing data encodes a genome wide single nucleotide polymorphism (SNP) backbone.
[0013] In some implementations, the DNA sequencing data encodes a plurality of polymorphisms in a genome wide single nucleotide polymorphism (SNP) backbone.
[0014] In some implementations, the computer-implemented method further includes: in response to detecting the solid tumor or neoplasm in the subject's blood sample, determining a presence or absence of cancer for the subject.
[0015] In some implementations, the computer-implemented method further includes: generating a report including the DNA sequencing data and an indication of the detected solid tumor or neoplasm, and / or the presence or absence of cancer for the subject.
[0016] In some implementations, the computer-implemented method further includes: generating display data for the report.
[0017] In some implementations, the computer-implemented method further includes: transmitting the report over a network.
[0018] In some implementations, the computer-implemented method further includes: in response to determining the presence or absence of cancer for the subject, determining a prognosis for the subject.
[0019] In some implementations, the computer-implemented method further includes: in response to determining the presence or absence of cancer for the subject, determining a response to treatment for the subject.
[0020] In some implementations, the computer-implemented method further includes: in response to determining the presence or absence of cancer for the subject, providing a determination of minimal or measurable residual disease for the subject.
[0021] In some implementations, the computer-implemented method further includes: sequencing the DNA to obtain the DNA sequencing data.
[0022] In some implementations, the DNA is sequenced using a next generation sequencing device (NGS).
[0023] In some implementations, the computer-implemented method further includes: sequencing the DNA or analyzing the plurality of chromosomal variations using a parallel processing operation.
[0024] In some implementations, a method for detecting solid tumors or neoplasms in a blood sample is provided. The method can include: detecting, using a computing device, the solid tumor or neoplasm in the subject's blood sample; and in response to detecting the solid tumor or neoplasm, providing a treatment to the subject.
[0025] In some implementations, the method further includes: obtaining a sample from the subject; isolating cell free (cf) deoxyribonucleic acid (DNAj(cfDNA) from the sample; and sequencing the cfDNA to obtain the DNA sequencing data.
[0026] In some implementations, the sample includes at least one of whole blood, peripheral blood, or plasma.
[0027] In some implementations, a system for detecting solid tumors or neoplasms is described herein. The system can include: at least one processor; and a memory operably coupled to the at least one processor, wherein the memory has computer executable instructions stored thereon that, when executed by the at least one processor, cause the at least one processor to: receive DNA sequencing data for a subject's blood sample; identify a plurality of chromosomal variations in the DNA sequencing data, wherein the plurality of variations includes a plurality of chromosomal instabilities (CINs); detect a solid tumor or neoplasm in the subject's blood sample based, at least in part, on analysis of the plurality of chromosomal variations.
[0028] In some implementations, memory has further computer executable instructions stored thereon that, when executed by the at least one processor, cause the at least one processor to: in response to detecting the solid tumor or neoplasm in the subject's blood sample, determine a presence or absence of cancer for the subject.
[0029] In some implementations, the memory has further computer executable instructions stored thereon that, when executed by the at least one processor, cause the at least one processor to: generate a report including the DNA sequencing data and an indication of the detected solid tumor or neoplasm in the subject's blood sample and / or the presence or absence of cancer.
[0030] In some implementations, the memory has further computer-executable instructions stored thereon that, when executed by the at least one processor, cause the at least one processor to: generate display data for the report.
[0031] In some implementations, the memory has further computer-executable instructions stored thereon that, when executed by the at least one processor, cause the at least one processor to: transmit the report over a network.
[0032] In some implementations, the memory has further computer executable instructions stored thereon that, when executed by the at least one processor, cause the at least one processor to: in response to determining the presence or absence of cancer, determine a prognosis for the subject.
[0033] In some implementations, the memory has further computer executable instructions stored thereon that, when executed by the at least one processor, cause the at least one processor to: in response to determining the presence or absence of cancer, determine a response to treatment for the subject.
[0034] In some implementations, the memory has further computer executable instructions stored thereon that, when executed by the at least one processor, cause the at least one processor to: in response to determining the presence or absence of cancer, determine a minimal or measurable residual disease for the subject.
[0035] In some implementations, a computer-implemented method for detecting presence of solid tumors or neoplasms in a subject is provided. The computer-implemented method can include: receiving Deoxyribonucleic acid (DNA) sequencing data for a subject's liquid biopsy sample; identifying a plurality of chromosomal variations in the DNA sequencing data, wherein the plurality of chromosomal variations includes a plurality of different chromosomal instabilities (CINs); and detecting a solid tumor or neoplasm in the subject based, at least in part, on analysis of the plurality of chromosomal variations.
[0036] In some implementations, the subject's liquid biopsy sample is a blood sample where the blood sample includes at least one of whole blood, peripheral blood, or plasma.
[0037] In some implementations, the subject's liquid biopsy sample is serum, saliva, sputum, cerebral spinal fluid, urine, lymph, or lacrimal fluid.
[0038] In some implementations, a computer-implemented method for determining a response to therapy is provided. The computer-implemented method can include: receiving first Deoxyribonucleic acid (DNA) sequencing data for a subject's first liquid biopsy sample; identifying a first plurality of chromosomal variations in the first DNA sequencing data; receiving second DNA sequencing data for the subject's second liquid biopsy sample; identifying a second plurality of chromosomal variations in the second DNA sequencing data; and determining the subject's response to therapy based, at least in part, on a comparison of the first plurality of chromosomal variations in the first DNA sequencing data and the second plurality of chromosomal variations in the second DNA sequencing data, wherein each plurality of chromosomal variations includes a plurality of different chromosomal instabilities (CINs).
[0039] In some implementations, each of the subject's first and second liquid biopsy samples is a blood sample, and the blood sample includes at least one of whole blood, peripheral blood, plasma, or serum.
[0040] In some implementations, each of the subject's first and second liquid biopsy samples is serum, saliva, sputum, cerebral spinal fluid, urine, lymph, or lacrimal fluid.
[0041] In some implementations, the subject's first liquid biopsy sample is obtained at a first time point and the subject's second liquid biopsy sample is obtained at a second time point, the second time point being subsequent in time relative to the first time point.
[0042] In some implementations, the subject's first liquid biopsy sample is obtained before or after treatment begins.
[0043] In some implementations, the subject's second liquid biopsy sample is obtained after treatment begins.
[0044] In some implementations, methods for detecting presence of solid tumors or neoplasms in a sample are described herein. The method can include: receiving DNA sequencing data for a subject's sample; identifying a plurality of chromosomal variations in the DNA sequencing data, wherein the plurality of chromosomal variations includes a plurality of different CINs; and detecting a solid tumor or neoplasm in the subject's sample based, at least in part, on analysis of the plurality of chromosomal variations.
[0045] In one embodiment, disclosed herein are methods for detecting a cancer, cancer recurrence, and / or metastasis (including, but not limited to cancer recurrence in a subject previouslytreated for cancer with no clinically apparent cancer, i.e., MRD or response to therapy or treatment, used interchangeably herein) in a subject comprising: a) obtaining a first sample (such as, for example a liquid biopsy including, but not limited to a liquid biopsy comprising whole blood, peripheral blood, plasma, serum, saliva, sputum, cerebral spinal fluid, urine, lymph, or lacrimal fluid) from a subject at a first timepoint; wherein the first time point occurs following treatment for a cancer; b)isolating cell free (cf) deoxyribonucleic acid (DNA)(cfDNA) from the first sample; c) measuring regions of chromosomal heterozygosity of polymorphic sites in the first sample (including, but not limited to measuring by next generation sequencing (NGS), allelic-specific hybridization, primer extension, oligonucleotide ligation, and / or invasive cleavage) to obtain an allele ratio thereby creating an internal reference measurement; d) obtaining a second sample (e.g., a liquid biopsy) from the subject at a second timepoint; wherein the second time point occurs after the first time point; e) isolating cell free cfDNA from the second sample; and f) measuring regions of chromosomal heterozygosity change of polymorphic sites in the second sample relative to the measurements obtain with the first sample; wherein when the first sample is obtained following treatment for a cancer, a change in heterozygosity of the cfDNA in the second sample away from the relative heterozygosity of the cfDNA in the first sample indicates the presence of contaminating circulating tumor (ct) DNA (ctDNA) and therefore the presence of a recurrent cancer and / or metastasis; and wherein no change in the heterozygosity of the cfDNA between the first and second samples indicates no recurrent cancer and / or metastasis. Exemplary systems and methods are described in PCT Application No. US2023 / 022361, titled "EARLY DETECTION OF CANCER USING ALLELE DRIFT AND CHROMOSOMAL INSTABILITY," filed on May 16, 2023, the content of which is hereby incorporated by reference in its entirety.
[0046] It should be understood that the above-described subject matter may also be implemented as a computer-controlled apparatus, a computer process, a computing system, or an article of manufacture, such as a computer-readable storage medium.
[0047] Other systems, methods, features and / or advantages will be or may become apparent to one with skill in the art upon examination of the following drawings and detailed description. It is intended that all such additional systems, methods, features and / or advantages be included within this description and be protected by the accompanying claims.BRIEF DESCRIPTION OF THE DRAWINGS
[0048] The components in the drawings are not necessarily to scale relative to each other. Like reference numerals designate corresponding parts throughout the several views.
[0049] FIGURE 1 is a flowchart of an example computer-implemented method for detecting a presence of a solid tumor cancer or neoplasm in a subject's blood in accordance with certain embodiments described herein.
[0050] FIGURE 2A, FIGURE 2B, and FIGURE 2C are a flowchart diagram depicting an example method for detecting a presence of a solid tumor cancer or neoplasm in a subject's blood in accordance with certain embodiments described herein.
[0051] FIGURE 3A shows a representation of 27 base pairs of DNA from white blood cells for both the maternal and paternal chromosomes as represented by the positive strand, of a double stranded helix. Of the 27 base pairs of DNA a total of 5 are known polymorphic sites, i.e. single nucleotide polymorphisms (SNPs).
[0052] FIGURE 3B shows the theoretical results of a A / B allele ratio for the 5 are known potential polymorphic loci previously illustrated in figure 3A for cell free DNA (cfDNA) with no contamination by neoplastic cell DNA (ctDNA).
[0053] FIGURE 3C shows the measured results of a A / B allele ratio for the 5 known potential polymorphic loci previously illustrated in Figure 3A for cell free DNA (cfDNA) with no contamination by neoplastic cell DNA (ctDNA).
[0054] FIGURE 3D shows the results of a A / B allele ratio for the 5 known potential polymorphic loci previously illustrated in Figure 3A for cell free DNA (cfDNA) with no contamination by neoplastic cell DNA (ctDNA).
[0055] FIGURE 3E shows a representation of neoplastic tissue that has lost one of two chromosomes.
[0056] FIGURE 3F shows the results of a A / B allele ratio for the 5 known potential polymorphic loci previously illustrated in Figure 3A, but in neoplastic tissue that has lost one of two chromosomes as illustrated in Figure 3E, i.e., for cell free DNA (cfDNA) with 100% neoplastic tissue with "loss of heterozygosity" (LOH)
[0057] FIGURE 3G shows the results of a A / B allele ratio for the 5 known potential polymorphic loci previously illustrated in Figure 3A, but in neoplastic tissue that has lost one of two chromosomes as illustrated in Figure 3E, i.e., cell free DNA (cfDNA) contaminated with neoplastic cell DNA (ctDNA) that have lost either the maternal or paternal chromosome (i.e. LOH).
[0058] FIGURE 3H shows a representation of where the paternal copy of one chromosome is lost.
[0059] FIGURE 31 shows the results of a A / B allele ratio for the 5 known potential polymorphic loci previously illustrated in Figure 3A for cell free DNA (cfDNA) contaminated with ahigh number of neoplastic cell DNA (ctDNA) with known allelic imbalance for a specific region of a chromosome in a patient undergoing treatment and being responsive to treatment.
[0060] FIGURE 3J shows the results of a A / B allele ratio for the 5 known potential polymorphic loci previously illustrated in Figure 3A for cell free DNA (cfDNA) contaminated with a high number of neoplastic cell DNA (ctDNA) with known allelic imbalance for a specific region of a chromosome in a patient undergoing treatment, but nonresponsive to treatment.
[0061] FIGURE 3K shows the results of a A / B allele ratio for the 5 known potential polymorphic loci previously illustrated in Figure 3A for cell free DNA (cfDNA) contaminated with a very low number of neoplastic cell DNA (ctDNA) with known allelic imbalance for a specific region of a chromosome in a patient post-treatment with significant response to treatment with a follow-up for minimal residual disease (MRD) in a patient with recurrence.
[0062] FIGURE 3L shows the results of a A / B allele ratio for the 5 known potential polymorphic loci previously illustrated in Figure 3A for cell free DNA (cfDNA) contaminated with a very low number of neoplastic cell DNA (ctDNA) with known allelic imbalance for a specific region of a chromosome in a patient post-treatment with significant response to treatment with a follow-up for minimal residual disease (MRD) in a patient with without recurrence.
[0063] FIGURE 4A, FIGURE 4B, FIGURE 4C, FIGURE 4D, and FIGURE 4E depict representative data for detecting chromosomal instability (CIN) on chromosome 3 and identifying concordance between tumor and cfDNA samples from a study that was conducted.
[0064] FIGURE 5 is an example computing device.
[0065] FIGURE 6 is an example system.DETAILED DESCRIPTION
[0066] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. Methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present disclosure.
[0067] The following description of the disclosure is provided as an enabling teaching of the disclosure in its best, currently known embodiment(s). To this end, those skilled in the relevant art will recognize and appreciate that many changes can be made to the various embodiments of the present disclosure, while still obtaining the beneficial results of the present disclosure. It will also be apparent that some of the desired benefits of the present disclosure can be obtained by selecting some of the features of the present disclosure without utilizing other features. Accordingly, those who work in the art will recognize that many modifications and adaptations to the presentdisclosure are possible and can even be desirable in certain circumstances and are a part of the present disclosure. Thus, the following description is provided as illustrative of the principles of the present disclosure and not in limitation thereof.
[0068] Reference will now be made in detail to the embodiments of the present disclosure, examples of which are illustrated in the drawings and the examples. The present disclosure may, however, be embodied in many different forms and should not be construed as limited to the embodiments set forth herein.
[0069] Terminology
[0070] The term "comprising" and variations thereof as used herein is used synonymously with the term "including" and variations thereof and are open, non-limiting terms. Although the terms "comprising" and "including" have been used herein to describe various embodiments, the terms "consisting essentially of" and "consisting of" can be used in place of "comprising" and "including" to provide for more specific embodiments and are also disclosed. As used in this disclosure and in the appended claims, the singular forms "a", "an", "the", include plural referents unless the context clearly dictates otherwise.
[0071] The following definitions are provided for the full understanding of terms used in this specification.
[0072] The terms "about" and "approximately" are defined as being "close to" as understood by one of ordinary skill in the art. In one non-limiting embodiment the terms are defined to be within 10%. In another non-limiting embodiment, the terms are defined to be within 5%. In still another non-limiting embodiment, the terms are defined to be within 1%.
[0073] As used herein, the terms "may," "optionally," and "may optionally" are used interchangeably and are meant to include cases in which the condition occurs as well as cases in which the condition does not occur. Thus, for example, the statement that a formulation "may include an excipient" is meant to include cases in which the formulation includes an excipient as well as cases in which the formulation does not include an excipient.
[0074] "Composition" refers to any agent that has a beneficial biological effect. Beneficial biological effects include both therapeutic effects, e.g., treatment of a disorder or other undesirable physiological condition, and prophylactic effects, e.g., prevention of a disorder or other undesirable physiological condition. The terms also encompass pharmaceutically acceptable, pharmacologically active derivatives of beneficial agents specifically mentioned herein, including, but not limited to, a vector, polynucleotide, cells, salts, esters, amides, proagents, active metabolites, isomers, fragments, analogs, and the like. When the term "composition" is used, then, or when a particular composition is specifically identified, it is to be understood that the termincludes the composition per se as well as pharmaceutically acceptable, pharmacologically active vector, polynucleotide, salts, esters, amides, proagents, conjugates, active metabolites, isomers, fragments, analogs, etc.
[0075] The term "subject" refers to any individual who is the target of administration or treatment. The subject can be a vertebrate, for example, a mammal. In one embodiment, the subject can be human, non-human primate, bovine, equine, porcine, canine, or feline. The subject can also be a guinea pig, rat, hamster, rabbit, mouse, or mole. Thus, the subject can be a human or veterinary patient. The term "patient" refers to a subject under the treatment of a clinician, e.g., physician.
[0076] A "control" is an alternative subject or sample used in an experiment for comparison purposes. A control can be "positive" or "negative."
[0077] The term "detect" or "detecting" refers to an output signal released for the purpose of sensing of physical phenomenon. An event or change in environment is sensed and signal output released in the form of light, heat, or a color change (i.e., color change from red to blue, white to black, or vice versa).
[0078] The term "cancer" is used to address any neoplastic disease. Specifically, it is used here to describe hematological malignancies, including lymphomas, leukemias and myelomas.
[0079] The term "kit" describes a wide variety of bags, containers, carrying cases, and other portable enclosures which may be used to carry and store solid substances, liquid substances, and other accessories necessary to perform a next generation sequencing, deep sequencing, or high- throughput sequencing / screening assays. Such kits and their contents along with any applicable procedures may be used to provide access to variations detected in the hematologic neoplasms in accordance with the teachings of the present disclosure.
[0080] A "gene" refers to a polynucleotide containing at least one open reading frame that is capable of encoding a particular polypeptide or protein after being transcribed and translated. Any of the polynucleotide sequences described herein may be used to identify larger fragments or full-length coding sequences of the gene with which they are associated. A "gene panel" refers to a compilation, list, or grouping of genes for testing of variations, mutations, and changes in one or more genes.
[0081] "Profiling" refers to a collection of information relating to a specific gene, protein, disease, or disorder. Profiles can describe the risk of disease, preventive care assessments and utilization, and disease prevalence.
[0082] A "nucleic acid" is a chemical compound that serves as the primary informationcarrying molecules in cells and make up the cellular genetic material. Nucleic acids comprisenucleotides, which are the monomers made of a 5-carbon sugar (usually ribose or deoxyribose), a phosphate group, and a nitrogenous base. A nucleic acid can also be a deoxyribonucleic acid (DNA) or a ribonucleic acid ( R N A) . A chimeric nucleic acid comprises two or more of the same kind of nucleic acid fused together to form one compound comprising genetic material.
[0083] A "variant," "mutant," or "derivative" of a particular nucleic acid sequence may be defined as a nucleic acid sequence having at least 50% sequence identity to the particular nucleic acid sequence over a certain length of one of the nucleic acid sequences using blastn with the "BLAST 2 Sequences" tool available at the National Center for Biotechnology Information's website. (See Tatiana A. Tatusova, Thomas L. Madden (1999), "Blast 2 sequences— a new tool for comparing protein and nucleotide sequences", FEMS Microbiol Lett. 174:247-250). In some embodiments a variant polynucleotide may show, for example, at least 60%, at least 70%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% or greater sequence identity over a certain defined length relative to a reference polynucleotide.
[0084] As used herein, a "mutation" refers to changing the structure of a gene, resulting in a variant form that may be transmitted to later generations. A mutation is caused by the alteration of single nucleotides in DNA, or the deletion, insertion, or rearrangement of larger sections of genes. A mutation can lead to the expression of a protein that has been changed physically or functionally leading to lethality, non-lethal dysfunction effects, or no effects.
[0085] An "assay standard", "standard", or a "verification sample" refers to any molecule, compound, or composition of known quantity (concentration, volume, mass, etc.) used to determine the quantity of an unknown molecule, compound, or composition. A standard is usually used in an assay to quantify a final product of the assay.
[0086] Ranges can be expressed herein as from "about" one particular value, and / or to "about" another particular value. When such a range is expressed, another embodiment includes from the one particular value and / or to the other particular value. Similarly, when values are expressed as approximations, by use of the antecedent "about," it will be understood that the particular value forms another embodiment. It will be further understood that the endpoints of each of the ranges are significant both in relation to the other endpoint, and independently of the other endpoint. It is also understood that there are a number of values disclosed herein, and that each value is also herein disclosed as "about" that particular value in addition to the value itself. For example, if the value "10" is disclosed, then "about 10" is also disclosed. It is also understood that when a value is disclosed that "less than or equal to" the value, "greater than or equal to the value" and possible ranges between values are also disclosed, as appropriately understood by the skilledartisan. For example, if the value "10" is disclosed the "less than or equal to 10"as well as "greater than or equal to 10" is also disclosed. It is also understood that the throughout the application, data is provided in a number of different formats, and that this data represents endpoints, starting points, and ranges for any combination of the data points. For example, if a particular data point "10" and a particular data point "15" are disclosed, it is understood that greater than, greater than or equal to, less than, less than or equal to, and equal to 10 and 15 are considered disclosed as well as between 10 and 15. It is also understood that each unit between two particular units are also disclosed. For example, if 10 and 15 are disclosed, then 11, 12, 13, and 14 are also disclosed.
[0087] An "increase" can refer to any change that results in a larger amount of a symptom, disease, composition, condition, or activity. A substance is also understood to increase the genetic output of a gene when the genetic output of the gene product with the substance is greater relative to the output of the gene product without the substance. Also for example, an increase can be a change in the symptoms of a disorder such that the symptoms are more than previously observed. An increase can be any individual, median, or average increase in a condition, symptom, activity, composition in a statistically significant amount. Thus, the increase can be a 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100% increase so long as the increase is statistically significant or observable.
[0088] A "decrease" can refer to any change that results in a smaller amount of a symptom, disease, composition, condition, or activity. A substance is also understood to decrease the genetic output of a gene when the genetic output of the gene product with the substance is less relative to the output of the gene product without the substance. Also for example, a decrease can be a change in the symptoms of a disorder such that the symptoms are less than previously observed. A decrease can be any individual, median, or average decrease in a condition, symptom, activity, composition in a statistically significant amount. Thus, the decrease can be a 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100% decrease so long as the decrease is statistically significant or observable.
[0089] "Inhibit," "inhibiting," and "inhibition" mean to decrease an activity, response, condition, disease, or other biological parameter. This can include but is not limited to the complete ablation of the activity, response, condition, or disease. This may also include, for example, a 10% reduction in the activity, response, condition, or disease as compared to the native or control level. Thus, the reduction can be a 10, 20, 30, 40, 50, 60, 70, 80, 90, 100%, or any amount of reduction in between as compared to native or control levels.
[0090] By "reduce" or other forms of the word, such as "reducing" or "reduction," is meant lowering of an event or characteristic (e.g., tumor growth). It is understood that this istypically in relation to some standard or expected value, in other words it is relative, but that it is not always necessary for the standard or relative value to be referred to. For example, "reduces tumor growth" means reducing the rate of growth of a tumor relative to a standard or a control.
[0091] By "prevent" or other forms of the word, such as "preventing" or "prevention," is meant to stop a particular event or characteristic, to stabilize or delay the development or progression of a particular event or characteristic, or to minimize the chances that a particular event or characteristic will occur. Prevent does not require comparison to a control as it is typically more absolute than, for example, reduce. As used herein, something could be reduced but not prevented, but something that is reduced could also be prevented. Likewise, something could be prevented but not reduced, but something that is prevented could also be reduced. It is understood that where reduce or prevent are used, unless specifically indicated otherwise, the use of the other word is also expressly disclosed.
[0092] "Biocompatible" generally refers to a material and any metabolites or degradation products thereof that are generally non-toxic to the recipient and do not cause significant adverse effects to the subject.
[0093] "Comprising" is intended to mean that the compositions, methods, etc. include the recited elements, but do not exclude others. "Consisting essentially of" when used to define compositions and methods, shall mean including the recited elements, but excluding other elements of any essential significance to the combination. Thus, a composition consisting essentially of the elements as defined herein would not exclude trace contaminants from the isolation and purification method and pharmaceutically acceptable carriers, such as phosphate buffered saline, preservatives, and the like. "Consisting of" shall mean excluding more than trace elements of other ingredients and substantial method steps for administering the compositions provided and / or claimed in this disclosure. Embodiments defined by each of these transition terms are within the scope of this disclosure.
[0094] "Effective amount" of an agent refers to a sufficient amount of an agent to provide a desired effect. The amount of agent that is "effective" will vary from subject to subject, depending on many factors such as the age and general condition of the subject, the particular agent or agents, and the like. Thus, it is not always possible to specify a quantified "effective amount." However, an appropriate "effective amount" in any subject case may be determined by one of ordinary skill in the art using routine experimentation. Also, as used herein, and unless specifically stated otherwise, an "effective amount" of an agent can also refer to an amount covering both therapeutically effective amounts and prophylactica lly effective amounts. An "effective amount" of an agent necessary to achieve a therapeutic effect may vary according to factors such as the age,sex, and weight of the subject. Dosage regimens can be adjusted to provide the optimum therapeutic response. For example, several divided doses may be administered daily or the dose may be proportionally reduced as indicated by the exigencies of the therapeutic situation.
[0095] A "pharmaceutically acceptable" component can refer to a component that is not biologically or otherwise undesirable, i.e., the component may be incorporated into a pharmaceutical formulation provided by the disclosure and administered to a subject as described herein without causing significant undesirable biological effects or interacting in a deleterious manner with any of the other components of the formulation in which it is contained. When used in reference to administration to a human, the term generally implies the component has met the required standards of toxicological and manufacturing testing or that it is included on the Inactive Ingredient Guide prepared by the U.S. Food and Drug Administration.
[0096] "Pharmaceutically acceptable carrier" (sometimes referred to as a "carrier") means a carrier or excipient that is useful in preparing a pharmaceutical or therapeutic composition that is generally safe and non-toxic and includes a carrier that is acceptable for veterinary and / or human pharmaceutical or therapeutic use. The terms "carrier" or "pharmaceutically acceptable carrier" can include, but are not limited to, phosphate buffered saline solution, water, emulsions (such as an oil / water or water / oil emulsion) and / or various types of wetting agents. As used herein, the term "carrier" encompasses, but is not limited to, any excipient, diluent, filler, salt, buffer, stabilizer, solubilizer, lipid, stabilizer, or other material well known in the art for use in pharmaceutical formulations and as described further herein.
[0097] "Pharmacologically active" (or simply "active"), as in a "pharmacologically active" derivative or analog, can refer to a derivative or analog (e.g., a salt, ester, amide, conjugate, metabolite, isomer, fragment, etc.) having the same type of pharmacological activity as the parent compound and approximately equivalent in degree.
[0098] "Therapeutic agent" refers to any composition that has a beneficial biological effect. Beneficial biological effects include both therapeutic effects, e.g., treatment of a disorder or other undesirable physiological condition, and prophylactic effects, e.g., prevention of a disorder or other undesirable physiological condition (e.g., a non-immunogenic cancer). The terms also encompass pharmaceutically acceptable, pharmacologically active derivatives of beneficial agents specifically mentioned herein, including, but not limited to, salts, esters, amides, proagents, active metabolites, isomers, fragments, analogs, and the like. When the terms "therapeutic agent" is used, then, or when a particular agent is specifically identified, it is to be understood that the term includes the agent per se as well as pharmaceutically acceptable, pharmacologically active salts, esters, amides, proagents, conjugates, active metabolites, isomers, fragments, analogs, etc.
[0099] "Therapeutically effective amount" or "therapeutically effective dose" of a composition (e.g. a composition comprising an agent) refers to an amount that is effective to achieve a desired therapeutic result. Therapeutically effective amounts of a given therapeutic agent will typically vary with respect to factors such as the type and severity of the disorder or disease being treated and the age, gender, and weight of the subject. The term can also refer to an amount of a therapeutic agent, or a rate of delivery of a therapeutic agent (e.g., amount over time), effective to facilitate a desired therapeutic effect. The precise desired therapeutic effect will vary according to the condition to be treated, the tolerance of the subject, the agent and / or agent formulation to be administered (e.g., the potency of the therapeutic agent, the concentration of agent in the formulation, and the like), and a variety of other factors that are appreciated by those of ordinary skill in the art. In some instances, a desired biological or medical response is achieved following administration of multiple dosages of the composition to the subject over a period of days, weeks, or years.
[0100] The terms "treat," "treating," "treatment," and grammatical variations thereof as used herein, refer to the medical management of a patient with the intent to cure, ameliorate, stabilize, or prevent a disease, pathological condition, or disorder. This term includes active treatment, that is, treatment directed specifically toward the improvement of a disease, pathological condition, or disorder, and also includes causal treatment, that is, treatment directed toward removal of the cause of the associated disease, pathological condition, or disorder. In addition, this term includes palliative treatment, that is, treatment designed for the relief of symptoms rather than the curing of the disease, pathological condition, or disorder; preventative treatment, that is, treatment directed to minimizing or partially or completely inhibiting the development of the associated disease, pathological condition, or disorder; and supportive treatment, that is, treatment employed to supplement another specific therapy directed toward the improvement of the associated disease, pathological condition, or disorder. Treatments according to the present disclosure may be applied preventively, prophylactica lly, pa lliatively or remedially. Prophylactic treatments are administered to a subject prior to onset (e.g., before obvious signs of cancer), during early onset (e.g., upon initial signs and symptoms of cancer), or after an established development of cancer. Prophylactic administration can occur for day(s) to years prior to the manifestation of symptoms of an infection.
[0101] The terms "Next Generation Sequencing (NGS) Assays" or "next generation sequencing (NGS)" refer to a massive parallel sequencing technique wherein isolated DNA or RNA samples are amplified, sequenced, and analyzed for complete genomic or transcriptomic overview. This method of sequencing requires smaller sample sizes, less reagents and time, and lower reagentcosts while providing whole genome data from a single subject. NGS is described in detail in PCT / US2023 / 011770, filed January 27, 2023, the disclosure of which is expressly incorporated herein by reference in its entirety.
[0102] In some embodiments, a duplication comprises repeating a gene at least one time. In some embodiments, a duplication comprises repeating a gene 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more.
[0103] A "single nucleotide variant", "SNV", or "single nucleotide polymorphism" refers to a nucleic acid sequence variation wherein a single nucleotide (adenine, thymine, cytosine, or guanine) in the genome sequence is altered into another nucleotide. In some embodiments, the nucleic acid sequence comprises a single nucleotide variant.
[0104] A "structural variant" refers to a nucleic acid, usually 1000 base pairs, or 1 kilobases (kb), in size, wherein the overall chromosomal structure is altered. "Structural variations" can further comprise other variations, including, but not limited to deletions, insertions, indels, copy number variations, duplications, inversion, substitutions, and derivatives thereof. In some embodiments, the nucleic acid comprises a structural variant.
[0105] The term "chromosomal instability" refers to a condition in which a cell of a subject with cancer loses or gains an entire or partial chromosome to aid in cancer cell survival. Chromosomal instability also refers to ongoing chromosome segregation errors throughout consecutive cell divisions that may result in various pathological phenotypes including copy number gains or losses (e.g., altered genome), increased copy number heterogeneity, and increased presence of micronuclei. When this occurs, heterozygous SNPs in this area of chromosomal loss or gain, so-called chromosomal instability (CIN), then revert to a state of homozygosity. Just as the presence of SNPs with two matched but different base pairs at given location can be used as a fingerprint in forensics, the comparison of SNPs in matched normal and tumor tissue can be used in a similar manner to investigate the presence or absence of cancer in cfDNA. As change in homozygosity in tumors almost always involved hundreds of thousands to millions of consecutive base pairs of DNA, there can be many unique heterozygous SNPs in that same area that now revert to a homozygous state. These SNPs affected by loss or gain of a specific piece of a chromosome are physically linked and the result is multiple alleles or loci changing from a heterozygous to homozygous state as one event that is referred to herein as "allele drift". In a comparable fashion complete or partial loss of a chromosome in a tumor can result in a copy number change that may be referred to as "copy number drift". Tumors can lose a chromosome and then duplicate the remaining copy. In these instances, there will be allele drift and is why the former is a preferred method of analysis for response to therapy or MRD. Elevated chromosomal instability, resulting invarious instances of chromosomal instability pathology, has been correlated with poor prognosis in several cancers. Thus, the present application provides machine-learning models to analyze biological samples to characterize chromosomal instability in a patient sample (e.g., a blood, plasma or FFPE sample) using high throughput machine-learning models.
[0106] A "change in zygosity" or "CIZ" refers to nucleic acid sequence wherein the number of heterozygous copies of a specific DNA segment within a genome is much less than expected. A "CIZ" can involve the deletion (loss) of genetic material or the duplication (gain) of additional copies, or no change at all in the copies of genetic material. A "CIZ" is a specific type of "genomic variant". In some embodiments, the nucleic acid sequence comprises a CIZ .
[0107] In some embodiments, a CIZ is a chromosomal variant involving a chromosomal arm CIZ, or chromosomal CIZ. Chromosomal CIZ refers to change in zygosity of an entire chromosomal arm or a significant portion of it. Each chromosome has two arms, referred to as the "p" arm (short arm) and the "q" arm (long arm), separated by the centromere. Chromosomal arm CIZ can occur in both somatic cells (non-germline cells) and germline cells (cells that give rise to eggs or sperm). The CIZ of a chromosomal arm is typically caused by genetic events such as chromosomal deletions, translocations, or large-scale structural rearrangements. These events can lead to the loss of diversity of genetic material and disrupt the normal balance of gene dosage, potentially affecting the function and regulation of genes located on the affected arm. Chromosomal arm CIZ can have significant consequences on an individual's health and development, depending on the specific chromosome involved and the genes located on the affected arm. The loss of essential genes or regulatory elements can result in developmental abnormalities, intellectual disabilities, and an increased risk of certain genetic disorders or diseases.
[0108] Method of detecting and treating cancer
[0109] There are currently many methods of characterizing response to one or more therapies in patients with clinically apparent cancer, or recurrence of disease in treated cancer patients with no clinically apparent cancer, i.e., minimal residual disease (MRD). Most recently, the technological advance of next generation sequencing (NGS) has created a prior state of the art in this field. The prior state of the art using NGS requires (i) obtaining a sample, such as but not limited to a tumor sample or liquid biopsy sample from the patient, (ii) measuring DNA mutations at nonpolymorphic loci in the sample, (iii) obtaining future blood samples from the patient during treatment for evaluating response to therapy, or post-treatment for evaluating recurrence of tumor after treatment (iv) isolating cell free DNA (cfDNA) from blood samples (liquid biopsy) that likely contain some level of circulating tumor cell free DNA (catena), and (v) predicting from these cfDNAsamples, based upon mutational analysis of nonpolymorphic loci in DNA coding regions, response of the tumor to treatment or recurrence of tumor after treatment.
[0110] However, response to therapy or MRD using this prior art of measuring DNA mutations at nonpolymorphic loci in a liquid biopsy lacks sensitivity and specificity. Lack of specificity is primarily due to inherent errors in NGS that impact the probability of a true event at a single nonpolymorphic loci. Although prior methodologies have improved errors in NGS, the prior art of measuring DNA mutations at nonpolymorphic loci still has significant problems with the probability of a true event at low levels of ctDNA. The lack of sensitivity of the prior art of analysis of nonpolymorphic loci to detect response to therapy or MRD relies on the probability of distinguishing multiple mutational independent events from sequencing errors from true changes whereby ctDNA is present in cfDNA. These multiple mutational events have no outcome in common and are called disjoint in regard to the probability of detecting a true event. For the prior art the basic rule of probability is if two events A and B are disjoint, then the probability of either event is the sum of the probabilities of the two events: P(A or B) = P(A) + P(B). The result of this regarding sensitivity is that up to one-fourth of all true events are not detected. Regarding specificity, the prior art of measuring DNA mutations at nonpolymorphic loci in a liquid biopsy is complicated by clonal hematopoiesis. The latter is the presence of mutations in cells in the blood (i.e., white blood cells) that release DNA into the cfDNA component of a liquid biopsy. Clonal hematopoiesis occurs in a large percentage of older patients with no clinical evidence of cancer. Although prior methodologies have improved the distinction of mutations in cfDNA due to clonal hematopoiesis (false positives) versus true positives of tumor recurrence, the results are less than optimal with up to 50% of false positive calls.
[0111] Embodiments of the present disclosure are directed to a computer- implemented method for characterizing response to therapy and MRD based upon Al and CIN . All human cells contain 46 chromosomes, 23 of maternal and 23 of paternal origin, referred to as autosomal chromosomes 1 through 22 and the sex chromosomes X or Y. When an entire or partial area of a chromosome is lost or gained through cell replication or other processes the loss of the contribution of one parent results in a change of allele balance.
[0112] For example, it is expected that a patient's cancer can be detected using cfDNA analysis (liquid biopsy) techniques described herein. Subsequent to treatment(s), the cfDNA of the patient relative to cfDNA is expected to decline if the treatment(s) are successful. Accordingly, embodiments of the present disclosure facilitate obtaining multiple measurements for a given patient over the course of multiple treatment events for specific therapies to measure the patient's response to treatment or lack thereof based, at least in part, on a relative decrease or increase, respectively of cfDNA.
[0113] Chromosomes are made up of deoxyribonucleic acid (DNA) and histone proteins with the former being that component that is transmitted from parent to child with rigor. DNA is made up four basic building blocks of adenine (A), cytosine (C), guanine (G), and thymine (T) for which there are approximately 3 billion such blocks across all 46 chromosomes. Since DNA is double stranded with a positive and negative strand there are always 2 of these blocks at each of these 3 billion positions that is referred to as base pairs. Close to 99% of these base pair positions are identical for all humans, referred to as nonpolymorphic loci or alleles, and the remaining 1% of positions that are different in greater than 95% of the population are called polymorphisms, or more specifically single nucleotide polymorphisms (SNPs). There are about 14 million SNPs in human DNA, or roughly one every 200 base pairs across the total 3 billion base pair positions. There is a great deal of variability from one person to the next for specific SNPs in their DNA and can be used as a molecular fingerprint in forensics for the investigation of crimes and other applications. As prior technologies only refer to the plus strand of DNA, for which everyone has two copies of for each chromosome, there are n to the fourth combinations of A, C, G, and T at each polymorphic loci. When both plus strands contain the same DNA molecule such as A and A, C and C, G and G, or T and T, that condition is referred to as homozygous. When both plus strands contain a different DNA molecule such as A and T, A and C, A and G, A and A, C and T, etc. then that condition is referred to as heterozygous with the more common nucleotide in the general population referred to as the "A allele" and the less common as the "B allele". On any positive strand of DNA with multiple linked and aligned heterozygous loci there can be any combination of "A" and "B" alleles and with the opposite designation on the other minus strand.
[0114] As discussed above, cancers (e.g., cancerous cells) often lose or gain an entire or partial area of a chromosome to be more efficient in cell replication. When this occurs heterozygous SNPs in this area of chromosomal loss or gain, so-called chromosomal instability (CIN), then revert to a state of homozygosity. Just as the presence of SNPs with two matched but different base pairs at a given location can be used as a fingerprint in forensics, the comparison of SNPs in matched normal and tumor tissue can be used in a similar manner to investigate the presence or absence of cancer in cfDNA. As change in homozygosity in tumors almost always involved hundreds of thousands to millions of consecutive base pairs of DNA, there can be many unique heterozygous SNPs in that same area that now revert to a homozygous state. These SNPs affected by loss or gain of a specific piece of a chromosome are physically linked and the result is multiple alleles or loci changing from a heterozygous to homozygous state as one event that is referred to herein as "allele drift". In a comparable fashion complete or partial loss of a maternal or paternal chromosome in a tumor can result in a copy number change that is referred to herein as "copy number drift". Astumors can lose either a maternal or paternal chromosome and then duplicate the remaining copy there are instances where there is copy number change, but there is no copy number drift. In these instances, there will be allele drift when there is no copy number drift and is why the former is a preferred method of analysis for response to therapy or MRD.
[0115] As compared to the prior art of looking for change in disjointed nonpolymorphic loci as the sum of independent events to detect cancer in a liquid biopsy, embodiments of the present disclosure allow for the identification of change in multiple joined polymorphic loci or regions of allele drift. When two events are not independent but rather joined as in heterozygous loci in a region of change the result is a conditional probability. If change in one heterozygous loci is detected in a region of change, then it is more likely that change will be detected at a second joined loci, etc. Since events A and B are not independent, then the probability of the intersection of A and B (the probability that both events occur) is defined by P(A and B) = P(A)P(B | A). Given that any region of change will have many joined heterozygous loci the conditional probabilities of any one loci is dependent on the consideration of the detection of change at any preceding loci. As compared to the disjointed probability of detecting change in the prior art, with this methodology in this embodiment the likelihood of detecting contamination in cfDNA with neoplastic cells (ctDNA) is based upon the conditional probability of detecting change at multiple adjacent heterozygous loci in a region of known change, i.e., allele drift.
[0116] In one embodiment, a computer-implemented method for detecting a presence of a solid tumor or neoplasm in a blood sample is provided. The computer-implemented method can comprise: receiving Deoxyribonucleic acid (DNA) sequencing data for a subject's blood sample; identifying a plurality of chromosomal variations in the DNA sequencing data, wherein the plurality of chromosomal variations comprises a plurality of chromosomal instabilities (CINs); and detecting the solid tumor or neoplasm based, at least in part, on analysis of the plurality of chromosomal variations.
[0117] In one embodiment, disclosed herein are computer-implemented methods for detecting a cancer, cancer recurrence, and / or metastasis (including, but not limited to cancer recurrence in a subject previously treated for cancer with no clinically apparent cancer, i.e., MRD) in a subject comprising: a) obtaining a first sample (such as, for example a liquid biopsy including, but not limited to a liquid biopsy comprising whole blood, peripheral blood, plasma, serum, saliva, sputum, cerebral spinal fluid, urine, lymph, or lacrimal fluid) from a subject at a first timepoint; wherein the first time point occurs following treatment for a cancer; b) isolating cell free (cf) deoxyribonucleic acid (DNA)(cfDNA) from the first tissue sample; c) measuring regions of chromosomal heterozygosity of polymorphic sites in the first sample (including, but not limited tomeasuring by next generation sequencing ( NGS), allelic-specific hybridization, primer extension, oligonucleotide ligation, and / or invasive cleavage) to obtain an allele ratio thereby creating an internal reference measurement; d) obtaining a second sample (e.g., liquid biopsy) from the subject at a second timepoint; wherein the second time point occurs after the first time point; e) isolating cell free cfDNA from the second sample; and f) measuring regions of chromosomal heterozygosity change of polymorphic sites in the second sample relative to the measurements obtain with the first sample; wherein when the first sample is obtained following treatment for a cancer, a change in heterozygosity of the cfDNA in the second sample away from the relative heterozygosity of the cfDNA in the first sample indicates the presence of contaminating circulating tumor (ct) DNA (ctDNA) and therefore the presence of a recurrent cancer and / or metastasis; and wherein no change in the heterozygosity of the cfDNA between the first and second samples indicates no recurrent cancer and / or metastasis.
[0118] The disclosed methods for detecting a cancer, cancer recurrence, or metastasis (including, but not limited to cancer recurrence in a subject previously treated for cancer with no clinically apparent cancer, i.e., MRD) were developed in recognition that circulating tumor DNA can be present in cfDNA samples where cancer remains, a recurrent cancer is developing or has formed, and / or a metastasis has occurred. Thus, in one embodiment, disclosed herein are methods for detecting a cancer, wherein the cfDNA in the first or second sample comprises circulating tumor (ct) DNA (ctDNA).
[0119] Detecting and / or measuring homozygosity / heterozygosity via allelic imbalance, and / or chromosomal instability at polymorphic sites in the tissue samples can be obtained by any means known in the art, including, but not limited to next generation sequencing (NGS), allelic-specific hybridization, primer extension, oligonucleotide ligation, and / or invasive cleavage.
[0120] A custom chromosomal instability (CIN) NGS caller was developed to identify concordance of tumor and cfDNA samples. Prior to analysis a quality control (QC) at the individual SNP level for each subject is performed. The total number of reads for all polymorphic (informative) SNPs across all somatic chromosomes in tumor and cfDNA must sum to 50 or greater to pass for further analysis. For those qualified SNPs the frequency of each qualified A and B allele for polymorphic (informative) SNPs is used to calculate the A / total allele ratio in an individual's peripheral blood mononuclear cell (PBMC) sample. Values for the PBMC Max Mapped Allele Count / Total Allele Count greater than 0.8 is considered homozygous:1) Max Mapped Allele Count / totalallelePBMc ^0.8 = polymorphic SNP
[0121] The absolute SNP difference for each polymorphic SNP using A / total allele ratio for PBMC to tumor is calculated and recorded as CIN SNP difference for each individual SNP:2) A:totalallelespBMc | - | A:totalallelesTuMOR = ASNP
[0122] Normal chromosomes are identified using a percentage of SNP differences that are greater than a threshold of 0.1 based upon a minimum of 20% neoplastic content for a presumed diploid state. Far extremes of total chromosomal allelic imbalance (Al) versus normal chromosome were identified using percent SNP differences for that chromosome. If 90% or greater of all SNP differences for any chromosome are greater than the threshold of 0.1 then this is considered as whole chromosome Al. Where n = polymorphic SNP and k = threshold A SNR:3) n>o.i^90° / ° ^wholeChrAi
[0123] If 10% or less of all SNP differences for any chromosome were not greater than 0.1 then this was considered as a normal chromosome with no Al:4) £ >0 i< 10% => NormChr
[0124] To identify partial chromosome Al, divide all SNPs into consecutive groups of 50 (i.e. SNP-group) for each autosomal chromosome. Among groups of 50 SNPs calculate percentage above or below baseline threshold of B allele frequency determined by healthy controls. If 70% of a group of 50 SNPs is above or below baseline threshold of B allele frequency determined by healthy controls, then group is considered candidate for partial chromosome Al. To be classified as partial chromosome Al two or more consecutive SNP-groups must be above or below baseline threshold of B allele frequency determined by healthy controls.
[0125] Further analysis at this point requires identification of breakpoints in chromosomes with partial Al and excludes those with whole chromosomal Al and normal chromosomes. Breakpoints are identified by finding 10 or more consecutive (by nucleotide position) polymorphic SNPs starting at the p-terminus with a SNP difference greater than 1.5x of SD of the absolute SNP differences for normal chromosomes for that case. Once a breakpoint is identified that defines the beginning of a region of Al for that specific chromosome. Termination of a region of Al requires the encounter of one absolute SNP difference less than 1.5x of SD of the absolute SNP differences for normal chromosomes for that case. Where n = polymorphic SNP and j is standard deviation of NormChr:5) PartialChrAi 2fc>i s*; GASNP n+l...n+10...nn<k=>RegionAi ,6) PartialChrAi - RegioriAi=> NOAI
[0126] After identifying whole and partial chromosomal changes allele with greater count is referred to as the retained allele ( RA) for each SNP and the lesser as the lost allele (LA). A tumor retained allele fraction (tRAF) is then calculated for each SNP by dividing the tumor retained allele count by the tumor lost allele count:7) tRAF = tRAC / tLAC
[0127] RA and LA designations from the tumor are then used to classify each SNP in the germline as retained or lost. A germline retained allele fraction (gRAF) is then calculated for each SNP by dividing the germline retained allele count (gRAC) by the germline lost allele count (gLAC):8) gRAF = gRAC / gLAC
[0128] A CIN score for the tumor (tCIN) is then calculated using the difference of sum of tRAF and gRAF in a region of CIN:9) tCIN score = sum(tRAF)-sum(gRAF)
[0129] RA and LA designations from the tumor are then used to classify each SNP in the cfDNA as retained or lost. A cfDNA retained allele fraction (cfRAF) is then calculated for each SNP by dividing the cfDNA retained allele count (cfRAC) by the cfDNA lost allele count (cfLAC):10) cfRAF = cfRAC / cfLAC
[0130] A CIN score for the subject cfDNA (cfCIN) is then calculated using the difference of sum of cfRAF and gRAF in a region of CIN in an individual with or prior evidence of cancer.11) cfCIN score = sum(cfRAF)-sum(gRAF)
[0131] A CIN score for the healthy control cfDNA (HCcfCIN) is then calculated using the difference of sum of cfRAF and gRAF in a region of CIN taken from a combined group of subjects with no evidence of cancer.12) HCcfCIN score = sum(HCcfRAF)-sum(HCgRAF)
[0132] Identification of circulating tumor DNA in subject cfDNA, or positive score, requires cfCIN to be greater than HCcfCIN.13) Positive ctCIN score = cfCIN > HCcfCIN
[0133] The systems and methods described herein address various challenges above. In particular, the systems and methods analyze DNA sequencing data for a subject in order to detect a presence or non-presence of a solid tumor or neoplasm. The DNA sequencing data encodes a genome wide single nucleotide polymorphism (SNP) backbone, or a plurality of polymorphisms in a genome wide single nucleotide polymorphism (SNP) backbone. The techniques described herein include identifying a plurality of chromosomal variations and / or CINs, particularly a plurality of different unique CINs, in the DNA sequencing data. At least one of the plurality of chromosomal instabilities can involve partial or whole instability of chromosome 1, chromosome 2, chromosome 3, chromosome 4, chromosome 5, chromosome 6, chromosome 7, chromosome 8, chromosome 9, chromosome 10, chromosome 11, chromosome 12, chromosome 13, chromosome 14, chromosome 15, chromosome 16, chromosome 17, chromosome 18, chromosome 19, chromosome 20, chromosome 21, chromosome 22, chromosome X, or chromosome Y. It should be understood that conventional technologies have yet to use CINs as a basis for detecting solid tumor or neoplasm. Optionally, the techniques described herein further include providing and / or modifying a diagnosis, prognosis, and / or treatment based on the detection.
[0134] FIG. 1 is a flowchart of an example computer-implemented method 100 for detecting a presence of a solid tumor or neoplasm in a subject's blood. In some implementations, the method 100 can be performed by a processing circuitry (for example, but not limited to, an application-specific integrated circuit (ASIC), or a central processing unit (CPU)). In some examples, the processing circuitry may be electrically coupled to and / or in electronic communication with other circuitries of an example computing device, such as, but not limited to, the example computing device 500 described above in connection with FIG. 5. In some examples, embodiments may take the form of a computer program product on a non-transitory computer-readable storage medium storing computer-readable program instruction (e.g., computer software). Any suitable computer-readable storage medium may be utilized, including non-transitory hard disks, CD-ROMs, flash memory, optical storage devices, or magnetic storage devices.
[0135] Referring now to FIG. 1, a flowchart illustrating example operations 100 for detecting a presence of a solid tumor or neoplasm in a subject's blood is shown. This disclosure contemplates that the example operations can be performed using one or more computing devices (e.g., at least the basic configuration illustrated in FIG. 5 by box 502). The method for detecting a presence of a solid tumor or neoplasm in a subject's blood according to the method of FIG. 1 includes identifying chromosomal variations. Conventional technologies do not use CINs to detect solid tumors or neoplasms in blood samples. Additionally, detection of solid tumor or neoplasm using the method shown in FIG. 1 can be completed more quickly (e.g., 1-3 days) than conventionaltechnologies, which typically have a turn-around-time of weeks for genomic and / or transcriptomic analysis.
[0136] At step 110, the method includes receiving DNA sequencing data for a subject. The DNA sequencing data is obtained from a sample, for example a liquid biopsy. Optionally, in some implementations, the liquid biopsy is a blood sample. The blood sample may include whole blood, peripheral blood (e.g., peripheral blood mononuclear cells (PBMCs), plasma, serum, or combinations thereof. As described herein, the DNA sequencing data includes a plurality of chromosomal variations, where the plurality of chromosomal variations include a plurality of different unique chromosomal instabilities (CINs), such as a partial chromosomal instability or a whole chromosomal instability. In some implementations, each of the plurality of chromosomal variations is a unique CIN. Additionally, each of the plurality of chromosomal variations can further comprise at least one of a plurality of candidate CINs for the solid tumor or neoplasm in the blood sample. In some implementations, step 110 includes the steps of (a) obtaining a blood sample from the subject; (b) extracting DNA from the blood sample, which includes isolating cell free (cf) deoxyribonucleic acid (DNAj(cfDNA) from the blood sample; and (c) sequencing the cfDNA to obtain the DNA sequencing data. As described herein, the blood sample may include peripheral blood mononuclear cells (PBMCs), plasma, or both PBMCs and plasma. In some implementations, the DNA is sequenced using a parallel processing operation, as discussed in more detail below in relation to FIG. 6.
[0137] In some implementations, the DNA sequencing data encodes a genome wide single nucleotide polymorphism (SNP) backbone. The term "single nucleotide polymorphism backbone" or "SNP backbone" refers to a core set of SNPs used as a reference or framework for genotyping or genetic analysis. SNPs are the most common type of genetic variation in the human genome, where a single nucleotide (A, T, C, or G) is altered at a specific position in the DNA sequence. SNPs can occur throughout the genome and are typically used as genetic markers to study genetic diversity, population genetics, and the association of specific variations with traits or diseases. The SNP backbone typically includes a representative set of SNPs that are informative for the population or specific research objectives. These SNPs are genotyped or analyzed to identify and compare genetic variations among individuals or populations. The selection of SNPs for the backbone often takes into consideration factors such as their frequency in the population, functional relevance, coverage of different genomic regions, and linkage disequilibrium patterns (correlation of variations at nearby positions). The chosen SNPs serve as representative markers to capture the genetic variation present in the studied population. As used herein, a genome wide SNP backbone includes a plurality (e.g., hundreds of thousands or millions) of SNPs that identify genetic variantsassociated with hematologic neoplasms. Alternatively, in some implementations, the DNA sequencing data encodes a plurality of polymorphisms in a genome wide single nucleotide polymorphism.
[0138] The DNA sequencing data can be obtained using a DNA sequencing device such as an NGS instrument that reads DNA strands and outputs an electrical waveform. A basecalling module (e.g., hardware, software, or combination thereof) of the NGS instrument transforms the electrical waveform into nucleobase sequences. In other words, the basecalling module receives the electrical waveform output by the sequencing device and outputs sequencing reads (A, T, C, G). The sequencing reads are optionally in text-based format such as FASTQ format. NGS systems are known in the art. For example, the iS EQ. 100 NGS sequencing system from Illumina, Inc. of San Diego, California is a sequencing platform that can be used with the systems and methods described herein. It should be understood that NGS systems are provided only as an example. This disclosure contemplates using other high-throughput sequencing systems including, but not limited to, third- generation sequencing systems.
[0139] DNA sequencing is a laboratory technique used to determine the precise order of nucleotides (A, T, C, and G) in a DNA molecule. To sequence the DNA in a tissue sample, one or more of the following steps are performed:
[0140] Sample collection: Collect a tissue sample from the subject. In some implementations described herein, the tissue sample is PBMC, Formalin-Fixed Paraffin-Embedded (FFPE) tissue, or plasma. Additionally, the subject is a mammal. In some embodiments, the subject is a human.
[0141] DNA extraction and purification: Isolate the DNA from the tissue sample. DNA extraction kits and protocols for isolating DNA are known in the art.
[0142] Library preparation: Prepare a DNA library for sequencing. This includes fragmenting the DNA into smaller pieces, adding adapters to the fragments, and amplifying them using polymerase chain reaction (PCR). The adapters allow the DNA fragments to be attached to a solid surface (e.g., a sequencing flow cell or microarray) and enable the sequencing reaction to occur.
[0143] Sequencing: Perform the DNA sequencing using a suitable sequencing platform or technology. As described herein, a high-throughput DNA sequencing system such as NGS device can be used. In some implementations, sequencing the DNA to obtain the DNA sequencing data includes extracting the DNA from at least one of the blood sample, blood mononuclear cells (PBMCs), plasma, or Formalin-Fixed Paraffin-Embedded (FFPE) tissue.
[0144] Data analysis-. Once the sequencing is complete, process and analyze the resulting raw data. This may include converting the raw sequencing data into a readable format, aligning the sequences to a reference genome (if available), identifying variants, and interpreting the genetic information obtained.
[0145] At step 120, the method includes identifying a plurality of chromosomal variations in the DNA sequencing data for the subject. In some implementations, one or more different unique CINs are identified. For example, one or more partial chromosomal instabilities and / or one or more whole chromosomal instabilities can be identified. The plurality of chromosomal instabilities can involve whole instability or partial instability of chromosome 1, chromosome 2, chromosome 3, chromosome 4, chromosome 5, chromosome 6, chromosome 7, chromosome 8, chromosome 9, chromosome 10, chromosome 11, chromosome 12, chromosome 13, chromosome 14, chromosome 15, chromosome 16, chromosome 17, chromosome 18, chromosome 19, chromosome 20, chromosome 21, chromosome 22, chromosome X, or chromosome Y.
[0146] At step 130, the method includes detecting a solid tumor or neoplasm based, at least in part, on analysis of the plurality of chromosomal variations. For example, the analysis can comprise comparing the DNA sequencing data to a healthy control DNA sequencing data set to identify one or more differences. The solid tumor or neoplasm is in one of a plurality of solid tumor primary sites including lung, trachea, breast, uterine, ovarian, skin, bladder, kidney, liver, pancreas, soft tissue, bone, mouth, salivary glands, bowel, rectum, gallbladder, urethra, fallopian tube, vulva, testicle, prostate, penis, thyroid, adrenal, parathyroid, heart, vessels, thymus, nerve, eye, ear, and subcutaneous tissue. In some implementations, analyzing the plurality of chromosomal variations comprises implementing a parallel processing operation, as discussed in more detail below in relation to FIG. 6.
[0147] Optionally, at step 140, the method includes determining the presence of absence of cancer for the subject in response to (or subsequent to) detecting the solid tumor or neoplasm. In some implementations, the method includes determining a prognosis for the subject and / or determining a response to treatment for the subject. Additionally, the method can include providing a determination of minimal or measurable residual disease for the subject or providing a treatment to the subject.
[0148] Optionally, in some implementations, the method further includes generating a report including the DNA sequencing data and the detected solid tumor or neoplasm in the subject, and / or the presence or absence of cancer for the subject. Optionally, the report is integrated into the subject's electronic health record (EHR). Alternatively or additionally, the method optionally further includes generating display data for the report. Alternatively or additionally, theT1method optionally further includes transmitting the report over a network. This disclosure contemplates that operations related to generation of the report can be performed using one or more computing devices (e.g., at least the basic configuration illustrated in FIG. 5 by box 502).
[0149] In some implementations, the method optionally further includes, in response to detecting the solid tumor or neoplasm in the subject, providing a diagnosis for the subject. The diagnosis can be the type of cancer and / or location of a solid tumor primary site (e.g., lung, trachea, breast). As described herein, the CINs described herein can be used to diagnosis the type of solid tumor cancer. In some embodiments, this disclosure contemplates that the CINs described herein are the only information used to make the diagnosis. Optionally, in other embodiments, the CINs described herein are used in combination with other test results (e.g., clinical evaluation, blood tests, bone marrow biopsy, imaging studies, etc.) to make the diagnosis. Additionally, the method optionally further includes, in response to detecting the solid tumor or neoplasm in the subject, providing a prognosis for the subject. Alternatively or additionally, the method optionally further includes recommending a treatment for the subject. Treatment approaches can vary depending on the specific subtype, stage, and patient factors but may include, but are not limited to, chemotherapy, radiation therapy, immunotherapy, targeted therapies, and stem cell transplantation. This disclosure contemplates that the operations related to providing diagnosis, prognosis, and / or treatment options can be performed using one or more computing devices (e.g., at least the basic configuration illustrated in FIG. 5 by box 502).
[0150] Optionally, in some implementations, the method further includes administering the recommended treatment to the subject.
[0151] FIG. 2A, FIG. 2B, and FIG. 2C are a flowchart depicting an example method for detecting a presence of a solid tumor or neoplasm in a subject's blood in accordance with certain embodiments described herein. The method can be at least partially implemented by a computing device 500 described in connection with FIG. 5 and / or the system 600 described in connection with FIG. 6.
[0152] With reference to FIG. 2A, at step 201A, the method includes identifying informative and non-informative SNPs from an individual's PBMC sample. For example, based on the following equation:Max Mapped Allele Count / totalallelepBMc ^0.8 = polymorphic SNP
[0153] At step 201B, the method includes removing the non-informative SNPs from analysis. For example, based on the following equation:Fraction [A / Total] > 0
[0154] At step 201C, the method includes removing informative SNPs with an inadequate read count. For example, based on the following equation:Total Read Count >= 50
[0155] At step 201D, the method includes removing informative SNPs with an abnormal fraction. For example, based on the following equation:Fraction [A / Total] < 0.56 and > 0.46
[0156] At step 203A, the method includes removing informative SNPs with an inadequate read count from a plasma sample. For example, based on the following equation:Total Read Count >= 50
[0157] At step 202A, the method includes removing informative SNPs with an inadequate read count from a tumor sample. For example based on the following equation:Total Read Count >= 50
[0158] At step 202B, the method includes determining a subset on Autosomes (i.e., Chromosome 1 through Chromosome 22).
[0159] At step 202C, the method includes calculating a B Allele frequency. For example based on the following equation:1-Fraction [A / Total]
[0160] At step 202D, the method includes identifying full chromosome instabilities (WholeChrAi) and full normal chromosomes (NormChr). For example based on the following algorithm:
[0161] (i) For each chromosome, calculate the fraction of informative SNPs with an absValDiff (i.e., abs[ (Tumor A / Total) - Germline A / Total)]) >= 0.1.
[0162] (ii) If the fraction of SNPs on a specific chromosome is greater than or equal to 0.9, the chromosome is considered a complete instability.
[0163] (iii) If the fraction of SNPs on a specific chromosome is less than or equal to 0.1, the chromosome is considered a completely normal chromosome.
[0164] At step 202E, the method includes identifying partial chromosome instabilities (PartialChr i) and partial normal regions (Pt. 1). For example based on the following algorithm:
[0165] (i) For every 50 consecutive SNPs on a chromosome, calculate the fraction ofSNPs with a B allele frequency greater than 0.5 or a B allele frequency less than 0.42.
[0166] (ii) If a fraction of SNPs in a 50 consecutive SNP group is greater than or equal to 0.7, then that group of SNPs is a partial instability.
[0167] (iii) Need two consecutive groups of 50-SNP-groups in which 70 percent or more of the SNPs meet the criteria in step (i) to be considered a partial instability.
[0168] (iv) The remainer of the SNPs not belonging to the identified partial CINs are considered partial chromosome normal regions.
[0169] With reference to FIG. 2B, at step 202F, the method includes identifying partial chromosome instabilities ( Pa rtialCh TAI) and partial normal regions (Pt 2). For example based on the following algorithm:
[0170] (i) Calculate the standard deviation of the absValDiff values across all SNPs among full normal chromosomes (i.e., sd_normal).
[0171] (ii) Calculate the standard deviation cut point (mean of the absValDiff values across all informative SNPs in all full normal chromosomes + [sd_normal*1.5]).
[0172] (iii) Within a partial chromosome instability (Cl N), if 10 or more consecutive SNPs have an absValDiff value less than the standard deviation cut point, the region is considered normal.
[0173] (iv) Within a partial chromosome normal region, the region is considered CIN if there are 5 or more consecutive SNPs with an absValDiff value greater than or equal to the standard deviation cut point.
[0174] At step 202G, the method includes identifying the Retained Allele (RA) and Lost Allele (LA) in a region of CIN. For example, based on the following algorithm:
[0175] (i) Identify the allele (i.e., T, A, C, or G) that has the largest mapped read count for an informative SNP (i.e., retained allele).
[0176] (ii) Identify the allele (i.e., T, A, C, or G) that has the second highest mapped read count for an informative SNP (i.e., lost allele).
[0177] At step 205A, in relation to data / information from one or more healthy control databases 210, the method includes determining a subset on informative SNPs that met quality control (Q ) metrics in all three tissue types isolated from a cancer subject (i.e., PBMC, FFPE, and Plasma).
[0178] At step 205B , the method includes calculating a healthy control germline retained fraction (HCgRAF) in a region of CIN. For example, by taking the mapped read count for the retained allele (HCgRAC) identified in the tumor and dividing it by the mapped read count for the loss allele (HCgLAC) identified in the tumor.
[0179] At step 205C, the method includes calculating a healthy control streck retained fraction (HCcfRAF) in a region of CIN. For example, by taking the mapped read count for theretained allele (HCcfRAC) identified in the tumor and dividing it by the mapped read count for the loss allele (HCcfLAC) identified in the tumor.
[0180] At step 205D, the method includes calculating a CIN score in healthy control streck sample(s) (HCcfCIN). For example based on the following algorithm:
[0181] (i) For each informative SNP, calculate the difference between the streck retained fraction (HCcfRAF) in the healthy control database and the germline retained fraction (HCgRAF) in the healthy control database.
[0182] (ii) Sum all those values up.
[0183] At step 205E, the method includes normalizing the Healthy Control CIN Score (HCcfCIN) to zero (i.e., if the CIN score is positive, subtract it by an equivalent value to get to zero. However, if the CIN score is negative, add an equivalent value to get to zero.
[0184] At step 202H, subsequent to identifying the Retained Allele (RA) and Lost Allele (LA) in a region of CIN at step 202G, the method includes calculating a germline retained fraction (gRAF) in a region of CIN . For example, by taking the mapped read count for the retained allele (gRAC) identified in the tumor and dividing it by the mapped read count for the loss allele (gLAC) identified in the tumor.
[0185] At step 2021, the method includes calculating a streck retained fraction (cfRAF) in a region of CIN. For example, by taking the mapped read count for the retained allele (cfRAC) identified in the tumor and dividing it by the mapped read count for the loss allele (cfLAC) identified in the tumor.
[0186] At step 202J, the method includes calculating a tumor retained fraction (tRAF) in a region of CIN. For example, by taking the mapped read count for the retained allele (tRAC) identified in the tumor and dividing it by the mapped read count for the loss allele (tLAC) identified in the tumor.
[0187] With reference to FIG. 2C, at step 202K, the method includes, for each unique CIN, calculating a tumor CIN score (tCIN) from the tumor sample. For example, based on the following algorithm:
[0188] (i) For each informative SNP, calculate the difference between tumor retained fraction (tRAF) in the cancer subject and the germline retained fraction (gRAF) in the cancer subject.
[0189] (ii) Sum all those values up.
[0190] At step 202L, the method includes, for each unique CIN, calculating a streckCIN score (cfCIN) from a cancer streck sample. For example, based on the following algorithm:
[0191] (i) For each informative SNP, calculate the difference between streck retained fraction (cfRAF) in the cancer subject and the germline retained fraction (gRAF) in the cancer subject.
[0192] (ii) Sum all those values up.
[0193] At step 206, the method includes, for each unique CIN, calculate the final CIN score in the streck cancer subject and determine if CIN is detected (ctCIN Score). For example, based on the following algorithm:
[0194] (i) Whatever the value is in the healthy control database 210 (HCcfCIN), add or subtract that value from the calculated CIN score in the cancer streck sample (cfCIN). For example, if the CIN score in the healthy control database 210 is 2.0, you would subtract 2.0 from the calculated CIN score in the cancer streck sample. However, if the CIN score in the healthy control database 210 is -2.0, you would add 2.0 to the calculated CIN score in the cancer streck sample (i.e., normalizing to zero).
[0195] (ii) If the final CIN score exceeds 0, then CIN is detected. If the final CIN score is less than or equal to 0, then the CIN is undetected in the cancer subject's streck sample.
[0196] Examples
[0197] The following examples are put forth so as to provide those of ordinary skill in the art with a complete disclosure and description of how the compounds, compositions, articles, devices and / or methods claimed herein are made and evaluated, and are intended to be purely exemplary and are not intended to limit the disclosure. Efforts have been made to ensure accuracy with respect to numbers (e.g., amounts, temperature, etc.), but some errors and deviations should be accounted for. Unless indicated otherwise, parts are parts by weight, temperature is in 0C or is at ambient temperature, and pressure is at or near atmospheric.
[0198] Example 1
[0199] FIG. 3A is an example of 27 base pairs of DNA from white blood cells for both the maternal and paternal chromosomes as represented by one strand, i.e. the positive strand, of a double stranded helix, in accordance with an embodiment. Of the 27 base pairs of DNA a total of 5 are known polymorphic sites, or specific nucleotides of DNA that can vary from one person to the next, i.e. single nucleotide polymorphisms (SNPs). When the same nucleotide exists on both the maternal and paternal chromosomes at one of these sites the state is referred to as homozygous. In this example there are 4 nonpolymorphic sites that exist in a homozygous state and for which there is no value for early detection of cancer. Three polymorphic sites have a different nucleotide on both the maternal and paternal chromosomes and exist in a state referred to as heterozygous and which have value for early detection of cancer. For those sites in a heterozygous state the specificnucleotide that is more common in the general population is called the "A" allele and can be on either the maternal or paternal chromosome. The less common nucleotide in the general population is called the "B" allele.
[0200] Example 2: Theoretical cfDNA with no ctDNA contamination
[0201] FIG. 3B shows the results of a A / B allele ratio for the 7 are known potential polymorphic loci previously illustrated in Figure 3A for cell free DNA (cfDNA) with no contamination by neoplastic cell DNA (ctDNA). In embodiments described herein, the A / B allele ratio is measured by counting the number of times the A allele nucleotide at a specific DNA position is observed in a single next generation sequencing evaluation of a single sample divided by the value for the respective B allele. In this example measurement of the A / B allele ratio for the 4 homozygous polymorphic loci would result in a theoretical value of either 0 or 1. Measurement of the A / B allele ratio for the 3 heterozygous polymorphic loci would result in a theoretical value of 0.50.
[0202] Example 3: Measured cfDNA with no ctDNA contamination
[0203] FIG. 3C shows the results of a A / B allele ratio for the 7 known potential polymorphic loci previously illustrated in Figure 3A for cell free DNA (cfDNA) with no contamination by neoplastic cell DNA (ctDNA). The actual measured values for the A / B allele ratios for each loci approximate the theoretical values but are not the same exact value due to technical variability. The degree of technical variability at each polymorphic loci can and does vary per various technologies used to measure allele ratios. This technical variability results in A / B allele ratio values slightly less than 1 or slightly more than 0 for homozygous loci. For heterozygous loci technical variability results in A / B allele ratio values that are slightly greater or less than 0.50.
[0204] Example 4: cfDNA with ctDNA no contamination
[0205] FIG. 3D shows the results of a A / B allele ratio for the 5 known potential polymorphic loci previously illustrated in Figure 3A for cell free DNA (cfDNA) with no contamination by neoplastic cell DNA (ctDNA). In this example the previous analysis of the neoplastic cells did not show allelic imbalance for these 7 polymorphic loci and there is no impact on the A / B allele ratio values from the theoretical values minus the technical variability.
[0206] Example 5: Loss of chromosome
[0207] FIG. 3E shows neoplastic tissue that has lost one of two chromosomes, in this instance the paternal copy, and illustrates that alleles that were previously heterozygous now become homozygous.
[0208] Example 6: ctDNA that has loss of homozygosity (LOH)
[0209] FIG. 3F shows 100% neoplastic tissue with "loss of heterozygosity" (LOH) that results in previously heterozygous loci with an A / B allele ratio equal to 0.50 now having a value of 0 or 1 minus technical variability. The result is all loci now appear to be homozygous.
[0210] Example 7: cfDNA comprising contamination with ctDNA that has loss of homozygosity (LOH)
[0211] FIG. 3G shows cell free DNA (cfDNA) contaminated with neoplastic cell DNA (ctDNA) that have lost either the maternal or paternal chromosome (i.e. LOH) and the A / B allele ratio of homozygous loci (Ho) remains at 0 or 100 minus technical variability, while the value for heterozygous loci move further from the theoretical value of 50% dependent on the amount of contamination.
[0212] Example 8: Loss of one copy of a chromosome
[0213] FIG. 3H shows an example where the paternal copy of one chromosome is lost. For this region of DNA, the direction of change of heterozygous loci from the theoretical value of 50% can be predicted after analysis is completed. In this example for the three polymorphic loci from left to right, or 5-prime to 3-prime, the retained alleles are A, B, and A. This results in an increase of the A / B allele ratio when the retained allele is A, and a decrease when the retained allele is B. The increase or decrease of the A / B allele ratio of these loci are joined and move together in a relative fashion, i.e., allele drift, allowing for a conditional probability of detection for this event.
[0214] Example 9: cfDNA showing responsive to Treatment
[0215] FIG. 31 shows cell free DNA (cfDNA) contaminated with a high number of neoplastic cell DNA (ctDNA) with known allelic imbalance for a specific region of a chromosome in a patient undergoing treatment. In this example three polymorphic loci that are heterozygous in normal cfDNA that would typically have an A / B allele theoretical ratio of 0.50 have values closer to a homozygous loci due to contamination of ctDNA. In this example the three loci have lost from left to right, or 5-prime to 3-prime, the B, A, and B allele. With response to therapy and diminishing ctDNA the direction of change of heterozygous (He) loci from the observed values can be predicted together in a relative fashion, i.e., allele drift, allowing for a conditional probability of detection for this event.
[0216] Example 10: cfDNA that is non-responsive to Treatment
[0217] FIG. 3J shows cell free DNA (cfDNA) contaminated with a high number of neoplastic cell DNA (ctDNA) with known allelic imbalance for a specific region of a chromosome in a patient undergoing treatment. In this example three polymorphic loci that are heterozygous in normal cfDNA that would typically have an A / B allele theoretical ratio of 0.50 have values closer to a homozygous loci due to contamination of ctDNA. In this example the three loci have lost from left toright, or 5-prime to 3-prime, the B, A, and B allele. When there is no response to therapy and no decrease in ctDNA the direction of change of heterozygous (He) loci from the observed values cannot be predicted together in a relative fashion, i.e., allele drift, not allowing for a conditional probability of detection for this event.
[0218] Example 11: cfDNA that is post treatment with recurrence
[0219] FIG. 3K shows cell free DNA (cfDNA) contaminated with a very low number of neoplastic cell DNA (ctDNA) with known allelic imbalance for a specific region of a chromosome in a patient post-treatment with significant response to treatment. In this example three polymorphic loci that are heterozygous in normal cfDNA have A / B allele ratios close to the theoretical ratio of 0.50 as would be expected with minimal contamination of ctDNA and response to treatment. In this example prior analysis of the tumor showed that the three loci have lost from left to right, or 5- prime to 3-prime, the B, A, and B allele. With follow-up for minimal residual disease (MRD) in a patient with recurrence and increasing ctDNA the direction of change of heterozygous (He) loci from the observed values can be predicted together in a relative fashion, i.e., allele drift, allowing for a conditional probability of detection for this event.
[0220] Example 12: cfDNA that is post treatment and no recurrence
[0221] FIG. 3L shows cell free DNA (cfDNA) contaminated with a very low number of neoplastic cell DNA (ctDNA) with known allelic imbalance for a specific region of a chromosome in a patient post-treatment with significant response to treatment. In this example three polymorphic loci that are heterozygous in normal cfDNA have A / B allele ratios close to the theoretical ratio of 0.50 as would be expected with minimal contamination of ctDNA and response to treatment. In this example prior analysis of the tumor showed that the three loci have lost from left to right, or 5- prime to 3-prime, the B, A, and B allele. With follow-up for minimal residual disease (MRD) in a patient without recurrence and no change in ctDNA the direction of change of heterozygous (He) loci from the observed values cannot be predicted together in a relative fashion, i.e., allele drift, not allowing for a conditional probability of detection for this event.
[0222] Experimental Results
[0223] FIGS. 4A-4E depict representative data for detecting chromosomal instability (CIN) on chromosome 3 and identifying concordance between tumor and cfDNA samples from a study that was conducted.
[0224] FIG. 4A shows representative histograms for three unique subjects (MRD- 020, MRD-027, and MRD-029) showing the frequency (y-axis) of the A / total allele ratio (x-axis) for qualified SNPs on chromosome 3 in tumor samples with no evidence of chromosomal instability (CIN).
[0225] FIG. 4B shows representative circos plots for the same subjects (MRD-020, MRD-027, and MRD-029) indicating the algorithm detected no whole-chromosome allelic instabilities (WholeChrAI) or partial chromosome allelic instabilities ( PartialCh rAI), identifying normal chromosomes (gray regions).
[0226] FIG. 4C shows representative histograms for three different subjects (MRD- 014, MRD-016, and MRD-018) showing the frequency (y-axis) of the A / total allele ratio (x-axis) for qualified SNPs on chromosome 3 in tumor samples with evidence of chromosomal instability (CIN).
[0227] FIG. 4D shows representative circos plots for the same subjects (MRD-014, MRD-016, and MRD-018), highlighting algorithm-detected chromosomal allelic instabilities on chromosome 3 (blue regions).
[0228] FIG. 4E is a table displaying sample IDs and the associated cfCIN scores (ctCIN scores), calculated in samples where chromosomal instability was detected on chromosome 3.
[0229] Computing Devices and Methods of Use
[0230] It should be appreciated that the logical operations described herein with respect to the various figures may be implemented (1) as a sequence of computer-implemented acts or program modules (i.e., software) running on a computing device (e.g., the computing device described in FIG. 5), (2) as interconnected machine logic circuits or circuit modules (i.e., hardware) within the computing device and / or (3) a combination of software and hardware of the computing device. Thus, the logical operations discussed herein are not limited to any specific combination of hardware and software. The implementation is a matter of choice dependent on the performance and other requirements of the computing device. Accordingly, the logical operations described herein are referred to variously as operations, structural devices, acts, or modules. These operations, structural devices, acts and modules may be implemented in software, in firmware, in specialpurpose digital logic, and any combination thereof. It should also be appreciated that more or fewer operations may be performed than shown in the figures and described herein. These operations may also be performed in a different order than those described herein.
[0231] Referring to FIG. 5, an example computing device 500 upon which embodiments of the present disclosure may be implemented is illustrated. It should be understood that the example computing device 500 is only one example of a suitable computing environment upon which embodiments of the present disclosure may be implemented. Optionally, the computing device 500 can be a well-known computing system including, but not limited to, personal computers, servers, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, personal network computers (PCs), minicomputers, mainframe computers, embedded systems,and / or distributed computing environments including a plurality of any of the above systems or devices. Distributed computing environments enable remote computing devices, which are connected to a communication network or other data transmission medium, to perform various tasks. In the distributed computing environment, the program modules, applications, and other data may be stored on local and / or remote computer storage media.
[0232] In its most basic configuration, the computing device 500 typically includes at least one processing unit 506 and system memory 504. Depending on the exact configuration and type of computing device, system memory 504 may be volatile (such as random-access memory (RAM)), non-volatile (such as read-only memory (ROM), flash memory, etc.), or some combination of the two. This most basic configuration is illustrated in FIG. 5 by the dashed line 502. The processing unit 506 may be a standard programmable processor that performs arithmetic and logic operations necessary for the operation of the computing device 500. The computing device 500 may also include a bus or other communication mechanism for communicating information among various components of the computing device 500.
[0233] Computing device 500 may have additional features / functionality. For example, the computing device 500 may include additional storage such as removable storage 508 and non-removable storage 510 including, but not limited to magnetic or optical disks or tapes. Computing device 500 may also contain network connection(s) 516 that allow the device to communicate with other devices. Computing device 500 may also have input device(s) 514 such as a keyboard, mouse, touch screen, etc. Output device(s) 512, such as a display, speakers, printer, etc., may also be included. The additional devices may be connected to the bus in order to facilitate communication of data among the components of the computing device 500. All these devices are well-known in the art and need not be discussed at length here.
[0234] The processing unit 506 may be configured to execute program code encoded in tangible, computer-readable media. Tangible, computer-readable media refers to any media that is capable of providing data that causes the computing device 500 (i.e., a machine) to operate in a particular fashion. Various computer-readable media may be utilized to provide instructions to the processing unit 506 for execution. Example of tangible, computer-readable media may include but is not limited to, volatile media, non-volatile media, removable media and nonremovable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules or other data. System memory 504, removable storage 508, and non-removable storage 510 are all examples of tangible computer storage media. Examples of tangible, computer-readable recording media include but are not limited to, an integrated circuit (e.g., field-programmable gate array or application-specific IC), a hard disk,an optical disk, a magneto-optical disk, a floppy disk, a magnetic tape, a holographic storage medium, a solid-state device, RAM, ROM, electrically erasable program read-only memory (EEPROM), flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices.
[0235] In an example implementation, the processing unit 506 may execute program code stored in the system memory 504. For example, the bus may carry data to the system memory 504, from which the processing unit 506 receives and executes instructions. The data received by the system memory 504 may optionally be stored on the removable storage 508 or the non-removable storage 510 before or after execution by the processing unit 506.
[0236] It should be understood that the various techniques described herein may be implemented in connection with hardware or software or, where appropriate, with a combination thereof. Thus, the methods and apparatuses of the presently disclosed subject matter, or certain aspects or portions thereof, may take the form of program code (i.e., instructions) embodied in tangible media, such as floppy diskettes, CD-ROMs, hard drives, or any other machine- readable storage medium wherein, when the program code is loaded into and executed by a machine, such as a computing device, the machine becomes an apparatus for practicing the presently disclosed subject matter. In the case of program code execution on programmable computers, the computing device generally includes a processor, a storage medium readable by the processor (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device. One or more programs may implement or utilize the processes described in connection with the presently disclosed subject matter, for example, through the use of an application programming interface (API), reusable controls, or the like. Such programs may be implemented in a high-level procedural or object-oriented programming language to communicate with a computer system. However, the program(s) can be implemented in assembly or machine language if desired. In any case, the language may be a compiled or interpreted language, and it may be combined with hardware implementations.
[0237] In one embodiment, disclosed herein is a non-transitory computer-readable storage medium comprising instructions that, when executed, cause at least one processor to perform the method of any preceding embodiments.
[0238] FIG. 6 is an example system 600 in accordance with certain embodiments of the present disclosure. As shown in FIG. 6, the system 600 includes a processing system 601 configured to communicate with one or more computing devices 603. In various implementations, the processing system 601 and the computing device(s) 603 are configured to transmit data to andreceive data from one another over a network 602. The system 600 can include one or more databases, data stores, repositories, and the like. As shown, the system 600 includes healthy control database(s) 615 in communication with the computing device(s) 603 and the processing system 601. In some implementations, the healthy control database(s) 615 can be hosted by the processing system 601.
[0239] In some embodiments, the processing system 601 is configured to sequence DNA and / or analyze a plurality of chromosomal variations or CINs using a parallel processing operation. For example, as shown, the processing system 601 comprises a plurality of networked devices that are operatively coupled to one another that are configured to process data / information in parallel.
[0240] In the example shown in FIG. 6, the processing system 601 comprises a first processing device 610A, a second processing device 610B, and third processing device 610C...an n-th processing device 610N. In some implementations, at least some of the processing devices 610A-N are or comprise an NGS device that is configured to sequence a plurality of DNA fragments in parallel, thereby facilitating fast and efficient sequencing. An example NGS device can include instrumentation, a fluidics system, detection component(s) (e.g., fluorescence-based or semiconductor-based detection components, and the like). Additionally or alternatively, each processing device 610A-N can include a plurality of processing components. For example, as shown, the first processing device comprises a DNA sequencer 220 and a DNA analyzer 230.
[0241] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.
Claims
WHAT IS CLAIMED:
1. A computer-implemented method for detecting presence of solid tumors or neoplasms in a blood sample, the computer-implemented method comprising: receiving Deoxyribonucleic acid (DNA) sequencing data for a subject's blood sample; identifying a plurality of chromosomal variations in the DNA sequencing data, wherein the plurality of chromosomal variations comprises a plurality of different chromosomal instabilities (CINs); and detecting a solid tumor or neoplasm in the subject's blood sample based, at least in part, on analysis of the plurality of chromosomal variations.
2. The computer-implemented method of claim 1, wherein each of the plurality of different CINs is a unique CIN.
3. The computer-implemented method of claim 1 or 2, wherein each of the plurality of different CINs is a partial chromosomal instability or whole chromosomal instability.
4. The computer-implemented method of any one of claims 1-3, wherein the analysis of the plurality of chromosomal variations comprises comparing the DNA sequencing data to a healthy control DNA sequencing dataset.
5. The computer-implemented method of claim 4, wherein the analysis of the plurality of chromosomal variations reveals that the plurality of different CINs comprise at least one of a plurality of candidate CINs for the solid tumor or neoplasm in the subject's blood sample.
6. The computer-implemented method of any one of claims 1-5, wherein the solid tumor or neoplasm is in one of a plurality of solid tumor primary sites including lung, trachea, breast, uterine, ovarian, skin, bladder, kidney, liver, pancreas, soft tissue, bone, mouth, salivary glands, bowel, rectum, gallbladder, urethra, fallopian tube, vulva, testicle, prostate, penis, thyroid, adrenal, parathyroid, heart, vessels, thymus, nerve, eye, ear, and subcutaneous tissue.
7. The computer-implemented method of any one of claims 1-5, wherein at least one of the plurality of chromosomal variations involves partial instability of chromosome 1, chromosome 2, chromosome 3, chromosome 4, chromosome 5, chromosome 6, chromosome 7, chromosome 8, chromosome 9, chromosome 10, chromosome 11, chromosome 12, chromosome 13, chromosome 14, chromosome 15, chromosome 16, chromosome 17, chromosome 18, chromosome 19,chromosome 20, chromosome 21, chromosome 22, chromosome X, or chromosome Y.
8. The computer-implemented method of any one of claims 1-5, wherein at least one of the plurality of chromosomal variations involves whole instability of chromosome 1, chromosome 2, chromosome 3, chromosome 4, chromosome 5, chromosome 6, chromosome 7, chromosome 8, chromosome 9, chromosome 10, chromosome 11, chromosome 12, chromosome 13, chromosome 14, chromosome 15, chromosome 16, chromosome 17, chromosome 18, chromosome 19, chromosome 20, chromosome 21, chromosome 22, chromosome X, or chromosome Y.
9. The computer-implemented method of any one of claims 1-8, wherein the DNA sequencing data encodes a genome wide single nucleotide polymorphism (SNP) backbone.
10. The computer-implemented method of any one of claims 1-8, wherein the DNA sequencing data encodes a plurality of polymorphisms in a genome wide single nucleotide polymorphism (SNP) backbone.
11. The computer-implemented method of any one of claims 1-10, further comprising: in response to detecting the solid tumor or neoplasm in the subject's blood sample, determining a presence or absence of cancer for the subject.
12. The computer-implemented method of claim 11, further comprising: generating a report comprising the DNA sequencing data and an indication of the detected solid tumor or neoplasm, and / or the presence or absence of cancer for the subject.
13. The computer-implemented method of claim 12, further comprising generating display data for the report.
14. The computer-implemented method of any one of claims 12-13, further comprising: transmitting the report over a network.
15. The computer-implemented method of any one of claims 11-14, further comprising, in response to determining the presence or absence of cancer for the subject, determining a prognosis for the subject.
16. The computer-implemented method of any one of claims 11-15, further comprising: in response to determining the presence or absence of cancer for the subject, determining a response to treatment for the subject.
17. The computer-implemented method of any one of claims 11-16, further comprising: in response to determining the presence or absence of cancer for the subject, providing a determination of minimal or measurable residual disease for the subject.
18. The computer-implemented method of any one of claims 1-17, further comprising: sequencing the DNA to obtain the DNA sequencing data.
19. The computer-implemented method of claim 18, wherein the DNA is sequenced using a next generation sequencing device (NGS).
20. The computer-implemented method of claim 18 or 19, further comprising: sequencing the DNA or analyzing the plurality of chromosomal variations using a parallel processing operation.
21. A method for detecting solid tumors or neoplasms in a blood sample, the method comprising: detecting, using a computing device, the solid tumor or neoplasm in the subject's blood sample according to any one of claims 1-20; and in response to detecting the solid tumor or neoplasm, providing a treatment to the subject.
22. The method of claim 20, further comprising: obtaining a sample from the subject; isolating cell free (cf) deoxyribonucleic acid (DNAj(cfDNA) from the sample; and sequencing the cfDNA to obtain the DNA sequencing data.
23. The method of claim 22, wherein the sample comprises at least one of whole blood, peripheral blood, or plasma.
24. A system for detecting solid tumors or neoplasms, the system comprising: at least one processor; anda memory operably coupled to the at least one processor, wherein the memory has computer executable instructions stored thereon that, when executed by the at least one processor, cause the at least one processor to: receive DNA sequencing data for a subject's blood sample; identify a plurality of chromosomal variations in the DNA sequencing data, wherein the plurality of variations comprises a plurality of chromosomal instabilities (CINs); detect a solid tumor or neoplasm in the subject's blood sample based, at least in part, on analysis of the plurality of chromosomal variations.
25. The system of claim 24, wherein the memory has further computer executable instructions stored thereon that, when executed by the at least one processor, cause the at least one processor to: in response to detecting the solid tumor or neoplasm in the subject's blood sample, determine a presence or absence of cancer for the subject.
26. The system of claim 25, wherein the memory has further computer executable instructions stored thereon that, when executed by the at least one processor, cause the at least one processor to: generate a report comprising the DNA sequencing data and an indication of the detected solid tumor or neoplasm in the subject's blood sample and / or the presence or absence of cancer.
27. The system of claim 26, wherein the memory has further computer-executable instructions stored thereon that, when executed by the at least one processor, cause the at least one processor to: generate display data for the report.
28. The system of any one of claims 26-27, wherein the memory has further computerexecutable instructions stored thereon that, when executed by the at least one processor, cause the at least one processor to: transmit the report over a network.
29. The system of any one of claims 25-28, wherein the memory has further computer executable instructions stored thereon that, when executed by the at least one processor, cause the at least one processor to:in response to determining the presence or absence of cancer, determine a prognosis for the subject.
30. The system of any one of claims 25-29, wherein the memory has further computer executable instructions stored thereon that, when executed by the at least one processor, cause the at least one processor to: in response to determining the presence or absence of cancer, determine a response to treatment for the subject.
31. The system of any one of claims 25-30, wherein the memory has further computer executable instructions stored thereon that, when executed by the at least one processor, cause the at least one processor to: in response to determining the presence or absence of cancer, determine a minimal or measurable residual disease for the subject.
32. A computer-implemented method for detecting presence of solid tumors or neoplasms in a subject, the computer-implemented method comprising: receiving Deoxyribonucleic acid (DNA) sequencing data for a subject's liquid biopsy sample; identifying a plurality of chromosomal variations in the DNA sequencing data, wherein the plurality of chromosomal variations comprises a plurality of different chromosomal instabilities (CINs); and detecting a solid tumor or neoplasm in the subject based, at least in part, on analysis of the plurality of chromosomal variations.
33. The computer-implemented method of claim 32, wherein the subject's liquid biopsy sample is a blood sample, the blood sample comprises at least one of whole blood, peripheral blood, plasma, or serum.
34. The computer-implemented method of claim 32, wherein the subject's liquid biopsy sample is serum, saliva, sputum, cerebral spinal fluid, urine, lymph, or lacrimal fluid.
35. A computer-implemented method for determining a response to therapy, the computer- implemented method comprising:receiving first Deoxyribonucleic acid (DNA) sequencing data for a subject's first liquid biopsy sample; identifying a first plurality of chromosomal variations in the first DNA sequencing data; receiving second DNA sequencing data for the subject's second liquid biopsy sample; identifying a second plurality of chromosomal variations in the second DNA sequencing data; and determining the subject's response to therapy based, at least in part, on a comparison of the first plurality of chromosomal variations in the first DNA sequencing data and the second plurality of chromosomal variations in the second DNA sequencing data, wherein each plurality of chromosomal variations comprises a plurality of different chromosomal instabilities (CINs).
36. The computer-implemented method of claim 35, wherein each of the subject's first and second liquid biopsy samples is a blood sample, and wherein the blood sample comprises at least one of whole blood, peripheral blood, plasma, or serum.
37. The computer-implemented method of claim 35, wherein each of the subject's first and second liquid biopsy samples is serum, saliva, sputum, cerebral spinal fluid, urine, lymph, or lacrimal fluid.
38. The computer-implemented method of any one of claims 35-37, wherein the subject's first liquid biopsy sample is obtained at a first time point and the subject's second liquid biopsy sample is obtained at a second time point, the second time point being subsequent in time relative to the first time point.
39. The computer-implemented method of claim 38, wherein the subject's first liquid biopsy sample is obtained before or after treatment begins.
40. The computer-implemented method of claim 38 or 39, wherein the subject's second liquid biopsy sample is obtained after treatment begins.
Citation Information
Patent Citations
Disease Detection in Liquid Biopsies
US20230042332A1
Cancer detection methods
WO2017048932A1
Early detection of cancer using allele drift and chromosomal instability
WO2023224978A1