Fragment Size Characterization of Cell-Free DNA Mutations Derived from Clonal Hematopoiesis

By profiling cfDNA fragments for variant allele frequency and size distribution, the method distinguishes cancer variants from hematopoietic cell variants, enhancing tumor mutation burden assessment and clinical decision-making.

JP7705797B2Active Publication Date: 2025-07-10ILLUMINA INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2021537785
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-10-08
Filing Date
2020-09-16
Publication Date
2025-07-10
Estimated Expiration
2040-09-16

AI Technical Summary

Technical Problem

Current methods struggle to distinguish somatic variants derived from hematopoietic cells from cancer variants in cell-free DNA (cfDNA) samples, leading to false-positive mutations that affect clinical decisions.

Method used

A method involving molecular profiling of cfDNA fragments to determine variant allele frequency (VAF) and generate a fragment size distribution profile to identify and remove hematopoietic cell variants, allowing for the distinction of cancer variants.

Benefits of technology

This approach enhances the accuracy of tumor mutation burden determination by effectively filtering out hematopoietic cell variants, improving the reliability of clinical decisions and enabling targeted therapies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007705797000003
    Figure 0007705797000003
  • Figure 0007705797000004
    Figure 0007705797000004
  • Figure 0007705797000005
    Figure 0007705797000005
Patent Text Reader

Abstract

Methods and systems are provided for distinguishing between cancer variants and somatic variants derived from hematopoietic cells in a cell-free DNA sample. In some embodiments, cancer variants can be distinguished from somatic variants derived from hematopoietic cells based on fragment size distribution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Some embodiments of the methods and systems provided herein relate to variant calling from sequence data obtained from cell-free DNA (cfDNA) samples. In some embodiments, somatic variants derived from hematopoietic cells can be distinguished from cancer variants based on the fragment size distribution of multiple variants.

Background Art

[0002] Mutations in human DNA are known to cause cancer, and these mutations are currently the focus of cancer research and treatment. Circulating tumor DNA (ctDNA) is a non-invasive real-time biomarker that can provide diagnostic and prognostic information for cancer patients before and after treatment. However, cell-free DNA (cfDNA) derived from tumor cells is only a small part, and most of the fragments originate from hematopoietic cells. Somatic mutations carried by hematopoietic cells can be a major cause of false-positive mutations in cfDNA that affect clinical decisions.

Summary of the Invention

Means for Solving the Problems

[0003] The present disclosure relates to methods and systems for distinguishing cancer variants and somatic variants derived from hematopoietic cells from cfDNA samples.

[0004] Some embodiments provided herein relate to methods for discriminating cancer variants in circulating tumor DNA (ctDNA) samples from hematopoietic cell variants. In some embodiments, the method comprises: (a) obtaining a ctDNA sample comprising a plurality of cell-free DNA (cfDNA) fragments, or having an obtained sample; (b) extracting cfDNA fragments comprising a plurality of variants from the sample; (c) performing molecular profiling for each of the plurality of variants; and (d) identifying cancer variants by removing the identified hematopoietic cell variants. In some embodiments, performing molecular profiling for each of the plurality of variants comprises: (i) determining a variant allele frequency (VAF) for each of the plurality of variants, including cancer variants and hematopoietic cell variants; and (ii) generating a fragment size distribution profile to identify hematopoietic cell variants.

[0005] Some embodiments provided herein relate to methods for determining the tumor mutation burden of a tumor. In some embodiments, the method comprises obtaining sequence data from a biological sample comprising tumor cells, determining a plurality of variants from the sequence data, and determining the number of cancer variants in the plurality of variants using any of the methods described herein, wherein the number of cancer variants is equal to the tumor mutation burden of the tumor.

[0006] Some embodiments provided herein relate to methods for treating a tumor. In some embodiments, the method comprises determining a tumor having a tumor mutation burden of 10 or more cancer variants by any of the methods described herein, and treating the tumor by administering an effective amount of a checkpoint inhibitor.

[0007] Some embodiments provided herein relate to an electronic system for analyzing genetic polymorphism data. In some embodiments, the system is an informatics module executed on a processor and adapted to identify a plurality of variants from sequence data from a cfDNA sample, the plurality of variants including cancer variants and hematopoietic cell variants; an analyzer for performing molecular profiling for each of the plurality of variants, configured to determine a variant allele frequency (VAF) for each of the plurality of variants and further configured to generate a fragment size distribution profile; an analyzer for identifying cancer variants by removing the identified hematopoietic cell variants; and a display module adapted to return variants that are not removed from the plurality of variants.

Brief Description of the Drawings

[0008]

Figure 1

[0009]

Figure 2

[0010]

Figure 3

[0011]

Figure 4

[0012]

Figure 5

[0013]

Figure 6

[0014]

Figure 7A

Figure 7B

Best Mode for Carrying Out the Invention

[0015] In the following detailed description, reference is made to the accompanying drawings, which form a part of this specification. In the drawings, like reference numerals typically identify like components, unless the context specifically indicates otherwise. The exemplary embodiments described in the detailed description, drawings, and claims are not intended to be limiting. Other embodiments may be utilized and other changes may be made without departing from the spirit or scope of the subject matter presented herein. It is readily understood that aspects of the present disclosure can be arranged, substituted, combined, separated, and designed in a variety of different configurations as generally described herein and illustrated in the drawings, and all of these are explicitly contemplated herein.

[0016] Embodiments of the systems, methods, and compositions provided herein relate to methods and systems for determining nucleic acid variants (“variant calls”) from sequence data obtained from a cell-free DNA (cfDNA) sample taken from a user or patient. In some embodiments, the methods and systems can distinguish somatic mutations from different cellular origins not related to cancer from tumor mutations based on fragment size distribution. In some embodiments, somatic variants derived from hematopoietic cells can both be distinguished from mutations derived from tumor cells obtained from a cfDNA sample based on the fragment size distribution of the variants. The cfDNA sample contains DNA fragments derived from tumor cells and other sources, such as those derived from clonal hematopoiesis. Since the fragment sizes of DNA derived from tumor cells are different from those of hematopoietic cells, fragments from the cfDNA sample can be applied to a fragment size distribution profile to distinguish tumor cells from hematopoietic cells, thereby providing an improved determination of the tumor mutation burden in the sample. More specifically, in some embodiments, fragments carrying somatic mutations from solid tumors have smaller sizes compared to fragments carrying somatic mutations from clonal hematopoiesis or leukemia.

[0017] Unless otherwise defined herein, scientific and technical terms used in connection with this application shall have the ordinary meaning as illuminated by this specification and as understood by those of ordinary skill in the art to which this disclosure pertains. It is to be understood that the disclosure is not limited to the specific methodologies, protocols, and reagents etc. described herein and thus may vary. Definitions of common terms in immunology and molecular biology are provided in Diagnosis and Therapy, 20th Edition, published by Merck Sharp & Dohme Corp., 2018 (ISBN 0911910190, 978 - 0911910421); Robert S. Porter et al. (eds.), the Encyclopedia of Molecular Cell Biology and Molecular Medicine, published by Blackwell Science Ltd., 1999 - 2012 (ISBN 9783527600908); and Robert A. Meyers (ed.), Molecular Biology and Biotechnology: a Comprehensive Desk Reference, published by VCH Publishers, Inc., 1995 (ISBN 1 - 56081 - 569 - 8); Immunology by werner Luttmann, published by Elsevier, 2006; Janeway’s Immunobiology, Kenneth Murphy, Allan Mowat, Casey weaver (eds.), W.W. Norton & Company, 2016 (ISBN 0815345054, 978 - 0815345053); Lewin’s Genes XI, published by Jones & Bartlett Publishers, 2014 (ISBN - 1449659055); Michael Richard Green and Joseph Sambrook, Molecular Cloning: A Laboratory Manual, 4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y., USA (2012) (ISBN 1936113414); Davis et al., Basic Methods in Molecular Biology, Elsevier Science Publishing, Inc., New York, USA (2012) (ISBN 044460149X); Laboratory Methods in Enzymology: DNA, Jon Lorsch (ed.) Elsevier, 2013 (ISBN 0124199542); Current Protocols in Molecular Biology (CPMB), Frederick M. Ausubel (ed.), John Wiley and Sons, 2014 (ISBN 047150338X, 9780471503385), Current Protocols in Protein Science (CPPS), John E. Coligan (ed.), John Wiley and Sons, Inc., 2005; and Current Protocols in Immunology (CPI) (John E. Coligan, ADA M Kruisbeek, David H Margulies, Ethan M Shevach, Warren Strobe, (eds.) John Wiley and Sons, Inc., 2003 (ISBN 0471142735, 9780471142737), the contents of each of which are hereby incorporated by reference in their entirety.

[0018] As used herein, "cell-free DNA" or "cfDNA" has its ordinary meaning as understood in the context of this specification and refers to DNA that freely circulates in the bloodstream, but is not necessarily derived from a tumor. CfDNA can be released from cells as a result of various processes, including both normal and abnormal apoptotic events, cell excretion, and necrosis. Specific forms of cfDNA can be present within the circulatory system as a result of various medical conditions, disease states, or pregnancy. Solid tissues, including cancer, also contribute to the plasma cfDNA pool. CfDNA is characterized by the length of the nucleic acid fragments due to fragmentation within nucleosomes, and the fragments can be approximately 100-200 bp in length, for example, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 bp in length, or within a range defined by any two of the foregoing values. In some embodiments, the fragments are 166 bp in length.

[0019] As used herein, "circulating tumor DNA" or "ctDNA" has its ordinary meaning as understood in the context of this specification and refers to fragmented DNA of tumor origin that is not associated with cells. CtDNA may be derived from a portion of the cfDNA found in plasma or serum and may be derived from a tumor or circulating tumor cells. CtDNA has the molecular signature of the tumor cell genome. CtDNA can be used to sample clonal populations at both the primary and metastatic sites by perfusion sampling, as opposed to microdissection of tumor tissue to examine some aspects of the genetic diversity within the tumor. However, due to dilution by abundant normal cfDNA, ctDNA may be present at low allele frequencies. In some embodiments, the low allele frequency is an amount less than 5%, less than 4%, less than 3%, less than 2%, less than 1%, less than 0.9%, less than 0.8%, less than 0.7%, less than 0.6%, less than 0.5%, less than 0.4%, less than 0.3%, or less than 0.2%, or within a range defined by any two of the foregoing values.

[0020] In some embodiments, the methods and systems described herein can distinguish between somatic polymorphisms derived from hematopoietic cells and mutations derived from tumor cells.

[0021] As used herein, a "variant" can include a polymorphism within a nucleic acid molecule. Polymorphisms can include insertions, deletions, variable length tandem repeats, single nucleotide mutations, and structural variants such as translocations, copy number polymorphisms, or combinations thereof. Variants can sometimes include germline variants or somatic variants. As used herein, "germline variant" can sometimes include variants present in germ cells and all cells of an individual, and may be inherited by offspring. As used herein, "somatic variant" can sometimes include variants present within tumor cells or carried by hematopoietic cells, and may not be present in other cells of the individual and may not be inherited.

[0022] Analysis of gene mutations can provide useful information in the study of various phenotypes, including certain somatic diseases such as genetic disorders and cancer. Variant alleles can include the mutant form of a gene at a specific position in its DNA sequence. Some parts of the gene sequence are different for each individual and have no effect as a result, while others result in dramatically different phenotypes. For example, a single mutation in a DNA sequence can switch a gene on / off or change the function of a protein in a metabolic chain. Gene data across populations with genetic diversity can provide insights not only into the relationship between genes and phenotypes, but also into the evolutionary history of phenotypes associated with variants. For example, changes in body organs or systems that occur over time, such as changes in the kidneys, hair, or muscle tissue, can be associated with somatic mutations.

[0023] As used herein, "variant allele frequency" or "VAF" has its ordinary meaning as understood in the context of this specification and refers to the ratio of the sequenced reads that match the variant divided by the overall coverage at the target position. VAF may include a measure of the proportion of the sequenced reads that carry the variant.

[0024] "Hematopoietic cells" has its ordinary meaning as understood in the context of this specification and refers to any type of cell in the hematopoietic system. Examples include, but are not limited to, undifferentiated cells such as hematopoietic stem cells and progenitor cells (HSPC), and differentiated cells such as megakaryocytes, platelets, erythrocytes, leukocytes, granulocytes, monocytes, lymphocytes, and natural killer (NK) cells. As used herein, "clonal hematopoiesis" has its ordinary meaning as understood in the context of this specification and refers to the clonal expansion of a subpopulation of hematopoietic cells that have one or more somatic mutations. Clonal hematopoiesis (CH) can be a major cause of false positive mutations identified in cfDNA and can thus affect clinical decisions. Accordingly, the present disclosure relates to methods and systems for determining whether somatic mutations are derived from CH or tumor cells.

[0025] Clonal hematopoiesis of indeterminate potential (CHIP) can be a common aging-related phenomenon, in which hematopoietic stem cells (HSCs) or other early blood cell progenitor cells contribute to the formation of genetically distinct subpopulations of blood cells. In some embodiments, the tumor mutation burden (TMB) of a tumor can be indicated by determining the origin of somatic variants. In some embodiments, the determination of the origin of somatic variants can be used for the determination of targeted therapies.

[0026] As used herein, "tumor mutational burden" or "TMB" has its ordinary meaning as understood in light of this specification and refers to a measure of the mutations carried by tumor cells. After recent studies have shown a correlation between TMB and the effectiveness of checkpoint inhibitor immunotherapy, TMB has emerged as an important biomarker for the selection of cancer therapies. When calculating TMB, it may be useful to identify and filter germline variants. Germline variants may include variants that an individual was born with (or that are shared between tumor and normal cells), but these are detected as variants when compared to the reference genome. Since these variants do not contribute to distinguishing tumor cells from normal cells, failure to filter them correctly can lead to an overestimation of TMB. Additionally, somatic variants derived from hematopoietic cells (e.g., clonal hematopoiesis) can be filtered to distinguish tumor cells from clonal hematopoiesis. Embodiments include determining the TMB of a cfDNA sample, selecting a treatment for a tumor according to the TMB, and treating a subject in need thereof.

[0027] In some embodiments, TMB can be calculated by dividing by the effective panel size to obtain the eligible variants. Eligible variants include, for example, variants in coding regions, variants that do not occur in low-confidence regions, variants having a frequency greater than 0.4% and less than 40%, variants having a coverage greater than 500-fold, single nucleotide variants (excluding multiple nucleotide variants), insertion and deletion variants (Indels), non-synonymous and synonymous variants, and variants having a COSMIC (Catalog of Somatic Mutations in Cancer) number greater than 50 are excluded, and / or variants having mutations in genes affected by clonal hematopoiesis, such as Tet methylcytosine dioxygenase 2 (TET2), tumor protein p53 (TP53), DNA (cytosine-5)-methyltransferase 3A (DNMT3A), and / or Casitas B-lineage lymphoma (CBL) are excluded. The effective panel size can include, for example, the total coding region having a coverage greater than 500-fold. Method

[0028] Some embodiments provided herein relate to methods for determining the origin of somatic variants. In some embodiments, the method includes distinguishing DNA mutations derived from clonal hematopoiesis (CH) from DNA mutations indicative of tumor variants in a cfDNA sample. In some embodiments, CH can be distinguished from tumor variants by analyzing the fragment size distribution of DNA fragments in cfDNA.

[0029] As used herein, "fragment size distribution" has its ordinary meaning in light of this specification and refers to fragments of cfDNA that are distributed in size to generate a fragment size profile. The generated fragment size profile can be used to distinguish somatic mutations from different cell origins.

[0030] An exemplary method for distinguishing somatic mutations from different cell origins is schematically shown in FIG. 1. Method 100 includes step 105 of obtaining a sample or having a sample that has been obtained. In some embodiments, the sample is a biological sample. In some embodiments, the biological sample can include tumor cells. In some embodiments, the biological sample can include a serum sample, a fecal sample, a blood sample, and a tumor sample. In some embodiments, the biological sample is fixed. In some embodiments, the sample includes cfDNA. In some embodiments, the sample includes ctDNA. In some embodiments, the sample includes a plurality of variants, such as somatic variants and germline variants. In some embodiments, the method includes removing germline variants.

[0031] The amount of the biological sample has no particular requirements as long as the biological sample contains sufficient nucleic acids for analysis. Thus, the amount of the biological sample may include an amount in the range of about 1 μL to about 500 μL, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, or 500 μL, or an amount within the range defined by any two of the foregoing values.

[0032] In some embodiments, the method includes obtaining a sample from a subject. In some embodiments, the method includes having a sample obtained from a subject. In some embodiments, the subject can provide a biological sample, or a separate entity can provide a biological sample. The biological sample can be any substance produced by the subject. Generally, the biological sample can be any tissue collected from the subject, or any substance produced by the subject. Examples of biological samples can include blood, plasma, saliva, cerebrospinal fluid (CSF), buccal tissue, urine, feces, skin, hair, and organ tissue. In some embodiments, the biological sample is a solid tumor or a biopsy of a solid tumor. In some embodiments, the biological sample is a formalin-fixed paraffin-embedded (FFPE) tissue sample. The biological sample can be any biological sample containing nucleic acids. The biological sample may be derived from a subject. The subject can be a mammal, reptile, amphibian, bird, or fish. In some embodiments, the subject is a human. In some embodiments, the method further includes obtaining a matched tumor sample. By matching the results of the tumor and cfDNA variants, a fragment size profile derived from tumor hematopoietic cells, healthy hematopoietic cells, and abnormal hematopoietic cells can be constructed.

[0033] In some embodiments, method 100 includes step 110 of extracting DNA from a sample. DNA of a biological sample can be extracted by any suitable extraction method. Methods for achieving this are well known to those skilled in the art and include, for example, phenol / chloroform extraction, ethanol precipitation, cesium chloride gradient, CHELEX or silica column, or bead method. DNA can be extracted from cells using methods known in the art and / or commercially available kits, for example, by using the QIAamp DNA blood Mini Kit or DNeasy Blood&Tissue Kit supplied by QIAGEN.

[0034] In some embodiments, method 100 includes step 115 of library preparation and enrichment. Library preparation and enrichment can be carried out according to methods known in the art. For example, library preparation and enrichment methods may include standard protocols that include steps of end repair and A-tailing, adapter ligation, ligation cleanup, index PCR, first hybridization, first target capture, second hybridization, second target capture, library amplification, amplified library cleanup, library quantification, and / or library normalization.

[0035] In some embodiments, method 100 further includes step 120 of sequencing. Sequencing of the DNA library can be performed, for example, using HiSeq. HiSeq can be performed using 151bp paired-end reads. Paired-end sequencing provides high-quality alignment across DNA regions containing repetitive sequences and generates long contigs for de novo sequencing by filling in gaps in the consensus sequence. Paired-end DNA sequencing also detects common DNA rearrangements such as insertions, deletions, and inversions. In some embodiments, sequencing includes molecular profiling using unique molecular identifiers (UMIs).

[0036] In some embodiments, method 100 further includes step 125 of variant allele frequency (VAF) analysis. The VAF analysis can be performed according to methods established in the art, and the proportion of reads at the site containing the variant allele is determined. In cfDNA, since the proportion of the tumor is low (typically less than 20% by volume), the VAF may vary significantly between the germline and somatic lineages. ctDNA may include highly sensitive detection of low-VAF variants in amounts of 0.2% to 0.4%.

[0037] Variant frequency analysis may include extracting variant data from sequence data collected from a sequencer. For example, germline variants can be removed by applying a filter to data representing multiple variants, such as a database filter or a proximity filter. A database filter can be used to identify variants as germline variants and extract the variants from data representing multiple variants in a sample. The database filter can be related to the allele count of the corresponding variant in the database for a particular variant among multiple variants. The proximity filter can be related to the proximity of the allele frequency of a particular variant among multiple variants, the position of the variant in a genomic region, and the allele frequency of the identified germline variant in the same genomic region. In some embodiments, applying the database filter includes determining a first germline variant among multiple variants, each of the first germline variants having an allele count in a first reference set of variants that is greater than or equal to a threshold allele count. In some embodiments, applying the proximity filter includes: (i) binning the variants of the multiple variants into multiple bins, where variants located in the same genomic region are binned into the same bin; (ii) determining database variants among the multiple variants, where the database variants are present within a second reference set of variants; and / or (iii) determining a second germline variant among the multiple variants, each of the second germline variants having an allele frequency within a proximity range of the allele frequency of at least one database variant in the same bin as the second germline variant.

[0038] In some embodiments, by classifying or binning variants of a plurality of variants into a plurality of bins, variants located in the same region of the genome can be classified or binned into the same bin. In some embodiments, the same region of the genome can be within the same chromosome, within the same arm of the chromosome, within the same chromosomal cytoband. In some embodiments, the same region of the genome can be within the same contiguous 100 Mb, 50 Mb, 40 Mb, 30 Mb, 20 Mb, 10 Mb, 5 Mb, 1 Mb, or within any range between any two of the foregoing numbers.

[0039] In some embodiments, the proximity filter also includes instructions or commands for determining which of the binned variants can be readily identified as germline variants. For example, the binned variants can have corresponding variants present in one or more reference databases and can be identified as germline variants.

[0040] In some embodiments, the proximity filter includes instructions for determining that variants having an allele frequency above a threshold frequency in a sample are germline variants. In some such embodiments, variants having an allele frequency of 0.7, 0.8, 0.9, or 1.0 or greater can be identified as germline variants, although it should be understood that higher or lower allele frequencies are still within the scope of the present disclosure.

[0041] In some embodiments, the proximity filter includes instructions for determining a proximity range of the allele frequency of variants that have not been identified as germline variants. The proximity range of the allele frequency of the variant can include the range of allele frequencies above and below the allele frequency of the variant. In some embodiments, the proximity range is a range having a maximum and a minimum from the allele frequencies of any number of variants within a range of 0.01, 0.02, 0.03, 0.04, 0.05, 0.06, 0.07, 0.08, 0.09, or any number between any two of the aforementioned numbers. For example, for a variant having an allele frequency of 0.2 and a proximity range of 0.05, the minimum and maximum of the proximity range are allele frequencies of 0.15 and 0.25, respectively.

[0042] In some embodiments, the proximity range is determined by the values of two (n) standard deviations of the binomial distribution, assuming that the reinforcing evidence for a given variant arises from a binomial process. For example, for a variant having an allele frequency (x) and coverage (y), the proximity range (z) is as follows: z = n * sqrt(fy * x * (1 - x)) / y

[0043] For example, for a variant with an allele frequency of 0.2 and a coverage / sequencing depth of 100, the proximity range is 0.08, and the minimum and maximum of the proximity range are allele frequencies of 0.12 and 0.28, respectively. In some embodiments, the proximity range is the higher of 0.05 or 2(n) standard deviations from the binomial distribution of the allele frequency of the variant above and below the allele frequency of the variant.

[0044] In some embodiments, if a variant has an allele frequency within the proximity of one or more identified germline variants within the same bin as the variant, the variant can be identified as a germline variant. In some embodiments, if a variant has an allele frequency within the proximity of more than 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 identified germline variants within the same bin as the variant, the variant can be identified as a germline variant. In some embodiments, if a variant has an allele frequency within the proximity of more than five identified germline variants within the same bin as the variant, the variant can be identified as a germline variant. For example, in embodiments where a variant has an allele frequency within the proximity of more than five identified germline variants within the same bin as the variant and the variant is identified as a germline variant: the allele frequency is 0.2, the proximity range is 0.05, thus having a minimum range of 0.15 and a maximum range of 0.25, and the variant binned within the bin representing chromosome 7 is identified as a germline variant, and more than five identified germline variants have allele frequencies within the proximity of the variant and are binned in the bin representing chromosome 7.

[0045] In some embodiments, the proximity filter identifies somatic variants that are variants not identified as germline variants. In some embodiments, the number of somatic variants obtained from sequencing data from a tumor is the tumor mutation burden of the tumor.

[0046] In some embodiments, a database filter or a proximity filter can be applied to a plurality of variants to identify and remove germline variants from the plurality of variants. In some embodiments, the database filter and the proximity filter can be applied sequentially. For example, the output of such a database filter can be used as the input to the proximity filter. Conversely, the output of the proximity filter can be used as the input to the database filter.

[0047] In some embodiments, after performing variant allele frequency analysis in step 125, method 100 further includes step 130 of fragment size distribution. Using genomic coordinates, the fragment size can be inferred using the consensus sequence after read collapse. In some embodiments, the fragment size distribution includes generating a profile of fragment sizes based on variant types of different cell origins, such that different cell origins or different variant types generate distinct fragment size profiles. In some embodiments, the provided fragment sizes are cell lineage dependent.

[0048] In some embodiments, method 100 includes step 135 of identifying cancer variants. Cancer variants can be identified by analyzing the fragment size distribution and extracting the fragment size distribution known to be associated with CH. In some embodiments, the identification of cancer variants includes fitting the fragment size distribution to a likelihood model. In some embodiments, the matched tumor samples are analyzed using method 100 described in FIG. 1, and by matching the results of tumor and cfDNA variants, it enables the construction of fragment size profiles derived from tumor hematopoietic cells, healthy hematopoietic cells, and abnormal hematopoietic cells. To identify somatic mutations from CH, a likelihood ratio test can be performed to fit the observed fragment sizes of different cell origins. In some embodiments, the identification of cancer variants is performed with a sensitivity greater than 75%, for example, 75, 80, 85, 90, 95, 96, 97, 98 or greater than 99%, or within a range defined by any two of the above values. Treatment method

[0049] Some embodiments of the methods and systems include methods of treating a subject having a tumor or a subject suspected of having a tumor. In some such embodiments, the number of cancer variants present in a cfDNA sample can be determined by the methods and systems provided herein. For example, sequence data can be obtained from a cfDNA sample, multiple variants can be identified from the sequence data, a fragment size distribution profile can be established to identify and characterize CH from cancer variants, thereby identifying cancer variants among the multiple variants. In some embodiments, the number of cancer variants obtained from sequencing data from a cfDNA sample is the TMB. In some embodiments, the TMB is calculated as the average number of cancer variants per genomic region, such as variants per 50 kb, 100 kb, 1 Mb, 10 Mb, 100 Mb. The TMB can be sampled by sequencing the entire genome or a portion thereof. For example, a portion of the genome can be sequenced by enriching one or more genomic regions of interest, such as a tumor gene panel, a complete exon, a partial exon, etc.

[0050] Some embodiments of treating a subject having a tumor or a subject suspected of having a tumor can include determining that a cfDNA sample has a TMB that is greater than or equal to a TMB threshold and contacting the tumor with an effective amount of a therapeutic agent. Some embodiments can include treating a subject having a tumor, determining that a cfDNA sample has a TMB that is greater than or equal to a TMB threshold, and administering an effective amount of a therapeutic agent to the subject. In some embodiments, the TMB threshold can be 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or any number within the range between any two of the foregoing numerical values.

[0051] In some embodiments, TMB is calculated by dividing the eligible variants by the effective panel size. Eligible variants include, for example, variants in the coding region, variants not in low-confidence regions, variants having a frequency greater than 0.4% and less than 40%, variants having a coverage greater than 500-fold, single nucleotide variants (excluding multiple nucleotide variants), insertion and deletion variants (Indels), non-synonymous and synonymous variants, and variants having a COSMIC number greater than 50 are excluded, and / or variants having mutations in ‘I’E T2, TP53, DNMT3A, and / or CBL are excluded. The effective panel size can include, for example, the total coding region having a coverage greater than 500-fold.

[0052] Examples of therapeutic agents include chemotherapeutic agents. In some embodiments, the therapeutic agent can include checkpoint inhibitors. Examples of checkpoint inhibitors include CTLA-4 inhibitors, PD-1 inhibitors, and PD-L1 inhibitors. In some embodiments, checkpoint inhibitors include ipilimumab, nivolumab, pembrolizumab, sparta Li zumab (Spartalizumab), atezo Lith mab (Atezolizumab), avelumab, and Du ruva R mab (Durvalumab) can be mentioned. Examples of tumors include large tumors, lung tumors, endometrial tumors, uterine tumors, gastric tumors, melanomas, breast tumors, pancreatic tumors, kidney tumors, bladder tumors, and brain tumors. Further examples of cancers that can be treated with the methods and systems included herein are listed in US Patent Application Publication No. 2018 / 0218789, which is hereby expressly incorporated by reference in its entirety.

[0053] Some embodiments include a computer-based system and a computer-implemented method for performing the methods described herein. In some embodiments, the system can be used to determine a fragment size distribution profile for discriminating CH variants and cancer variants. In some embodiments, the system further comprises a database filter and / or a proximity filter applied to the polymorphism data for identifying and removing germline variants. Some embodiments of the methods and systems provided herein include an electronic system for analyzing polymorphism data. In some such embodiments, the system and computer-implemented method include an analyzer of variant allele frequency and fragment size distribution. Some embodiments can include an informatics module, executed on a processor, adapted to identify a plurality of variants from sequence data from a biological sample, the plurality of variants including CH and cancer variants. Some embodiments provided herein include a computer-implemented method for identifying CH in a plurality of variants. Some such embodiments can include receiving a plurality of variants from sequence data from a biological sample, the plurality of variants possibly including CH and cancer variants. Some embodiments include matching the results of tumor and cfDNA variants to construct fragment size profiles derived from tumor hematopoietic cells, healthy hematopoietic cells, and abnormal hematopoietic cells. In some embodiments, the tumor variants obtained from sequencing data from a cfDNA sample are TMB.

[0054] The system may comprise one or more client components. The one or more client components may include a user interface. The system may comprise one or more server components. The server components may include one or more memory locations. The one or more memory locations can be configured to receive data inputs. The data inputs may include sequencing data. The sequencing data can be generated from a nucleic acid sample from a subject. The system may further comprise one or more computer processors. The one or more computer processors can be operably coupled to the one or more memory locations. The one or more computer processors can be programmed to map the sequencing data to a reference sequence. The one or more computer processors can be further programmed to determine the presence or absence of a plurality of variants from the sequencing data. The one or more computer processors can be further programmed to determine the variant allele frequency. The one or more computer processors can be further programmed to determine a fragment size distribution profile. The one or more computer processors can be further programmed to determine the classification of variants of different origins based on the fragment size distribution. The one or more computer processors can be further programmed to generate an output for display on a screen. The output may include one or more reports identifying CH and / or cancer variants.

[0055] Some embodiments of the method and system may include one or more client components. The one or more client components may include one or more software components, one or more hardware components, or a combination thereof. The one or more client components can access one or more services via one or more server components. The one or more services can be accessed by one or more client components via a network. As used herein, "service" is used to refer to any product, method, function, or use of the system. For example, a user can apply for a genetic test. The application can be made via one or more client components of the system, and a request can be sent to one or more server components of the system via a network. The network can be the Internet, the Internet and / or an extranet, or an intranet and / or an extranet communicating with the Internet. In some cases, the network is a telecommunications and / or data network. The network can include one or more computer servers that enable distributed computing such as cloud computing. The network can, in some cases, implement a peer-to-peer network with the assistance of a computer system, and this peer-to-peer network can enable devices coupled to the computer system to operate as clients or servers.

[0056] Some embodiments of the system may include one or more memory locations such as random access memory, read-only memory, flash memory; an electronic storage unit such as a hard disk; a communication interface such as a network adapter for communicating with one or more other systems; and / or peripheral devices such as a cache, other memory, data storage and / or an electronic display adapter. The memory, storage unit, interface, and / or peripheral devices can communicate with the CPU via a communication bus such as a motherboard. The storage unit can be a data storage unit or a data repository for storing data. In one example, one or more memory locations can store the received sequencing data.

[0057] Some embodiments of the method and system may include one or more computer processors. The one or more computer processors may be operably coupled to one or more memory locations and, for example, can access the stored sequencing data. The one or more computer processors can implement machine-executable code for performing the methods described herein. For example, the one or more computer processors can execute the machine-readable code to map the sequencing data input to a reference sequence and / or identify CH variants and / or cancer variants.

[0058] Some embodiments of the methods and systems provided herein may include machine-executable code or machine-readable code. In some such embodiments, the machine-executable code or machine-readable code may be provided in the form of software. In use, the code may be executed by a processor. Optionally, the code may be retrieved from a storage unit and stored in a memory that is ready for access by the processor. In some embodiments, the electronic storage unit may be excluded, and the machine-executable instructions may be stored in the memory. The code can be pre-compiled and configured for use on a machine having a processor adapted to execute the code, and can also be compiled or interpreted at run-time. The code can be provided in a programming language selected to enable the code to be executed in a pre-compiled form, a form that is compiled or interpreted at run-time.

[0059] Some embodiments of the systems and methods provided herein, such as computer systems, can be implemented in programming. Various aspects of the technology can typically be regarded as a "product" or "article of manufacture" in the form of machine (or processor) executable code and / or associated data held or incorporated in some form of machine-readable medium. The machine-executable code can be stored in an electronic storage unit, such as a memory or a hard disk. The "storage" type of medium can include tangible memory of a computer, a processor, etc., or their associated modules, such as any or all of various semiconductor memories, tape drives, disk drives, etc., which can always provide a persistent storage area for software programming. All or part of the software may, in some cases, be communicated via the Internet or various other electrical communication networks. Such communication can, for example, enable the loading of software from one computer or processor to another, such as from a management server or host computer to an application server's computer platform. Thus, another type of medium that can have software elements includes light waves, radio waves, and electromagnetic waves, such as those used via the physical interface between local devices, wired and optical terrestrial links networks, and various air links. Physical elements that carry waves, such as wired or wireless links, optical links, etc., can also be regarded as media having software. As used herein, unless limited to a fixed tangible "storage" medium, terms such as computer or machine "readable medium" refer to any medium involved in providing instructions to a processor for execution.

[0060] Some embodiments of the methods and systems disclosed herein may comprise or be communicable with one or more electronic displays. The electronic display may be part of a computer system or may be coupled to the computer system directly or via a network. The computer system may comprise a user interface (UI) for providing the various features and functions disclosed herein. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces. The UI may provide an interactive tool for using the methods and systems described herein. As an example, a UI contemplated herein may be a web-based tool through which a healthcare provider may initiate a genetic test, customize a list of genetic variants to be tested, and receive and view a biomedical report.

[0061] Some embodiments of the methods and systems disclosed herein may include a biomedical database, a genomic database, a biomedical report, a disease report, a case management analysis, and a rare variant identification analysis based on one or more databases, one or more assays, data and / or information from one or more data or results, one or more outputs based on or derived from one or more assays, one or more outputs based on or derived from one or more data or results, or a combination thereof.

Example

[0062] Embodiments of the present invention are further specified in the following examples. It should be understood that these examples are given by way of illustration only. From the above considerations and these examples, those skilled in the art can identify the essential features of the present invention and make various changes and modifications to the embodiments of the present invention to adapt to various applications and conditions without departing from its spirit and scope. Therefore, in addition to what is shown and described in this specification, various changes to the embodiments of the present invention will be apparent to those skilled in the art from the foregoing description. Such changes are also intended to be included within the scope of the appended claims. The disclosure of each reference described in this specification is hereby incorporated by reference in its entirety into this specification, and the disclosure referred to in this specification is also incorporated. Example 1 Determination of Variant Allele Frequencies in FFPE vs Plasma Samples

[0063] Sequence data were obtained from cell-free DNA (cfDNA) and matched tumor samples. Samples such as solid tumors and leukemias were collected across four original tissue types by various tumor stages. As shown in Table 1, a total of 85 plasma samples across four tissue types were analyzed, and 15 bladder samples and 32 lung samples matched to FFPE tissues were analyzed. [Table 1]

[0064] Figure 2 shows the determination of variant allele frequencies between FFPE and plasma samples. As shown in Figure 2, among 47 samples with matched FFPE and plasma, 33 COSMIC hot spot variants were detected in plasma. Among the 33 variants, 17 variants were detected in FFPE, with VAF over 3%, 6 with VAF below 3%, and 10 wild-type FFPE were detected. As shown, most of the mutations found only within plasma samples rather than FFPE samples gathered in TP53, DNMT3A, TET2, SF3B1, and CBL, which are known to be related to clonal hematopoiesis (CH). CH mutations were also detected in FFPE samples, and the variant allele frequencies were low.

[0065] Figure 3 shows the comparison of VAF between somatic mutations and CH mutations. As shown in Figure 3, the VAF of somatic mutations was significantly higher in FFPE (p = 2e -5 ), which is likely due to tumor excretion, while the VAF of CH mutations was significantly higher in plasma samples (p = 0.01). Example 2 Fragment size distribution

[0066] Using the VAF determination shown in Example 1, a fragment size profile was constructed. The fragment size profile is derived from tumor hematopoietic cells, healthy hematopoietic cells, and abnormal hematopoietic cells. The three main types of variants are present in plasma, somatic, CH, and germline. These are derived from different tissue origins, as shown in Table 2.

Table 2

[0067] Differences in fragment sizes between variants and different tissue origins were determined by extracting fragments from sequencing data carrying the mutant alleles. The results were tabulated across all samples. As shown in Figure 4, the fragment size distribution of the mutations was found to differ depending on the origin. The size distribution of fragments carrying somatic mutations from solid tumors (peak at 138 bp) was shifted compared to fragments carrying somatic mutations from CH or leukemia (peak at 166 bp). No significant difference in size distribution was observed between fragments carrying somatic mutations and healthy hematopoietic cells (p-value = 0.86).

[0068] As shown in Figure 5, mutations of different origins were classified by fragment size distribution. By mixing fragments of different origins, 10,000 CH mutations or somatic mutations of different VAFs were simulated in silico at 2000-fold coverage. By fitting the fragment size distribution with a likelihood model, sensitivities of 81.5%, 92.5%, 98.3%, and 99.8% were achieved, and specificities of 82%, 92.5%, 97.5%, and 99.9% were achieved for CH mutations of 1%, 2.5%, 5%, and 10% respectively.

[0069] These examples demonstrate that the fragment size distribution of cfDNA released by malignant or healthy hematopoietic cells is different from that of cfDNA released by solid tumors. In addition, the fragment size distribution can be used to distinguish somatic mutations from different cell origins. Example 3 Clonal Hematopoietic Variants in cfDNA

[0070] Forty pairs of cfDNA and buffy coat (white blood cell) DNA were profiled using the method described in Figure 1. Variants were observed as non-germline lineages (with low VAF) in both cfDNA and buffy coat. The results included 106 variants, of which 92 were non-synonymous and 14 were synonymous. As shown in Figure 6, the VAF determined for cfDNA correlates with the VAF determined for the buffy coat. Example 4 Measurement of Tumor Mutation Burden

[0071] Using the samples analyzed in Example 3, the tumor mutation burden was measured. The samples included 40 pairs of cfDNA and buffy coat DNA, which were profiled using the method described in FIG. 1.

[0072] The untreated TMB was calculated by dividing the number of eligible variants by the effective panel size. Eligible variants included variants in the coding region, variants not in low-confidence regions, variants with a frequency greater than 0.4% and less than 40%, variants with coverage greater than 500-fold, single nucleotide variants (SNVs), insertion and deletion variants (Indels), non-synonymous and synonymous variants. Variants with a COSMIC number greater than 50 were excluded, multiple nucleotide variants (MNVs) were excluded, and variants having mutations in 'I'E T2, TP53, DNMT3A, and / or CBL were excluded. The effective panel size included the total coding region with coverage greater than 500-fold. In this example, the total number of variants included 1025, the number of variants after germline filtering was 121, the number of variants in the eligible region was 86, the number of SNVs and Indels in the eligible region was 81, the number of variants after COSMIC removal was 80, the number of variants with approximately 0.4% was 78, and the number of variants excluding the TET2, TP53, DNMT3A, and CBL genes was 75. Therefore, the total number of eligible variants was 76. The effective panel size was 1.307291 Mb. The untreated TMB was 76 / 1.30729 = 57.4 mutations / Mb. The adjusted TMB was (57.37055 - 1.5) / 0.91 = 61.4.

[0073] As shown in FIG. 7A, compared to the total blood cell TMB (T / N TMB), the TMB of the tumor-limited TMB (T-limited TMB) had an R of 0.91 2correlates, and the tumor - restricted TMB is higher than the normal TMB of the tumor due to the CH variant. As shown in Figure 7B, compared with the T - restricted TMB after clonal hematopoiesis adjustment, the TMB in the T / N TMB has an R of 0.934 2 correlates, and the tumor - restricted TMB is similar to the normal TMB of the tumor.

[0074] As used herein, the term "comprising" is synonymous with "including," "containing," or "characterized by," is inclusive or non - limiting, and does not exclude additional, unrecited elements or method steps.

[0075] The foregoing description discloses some methods and materials of the present invention. The present invention is amenable to modifications of methods and materials, as well as changes in manufacturing methods and equipment. Such modifications will be apparent to those skilled in the art in view of the present disclosure or the practice of the invention disclosed herein. Accordingly, the present invention is not intended to be limited to the specific embodiments disclosed herein, but rather is intended to cover all modifications and alternatives falling within the true scope and spirit of the invention.

[0076] All references cited herein, including but not limited to published and unpublished applications, patents, and literature references, are hereby incorporated by reference in their entirety and form a part of this specification. If the publications and patents or patent applications incorporated by reference conflict with the disclosure herein, this specification is intended to supersede and / or be superior to such conflicting materials. In certain embodiments, for example, the following items are provided. (Item 1) A method for distinguishing cancer variants from hematopoietic cell variants in a circulating tumor DNA (ctDNA) sample, comprising: (a) obtaining a ctDNA sample comprising a plurality of cell - free DNA (cfDNA) fragments, or having said obtained sample; (b) extracting cfDNA fragments containing a plurality of variants from said sample; (c) Performing molecular profiling for each of the plurality of variants, (i) determining a variant allele frequency (VAF) for each of the plurality of variants, including cancer variants and hematopoietic cell variants; (ii) generating a fragment size distribution profile to identify hematopoietic cell variants; and (d) identifying cancer variants by removing the identified hematopoietic cell variants. (Item 2) The method according to Item 1, further comprising removing germline variants from the plurality of variants. (Item 3) The method according to Item 2, wherein the germline variants are removed by applying a database filter or a proximity filter to the plurality of variants. (Item 4) The method according to Item 1, further comprising sequencing the cfDNA fragments to obtain sequence data. (Item 5) The method according to Item 4, further comprising aligning the sequence data with a reference sequence and identifying variants in the sequence data. (Item 6) The method according to Item 1, wherein the ctDNA sample is derived from a solid sample or a plasma sample. (Item 7) The method according to Item 6, wherein the solid sample is fixed. (Item 8) The method according to Item 6, wherein the sample contains tumor cells. (Item 9) The method according to Item 6, wherein the sample contains a serum sample, a fecal sample, a blood sample, or a tumor sample. (Item 10) The method according to Item 1, wherein the method is a computer-implemented method. (Item 11) A method for determining the tumor mutation burden of a tumor, Obtaining sequence data from a biological sample containing tumor cells, Determining a plurality of variants from the sequence data, Determining the number of cancer variants among the plurality of variants by the method according to item 1, wherein the number of cancer variants is equal to the tumor mutation burden of the tumor, and determining the number of cancer variants, a method comprising. (Item 12) A method for treating a tumor, Determining a tumor having a tumor mutation burden of 10 or more cancer variants by the method according to item 11, Treating the tumor by administering an effective amount of a checkpoint inhibitor, a method comprising. (Item 13) The method according to item 12, wherein the tumor is selected from the group consisting of a large tumor, a lung tumor, an endometrial tumor, a uterine tumor, a gastric tumor, a melanoma, a breast tumor, a pancreatic tumor, a kidney tumor, a bladder tumor, and a brain tumor. (Item 14) The method according to item 12, wherein the checkpoint inhibitor is selected from the group consisting of a CTLA-4 inhibitor, a PD-1 inhibitor, and a PD-L1 inhibitor. (Item 15) The checkpoint inhibitor is ipilimumab, nivolumab, pembrolizumab, sparta Li zumab, atezo Lith mab, avelumab, and Du luba R The method according to item 12, selected from the group consisting of mab. (Item 16) An electronic system for analyzing genetic polymorphism data, An informatics module adapted to identify a plurality of variants from sequence data from a cfDNA sample, which is executed on a processor, wherein the plurality of variants include cancer variants and hematopoietic cell variants, an informatics module; An analyzer for performing molecular profiling for each of the plurality of variants, configured to determine a variant allele frequency (VAF) for each of the plurality of variants and further configured to generate a fragment size distribution profile, and an analyzer An analyzer for identifying cancer variants by removing identified hematopoietic cell variants A display module adapted to return variants that are not removed from the plurality of variants. An electronic system comprising (Item 17) The system according to item 16, further comprising a database filter module or a proximity filter module configured to remove germline variants from the plurality of variants.

Claims

**Claim 1** A method for distinguishing somatic variants derived from solid tumor cells from somatic variants derived from hematopoietic cells related to clonal hematopoiesis in a circulating tumor DNA (ctDNA) sample, comprising: (a) extracting cfDNA fragments containing a plurality of variants from a ctDNA sample containing a plurality of cell-free DNA (cfDNA) fragments; (b) excluding variants having a catalog of somatic mutations in more than 50 cancers (COSMIC) from the plurality of variants; (c) performing molecular profiling for each of the plurality of variants based on sequence data from the ctDNA sample containing the plurality of cell-free DNA (cfDNA) fragments of (a), comprising: (i) determining a variant allele frequency (VAF) for each of the plurality of variants, including the somatic variants derived from the solid tumor cells and the somatic variants derived from hematopoietic cells related to clonal hematopoiesis; (ii) generating a fragment size distribution profile based on the determined VAF and identifying the somatic variants derived from hematopoietic cells related to the clonal hematopoiesis; and (d) identifying the somatic variants derived from the solid tumor cells by removing the sequence data of the somatic variants derived from hematopoietic cells related to the clonal hematopoiesis from the sequence data of the plurality of variants, wherein the somatic variants derived from hematopoietic cells related to the clonal hematopoiesis include mutations in TET2, TP53, DNMT3A, and / or CBL. A method comprising the above steps. **Claim 2** The method according to claim 1, further comprising removing sequence data of germline variants from the sequence data of the plurality of variants. **Claim 3** The method according to claim 2, wherein the sequence data of the germline variants is removed by applying a database filter or a proximity filter to the sequence data of the plurality of variants. **Claim 4** The method according to claim 1, further comprising sequencing the cfDNA fragments to obtain the sequence data. **Claim 5** The method according to claim 4, further comprising aligning the sequence data with a reference sequence and identifying variants in the sequence data. **Claim 6** The method according to claim 1, wherein the cDNA sample is derived from a solid sample or a plasma sample.

7. The method according to claim 6, wherein the solid sample is fixed.

8. The method according to claim 6, wherein the sample contains tumor cells.

9. The method according to claim 6, wherein the sample contains a serum sample, a fecal sample, a blood sample, or a tumor sample.

10. The method according to claim 1, wherein the method is a computer-implemented method.

11. A method for determining the tumor mutation burden of a tumor, comprising: obtaining sequence data from a biological sample containing tumor cells; determining a plurality of variants from the sequence data; determining the number of somatic variants derived from the solid tumor cells in the plurality of variants by the method according to claim 1, wherein the number of somatic variants derived from the solid tumor cells is equal to the tumor mutation burden of the tumor; A method comprising the above steps.

Citation Information

Patent Citations

  • Systems and methods for analyzing nucleic acids

    JP2018513508A

  • Methods and systems for assessing tumor mutation burden

    JP2019512218A

  • Ultra-sensitive detection of circulating tumor DNA through genome-wide integration

    WO2019169042A1