Cell-free DNA methylation testing for breast cancer
A cell-free DNA methylation-based method using machine learning creates an MRD signature to enhance the detection of metastatic breast cancer recurrence risk, addressing the limitations of existing tests by improving sensitivity and specificity.
Patent Information
- Application Number
- JP2025528913
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-22
- Filing Date
- 2023-11-22
- Publication Date
- 2025-12-16
AI Technical Summary
Current diagnostic methods for metastatic breast cancer (MBC) and minimal residual disease (MRD) are not sensitive or specific enough, and existing tests like the 21-gene Oncotype DX, 70-gene Mamaprint, and 76-gene Rotterdam signatures fail to accurately predict disease recurrence, especially in high-risk subtypes such as HER2-positive and triple-negative breast cancer.
Development of a method using cell-free DNA methylation patterns to create a minimal residual disease (MRD) signature through machine learning, trained on cancer and non-cancer samples, to identify patients at high risk of relapse by analyzing methylation patterns in cell-free DNA.
The method provides a more sensitive and specific means to detect MRD prior to clinical relapse, allowing for early identification of patients at high risk of recurrence using minimally invasive techniques.
Smart Images

Figure 2025540676000001_ABST
Abstract
Description
[Technical Field]
[0001] Related Applications This application claims priority under 35 U.S.C. § 119(e) to U.S. Provisional Patent Application No. 63 / 384,731, filed November 22, 2022, which is incorporated herein by reference.
[0002] Federal Funding Statement This invention was made with government support under Grant No. CA201352 awarded by the National Institutes of Health. The government has certain rights in this invention. [Background technology]
[0003] Despite recent advances in breast cancer (BC) screening, diagnosis, and treatment, some patients still develop metastasis and die from the disease. The 5-year overall survival (OS) rate for patients with metastatic breast cancer is less than 25%. The risk of metastasis increases with tumor size, lymph node involvement, lack of estrogen receptor (ER) expression, overexpression of human epidermal growth factor receptor 2 (HER2), and high histologic differentiation (grade).
[0004] Clinical tests based on molecular profiles have improved our ability to stratify patients based on their risk of recurrence. The most widely used multigene predictive classifier is the 21-gene Oncotype DX signature (Exact Sciences, USA), while others include the 70-gene Mamaprint signature (Agendia, Netherlands), the 76-gene Rotterdam signature, and the PAM50 specific classifier (NanoString, USA). Although these tests are based on large sets of specifically selected genes, none of them can accurately predict disease recurrence. Furthermore, these tests reflect the instantaneous state of each individual cancer at the time of biopsy and are not intended to be used to monitor changes in the cancer's molecular profile over time. All of these tests are recommended for ER-positive tumors but are not approved for high-risk BC subtypes such as HER2-positive and triple-negative breast cancer (TNBC). An alternative strategy to stratify patients at high risk of systemic recurrence is response to neoadjuvant therapy. However, meta-analyses have not demonstrated a correlation between pathologic complete response to treatment and disease-free survival (DFS) or OS.
[0005] Metastatic breast cancer (MBC) is an incurable disease affecting 10–15% of breast cancer patients. MBC arises from disseminated cells derived from the primary tumor mass before treatment and / or from minimal residual disease remaining after treatment. If these cells persist after systemic chemotherapy (adjuvant or neoadjuvant), they can lead to recurrence months or even years after primary treatment. Historically, the only way to detect recurrence was through local recurrence or metastatic nodules. While whole-body CT scans may be indicated in high-risk patients, in most MBC patients, the first indication of recurrence is symptoms caused by organ damage due to local metastatic growth. Such metastases are well-documented and often difficult to treat, even with high-dose chemotherapy and surgical intervention. Recently, personalized tumor-based circulating tumor (ct) DNA molecular residual disease (MRD) testing in breast cancer has been developed to support critical treatment decisions. The Signatera™ Residual Disease Test is a custom-tailored blood test for individuals diagnosed with breast cancer or other solid tumors. Signatera™ can detect molecular residual disease (MRD) in the form of circulating tumor DNA, although the test uses tumor-informed mutation data and remains limited in sensitivity and specificity.
[0006] Therefore, there is a need for new diagnostic methods for MBC and MRD that are more sensitive and specific than current methods. The present disclosure meets these needs. Summary of the Invention
[0007] Disclosed herein are tools and assays specifically designed to detect evidence of MRD prior to clinical relapse using cost-effective, minimally invasive techniques that can be applied after treatment completion, at diagnosis, and repeatedly during treatment. Embodiments of the present disclosure can be used to identify and measure methylation patterns in cell-free (cf) DNA to develop an MRD signature. This signature identifies patients at highest risk of relapse.
[0008] Monitoring MRD using cfDNA has emerged as a promising blood-based biomarker strategy. Evidence of MRD is considered a prognostic marker for identifying individuals at high risk of recurrence. cfDNA is an excellent analytical substrate for MRD monitoring because it: 1) contains rich information from multiple tissue types; 2) is minimally invasive to the patient, requiring only standard venous blood sampling; and 3) can be easily repeated over time. Furthermore, cfDNA may more accurately reflect the tissue of origin, whereas traditional biopsies can be biased by subclonal and tumor heterogeneity.
[0009] Therefore, the present disclosure provides panel assays and methods for using them.In some embodiments, a method is provided for determining whether a subject has minimal residual disease (MRD), the method comprises: a) training a machine learning model to develop MRD signature, the machine learning program is trained using target regions from cancer samples and corresponding target regions from non-cancer samples, and the MRD signature is based on the comparison of the methylation pattern of the target regions of cancer samples and the methylation pattern of corresponding target regions of non-cancer samples; b) determining the methylation pattern of the target regions of cell-free deoxyribonucleic acid (cfDNA) samples obtained from the subject; c) applying the MRD signature to the methylation pattern of the target regions of the cfDNA obtained from the subject; And d) determining whether the subject has MRD based on the MRD signature.
[0010] These and other features and advantages of the present invention will be more fully understood by reference to the following detailed description of the invention in conjunction with the appended claims, which should be understood to be defined by the description and not by any specific consideration of the features and advantages set forth herein.
[0011] The following drawings form part of the specification and are included to further demonstrate certain embodiments or various aspects of the present invention. In some cases, embodiments of the present invention can be best understood by reference to the accompanying drawings in conjunction with the detailed description set forth herein. The description and accompanying drawings may emphasize particular examples or aspects of the present invention. However, one of ordinary skill in the art will understand that some of the examples or aspects of the present invention can be used in combination with other examples or aspects. [Brief explanation of the drawings]
[0012] [Figure 1] An example of CpG methylation status in a hypothetical genomic region is shown, where solid dots represent methyl groups and hollow dots indicate the absence of methyl groups. [Figure 2] FIG. 1 shows coverage and beta values evaluated by FLAME and BSSEQ. [Figure 3] FIG. 1 shows a comparison of the number of synthetic methylated fragments counted by FLAME and the expected frequency. [Figure 4] CSF detected by in silico spiking-in of MBC cfDNA into healthy cfDNA. [Figure 5A-B] WGBS demonstrates that the methylation profile of MBC differs from that of DFS and healthy. A) Heat scatter plot of methylation rates for the three study groups by pairwise comparison. The numbers in the upper right corner indicate Pearson correlation coefficients. The histogram on the diagonal shows the frequency of methylation rates per CpG per pool. MBC is shifted to the left compared to the DFS and healthy groups, indicating genome-wide hypomethylation. B) Principal component analysis (PCA) of the methylation profile of each cfDNA pool. Samples that are close to each other in clustering or PCA have similar methylation profiles. (See, e.g., Legendre et al. Clin Epigenetics 2015 Sep 16;7(1):100. doi:10.1186 / s13148-015-0135-8) [Figure 6]Receiver operating characteristic (ROC) curves showing the performance of the random forest classifier model on a training set of 30 samples demonstrate high sensitivity and specificity in classifying MBC from healthy patients using cfDNA. Area under the curve (AUC) is annotated. [Figure 7] Evidence of MRD in cfDNA collected after neoadjuvant therapy and postoperatively (color-coded) is shown. Each plot is subdivided by patient outcome: DFS (disease-free survivor), REC (relapse), and NDF (disease never went into remission). A) Probability scores assessed with the RF model (Figure 3) show little change between time points and minimal differences between samples. B) The number of cancer-specific fragments (CSF) per sample shows a significant decrease in DFS, an increase in relapse samples, and a minor increase in samples from patients who never went into remission. DETAILED DESCRIPTION OF THE INVENTION
[0013] definition The following definitions are included for a clear and consistent understanding of the specification and claims. As used herein, the described terms have the following meanings: All other terms and expressions used herein have the ordinary meanings understood by those skilled in the art. These ordinary meanings can be obtained by consulting technical dictionaries (e.g., Hawley's Condensed Chemical Dictionary, 14th Edition, by R.J. Lewis, John Wiley & Sons, New York, NY, 2001 or Singleton, et al., Dictionary of Microbiology and Molecular Biology, 2nd ed., John Wiley and Sons, New York (1994), and Hale & Markham, The Harper Collins Dictionary of Biology, Harper Perennial, NY (1991). Common laboratory techniques (such as DNA extraction, RNA extraction, cloning, cell culture, etc.) are well known in the art and are described, for example, in Molecular Cloning: A Laboratory Manual, J. Sambrook et al., 4th edition, Cold Spring Harbor Laboratory Press, 2012.
[0014] References in the specification to "one embodiment," "an embodiment," or the like indicate that the described embodiment includes a particular aspect, feature, structure, moiety, or characteristic, but not all embodiments necessarily include that aspect, feature, structure, moiety, or characteristic. Moreover, such phrases may, but do not necessarily, refer to the same embodiment referenced elsewhere in the specification. Furthermore, when a particular aspect, feature, structure, moiety, or characteristic is described in connection with an embodiment, it is within the knowledge of one of ordinary skill in the art to associate or connect that aspect, feature, structure, moiety, or characteristic with other embodiments, whether or not explicitly described.
[0015] As used herein, where the term "comprising" is used, it is contemplated that the terms "consisting of" or "consisting essentially of" may be used instead. As used herein, "comprising" is synonymous with "including," "containing," or "characterized by" and is inclusive or open-ended, and does not exclude additional, unrecited elements or method steps. As used herein, "consisting of" excludes elements, steps, or ingredients not expressly recited in the embodiment. As used herein, "consisting essentially of" does not exclude materials or steps that do not materially affect the basic and novel characteristics of the embodiment. In any instance herein, any of the terms "comprising," "consisting essentially of," and "consisting of" may be substituted for either of the other two terms. The disclosure illustratively set forth herein may suitably be practiced in the absence of any element or limitation not specifically disclosed herein.
[0016] The singular forms "a," "an," and "the" include the plural unless the context clearly dictates otherwise. Thus, for example, a reference to a "compound" includes a plurality of such compounds, and reference to compound X includes a plurality of compounds X. Furthermore, it should be noted that the claims may be drafted to exclude any optional element. Thus, this statement is intended to serve as antecedent basis for using exclusive terminology (e.g., "solely," "only," etc.) in connection with any element described herein and / or in connection with the recitation of claim elements or the use of "negative" limitations.
[0017] The term "and / or" refers to any one of the items, a combination of items, or all of the items associated with the term. The terms "one or more" and "at least one" are readily understood by those skilled in the art when read in their context. For example, the term can mean 1, 2, 3, 4, 5, 6, 10, 100, or an upper limit of about 10, 100, or 1000 times the stated lower limit. For example, one or more substituents on a phenyl ring refers to one to five substituents on the ring.
[0018] As will be understood by those of ordinary skill in the art, all numerical values, including those expressing quantities of ingredients, molecular weights, reaction conditions, and like properties, are approximations and are understood to be modified in all cases as necessary by the use of the term "about." These values may vary depending upon the desired properties sought by the skilled artisan using the teachings set forth herein. It is also understood that such values contain inherent variation necessarily resulting from the standard deviation found in their respective testing measurements. When values are expressed as approximations, by use of the antecedent "about," it will be understood that the particular value without the modifier "about" also forms another embodiment.
[0019] The terms "about" and "approximately" are used interchangeably. Both terms can refer to a variation of ±5%, ±10%, ±20%, or ±25% of the specified value. For example, "about 50%" can, in some embodiments, include a variation of 45% to 55%, or a range otherwise defined in a particular claim. In the case of integer ranges, the term "about" can include integers one or two integers greater and / or less than the specified integers at both ends of the range. Unless otherwise stated herein, the terms "about" and "approximately" are intended to include values (e.g., weight percentages) near the specified range that are functionally equivalent for individual components, compositions, or embodiments. The terms "about" and "approximately" can also modify the endpoints of a stated range, as explained in the paragraph above.
[0020] As will be understood by those skilled in the art, for all purposes, particularly for purposes of providing a written description, all ranges described herein encompass all possible subranges and combinations of subranges, including the individual values (especially integer values) that make up the range. Thus, it is understood that each unit between two specific units is also disclosed. For example, if 10 to 15 is disclosed, 11, 12, 13, and 14 are also disclosed individually and as part of a range. A described range (e.g., weight percent or carbon group) includes each specific value, integer, decimal, or identity within the range. A described range is readily recognizable as fully describing and enabling the same range to be divided into at least equal halves, thirds, quarters, fifths, or tenths. As a non-limiting example, each range discussed herein can be readily divided into a lower third, middle third, upper third, etc. As one skilled in the art would understand, expressions such as "up to," "at least," "greater than," "less than," "more than," "or greater than or equal to," and the like, are inclusive of the stated numerical values, and these terms refer to ranges that can be divided into subranges as described above. For the same reason, all ratios described herein include all subratios that fall within the broader ratio. Accordingly, specific values recited for radicals, substituents, and ranges are for illustrative purposes only and do not exclude other defined or other values within the defined ranges of radicals and substituents. Furthermore, it is understood that the endpoints of each range are significant not only in relation to the other endpoint, but also independently of the other endpoint.
[0021] The present disclosure provides ranges, limits, and deviations for variables such as volume, mass, percentages, and ratios. Those skilled in the art will understand that a range from, for example, "number 1" to "number 2" refers to a continuous range of numbers, including integers and decimals. For example, 1 to 10 refers to 1, 2, 3, 4, 5, ..., 9, 10. It also refers to 1.0, 1.1, 1.2, 1.3, ..., 9.8, 9.9, 10.0, and further refers to 1.01, 1.02, 1.03, .... When a disclosed variable is a number less than "number 10," it refers to a continuous range, including integers and decimals less than number 10, as explained above. Similarly, when a disclosed variable is a number greater than "number 10," it refers to a continuous range, including integers and decimals greater than number 10. These ranges may be modified by the term "about," as explained above.
[0022] As used herein, the term "substantially" is a broad term and is used in its ordinary sense, without limitation, including most, but not necessarily all, of what is specified. For example, the term may refer to a numerical value that is not a full 100% numerical value. The full numerical value may be less than about 1%, about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, about 10%, about 15%, or about 20%.
[0023] As used herein, the term "portion" or "part thereof" refers to consecutive nucleotides of the sequence of the specific region. A portion according to the present invention comprises or consists of at least 15 or 20 consecutive nucleotides, preferably at least 100, 200, 300, 500, or 700 consecutive nucleotides, and more preferably at least 1, 2, 3, 4, or 5 consecutive kb of the specific region. For example, a portion may comprise or consist of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 consecutive kb of the specific region.
[0024] Those skilled in the art will readily recognize that where members are grouped in a common manner (e.g., in the case of Markush groups), the invention includes not only the entire group recited as a whole, but also each individual member of the group and all possible subgroups of the primary group. Moreover, for all purposes, the invention encompasses not only the primary group but also the primary group absent one or more group members. Thus, the invention contemplates the explicit exclusion of one or more of the stated group members. Thus, for any of the disclosed categories or embodiments, a condition may be applied that excludes one or more of the stated elements, species, or embodiments from such category or embodiment, e.g., as an explicit negative limitation.
[0025] "Contacting" means the act of touching, touching, or bringing into close proximity or proximity, including at the cellular or molecular level, e.g., in solution, in a reaction mixture, in vitro, or in vivo, to bring about, e.g., a physiological response, a chemical response, or a physical change.
[0026] An "effective amount" refers to an amount effective to treat a disease, disorder, and / or condition or to produce a described effect. For example, an effective amount can be an amount effective to reduce the progression or severity of the condition or symptom being treated. Determining a therapeutically effective amount is within the capabilities of one of ordinary skill in the art. The term "effective amount" is intended to encompass an amount of a compound described herein, or an amount of a combination of compounds described herein, e.g., an amount effective to treat or prevent a disease or disorder or treat a symptom of a disease or disorder in a host. Thus, an "effective amount" generally refers to an amount that produces a desired effect.
[0027] Alternatively, the term "effective amount" or "therapeutically effective amount," as used herein, refers to a sufficient quantity of an agent or composition or combination of compositions administered to relieve to some extent one or more of the symptoms of the disease or condition being treated. The result can be a reduction and / or alleviation of the signs, symptoms, or causes of a disease, or any other desired alteration of a biological system. For example, an "effective amount" for therapeutic use is an amount of a composition comprising a compound disclosed herein that results in a clinically significant reduction in a symptom of the disease. An appropriate "effective" amount in any individual case can be determined using techniques such as a dose escalation study. Doses can be administered in one or more administrations. However, the precise determination of an effective dose can be based on patient-specific factors, including, but not limited to, the patient's age, size, type or extent of disease, stage of disease, route of administration of the composition, type or extent of concurrent adjunctive therapy, ongoing disease process, and the type of treatment desired (e.g., active versus conventional therapy).
[0028] The terms "treating," "treat," and "treatment" include (i) preventing the occurrence of a disease, pathological, or medical condition (e.g., prophylaxis), (ii) inhibiting or halting the progression of a disease, pathological, or medical condition, (iii) relieving a disease, pathological, or medical condition, and / or (iv) alleviating symptoms associated with a disease, pathological, or medical condition. Thus, the terms "treat," "treatment," and "treating" extend to prophylaxis and can include preventing, prevention, preventing, lowering, stopping, or reversing the progression or severity of the condition or symptom being treated. Thus, the term "treatment" can include medical, therapeutic, and / or prophylactic administration, as appropriate.
[0029] As used herein, "subject" or "patient" refers to an individual who has or is at risk for a disease or other malignancy. A patient can be human or non-human and can include, for example, animal strains or species used as "model systems" for research purposes (e.g., the mouse model described herein). Similarly, a patient can include an adult or minor (e.g., a child). Furthermore, a patient can refer to any organism, preferably a mammal (e.g., human or non-human), that may benefit from the administration of a composition contemplated herein. Examples of mammals include any organism belonging to the mammalian class, including, but not limited to, humans, non-human primates such as chimpanzees, and other ape and monkey species; livestock (cows, horses, sheep, goats, pigs); pets (rabbits, dogs, cats); and laboratory animals (rodents such as rats, mice, and guinea pigs). Examples of non-mammals include, but are not limited to, birds, fish, and the like. In one embodiment of the methods provided herein, the mammal is a human.
[0030] As used herein, the terms "providing," "administering," and "introducing" are used interchangeably and refer to placing a compound of the present disclosure into a subject by a method or route that results in the compound being at least partially localized at a desired site. The compound can be administered by any suitable route that delivers it to the desired site in the subject.
[0031] The terms "inhibit," "inhibiting," and "inhibition" refer to slowing, stopping, or reversing the growth or progression of a disease, infection, condition, or group of cells. Inhibition can be, for example, greater than about 20%, 40%, 60%, 80%, 90%, 95%, or 99% compared to growth or progression in the absence of treatment or contact.
[0032] The term "amplicon" refers to a nucleic acid product resulting from amplification of a target nucleic acid sequence. Amplification is often performed by PCR. Amplicon sizes range from 20 base pairs to 15,000 base pairs for long-range PCR, but are typically 100 to 1,000 base pairs for bisulfite-treated DNA used in methylation analysis.
[0033] "Amplification" refers to an increase in the copy number of a nucleic acid molecule. The resulting amplification product is called an "amplicon." Amplification of a nucleic acid molecule (such as a DNA or RNA molecule) refers to the use of a technique to increase the copy number of the nucleic acid molecule in a sample. An example of amplification is polymerase chain reaction (PCR), which is a method in which a sample is contacted with a pair of oligonucleotide primers under conditions in which the primers hybridize to a nucleic acid template in the sample. Amplification products can be characterized by techniques such as electrophoresis, restriction enzyme cleavage patterns, oligonucleotide hybridization or ligation, and / or nucleic acid sequencing. In some embodiments, the methods provided herein can include producing amplified nucleic acids under isothermal or thermally alternating conditions.
[0034] The term "biological sample" refers to a sample obtained from an individual. As used herein, biological samples include all clinical samples containing genomic DNA (such as cell-free genomic DNA) useful for cancer diagnosis and prognosis, including, but not limited to, cells, tissues, and bodily fluids (blood, blood derivatives and fractions (such as serum or plasma), oral epithelium, saliva, urine, stool, bronchial aspirate, sputum, biopsies (such as tumor biopsies), and CVS samples). A "biological sample" obtained from or derived from an individual includes a sample that has been appropriately processed after being obtained from the individual (e.g., processed to separate genomic DNA for bisulfite treatment).
[0035] The term "bisulfite treatment" refers to treating DNA with bisulfite or its salts (e.g., sodium bisulfite (NaHSO3)). Bisulfite reacts readily with the 5,6-double bond of cytosine but poorly with methylated cytosine. Cytosine reacts with the bisulfite ion to form a sulfonated cytosine reaction intermediate, which then undergoes deamination to produce sulfonated uracil. The sulfonate group can be removed under alkaline conditions to form uracil. Uracil is recognized as thymine by polymerases, and amplification results in adenine-thymine base pairs rather than cytosine-guanine base pairs.
[0036] The term "cancer" refers to a biological state in which a malignant tumor or other neoplasm has undergone characteristic anaplasia characterized by loss of differentiation, increased rate of proliferation, invasion of surrounding tissues, and the ability to metastasize. A neoplasm is a new, abnormal growth, particularly a new growth of tissue or cells, where the growth is uncontrolled and progressive. A tumor is one example of a neoplasm. Non-limiting examples of types of cancer include lung cancer, stomach cancer, colon cancer, breast cancer, uterine cancer, bladder cancer, head and neck cancer, kidney cancer, liver cancer, ovarian cancer, pancreatic cancer, prostate cancer, and rectal cancer, among others.
[0037] The terms "polynucleotide" and "nucleic acid" are used interchangeably and refer to at least two or more ribo- or deoxyribonucleic acid base pairs (nucleotides) linked by a phosphate bond or equivalent bond. Nucleic acids include polynucleotides and polynucleosides. Nucleic acids include single molecules, double molecules, triplex molecules, circular molecules, or linear molecules. Examples of nucleic acids include, but are not limited to, RNA, DNA, cDNA, genomic nucleic acids, naturally occurring nucleic acids, and non-naturally occurring nucleic acids such as synthetic nucleic acids. Short nucleic acids and polynucleotides (e.g., 10-20, 20-30, 30-50, or 50-100 nucleotides) are typically referred to as single- or double-stranded DNA "oligonucleotides" or "probes."
[0038] The term "DNA (deoxyribonucleic acid)" refers to the long, chain-like polymer that comprises the genetic material of most living organisms. The repeating units of a DNA polymer are four different types of nucleotides, each containing one of four bases (adenine, guanine, cytosine, or thymine) attached to a deoxyribose sugar with a phosphate group attached. Triplets of nucleotides (called codons) code for each amino acid in a polypeptide or a stop signal. The term codon is also used for the corresponding (and complementary) sequence of three nucleotides in the mRNA into which the DNA sequence is transcribed.
[0039] "Cell-free DNA" refers to DNA that is not contained within intact cells, such as DNA found in plasma or serum.
[0040] A "target nucleic acid molecule" refers to a nucleic acid molecule whose detection, amplification, quantification, qualitative detection, or a combination thereof is desired. The nucleic acid molecule need not be in a purified form. A variety of other nucleic acid molecules may be present along with the target nucleic acid molecule. For example, the target nucleic acid molecule may be a specific nucleic acid molecule whose amplification and / or methylation status is desired to be assessed. If necessary, purification or isolation of the target nucleic acid molecule can be performed by methods well known to those skilled in the art, such as using commercially available purification kits.
[0041] "Methylation level" refers to the methylation status (methylated or unmethylated) of cytosine nucleotides at one or more CpG sites within a genomic sequence.
[0042] A "CpG site" refers to a dinucleotide DNA sequence containing a cytosine followed by a guanine in the 5' to 3' direction. The cytosine nucleotide at a CpG site in genomic DNA is targeted by intracellular methyltransferases and can have a methylation state of methylated or unmethylated. Reference to a "methylated CpG site" or similar term refers to a CpG site in genomic DNA that has a 5-methylcytosine nucleotide.
[0043] As used herein, "sequence identity" or "identity," in the context of two nucleic acid or polypeptide sequences, refers to a specified percentage of identical residues in two sequences when aligned for maximum matching within a specified comparison window, as determined by a sequence comparison algorithm or visual inspection. When percentage sequence identity is used with respect to proteins, it is recognized that non-identical residue positions typically differ by conservative amino acid substitutions (substitution of an amino acid residue with another that has similar chemical properties (e.g., charge or hydrophilicity), thus not altering the functional properties of the molecule). When sequences differ by conservative substitutions, the percentage of sequence identity may be adjusted upward to correct for the conservatism of the substitution. Sequences that differ by such conservative substitutions are said to have "sequence similarity" or "similarity." Methods for making this adjustment are well known to those of skill in the art. Typically, this involves scoring conservative substitutions as partial rather than complete mismatches, thereby increasing the percentage of sequence identity. For example, conservative substitutions are given a score between 0 and 1, where identical amino acids are assigned a score of 1 and non-conservative substitutions are assigned a score of 0. Scoring of conservative substitutions is calculated, for example, as implemented in the program PC / GENE (Intelligenetics, Mountain View, Calif.).
[0044] As used herein, "percent sequence identity" refers to a value determined by comparing two optimally aligned sequences within a comparison window, where the region of the polynucleotide sequence within the comparison window may contain additions or deletions (i.e., gaps) compared to the reference sequence (which does not contain additions or deletions) due to optimal alignment of the two sequences. The percentage is calculated by calculating the number of positions where the same nucleotide base or amino acid residue occurs in both sequences, dividing the number of matching positions by the total number of positions within the comparison window, and multiplying the result by 100 to calculate the percentage of sequence identity.
[0045] The term "substantial identity" in the context of peptides indicates that the peptide has at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, or 94%, or 95%, 96%, 97%, 98%, or 99% sequence identity with the reference sequence within a specified comparison window. In certain embodiments, optimal alignment is performed using the Needleman and Wunsch homology alignment algorithm (Needleman and Wunsch, JMB, 48, 443 (1970)). An indication that two peptide sequences are substantially identical is that one peptide immunologically reacts with an antibody raised against the other peptide. Thus, a peptide is substantially identical to another peptide where the two peptides differ only by conservative substitutions. Accordingly, embodiments of the present invention also provide nucleic acid molecules and peptides that are substantially identical to the nucleic acid molecules and peptides presented herein.
[0046] In sequence comparison, typically, one sequence serves as a reference sequence and is compared to a test sequence. When using a sequence comparison algorithm, the test and reference sequences are input into a computer, subsequence coordinates are designated, if necessary, and sequence algorithm program parameters are designated. The sequence comparison algorithm calculates the percent sequence identity of the test sequence relative to the reference sequence based on the designated program parameters.
[0047] As used herein, the term "primer" refers to a short polynucleotide that hybridizes to a target polynucleotide sequence and serves as a starting point for the synthesis of a new polynucleotide.
[0048] "Multiplex" refers to the use of two or more pairs of primers to simultaneously amplify multiple target gene fragments in a single tube. In this method, all primers can be contained in a single tube into which the sample is introduced or placed. Then, multiple forward and reverse primers in the tube amplify all desired influenza virus and control gene fragments.
[0049] As used herein, "complement" refers to the complementary sequence of a nucleic acid according to standard Watson-Crick base pairing rules. A complementary sequence is an RNA sequence complementary to a DNA sequence or its complementary sequence, and may also be cDNA. As used herein, "substantially complementary" means that two sequences hybridize under stringent hybridization conditions. Those skilled in the art will understand that a substantially complementary sequence need not hybridize over the entire length. In particular, a substantially complementary sequence includes a contiguous sequence of bases located 3' or 5' of a contiguous base sequence that hybridizes with a target or marker sequence under stringent hybridization conditions, but that does not hybridize with the target or marker sequence.
[0050] "Hybridization" refers to the reaction of one or more polynucleotides to form a complex stabilized by hydrogen bonds between the bases of the nucleotide residues. Hydrogen bonding can occur through Watson-Crick base pairing, Hoogstein binding, or other sequence-specific methods. This complex can involve two strands forming a duplex structure, three or more strands forming a multistranded complex, a single self-hybridizing strand, or any combination of these. A hybridization reaction can constitute a step in a more extensive process, such as the initiation of a PCR reaction or the enzymatic cleavage of a polynucleotide by a ribozyme.
[0051] Examples of stringent hybridization conditions include an incubation temperature of about 25°C to about 37°C, a hybridization buffer concentration of about 6xSSC to about 10xSSC, a formamide concentration of about 0% to about 25%, and a wash solution of about 4xSSC to about 8xSSC. Examples of moderate hybridization conditions include an incubation temperature of about 40°C to about 50°C, a buffer concentration of about 9xSSC to about 2xSSC, a formamide concentration of about 30% to about 50%, and a wash solution of about 5xSSC to about 2xSSC. Examples of high stringency conditions include an incubation temperature of about 55°C to about 68°C, a buffer concentration of about 1xSSC to about 0.1xSSC, a formamide concentration of about 55% to about 75%, and a wash solution of about 1xSSC, 0.1xSSC, or deionized water. Generally, hybridization incubation times range from 5 minutes to 24 hours, with one, two, or more wash steps, with wash incubation times of about 1, 2, or 15 minutes. SSC is 0.15 M NaCl and 15 mM citrate buffer. It will be understood that equivalent SSCs using other buffer systems can be used.
[0052] As used herein, the term "reference genome" refers to any particular known, sequenced, or characterized genome (whether partial or complete) that can be used to reference sequences identified from a subject. Representative reference genomes for human subjects and many other organisms are provided in online genome browsers operated by the National Center for Biotechnology Information (NCBI) or the University of California, Santa Cruz (UCSC). "Genome" refers to the complete genetic information of an organism or virus, represented in nucleic acid sequence. As used herein, a reference sequence or reference genome often refers to assembled or partially assembled genome sequences obtained from an individual or multiple individuals. In some embodiments, a reference genome refers to assembled or partially assembled genome sequences obtained from one or more human individuals. A reference genome can be considered a representative example of a gene set for a species. In some embodiments, a reference genome includes sequences assigned to chromosomes. An example of a human reference genome is GRCh37 (UCSC equivalent: hg19).
[0053] As used herein, the term "normal reference standard" refers to the control level, degree, or range of DNA methylation at a specific genomic region or gene in a sample not associated with cancer. The term "normal reference cutoff value" refers to the control threshold level of DNA methylation at a specific genomic region or gene, or the differential methylation value (DMV). In some embodiments, a DNA methylation level above the normal reference cutoff value is associated with having or developing cancer. In some embodiments, a DNA methylation level at or below the normal reference cutoff value is associated with not having or developing cancer.
[0054] "Detection" refers to determining the presence, absence, and / or degree of methylation in a nucleic acid of interest in a sample. Detection does not require a method to provide 100% sensitivity and / or 100% specificity.
[0055] "RT-PCR" refers to reverse transcription polymerase chain reaction, which is used to detect specific RNA (in this case, a specific gene fragment of the influenza virus genome) by transcribing the target RNA into its DNA complement using, for example, reverse transcriptase. The newly synthesized cDNA can be amplified using conventional PCR. In one embodiment, the RT-PCR provided herein is a one-step method in which the entire reaction, from cDNA synthesis to PCR amplification, is carried out in a single tube. Alternatively, the process described herein is compatible with two-step reactions in which the reverse transcriptase reaction and PCR amplification must be carried out in separate tubes. See Real-Time PCR: Current Technology and Applications, Logan, Edwards, and Saunders eds., Caister Academic Press, 2009; Bustin A-Z of Quantitative PCR (IUL Biotechnology, No. 5).
[0056] As used herein, a "fragment" of DNA refers to a fragment of about 10 bp, about 20 bp, about 30 bp, about 40 bp, about 50 bp, about 60 bp, about 70 bp, about 80 bp, about 90 bp, about 100 bp, about 110 bp, about 120 bp, about 130 bp, about 140 bp, about 150 bp, about 160 bp, 170 bp, about 180 bp, about 190 bp, about 200 bp, about 210 bp, about 220 bp, about 230 bp, about 240 bp, about 250 bp, about 260 bp, about 270 bp, about 280 bp, about 290 bp, about 300 bp, about 310 bp, about 320 bp, about 330 bp, about 340 bp, about 350 bp, about 360 bp, about 370 bp, about 380 bp, about 390 bp, about 400 bp, about 410 bp, about 420 bp, about 430 bp, about 440 bp, about 450 bp, about 460 bp, about 470 bp, about 480 bp, about 490 bp, about 500 bp, about 510 bp, about 520 bp, about 530 bp, about 540 bp, about 550 bp, about 560 bp, about 570 bp, about 580 bp, about 590 bp, about 600 bp, about 610 bp, about 620 bp, about 630 bp, This refers to cell-free DNA fragments that are 10 bp, about 220 bp, about 230 bp, 240 bp, about 250 bp, about 260 bp, about 270 bp, 280 bp, about 290 bp, about 300 bp, about 310 bp, about 320 bp, about 330 bp, about 340 bp, about 350 bp, about 360 bp, about 370 bp, about 380 bp, about 390 bp, or about 400 bp in length. Typically, DNA fragments are about 100 bp to about 200 bp, about 120 bp to about 180 bp, or about 140 bp to about 160 bp.
[0057] The term "neoadjuvant therapy" refers to treatment (such as chemotherapy or hormone therapy) given before primary cancer treatment (such as surgery) to improve the outcome of the primary treatment.
[0058] The term "chemotherapy" refers to the treatment of cancer with anti-tumor or chemotherapeutic agents as part of a standard regimen. Chemotherapy may be given with curative intent or with the intent to prolong survival or palliation of symptoms. It may be used in conjunction with other cancer treatments, such as radiation therapy or surgery.
[0059] The term "methylation" refers to the addition of a methyl group to the 5' carbon of the cytosine base of CpG deoxyribonucleic acid sequences in the genome.
[0060] The term "adjacent CpG sites" refers to a set of CpG sites that are present within a genomic feature or within a short genetic distance. Genomic features include promoters, enhancers, exons, introns, 5'-untranslated regions (UTRs), 3'-UTRs, gene bodies, stem cell-associated regions, CpG islands, CpG shelves, CpG shores, LINEs, SINEs, or LTRs. Short genetic distances are 10bp, 11bp, 12bp, 13bp, 14bp, 15bp, 16bp, 17bp, 18bp, 19bp, 20bp, 21bp, 22bp, 23bp, 24bp, 25bp, 26bp, 27bp, 28bp, 29bp, 30bp, 31bp, 32bp, 33bp, 34bp, 35bp, 36bp, 37bp, 38bp, 39bp, 40bp, 41bp, 42bp, 43bp, 44bp, 45bp, 46bp, 47bp, 48bp, 49bp, 50bp, 51bp, 52bp, 53bp, 54bp, 55bp, 56bp, 57bp, The length may be 58bp, 59bp, 60bp, 61bp, 62bp, 63bp, 64bp, 65bp, 66bp, 67bp, 68bp, 69bp, 70bp, 71bp, 72bp, 73bp, 74bp, 75bp, 76bp, 77bp, 78bp, 79bp, 80bp, 81bp, 82bp, 83bp, 84bp, 85bp, 86bp, 87bp, 88bp, 89bp, 90bp, 91bp, 92bp, 93bp, 94bp, 95bp, 96bp, 97bp, 98bp, 99bp, 100bp, 250bp, 500bp, 750bp, or 1,000bp. In some cases, adjacent CpG sites may be present within the sequencing read.
[0061] "Minimal residual disease" or "MRD" refers to cancer cells (e.g., breast cancer cells) that remain after treatment and cannot be detected by a scan or test to determine remission status (i.e., cancer-free status). MRD can occur in any of the cancer treatments described herein.
[0062] A "fragment assessment region" or "FAR" refers to a set of n CpG coordinates within a DMR call (target region) that are within one base pair of each other from a single CpG, and whose length is one base pair less than the expected fragment length (usually 160 bp).
[0063] An embodiment of the present invention. The present disclosure provides panel assays and various methods for detecting the difference in the methylation pattern of target regions of cfDNA.The difference in the methylation pattern of target regions of sample can indicate, for example, the presence or absence of breast cancer, the severity of breast cancer, susceptibility to breast cancer, recurrence or susceptibility to recurrence of breast cancer, the presence or absence of minimal residual disease (MRD), and susceptibility to MRD.The methylation pattern of target regions of cfDNA in sample can be analyzed using a machine learning algorithm that is trained using target regions of cfDNA of cancer samples (such as metastatic breast cancer) and non-cancer control samples, and develops an MRD signature that is used to detect MRD in test subjects.
[0064] Description of embodiments. Description 1. A method for determining whether a subject has minimal residual disease (MRD), comprising the steps of: a) training a machine learning program to develop an MRD signature, wherein the machine learning program is trained using a plurality of target regions from a cancer sample and a corresponding plurality of target regions from a non-cancerous sample, and the MRD signature is based on a comparison of the methylation pattern of the plurality of target regions in the cancer sample with the methylation pattern of the corresponding plurality of target regions in the non-cancerous sample; b) determining the methylation pattern of the plurality of target regions in a cell-free deoxyribonucleic acid (cfDNA) sample obtained from the subject; c) applying the MRD signature to the methylation pattern of the plurality of target regions in the cfDNA obtained from the subject; and d) determining whether the subject has MRD based on the MRD signature.
[0065] Description 2. The method of description 1, wherein the plurality of target regions in a cfDNA sample from the subject are identical to the plurality of target genomic regions in both the cancer and non-cancer samples used to develop the MRD signature.
[0066] Item 3. The method of items 1 or 2, wherein the methylation pattern of the plurality of target regions is determined using one or more of post whole genome library hybrid probe capture, enzymatic treatment, bisulfite amplicon sequencing (BSAS), bisulfite treatment of DNA, methylation-sensitive polymerase chain reaction, and a combination of bisulfite conversion and bisulfite restriction analysis.
[0067] Item 4. The method according to any one of items 1 to 3, wherein the methylation pattern of each of the plurality of target regions is determined using a hybrid probe capture method.
[0068] Statement 5. The method of any one of statements 1 to 4, wherein the hybrid probe capture method comprises using one or more hybrid capture probes comprising ribonucleic acid or deoxyribonucleic acid.
[0069] Statement 6. The method of any one of statements 1-5, wherein each of the one or more hybrid capture probes further comprises an affinity tag selected from the group consisting of biotin and streptavidin.
[0070] Item 7. The method of any one of items 1-6, wherein the plurality of target regions from the cancer sample and the non-cancer sample comprises about 60% to about 70% of the target regions in Table 1.
[0071] Statement 8. The method of any one of statements 1 to 7, wherein the plurality of target regions comprises about 70% to about 80% of the target regions of Table 1.
[0072] Item 9. The method of any one of items 1-8, wherein the plurality of target regions comprises about 80% to about 90% of the target regions of Table 1.
[0073] Item 10. The method of any one of items 1-9, wherein the plurality of target regions is greater than about 95% of the target regions of Table 1.
[0074] Item 11. The method of any one of items 1 to 10, wherein the cfDNA sample is extracted from whole blood, plasma, serum, or urine.
[0075] 12. The method of any one of claims 1 to 11, further comprising: e) combining adjacent CpGs of each of the plurality of target regions into n to m contiguous CpG blocks, where n is greater than or equal to 1 and m is less than the length of the corresponding target region; f) removing target regions having fewer than n CpG blocks and more than m CpG blocks; and g) filtering the target regions remaining after step f) using a k-means clustering function based on their adjacent CpGs to provide one or more fragment assessment regions (FARs).
[0076] 13. The method of any one of statements 1-12, further comprising aggregating the methylation state of each FAR according to: statement 13.h) identifying methylation patterns of all or substantially all possible CpGs; i) selecting all sequencing reads that overlap with the FARs; j) extracting the methylation state of each CpG in the sequencing reads that span the FARs; k) counting each different methylation pattern in the FARs to calculate the number of methylation states; and l) outputting the results of steps h) to k), wherein the output includes any one or more of the position of the FARs, the methylation pattern of the FARs, and the number of FARs.
[0077] Description 14. The method of any one of descriptions 1-13, comprising the steps of: combining the FAR counts; normalizing the FAR counts based on sequencing depth; and identifying differentially expressed FARs between the subject's cfDNA sample and the cancer and non-cancer samples.
[0078] Description 15. The method of any one of descriptions 1-14, wherein a trained machine learning program is used to determine the likelihood that a subject has or will develop metastatic breast cancer, a breast cancer recurrence, or both metastatic breast cancer and a breast cancer recurrence.
[0079] Item 16. The method of any one of items 1-15, wherein the machine learning program comprises one or more of a random forest, a support vector machine (SVM), a neural network, a generalized linear model (GLM), a gradient boosting model (GBM), extreme gradient boosting (XGB), or a deep learning algorithm.
[0080] Statement 17. The method of any one of statements 1-16, wherein the cancer sample and the non-cancerous sample comprise one or more of a breast cancer sample, a known metastatic breast cancer sample, a breast cancer recurrence sample, a sample from a subject who has completed a cancer treatment regimen, and a sample from a subject with no evidence of disease on standard of care.
[0081] Claim 18. The method of any one of claims 1 to 17, further comprising treating the subject with MRD, wherein the treatment comprises one or more of radiation therapy, surgery to remove the cancer, and administration of a therapeutic agent to the patient, thereby treating the MRD.
[0082] In general, embodiments of the present disclosure include bisulfite converting nucleic acids from a subject's cfDNA sample, e.g., using whole genome bisulfite sequencing (WGBS) or hybrid probe capture; next-generation sequencing the converted and enriched nucleic acids; collecting methylation data from target regions (e.g., the target regions listed in Table 1); and using a trained machine learning algorithm to determine, e.g., the presence or absence of breast cancer, the severity of breast cancer, the histological subtype of breast cancer, or susceptibility to breast cancer.
[0083] In some embodiments, methylation data is used to develop cancer signatures, such as minimal residual disease (MRD), breast cancer recurrence, or MBC signatures, which indicate the presence of MRD in a patient, or to identify patients at high risk for cancer recurrence or MRD development. Certain embodiments are used to detect evidence of MRD before clinical recurrence, and the non-invasive method can be easily repeated after completion of the primary treatment regimen. In one embodiment, a method for determining the presence of MRD includes analyzing the methylation pattern of a specific target region of cfDNA. Typically, a beta value (the ratio of the number of methylated CpGs at a specific locus to the total number of CpGs at the same locus) is used to develop a differentially methylated region score (or "DMR" score). The DMR score can be used to determine, for example, the presence or absence of cancer or MRD, the likelihood of developing MRD, or the likelihood of cancer recurrence based on a comparison of the DMR value of the tested subject with the DMR value of a healthy subject or control value. Because the source of the test sample is cfDNA, which may be derived from multiple tissue sources, the resulting beta value represents the average methylation status across multiple tissue sources. In contrast, methylation pattern analysis, or fragment-level techniques, aggregate all possible methylation states of adjacent CpGs, thereby preserving the context of each CpG island. As an example, fragments IV and V in Figure 1 have an average beta value of 0.5 (half of the CpGs are methylated) for the last four CpGs, resulting in a Δβ value of 0. However, these fragments have completely opposite methylation patterns, suggesting a different tissue of origin, with one of the fragments likely being tumor-derived. Therefore, this fragment-level methylation pattern analysis allows for Boolean feature classification, i.e., assessing the presence or absence of cancer-specific fragments of DNA (CSFs) in a given cfDNA sample. This technique may be more sensitive in situations of low tumor burden, such as MRD.
[0084] In one embodiment, a method for analyzing the methylation pattern of a specific target region includes the steps of CpG clustering, methylation aggregation, and fragment analysis. The CpG clustering step combines adjacent CpGs into distinct blocks containing n to m CpGs, where n and m are user-specified. Preferably, the length of these blocks is shorter than the length of the fragment. To evaluate the methylation pattern of CpGs within a selected block, sequencing reads must span these CpGs. More specifically, the CpG clustering step works in two stages: first, adjacent CpGs are combined into contiguous regions, and regions containing fewer than n CpGs are removed, while regions containing n to m CpGs and shorter than a maximum length are retained. Next, all other regions are recursively partitioned until a region meets user-specified constraints or is deemed inappropriate. Region refinement can be performed using k-means clustering based on the nearest neighboring CpGs. The remaining target regions are sometimes referred to as "fragment assessment regions (FARs)."
[0085] Generally, methylation aggregation can be performed after FARs are selected to count all possible methylation states within each fragment. For example, methylation aggregation can take two input files: a bedGraph file listing the genomic coordinates and number of CpGs for each FAR, and a bam file containing mapped sequencing reads. The bam file is filtered to retain only sequencing reads that overlap with the FARs. This is done to reduce execution time and memory usage in subsequent steps. For each sequencing read, the genomic coordinates, mapping information, and methylation state are recorded in a custom data structure. More specifically, methylation aggregation involves the following steps: (a) identifying all possible methylation patterns based on the number of CpGs within the FARs; (b) selecting all reads that overlap with the FARs; (c) extracting the methylation state of each CpG within the reads that span the FARs; and (d) counting each distinct methylation pattern within the fragment. If no reads overlap with the FARs, all values are assigned as NA. This produces a tabular output containing rows detailing, among other things, the location of the FAR, the methylation pattern, and the number or value of the methylation pattern.
[0086] The fragment analysis step combines fragment counts from multiple samples, normalizes them based on sequencing depth, and analyzes differentially expressed fragments between groups. Additionally, the output supports data visualization and may provide export functionality for further analysis in packages such as SAS, SPSS, and Microsoft Excel.
[0087] In some embodiments, a biological sample containing cfDNA for which methylation patterns can be analyzed is collected, for example, from a patient with a tumor or mass, or a patient suspected of having a tumor or mass. In some embodiments, a biological sample containing cfDNA for which methylation patterns can be analyzed is collected, for example, from a patient who has completed a cancer treatment regimen and is suspected of having MRD. In some embodiments, a biological sample containing cfDNA can be collected from a patient previously diagnosed with cancer and / or currently diagnosed in remission. In some embodiments, a biological sample containing cfDNA can be collected from a patient who has completed all or part of a regimen. Preferably, the biological sample is collected by standard biopsy or liquid biopsy. cfDNA can be collected from whole blood, plasma, serum, or urine. In some embodiments, the volume of the sample, such as whole blood, is about 50 μL to about 5 mL, about 100 μL to about 5 mL, about 150 μL to about 5 mL, about 200 μL to about 5 mL, about 250 μL to about 5 mL, about 300 μL to about 5 mL, about 350 μL to about 5 mL, about 400 μL to about 5 mL, about 450 μL to about 5 mL, about 500 μL to about 5 mL, or about 550 μL to about 5 mL. mL, about 600 μL to about 5 mL, about 700 μL to about 5 mL, about 750 μL to about 5 mL, about 800 μL to about 5 mL, about 850 μL to about 5 mL, about 900 μL to about 5 mL, about 950 μL to about 5 mL, about 1 mL to about 5 mL, about 1.5 mL to about 5 mL, about 2 mL to about 5 mL, about 2.5 mL to about 5 mL, or about 3 mL to about 5 mL. In another embodiment, the volume of the sample, such as whole blood, may comprise about 5 mL to about 10 mL.
[0088] Separation and extraction of cfDNA can be carried out by collecting bodily fluids using various techniques.In some cases, collection can include using a syringe to draw bodily fluids from a subject.In other cases, collection can include pipetting or collecting bodily fluids directly into collection containers.
[0089] After collecting body fluid, cfDNA can be separated and extracted by various techniques known to those skilled in the art.In some cases, commercially available kits such as Thermofisher MagMax cfDNA Kit or Qiagen Qiamp® Circulating Nucleic Acid Kit protocol can be used to separate, extract and prepare cell-free nucleic acid.In other examples, Qiagen Qubit™ dsDNA HS Assay kit protocol, Agilent™ DNA 1000 kit, or TruSeq™ Sequencing Library Preparation; Low-Throughput (LT) protocol, Roche KAPA Hyper Prep Kit, Swift Biosciences Methyl-Seq Library Prep Kit, Nugen Ultra-low Methyl-Seq Kit can be used.
[0090] Alternatively, cfDNA can be extracted and separated from body fluids through a separation step, which separates the cfDNA present in the solution of body fluids from cells and other insoluble components of body fluids.Separation can include, but is not limited to, techniques such as centrifugation or filtration.In other cases, cells are not first separated from cfDNA, but rather dissolved.For example, the genomic DNA of complete cells can be separated by selective precipitation.
[0091] In some embodiments, a method for determining the methylation pattern of one or more target nucleic acids includes methylation sequencing. For example, the methylation patterns of CpG sites within the target regions listed in Table 1 can be detected using DNA methylation sequencing. DNA methylation sequencing includes, for example, treating DNA from a sample with bisulfite to convert unmethylated cytosines to uracil, amplifying the target nucleic acids within the treated genomic DNA (e.g., by PCR amplification), and sequencing the resulting amplicons. Sequencing generates nucleotide reads, which can be aligned to a genomic reference sequence and used to quantify the methylation levels of all CpGs within the amplicon. Cytosines in non-CpG contexts can be used to track bisulfite conversion efficiency for each sample. This procedure is efficient in both time and cost, and multiple samples can be sequenced in parallel using 96-well plates, producing reproducible measurements of methylation when measured in independent experiments.
[0092] The nucleic acid molecule may be subjected (e.g., after extraction from the sample) to conditions sufficient to convert unmethylated cytosines within the nucleic acid molecule to uracil. For example, to detect DNA methylation, in certain embodiments, the DNA to be analyzed is first converted to convert unmethylated cytosines to uracil. In one embodiment, a chemical reagent may be used that selectively modifies either the methylated or unmethylated form of a CpG dinucleotide motif. Suitable chemical reagents include hydrazine and bisulfite ions. Preferably, the isolated DNA is treated with sodium bisulfite (NaHSO3) to convert unmethylated cytosines to uracil while preserving methylated cytosines. Without wishing to be bound by theory, sodium bisulfite has been found to readily react with the 5,6-double bond of cytosine but to be less reactive with methylated cytosine. Cytosine reacts with bisulfite ions to form a sulfonated cytosine reaction intermediate that is susceptible to deamination, producing sulfonated uracil. The sulfonated group can be removed under alkaline conditions to form uracil. This nucleotide conversion alters the original DNA sequence. It is well known that the resulting uracil exhibits the base-pairing behavior of thymine, which differs from that of cytosine. Therefore, uracil is recognized as thymine by DNA polymerases. Therefore, after PCR or sequencing, the resulting product contains cytosine only at positions where 5-methylcytosine was present in the starting template DNA. This allows for the differentiation of unmethylated and methylated cytosines.
[0093] Nucleic acid molecules may be subjected to further processing, including other derivatization processes (e.g., incorporating, modifying, and / or deleting one or more sequences, tags, or labels). In some cases, functional sequences (e.g., sequencing adapters, flow cell adapters, sequencing primers, etc.) may be added to nucleic acid molecules to facilitate nucleic acid sequencing. Thus, derivatives of nucleic acid molecules from a sample can include processed nucleic acid molecules, including bisulfite-modified nucleic acid molecules, reverse-transcribed nucleic acid molecules, tagged nucleic acid molecules, barcoded nucleic acid molecules, and other modified nucleic acid molecules.
[0094] In some embodiments, the methylation pattern of the target region may be determined using one or more of hybrid probe capture (Buckley et al., NAR Genom Bioinform. 2022 Dec 31;4(4):lqac099. doi: 10.1093 / nargab / lqac099), targeted bisulfite amplicon sequencing, bisulfite DNA treatment, WGBS, combined bisulfite conversion and bisulfite restriction analysis (COBRA), bisulfite PCR, bisulfite modification, bisulfite pyrosequencing, methylated CpG island amplification, CpG-binding column-based CpG island isolation, differential methylation hybridization of CpG island arrays, high-performance liquid chromatography, DNA methyltransferase assay, methylation-sensitive PCR, cloning of differentially methylated sequences, post-restriction methylation detection, restriction landmark genomic scanning, methylation-sensitive restriction fingerprinting, or Southern blot analysis.
[0095] In some embodiments, one or more hybrid capture probes hybridize to a plurality of target regions, each of the plurality of target regions comprising a uracil at each position corresponding to an unmethylated cytosine in the DNA molecule, and the one or more hybrid capture probes are complementary to one or more of the plurality of target regions. In some embodiments, one or more hybrid capture probes hybridize to a plurality of target regions, each of the plurality of target regions comprising a thymine at each position corresponding to an unmethylated cytosine in the DNA molecule.
[0096] In some embodiments, the one or more hybrid capture probes are configured to hybridize to any of the following: a) a nucleotide sequence of a plurality of target regions comprising a uracil at each position corresponding to a cytosine in a CpG site of the nucleic acid molecule; b) a nucleotide sequence of a plurality of target regions comprising a uracil at each position corresponding to a cytosine in a CpG site of the nucleic acid molecule; or c) a nucleotide sequence of a plurality of target regions comprising a cytosine at each position corresponding to a cytosine in a CpG site of the nucleic acid molecule.
[0097] In one embodiment, the method used to determine the methylation level of one or more target regions in cfDNA is WGBS (Cokus et al., 2008. Nature, 452(7184):215-219; Lister, et al., 2009. Nature, 462(7271):315-322; Harris, et al., 2010. Nat Biotechnol, 28(10):1097-1105).
[0098] Other methods for determining the methylation status of CpG sites can also be used. Numerous DNA methylation detection methods are known in the art, including hybrid probe capture (REF), methylation-specific enzyme digestion (Singer-Sam et al., Nucleic Acids Res. 18(3):687, 1990; Taylor et al., Leukemia 15(4):583-9, 2001), methylation-specific PCR (MSP or MSPCR) (Herman et al., Proc Natl Acad Sci USA 93(18):9821-6, 1996), methylation-sensitive single-nucleotide primer extension (MS-SnuPE) (Gonzalgo et al., Nucleic Acids Res. 25(12):2529-31, 1997), restriction landmark genomic scanning (RLGS) (Kawai, Mol Cell Biol. 14(11):7421-7, 1994; Akama, et al., Cancer Res. 57(15):3294-9, 1997), and differential methylation hybridization (DMH) (Huang et al., Hum Mol Genet. 8(3):459-70, 1999). In some embodiments, methylation levels may be determined using one or more DNA methylation sequencing assays, with or without bisulfite treatment of the DNA.
[0099] In one embodiment, Reduced Representation Bisulfite Sequencing (RRBS) is used to measure the methylation level of a target region. RRBS generally begins with bisulfite treatment of nucleic acids to convert all unmethylated cytosines to uracil, followed by restriction enzyme digestion (e.g., with an enzyme that recognizes sites containing CG sequences, e.g., MspI), ligation with an adaptor ligand, and full-fragment sequencing. The choice of restriction enzyme enriches for fragments in CpG-dense regions, reducing the number of redundant sequences that may map to multiple locations in a gene during analysis. Therefore, RRBS reduces the sample complexity of a nucleic acid sample by selecting a subset of restriction fragments for sequencing (e.g., size selection by preparative gel electrophoresis). In contrast to whole-genome sequencing using bisulfite, each fragment generated by restriction enzyme digestion contains DNA methylation information for at least one CpG dinucleotide. Thus, RRBS enriches samples for promoters, CpG islands, and other genomic features with a high frequency of restriction enzyme cleavage sites in these regions, thus providing an assay to assess the methylation status of one or more genomic loci.
[0100] A typical protocol for RRBS involves digesting the sample nucleic acid with a restriction enzyme such as Mspl, projecting and A-tailing, ligating adapters, converting with bisulfite, and performing PCR (see, for example, Gu et al., (2010), Nat Methods 7:133-6; Meissner et al., (2005), Nucleic Acids Res. 33:5868-77).
[0101] In some embodiments, for example, identifying the presence and / or severity of cancer, such as metastatic breast cancer, identifying breast cancer recurrence, identifying susceptibility to breast cancer recurrence, identifying MRD, identifying susceptibility to MRD, or identifying MRD after a subject has completed a cancer treatment regimen may involve using hybrid capture probes configured to selectively enrich for nucleic acid molecules (e.g., DNA or RNA molecules) or sequences thereof. Such probes may be pull-down probes (e.g., bait sets). The selectively enriched nucleic acid molecules or their sequences may correspond to one or more target regions in the methylation profile of the dataset. Specific sequences, modifications (e.g., methylation status), deletions, additions, single nucleotide polymorphisms, copy number variations, or other features in the selectively enriched nucleic acid molecules or their sequences may indicate, for example, the presence and / or severity of breast cancer, the presence or absence of MRD, or susceptibility to MRD, or the presence or absence of MRD or susceptibility to developing MRD during or after a cancer treatment regimen (e.g., adjuvant or neoadjuvant therapy). The probes may be selective for (i.e., complementary to) a subset of specific target regions and / or differentially methylated regions (e.g., CpG sites, CpA sites, CpT sites, and / or CpC sites) in Table 1 in a cfDNA sample. The probes may be configured to selectively enrich nucleic acid molecules (e.g., DNA or RNA molecules) or sequences thereof corresponding to multiple target nucleic acids in a target genomic sequence, such as a subset of one or more genomic regions and / or differentially methylated regions (e.g., CpG sites, CpA sites, CpT sites, and / or CpC sites) in a cell-free biological sample. The probes may be nucleic acid molecules (e.g., DNA or RNA molecules) having sequence complementarity with the target nucleic acid sequence. These nucleic acid molecules may be primers or enrichment sequences.Assaying nucleic acid molecules in a sample (e.g., acellular biological sample) using probes selected for target nucleic acid sequences can include array hybridization, polymerase chain reaction (PCR), or nucleic acid sequencing (e.g., DNA sequencing or RNA sequencing).The number of target nucleic acid sequences selectively enriched by such a scheme can include at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 50, at least 100, at least 150, at least 200, at least 300, at least 500, at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, or more than 5000 target nucleic acid sequences from different target genome regions.The use of such probes to enrich target nucleic acids can be referred to as "hybrid capture." The use of such hybrid capture probes can be performed before or after bisulfite conversion (if applicable). Examples of target nucleic acid sequences include sequences related to the target regions shown in Table 1.
[0102] In some embodiments, cfDNA samples can be collected from plasma samples of subjects with or suspected of having breast cancer, breast cancer recurrence, MBC, or MRD.The extracted cfDNA can be contacted with a bisulfite compound to undergo bisulfite conversion.Then, a library can be prepared from the bisulfite-converted nucleic acid.Then, a portion of the library can be hybridized with various capture probes that are complementary to one or more DNA strands of target region, or complementary to the target sequence that has been modified by bisulfite conversion, such as CpG islands.
[0103] Non-limiting examples of methods for preparing libraries include transposome-mediated protocols with dual indexing and / or kits (e.g., TruSeq Methyl Capture EPIC Library Prep Kit (Illumina, CA, USA), Kapa Hyper Prep Kit (Kapa Biosystems)). Indexing can be performed using adapters such as TruSeq DNA LT adapters (Illumina). Sequencing of libraries is performed using sequencer platforms (e.g., MiSeq, HiSeq, Illumina Roche KAPA Hyper Prep Kit, Swift Biosciences Methyl-Seq Library Prep Kit, Nugen Ultra-low Methyl-Seq Kit).
[0104] Preferably, the capture probe is a DNA probe or an RNA probe that is complementary to at least a portion of the nucleotide sequence of the target genome region, or complementary to at least a portion of the nucleotide sequence of the target genome region that has been modified by bisulfite conversion. In some embodiments, multiple capture probes that overlap one or more portions of each target genome region can be used (i.e., tiling). In this way, multiple capture probes can be used to saturate the target genome region, ensuring enrichment of the target genome region. Capture probes can be designed using publicly available software, or can be purchased commercially.
[0105] Generally, the target strand can be the "plus" strand (e.g., the strand that is transcribed into mRNA and then translated into protein) or its complementary "minus" strand. In some embodiments, the assay panel includes two probe sets, one probe that targets the plus strand of the target genomic region and the other probe that targets the minus strand.
[0106] In some embodiments, the capture probes may be tagged with biotin, streptavidin, digitonin, or other known affinity tags. After hybridization to the target genomic region, the biotinylated capture probes are "pulled down" from the library using streptavidin beads or other streptavidin-coated surfaces, thereby enriching the target genomic region. In other embodiments, the probes may be immobilized on an assay panel comprising a solid surface, such as a glass microarray slide. In some embodiments, an exemplary assay panel includes at least 1,000, 2,000, 2,500, 5,000, 10,000, 12,000, 15,000, 20,000, 25,000, 30,000, 35,000, or 40,000 hybrid capture probes complementary to the target regions disclosed in Table 1. In some embodiments, the assay panel comprises about 100, about 200, about 300, about 400, about 500, about 600, about 700, about 800, about 900, about 1,000, about 1,500, about 2,000, about 2,500, about 3,000, about 3,500, about 4,000, about 4,500, about 5,000, about 5,500, about 6,000, about 6,500, about 7,000, about 7,500, about 8,000, about 8,500, about 9,000, about 9,500, or about 10,000 pairs of hybrid capture probes complementary to target regions disclosed in Table 1. In some embodiments, each hybrid capture probe on the assay panel comprises less than 300, 250, 200, or 150 nucleotides. In some embodiments, each probe on the panel comprises 100-150 nucleotides.
[0107] The enriched target genomic regions can then be sequenced using next-generation sequencing technologies such as pyrosequencing, single-molecule real-time sequencing, sequencing by synthesis, sequencing by ligation (SOLID sequencing), and nanopore sequencing.
[0108] Nucleic acid molecules (e.g., extracted cfDNA) or their derivatives can be sequenced to provide multiple sequencing reads. The sequencing reads can be aligned with or analyzed relative to a reference genome. Based on at least a portion of the sequencing reads, the absolute or relative amount of nucleic acid molecules corresponding to one or more genomic regions (including the absolute or relative level of methylation within the molecules) can be measured. Alternatively, the sequencing reads may not be used to determine the amount or relative amount of nucleic acid molecules. A dataset including a genomic profile (e.g., a methylation profile) of one or more genomic regions of the sample is generated based on at least a portion of the sequencing reads. The sequencing reads can be processed to identify the methylation pattern of target regions of cfDNA in the sample.
[0109] Sequence identification can be performed by sequencing, array hybridization (e.g., Affymetrix), or nucleic acid amplification (e.g., PCR), etc. Sequencing can be performed by any suitable sequencing method, such as massively parallel sequencing (MPS), paired-end sequencing, high-throughput sequencing, next-generation sequencing (NGS), shotgun sequencing, single-molecule sequencing, nanopore sequencing, nanopore sequencing with direct detection or inference of methylation status, semiconductor sequencing, pyrosequencing, sequencing by synthesis (SBS), sequencing by ligation, sequencing by hybridization, and RNA-Seq (Illumina).
[0110] Sequencing and / or preparation of nucleic acid samples for sequencing may involve performing one or more nucleic acid reactions (e.g., one or more nucleic acid amplification processes (e.g., of DNA or RNA molecules)). Nucleic acid amplification may include, for example, reverse transcription, primer extension, asymmetric amplification, rolling circle amplification, ligase chain reaction, polymerase chain reaction (PCR), and multiple displacement amplification. Examples of PCR methods include digital PCR (dPCR), emulsion PCR (ePCR), quantitative PCR (qPCR), real-time PCR (RT-PCR), hot-start PCR, multiplex PCR, asymmetric PCR, nested PCR, and assembly PCR. An appropriate number of rounds of nucleic acid amplification (e.g., PCR (qPCR, RT-PCR, dPCR, etc.)) may be performed to amplify the initial amount of nucleic acid molecules (e.g., DNA molecules) or their derivatives to the desired amount for subsequent sequencing. In some cases, PCR may be used for global amplification of nucleic acid molecules. This may include first ligating adapter sequences to different molecules, followed by PCR amplification using universal primers. PCR can be performed using any of a number of commercially available kits from Life Technologies, Affymetrix, Promega, Qiagen, etc. In other cases, only specific target nucleic acids within a population of nucleic acids may be amplified. Specific primers (optionally in combination with adapter ligation) may be used to selectively amplify specific targets for downstream sequencing. In some cases, nested primers targeting specific genomic regions may be used. Nucleic acid amplification can include targeted amplification of one or more genetic loci, genomic regions, cfDNA target regions, or differentially methylated regions (e.g., CpG sites, CpA sites, CpT sites, and / or CpC sites), and particularly the target regions listed in Table 1 below. In some cases, nucleic acid amplification is performed after bisulfite conversion. Such procedures are sometimes referred to as targeted bisulfite amplicon sequencing (TBAS). Nucleic acid amplification can include the use of one or more primers, probes, enzymes (e.g., polymerases), buffers, and deoxyribonucleotides.Nucleic acid amplification can be isothermal or involve thermal cycling. Thermal cycling can involve temperature changes associated with various processes in nucleic acid amplification, such as initialization, denaturation, annealing, and extension. Sequencing can involve methods that perform reverse transcription (RT) and PCR simultaneously, such as the OneStep RT-PCR kit protocol from Qiagen, NEB, Thermo Fisher Scientific, or Bio-Rad.
[0111] Nucleic acid molecules (e.g., DNA or RNA molecules) or their derivatives can be labeled or tagged with distinguishable tags to allow for multiple sample multiplexing. For example, all nucleic acid molecules or their derivatives associated with a particular sample or subject can be tagged or labeled (e.g., with a barcode such as a nucleic acid barcode sequence or a fluorescent label). Nucleic acid molecules or their derivatives associated with other samples or subjects can be tagged or labeled with different tags or labels, thereby allowing the nucleic acid molecules or their derivatives to be associated with the sample or subject from which they originate. Such tagging or labeling facilitates multiplexing for simultaneously analyzing (e.g., sequencing) nucleic acid molecules or their derivatives from multiple samples and / or subjects. Any number of samples can be multiplexed. For example, a multiplexed reaction may contain at least about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, or more than 100 nucleic acid molecules or their derivatives from initial samples. These samples may be from the same or different subjects. For example, tagging multiple samples with sample barcodes (e.g., nucleic acid barcode sequences) allows each nucleic acid molecule (e.g., DNA molecule) or its derivative to be traced back to the sample (and / or subject) from which it was derived. Sample barcodes allow samples from multiple subjects to be distinguished from one another, thereby allowing sequences within such samples, such as in a pool, to be simultaneously identified. Tags, labels, and / or barcodes may be attached to nucleic acid molecules or their derivatives by ligation, primer extension, nucleic acid amplification, or other methods. In some cases, nucleic acid molecules or their derivatives of a particular sample are tagged, labeled, or barcoded with different tags, labels, or barcodes (e.g., unique molecular identifiers), thereby allowing different nucleic acid molecules or their derivatives from the same sample to be differentially tagged, labeled, or barcoded.In some cases, nucleic acid molecules or derivatives thereof from a given sample may be labeled with both different and identical labels, such that each nucleic acid molecule or derivative thereof associated with the sample contains both a unique label and a common label.
[0112] After nucleic acid molecules or their derivatives are subjected to sequencing, the sequencing reads can be subjected to suitable bioinformatics processes to generate a dataset comprising the methylation pattern of one or more target regions of cfDNA samples.For example, the sequencing reads can be aligned to one or more reference genomes (e.g., human genome).The aligned sequencing reads can be quantified at one or more genome loci or target regions to generate a dataset comprising the methylation pattern profile of one or more target regions of acellular biological samples.The quantification of sequences can be expressed as unnormalized or normalized values.
[0113] In some embodiments, alignment of bisulfite converted DNA is performed using a software program such as Bismark (Krueger et al. (2011) Bioinformatics, 27(11):157171). Bismark performs read mapping and methylation calling in a single step, and its output distinguishes between cytosines in CpG, CHG, and CHH contexts. Bismark is released under the GNU GPLv3+ license. The source code is freely available at bioinformatics.bbsrc.ac.uk / projects / bismark / . In some embodiments, methylation differences at specific loci / regions are calculated, for example, using one or more publicly available programs for analyzing and / or determining methylation levels or target polynucleotide regions. In some embodiments, methods used to analyze and / or determine the methylation level of a target polynucleotide region include Metilene (Juhling et al., Genome Res., 2016;26(2):256-262) or GenomeStudio Software available online from Illumina, Inc. Other methods for determining differential methylation of a target polynucleotide region are described in Hovestadt et al., 2014; Nature, 510(7506), 537-541.
[0114] In some embodiments, the target genomic regions tested to determine the presence or absence of breast cancer in a subject comprise at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% of the target regions listed in Table 1.
[0115] In some embodiments, the target regions examined to determine the severity of a breast cancer subject (i.e., stage I, stage II, stage III, or stage IV cancer) comprise at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% of the target regions listed in Table 1.
[0116] Some embodiments may be used to determine MBC, breast cancer recurrence, and / or the presence of minimal residual disease (MRD), the name given to the small number of cancer cells that remain in the body when a patient is in or considered to be in remission during or after treatment, which is the primary cause of cancer recurrence.
[0117] The target genomic regions to be examined to determine the presence of MBC in a subject include at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% of the target regions listed in Table 1.
[0118] The target genomic regions examined to determine breast cancer recurrence in a subject include at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% of the target regions listed in Table 1.
[0119] The target genomic regions tested at the time of diagnosis to determine a subject's susceptibility to breast cancer recurrence include at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% of the target regions listed in Table 1.
[0120] The target genomic regions tested to determine the presence of MRD in a subject include at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% of the target regions listed in Table 1.
[0121] The target genomic regions tested to determine a subject's MRD susceptibility include at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% of the target regions listed in Table 1.
[0122] The target genomic regions tested to determine the presence or absence of, or susceptibility to, MRD in a subject receiving a cancer treatment regimen include at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% of the target regions listed in Table 1.
[0123] The target genomic regions tested to determine the presence or absence of MRD or MRD susceptibility in a subject after completion of a cancer treatment regimen include at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% of the target regions listed in Table 1.
[0124] In some embodiments, for example, the target genomic regions tested to determine the presence of MBC, the presence of or susceptibility to MRD, or the presence of or susceptibility to breast cancer recurrence in a subject may include about 20% to about 30%, about 30% to about 40%, about 40% to about 50%, about 50% to about 60%, about 60% to about 70%, about 70% to about 80%, about 80% to about 90%, about 95%, about 96%, about 97%, about 98%, about 99%, or about 100% of the target regions listed in Table 1.
[0125] In some embodiments, for example, the target genomic regions to be tested to determine the presence of MBC, the presence of or susceptibility to MRD, or the presence of or susceptibility to breast cancer recurrence in a subject may include about 50, about 100, about 150, about 200, about 250, about 300, about 350, about 400, about 450, about 500, about 550, about 600, about 650, about 700, about 750, about 800, about 850, about 900, about 950, about 1000, about 1050, about 1100, about 1150, about 1200, about 1250, about 1300, about 1350, about 1400, or about 1450, about 1500, about 1550, or about 1564 of the target regions listed in Table 1. [Table 1] TIFF2025540676000003.tif234159TIFF2025540676000004.tif231159TIFF2025540676000005.tif234159TIFF2025540676000006.tif233159TIFF2025540676000007.tif234159TIFF2025540676000008.tif234159TIFF2025540676000009.tif234159TIFF2025540676000010.tif233159TIFF2025540676000011.tif234159TIFF2025540676000012.tif234159TIFF2025540676000013.tif233159TIFF2025540676000014.tif233159TIFF2025540676000015.tif232159TIFF2025540676000016.tif234159TIFF2025540676000017.tif233159TIFF2025540676000018.tif235159TIFF2025540676000019.tif233159TIFF2025540676000020.tif234159TIFF2025540676000021.tif233159TIFF2025540676000022.tif232159TIFF2025540676000023.tif233159TIFF2025540676000024.tif233159TIFF2025540676000025.tif233159TIFF2025540676000026.tif233159TIFF2025540676000027.tif233159TIFF2025540676000028.tif235159TIFF2025540676000029.tif234159TIFF2025540676000030.tif234159TIFF2025540676000031.tif234159TIFF2025540676000032.tif234159TIFF2025540676000033.tif236159TIFF2025540676000034.tif234159TIFF2025540676000035.tif234159TIFF2025540676000036.tif234159TIFF2025540676000037.tif235159TIFF2025540676000038.tif232159TIFF2025540676000039.tif233159TIFF2025540676000040.tif233159TIFF2025540676000041.tif233159TIFF2025540676000042.tif231159TIFF2025540676000043.tif232159TIFF2025540676000044.tif233159TIFF2025540676000045.tif232159TIFF2025540676000046.tif233159TIFF2025540676000047.tif231159TIFF2025540676000048.tif233159TIFF2025540676000049.tif232159TIFF2025540676000050.tif233159TIFF2025540676000051.tif234159TIFF2025540676000052.tif235159TIFF2025540676000053.tif235159TIFF2025540676000054.tif234159TIFF2025540676000055.tif234159TIFF2025540676000056.tif233159TIFF2025540676000057.tif233159TIFF2025540676000058.tif234159TIFF2025540676000059.tif235159TIFF2025540676000060.tif232159TIFF2025540676000061.tif234159TIFF2025540676000062.tif232159TIFF2025540676000063.tif233159TIFF2025540676000064.tif233159TIFF2025540676000065.tif235159TIFF2025540676000066.tif234159TIFF2025540676000067.tif235159TIFF2025540676000068.tif234159TIFF2025540676000069.tif235159TIFF2025540676000070.tif232159TIFF202554067600 0071.tif235159TIFF2025540676000072.tif234159TIFF2025540676000073.tif233159TIFF2025540 676000074.tif232159TIFF2025540676000075.tif235159TIFF2025540676000076.tif232159TIFF20 25540676000077.tif235159TIFF2025540676000078.tif234159TIFF2025540676000079.tif233159.
[0126] In some embodiments, the target genomic regions tested to determine the presence of MBC, MRD, or breast cancer recurrence in a subject may include about 700 to about 750 of the target regions listed in Table 1. In some embodiments, for example, the target genomic regions tested to determine the presence of MBC, the presence or susceptibility of MRD, or the presence or susceptibility of breast cancer recurrence in a subject include all of the target regions listed in Table 1.
[0127] In some embodiments, detecting cfDNA in sample further comprises aligning the DNA sequence obtained by next-generation sequencing to human reference genome.In certain embodiments, human reference genome GRCh37 (UCSC version hg19) is incorporated herein in its entirety.This genome assembly can be found, for example, at www.genome.ucsc.edu.
[0128] In some embodiments, the nucleotide sequence for analyzing nucleic acid methylation patterns includes a target region sequence listed in Table 1, and may further include 1 to 100, 1 to 150, 1 to 200, 1 to 300, 1 to 400, 1 to 500, 500 to 1000, 1000 to 1500, 1500 to 2000, 2000 to 2500, 2500 to 3000, 3000 to 3500, or 3500 to 4000 nucleotides immediately adjacent to or upstream or downstream of the target genomic region listed in Table 1.
[0129] In some embodiments, the methylation pattern of target region of cfDNA is determined in the region of selected gene or gene group.Non-limiting examples include the region in the untranslated region (UTR) of selected gene or gene group, the region within 1.5 kb upstream of the transcription start site of selected gene or gene group, and the region in the first exon of selected gene or gene group.In other embodiments, the target region of cfDNA is in the non-genic region of genomic DNA.
[0130] Embodiments of the methods described herein may also be used to determine the methylation patterns of specific target regions associated with various cancers, for example, to predict malignancy, stage of malignancy, susceptibility to cancer recurrence, and / or the presence or susceptibility to MRD. Exemplary cancers include leukemias (acute leukemias (e.g., 11q23-positive acute leukemia, acute lymphocytic leukemia, acute myeloid leukemia, acute myelogenous leukemia, and myeloblastic, myeloblastic, promyelocytic, myelomonocytic, monocytic, and erythroblastic leukemia), chronic leukemias (e.g., chronic myeloid (granulocytic) leukemia, chronic myelogenous leukemia, and chronic lymphocytic leukemia), polycythemia vera, lymphoma, Hodgkin's disease, non-Hodgkin's lymphoma (low-grade and high-grade forms), multiple myeloma, Waldenstrom's macroglobulinemia, heavy chain disease, myelodysplastic syndrome, hairy cell leukemia, and myelodysplasia. Other tumors include sarcomas and carcinomas, including fibrosarcoma, myxosarcoma, liposarcoma, chondrosarcoma, osteosarcoma, other sarcomas, synovioma, mesothelioma, and leukemia. These include tumors such as leiomyosarcoma, Ewing's tumor, leiomyosarcoma, rhabdomyosarcoma, colorectal cancer, lymphoid tumors, pancreatic cancer, breast cancer (basal, ductal, and lobular carcinoma of the breast), lung cancer, ovarian cancer, prostate cancer, hepatocellular carcinoma, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, medullary thyroid carcinoma, papillary thyroid carcinoma, pheochromocytoma, sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinoma, medullary carcinoma, bronchogenic carcinoma, renal cell carcinoma, hepatocellular carcinoma, cholangiocarcinoma, choriocarcinoma, Wilms' tumor, cervical cancer, testicular tumor, seminoma, bladder cancer, and central nervous system tumors (e.g., glioblastoma, astrocytoma, medulloblastoma, craniopharyngioma, ependymoma, pineal tumor, hemangioblastoma, acoustic neuroma, oligodendroglioma, meningioma, melanoma, neuroblastoma, and retinoblastoma). Any of the cancers listed above may develop MRD after treatment.
[0131] For example, by using the target regions set forth in Table 1, embodiments of the invention may have greater than 75% sensitivity in detecting breast cancer, breast cancer recurrence, MBC, or MRD; greater than 80% sensitivity in detecting breast cancer, breast cancer recurrence, MBC, or MRD; greater than 85% sensitivity in detecting breast cancer, breast cancer recurrence, MBC, or MRD; greater than 90% sensitivity in detecting breast cancer, breast cancer recurrence, MBC, or MRD; greater than 95% sensitivity in detecting breast cancer, breast cancer recurrence, MBC, or MRD; greater than 96% sensitivity in detecting breast cancer, breast cancer recurrence, MBC, or MRD; greater than 97% sensitivity in detecting breast cancer, breast cancer recurrence, MBC, or MRD; greater than 98% sensitivity in detecting breast cancer, breast cancer recurrence, MBC, or MRD; greater than 99% sensitivity in detecting breast cancer, breast cancer recurrence, MBC, or MRD; or 100% sensitivity in detecting breast cancer, breast cancer recurrence, MBC, or MRD.
[0132] In some embodiments, subjects may be tested for the presence or absence of MRD using the methods described herein at any time during cancer treatment or after completion of a cancer treatment regimen.
[0133] If a subject is identified as having a high risk of developing cancer (e.g., breast cancer), cancer recurrence (e.g., breast cancer), MBC, or MRD, the subject can be administered preventive treatment or therapy. For example, preventive measures include, but are not limited to, surgery, tamoxifen administration, raloxifene administration, etc. In the case of solid tumors, surgical resection can be performed.
[0134] If the subject is identified as having breast cancer, recurrence of breast cancer, MBC, or MRD, the subject can be administered a clinical procedure or cancer treatment. Exemplary therapies or procedures include, but are not limited to, surgery, radiation therapy, chemotherapy, hormone therapy, targeted therapy, and / or administering an effective amount of one or more of the following therapeutic agents: angiogenesis inhibitors, such as angiostatin K1-3, DL-α-difluoromethylornithine, endostatin, fumagillin, genistein, minocycline, staurosporine, and (±)-thalidomide; DNA intercalators / crosslinkers, such as bleomycin, carboplatin, carmustine, chlorambucil, cyclophosphamide, cis-diammineplatinum(II) dichloride (cisplatin), melphalan, mitoxantrone, oxaliplatin; DNA synthesis inhibitors, such as (±)-amethopterin (methotrexate), 3-amino-1,2,4-benzotriazine, 1,4-dioxide, aminopterin, cytosine β-D-arabinofuranosides, 5-fluoro-5′-deoxyuridine, 5-fluorouracil, ganciclovir, hydroxyurea, and mitomycin C; DNA-RNA transcription regulators, such as actinomycin D, daunorubicin, doxorubicin, homoharringtonine, and idarubicin; enzyme inhibitors, such as S(+)-camptothecin, curcumin, (−)-deguelin, 5,6-dichlorobenzimidazole 1-β-D-ribofuranoside, etoposide, formestane, fostriecin, hispidin, 2-imino-1-imidazolidineacetic acid (cyclocreatine), mevinolin, trichostatin A, tyrphostin AG 34, and tyrphostin AG 879; Gene modulators, such as 5-aza-2′-deoxycytidine, 5-azacytidine, cholecalciferol (vitamin D3), 4-hydroxytamoxifen, melatonin, mifepristone, raloxifene, all-trans-retinal (vitamin A aldehyde), retinoic acid, all-trans (vitamin A acid), 9-cis-retinoic acid, 13-cis-retinoic acid, retinol (vitamin A), tamoxifen, and troglitazone;Microtubule inhibitors, such as colchicine, dolastatin 15, nocodazole, paclitaxel, podophyllotoxin, rhizoxin, vinblastine, vincristine, vindesine, and vinorelbine (navelbine); and unclassified antitumor agents, such as 17-(allylamino)-17-demethoxygeldanamycin, 4-amino-1,8-naphthalimide, apigenin, brefeldin A, cimetidine, dichloromethylenediphosphonic acid, leuprolide (leuprorelin), luteinizing hormone-releasing hormone, pifithrin-α, rapamycin, sex hormone-binding globulin, thapsigallin, and urinary trypsin inhibitor fragment (bikunin). The antitumor agent may also be a neoantigen. Neoantigens are tumor-associated peptides that stimulate anti-tumor responses and function as active pharmaceutical ingredients in vaccine compositions, and are described in U.S. Patent Application Publication No. 2011 / 0293637, which is incorporated herein by reference in its entirety. Anti-tumor agents include monoclonal antibodies such as rituximab, alemtuzumab, ipilimumab, bevacizumab, cetuximab, panitumumab, and trastuzumab, vemurafenib, imatinib mesylate, erlotinib, gefitinib, vismodegib; 90 Y-ibritumomab tiuxetan, 131 l-tositumomab, ado-trastuzumab emtansine, lapatinib, pertuzumab, ado-trastuzumab emtansine, regorafenib, sunitinib, denosumab, sorafenib, pazopanib, axitinib, dasatinib, nilotinib, bosutinib, ofatumumab, obinutuzumab, ibrutinib, idelalisib, crizotinib, erlotinib (Tarceva®), afatinib dimaleate, ceritinib, tositumomab, and 131The anti-tumor agent may be I-tositumomab, ibritumomab tiucetan, brentuximab vedotin, bortezomib, siltuximab, trametinib, dabrafenib, pembrolizumab, carfilzomib, ramucirumab, cabozantinib, or vandetanib. The anti-tumor agent may be a cytokine such as interferon (INF), interleukin (IL), or hematopoietic growth factor. The anti-tumor agent may be INF-α, IL-2, aldesleukin, IL-2, erythropoietin, granulocyte-macrophage colony-stimulating factor (GM-CSF), or granulocyte colony-stimulating factor. Antitumor agents include toremifene, fulvestrant, anastrozole, exemestane, letrozole, dib-aflibercept, alitretinoin, temsirolimus, tretinoin, denileukin-diftitox, vorinostat, romidepsin, bexarotene, pralatrexate, lenalidomide, belinstat, pomalidomide, cabazitaxel, enzalutamide, abiraterone acetate, 223 The antitumor agent may be a targeted therapy agent such as radium chloride or everolimus. The antitumor agent may be a checkpoint inhibitor, such as an inhibitor of the programmed cell death-1 (PD-1) pathway, for example, an anti-PD-1 antibody (nivolumab). The inhibitor may be an anti-cytotoxic T-lymphocyte-associated antigen (CTLA-4) antibody. The inhibitor may target another member of the CD28 CTLA4 Ig superfamily, such as BTLA, LAG3, ICOS, PDL1, or KIR. The checkpoint inhibitor may target a member of the TNFR superfamily, such as CD40, OX40, CD137, GITR, CD27, or TIM-3. Furthermore, the antitumor agent may be an epigenetic targeting agent, such as an HDAC inhibitor, a kinase inhibitor, a DNA methyltransferase inhibitor, a histone demethylase inhibitor, or a histone methylation inhibitor. The epigenetic agent may be azacitidine, decitabine, vorinostat, romidepsin, or ruxolitinib.
[0135] In some embodiments, methods for treating cancer (e.g., breast cancer), cancer recurrence (e.g., breast cancer), MBC, or MRD may include administering, alone or in combination with a suitable carrier or vehicle, an effective amount of a suitable substance capable of targeting intracellular proteins, small molecules, or nucleic acid molecules, including, but not limited to, antibodies or functional fragments thereof (e.g., Fab′, F(ab′)2, Fab, Fv, rlgG, and scFv fragments and genetically engineered or otherwise modified forms of immunoglobulins, e.g., intracellular antibodies and chimeric antibodies), small molecule inhibitors of proteins, chimeric proteins or peptides, gene therapy agents for transcription inhibition, or RNA interference (RNAi)-related molecules or morpholino molecules that inhibit gene expression and / or translation. In one embodiment, the inhibitor is an RNAi-related molecule, such as an siRNA or shRNA, for inhibiting translation. RNA interference (RNAi) molecules are small nucleic acid molecules, such as small interfering RNA (siRNA), double-stranded RNA (dsRNA), microRNA (miRNA), or short hairpin RNA (shRNA) molecules, that bind complementary to a portion of a target gene or mRNA so as to reduce the expression level of the target.
[0136] Appropriate pharmaceutical compositions containing one or more agents described herein are administered and dosed according to principles of good medical practice, taking into account the individual patient's clinical condition, the site and method of administration, the administration schedule, the patient's age, sex, and weight, and other factors known to medical professionals. The therapeutically effective amount herein is therefore determined by considerations known to those skilled in the art. For example, an effective amount of a pharmaceutical composition is the amount necessary to therapeutically reduce the expression of a target gene. The amount of the pharmaceutical composition should be effective to achieve improvement, including, but not limited to, complete prevention, improved survival, more rapid recovery, or improvement or elimination of symptoms associated with the chronic inflammatory condition being treated, and other indicators selected as appropriate measures by those skilled in the art. In accordance with the technology of the present invention, an appropriate single dose is one that, when administered one or more times over an appropriate period of time, can prevent or alleviate (reduce or eliminate) symptoms in a patient. Those skilled in the art can easily determine an appropriate single dose for systemic administration based on the patient's size and route of administration.
[0137] Pharmaceutical compositions can be formulated according to known methods for preparing medicament-useful compositions. Furthermore, as used herein, "pharmaceutically acceptable carrier" refers to any standard pharmaceutically acceptable carrier. Pharmaceutically acceptable carriers include diluents, adjuvants, and vehicles, as well as implant carriers, and inert, non-toxic solid or liquid fillers, diluents, or encapsulating materials that do not react with the active ingredients of the technology. Examples include, but are not limited to, phosphate-buffered saline, saline, water, and emulsions such as oil / water emulsions. Carriers can be solvents or dispersion media containing, for example, ethanol, polyols (e.g., glycerol, propylene glycol, liquid polyethylene glycol, etc.), suitable mixtures thereof, and vegetable oils.
[0138] Pharmaceutically acceptable carrier compositions are well known to those skilled in the art and are described in numerous readily available sources. For example, Remington: The Science and Practice of Pharmacy (Gerbino, PP2005) Philadelphia, Pa., Lippincott Williams & Wilkins, 21st ed.) describes pharmaceutical formulations that can be used in connection with this technology. Suitable formulations for parenteral administration include sterile injectable solutions, which may contain, for example, antioxidants, buffers, bacteriostats, and solutes that render the formulation isotonic with the blood of the intended recipient, as well as aqueous and non-aqueous sterile suspensions, which may contain suspending agents and thickening agents. The formulations may be filled into single-dose or multi-dose containers (e.g., sealed ampoules and vials) and stored in a freeze-dried (lyophilized) state, requiring only the addition of sterile liquid carriers (e.g., water for injection) prior to use. Extemporaneous injection solutions and suspensions can be prepared from sterile powders, granules, tablets, etc. In addition to the ingredients particularly mentioned above, the formulations of the present technology can include other conventional agents, depending on the type of formulation.
[0139] In some embodiments, the methods described herein can also be implemented by using a computer system. For example, any of the above steps of evaluating sequencing reads to determine the methylation status of CpG sites can be performed by software components loaded onto a computer or other information appliance or digital device. When configured in this way, the computer, appliance, or device can then perform all or part of the above steps to assist in analyzing or compare values related to the methylation of one or more CpG sites. The above features can be implemented in one or more computer programs and performed by one or more computers running such programs.
[0140] Additionally, various aspects of the methods disclosed herein can be implemented using computer-based computations, machine learning (e.g., support vector machines (SVM), lasso, generalized linear models (GLM), gradient boosting models (GBM), extreme gradient boosting (XGB), elastic net regularized generalized linear models (Glmnet), random forests, gradient boosting (on random forests, C5.0 decision trees), and other software tools, or combinations thereof. For example, the methylation status of CpG sites can be assigned by a computer based on the underlying sequencing reads of amplicons from a sequencing assay. In another example, the methylation value for a DNA region or portion thereof can be compared by a computer to a threshold value, as described herein. Advantageously, these tools are provided in the form of computer programs that are executable on conventional general-purpose computer systems.
[0141] In some embodiments, methods used to analyze and / or determine the methylation level of a target polynucleotide region include Metilene (Juhling et al., Genome Res., 2016;26(2):256-262) or GenomeStudio Software available online from Illumina, Inc., or methods described in Hovestadt et al., 2014; Nature, 510(7506), 537-541.
[0142] In some embodiments, a method for identifying breast cancer, breast cancer severity, cancer recurrence, MBC, or MRD in a subject can include the use of a machine learning algorithm. The machine learning algorithm can be a trained algorithm. The machine learning algorithm is trained based on one or more features and used to process a dataset generated by analyzing nucleic acid molecules in a sample (e.g., an acellular biological sample). The dataset includes a methylation profile of one or more genomic regions of the acellular biological sample. Examples of the use and training of machine learning algorithms are described, for example, in WO / 2022 / 178108 (Salhia et al.).
[0143] In some embodiments, a computer with at least one processor can be configured to receive multiple sequencing results of DNA methylation sequencing reactions, including methylation patterns of one or more target regions disclosed herein, from, for example, a patient having a mass (e.g., a breast mass) or other tumor, or a patient suspected of having cancer, or a patient showing clinical signs of cancer. In some embodiments, the machine learning algorithm or program used to develop the MRD signature comprises comparing and analyzing the methylation patterns of multiple target regions of a cancer sample with the methylation patterns of multiple target regions of a non-cancerous sample. In some embodiments, the cancer sample is from a stage IV cancer sample, such as, for example, metastatic breast cancer. In some embodiments, an MRD signature is developed by determining and analyzing methylation patterns of multiple target regions in both cancer and non-cancer samples, where the multiple target regions comprise at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% of the target regions listed in Table 1. The methylation patterns of the cancer samples can then be compared to the methylation patterns of the non-cancer samples to develop an MRD signature, which is discussed in more detail below.
[0144] Any computer-readable medium herein may be non-transitory (e.g., volatile memory such as DRAM or SRAM, magnetic storage, optical storage, etc.) and / or tangible media. Store operations described herein may be implemented by storing on one or more computer-readable media (e.g., computer-readable storage media or other tangible media). Things described as "stored" (e.g., data created and used during implementation) can be stored on one or more computer-readable media (e.g., computer-readable storage media or other tangible media). Computer-readable media may be limited to embodiments not consisting of signals.
[0145] Results and Discussion Breast cancer is the second leading cause of cancer deaths among women in the United States. Metastatic breast cancer (MBC) accounts for 10–15% of breast cancer cases but has the poorest prognosis, with a 5-year overall survival rate (OS) of less than 25%. MBC arises from disseminated cells derived from the primary tumor mass before treatment and / or minimal residual disease (MRD) remaining after treatment. Molecular-based clinical tests have improved our ability to stratify patients based on recurrence risk using molecular profiles derived from primary tumor tissue. However, tumor tissue is not always available and provides only a momentary tumor status. Therefore, biomarkers that can be noninvasively and repeatedly monitored over a period of time to predict recurrence risk are needed. Cell-free (cf)DNA is a promising biomarker for detecting MRD. In this study, we use cfDNA methylation patterns as a marker of MRD. We developed a bioinformatic pipeline to extract methylation information from each cfDNA fragment, rather than the average methylation (beta) value. Our results indicate that this method may be sensitive for detecting cancer-specific methylation in low-disease environments, such as MRD. We compared this "fragment-level" approach and beta values with an in silico spike-in model, allowing us to determine the theoretical detection limits of each approach. The software can detect evidence of residual disease in a longitudinal cohort of women who relapsed after primary treatment and disease-free survivors (DFS). The cohort consisted of blood samples taken at four time points: pre-treatment, during treatment, and post-treatment. The test consists of 1,564 differentially methylated regions (DMRs), some or all of which can be used to detect MRD, breast cancer recurrence, or MBC, identifying women at high risk of recurrence who may benefit from additional therapy. This marks a major step toward developing a blood test to monitor and predict distant breast cancer recurrence.
[0146] Previous studies have shown that detectable differences in beta values exist between healthy individuals and MBC patients. However, recent studies have demonstrated that fragment-level methylation is a powerful analytical tool for detecting unique methylation states, especially in cfDNA (Klein et al. Ann Oncol 32, 1167-1177, (2021); Liu et al. Mol Cancer 20, 36, (2021); Moss et al. Nat Commun 9, 5068, (2018); Kang et al. Genome Biol 18, 53, (2017); Guo et al. Nat Genet 49, 635-642, (2017)). To date, there are no publicly available tools for analyzing fragment-level methylation data, nor have direct comparisons been made with beta-based analyses.
[0147] Beta values are the proportion of methylated CpGs at a particular locus (
number
[0148] A novel approach to methylation assessment evaluates each DNA fragment (sequencing read) to determine CpG methylation patterns. This fragment-level approach offers several potential advantages over beta values: 1) the status of neighboring CpGs is retained in the fragment-level data. When calculating region methylation, beta values for individual CpGs may be calculated, but these are averaged to provide a single index of differentially methylated regions (DMRs). In contrast, fragment-level approaches aggregate the methylation status of all neighboring CpGs and retain the status of the surrounding CpGs. This increased data resolution is expected to improve classifier performance in early detection of MRD because it captures biological variation across all possible DNA methylation states. For example, in fragments IV and V in Figure 1, the average beta value for the last four CpGs is 0.5, resulting in a delta-beta value of 0. However, these fragments have completely opposite methylation patterns, suggesting different tissue origins. 2) Fragment-level methylation allows for Boolean feature classification—that is, assessing whether a cancer-specific fragment of DNA (CSF) is present or absent in a given sample. This approach may be more sensitive in situations of low tumor burden, such as MRD.
[0149] We developed a computer-implemented package for assessing fragment-level DNA methylation—Fragment Level Assessment and Methylation Extraction (FLAME). FLAME consists of three main functions: CpG clustering, methylation aggregation, and fragment analysis.
[0150] FLAME's CpG clustering subroutine joins adjacent CpGs into discrete blocks containing n–m CpGs, where n and m are user-specified. Importantly, these blocks must be shorter than the fragment length. To assess the methylation patterns of CpGs within a block, reads must span these CpGs. The clustering algorithm works in two stages: 1) Join adjacent CpGs into contiguous regions. Regions containing fewer than n CpGs are removed, while regions containing n–m CpGs and shorter than a maximum length are retained. 2) Remaining regions are recursively split until they satisfy user-specified constraints or are deemed inappropriate. Region refinement is performed using k-means clustering based on the nearest neighboring CpGs. Regions that pass the filter are hereafter referred to as "fragment assessment regions (FARs)."
[0151] After FARs are selected, methylation aggregation can be performed to count all methylation states within each fragment. FLAME accepts two files as input: a bedGraph file listing the genomic coordinates and CpG counts for each FAR, and a bam file containing mapped reads. The program filters the bam file to retain only reads that overlap with FARs, thereby reducing execution time and memory usage in subsequent steps. For each read, the genomic coordinates, mapping information, and methylation state are recorded in a custom data structure. Methylation aggregation consists of the following steps: (a) identifying all possible methylation patterns based on the number of CpGs within the FARs; (b) selecting all reads that overlap with the FARs; (c) extracting the methylation state of each CpG within the reads that span the FARs; and (d) counting each unique methylation pattern within the fragment and returning a data structure similar to Table 2. If no reads overlap with the FARs, all values are assigned as NA. Finally, FLAME outputs a table detailing the FAR location, methylation pattern, and count value in each row. [Table 2]
[0152] Finally, FLAME combines fragment counts from multiple samples, normalizes them based on sequencing depth, and searches for differentially expressed fragments between groups (i.e., to compare and distinguish methylation patterns observed in cancer samples from those found in healthy control samples). Additionally, FLAME supports data visualization and provides export functionality for further analysis in packages such as SAS, SPSS, and Microsoft Excel. As used herein, "sequencing depth" refers to the number of times a genomic locus is covered by reads (e.g., 10 reads covering a locus = 10X sequencing depth). In some embodiments, FLAME includes at least the description of embodiments 11-14.
[0153] Testing the Method First, we conducted beta testing to ensure that FLAME worked as intended. We compared the region coverage assessed by FLAME with that of our standard methylation pipeline. We also compared beta values, which can be easily calculated using fragment-level outputs. We observed a close linear relationship between beta values calculated by FLAME and those calculated by bsseq, a widely used R package for bisulfite sequencing data analysis (Figure 2a). FLAME calculated a lower total coverage than bsseq, because FLAME only counts reads that completely span the FAR (Figure 2b).
[0154] In the absence of publicly available software that provides a similar comparison for fragment analysis, we used synthetic datasets with known fragment methylation patterns and counts. To generate these test datasets in silico, we focused on selected genome (hg19) sequences and generated synthetic reads by simulating "methylation" with C-to-T substitutions using predefined patterns and counts. Fastq files were generated from these sequences, and alignment and fragment-level aggregation were performed using the same software and parameters as for biological samples. This provided a means to compare the aggregated data to "ground truth." We found that the methylation patterns counted by FLAME matched our expected counts in these synthetic samples (Figure 3), except in cases where alignment was poor.
[0155] To assess BC-specific methylation signals with FLAME, we in silico spiked in cfDNA from MBC patients into cfDNA from healthy individuals, using it as a surrogate for cfDNA tumor burden, to determine the theoretical detection limit. We found cancer-specific fragments (CSFs) previously observed in MBC patients but not in healthy individuals. We observed CSFs in low simulated tumor fractions, demonstrating that this method can detect evidence of MRD even in low tumor burden scenarios (Figure 4).
[0156] To evaluate FLAME's ability to identify genome-wide DNA methylation changes, we compared genome-wide fragment-level data output with Metilene's DMR results. DMRs are based on beta value differences, and significant overlap was identified between the DMRs detected by beta values and regions detected by fragment-level analysis. We then compared the genomic loci of all significantly differentially methylated fragments with Metilene's DMRs and assessed the percentage of overlapping regions and regions gained. We then identified regions with significantly different DNA methylation patterns but not detected by Metilene due to low average delta-beta across the region. These regions represent methylation changes detectable only by fragment-level methods.
[0157] Statistical Planning and Machine Learning The output can be evaluated in two ways. First, we compare the sensitivity and specificity of a model built using beta values from cfDNA obtained from MBC patients with that of a machine learning (ML) model built using fragment-level data. Briefly, MBC and healthy samples are divided into a 70 / 30 train / test set. A matrix containing fragment-level data and beta value data is used to train ML models to predict MBC for healthy samples. These models are constructed using multiple algorithms, including but not limited to random forests (RF), support vector machines (SVM), neural networks, generalized linear models (GLM), gradient boosting models (GBM), extreme gradient boosting (XGB), or deep learning algorithms. As a precaution against overfitting the final model, the training may be repeatedly subdivided (iterative cross-validation) during the training process. The test set is evaluated with the final model, and the model's sensitivity and specificity are assessed using receiver operating characteristic (ROC) analysis. We construct ROC curves and calculate the area under the curve (AUC) to evaluate performance. We then directly compare the AUCs of the models built on the fragment-level data and the beta data. This process may be repeated multiple times to obtain an average AUC for both input formats.
[0158] Previously published results of WGBS on cfDNA obtained from plasma samples from three cohorts (40 individuals each): Cohort 1, MBC from various organs; Cohort 2, disease-free survivors; and Cohort 3, samples from healthy women with no history of cancer. Differential methylation analysis of the WGBS data observed relatively few differences between healthy subjects and disease-free survivors. This is indicated by the relatively small number of differentially methylated loci (n = 87,935) and the high Pearson correlation coefficient (0.83), as well as by hierarchical clustering and principal component analysis (Figure 5). On the other hand, there was a significant difference of approximately 5.0 x 10 between women with and without metastatic disease. 6Differentially methylated loci were detected, suggesting that cfDNA methylation patterns may be useful for monitoring treatment and detecting the presence of minimal residual disease.
[0159] As a continuation of this work, the R01 study was expanded to include the following subtype-specific pools: ER+ / HER2- (n=13), ER- / HER2+ (n=8), ER+ / HER2+ (n=9), TNBC (n=5), and TNBC-AA (African American-specific pool, n=20). Three additional pools of healthy individuals were created as controls: two from primarily European women (n=9, 10) and one from African American women (n=20). All pools were profiled by WGBS. After sequencing, paired-end reads were aligned to hg19 (GRCh37) using Bismark Bisulfite Read Mapper (Krueger et al., Bioinformatics 27, 1571–1572, doi:10.1093 / bioinformatics / btr167(2011)). DMRs were identified using the open-source software Metilene (Juhling et al. Genome Res 26, 256–262, doi:10.1101 / gr.196394.115(2016)). DMRs were filtered based on |Δβ|, FDR-corrected p-value, and sequencing depth to obtain a multisubtype signature consisting of 713 regions. To validate this MBC signature, we designed a targeted assay using hybrid probe capture, enabling cost-effective sequencing of multiple samples. We captured and sequenced 96 samples (64 MBC, 32 healthy) to validate the 713 DMRs previously identified as MBC signatures. A random forest model was constructed from the average β values across the 713 regions to distinguish MBC patients from healthy individuals. Thirty percent of the samples were excluded as a test set. These results demonstrated that these regions could distinguish healthy samples from MBC with high accuracy (Figure 6), with an AUC of 0.93, a sensitivity of 0.8, a specificity of 0.9, a positive predictive value of 0.94, and a negative predictive value of 0.69.
[0160] We next assessed the methylation profiles of BC patients before and after treatment to determine whether BC signals could detect evidence of MRD from beta or fragment-level data.We collected cfDNA from a cohort of 119 women with BC at four time points (before neoadjuvant therapy, after neoadjuvant therapy, after surgery, and 1 year after treatment) (Table 3) (hereafter referred to as the "Mayo cohort"). [Table 3]
[0161] To generate pilot data from this cohort, we analyzed 16 cfDNA plasma samples from four patients from this dataset. Two of these patients relapsed, one had never been in disease remission (according to the physician's judgment), and one had not relapsed at the time of analysis. To preliminarily evaluate the potential utility of using cfDNA methylation to detect evidence of relapse, we attempted to predict relapse by: 1) predicting disease status using the beta value and the RF model described in detail above, and 2) tabulating the number of CSF fragments in each sample using FLAME (Figure 7). Here, CSF fragments were defined as methylation patterns detected in at least 5% of the 64 stage IV samples and not detected in normal cfDNA samples. Fragment counts were tabulated using a proof-of-concept version of the software. Our results show that the RF model constructed using the beta value does not show significant change across time points, whereas CSF fragments show a clear decrease in signal in DFS, an increase in signal in relapse samples, and a slight increase in signal in samples that had never been in disease remission.
[0162] While particular embodiments have been described with reference to disclosed embodiments and examples, such embodiments are merely illustrative and do not limit the scope of the invention. Changes and modifications may be made in accordance with one of ordinary skill in the art without departing from the scope of the invention in its broader aspects, as defined in the following claims.
[0163] All publications, patents, and patent documents are incorporated by reference herein, as though individually incorporated by reference. These include Legendre et al. Clin Epigenetics 2015 Sep 16;7(1):100. doi:10.1186 / s13148-015-0135-8; Buckley et al., Clin Cancer Res. 2023 Oct 9. doi:10.1158 / 1078-0432. CCR-23-1197. PMID:37812492; U.S. Patent No. 10,525,148 (Salhia et al.); U.S. Patent No. 11,035,849 (Salhia et al.); U.S. Patent Application Publication No. 20200340062 (Salhia et al.); WO / 2020150258 (Olsen et al.); and WO / 2022 / 178108 (Salhia et al.). No limitations inconsistent with this disclosure should be understood therefrom. The invention has been described with reference to various specific and preferred embodiments and techniques. However, it should be understood that many variations and modifications are possible within the spirit and scope of the invention.
Claims
1. 1. A method for determining whether a subject has minimal residual disease (MRD), comprising: a) training a machine learning model to develop an MRD signature, wherein the machine learning program is trained using target regions from cancer samples and corresponding target regions from non-cancer samples, and the MRD signature is based on a comparison of the methylation patterns of the target regions in the cancer samples with the methylation patterns of the corresponding target regions in the non-cancer samples; b) determining the methylation pattern of the target region in a cell-free deoxyribonucleic acid (cfDNA) sample obtained from said subject; c) applying the MRD signature to the methylation pattern of the target region of the cfDNA obtained from the subject; and d) determining whether the subject has MRD based on the MRD signature.
2. 2. The method of claim 1, wherein the target region of the cfDNA sample from the subject is identical to a target region of both the cancer sample and the non-cancer sample used to develop the MRD signature.
3. 3. The method of claim 2, wherein the methylation pattern of the target region is determined using any one or more of whole genome library hybrid probe capture, enzymatic treatment, bisulfite amplicon sequencing (BSAS), bisulfite treatment of DNA, methylation-sensitive polymerase chain reaction, and a combination of bisulfite conversion and bisulfite restriction analysis.
4. 2. The method of claim 1, wherein the methylation pattern of each of the target regions is determined using a hybrid probe capture method.
5. 5. The method of claim 4, wherein the hybrid probe capture method comprises using one or more hybrid capture probes comprising ribonucleic acid or deoxyribonucleic acid.
6. 6. The method of claim 5, wherein each of the one or more hybrid capture probes further comprises an affinity tag selected from the group consisting of biotin and streptavidin.
7. 10. The method of claim 1, wherein the target regions from the cancer sample and the non-cancerous sample comprise about 60% to about 70% of the target regions of Table 1.
8. 8. The method of claim 7, wherein the target region comprises about 70% to about 80% of the target region of Table 1.
9. 9. The method of claim 8, wherein the target region comprises about 80% to about 90% of the target region of Table 1.
10. 10. The method of claim 9, wherein the target region is greater than about 95% of the target region of Table 1.
11. 10. The method of claim 1, wherein the cfDNA sample is extracted from whole blood, plasma, serum, or urine.
12. e) combining adjacent CpGs of each said target region into n to m consecutive CpG blocks, where n is equal to or greater than 1 and m is less than the length of the corresponding target region; f) removing target regions with fewer than n CpG blocks and target regions with more than m CpG blocks; and g) filtering the target regions remaining after step f) using a k-means clustering function based on neighboring CpGs to provide one or more Fragment Assessment Regions (FARs); The method of claim 1 further comprising:
13. h) identifying all or substantially all possible methylation patterns of CpGs within said FAR; i) selecting all sequencing reads that overlap with said FAR; j) extracting the methylation status of each CpG within the sequencing reads spanning the FAR; k) counting each different methylation pattern within the FAR to obtain a number of methylation states; and l) outputting the results of steps h) to k), wherein said output includes one or more of the location of said FARs, the methylation pattern of said FARs, and the number of said FARs; 13. The method of claim 12, further comprising summarizing the methylation status of each FAR according to:
14. adding up the respective numbers of the FARs; Normalizing the number of FARs based on sequencing depth; and identifying FARs that are differentially methylated between the subject's cfDNA sample and the cancer and non-cancer samples; 14. The method of claim 13, further comprising:
15. 10. The method of claim 1, wherein a trained machine learning model is used to determine whether the subject is likely to have or develop metastatic breast cancer, a breast cancer recurrence, or both metastatic breast cancer and a breast cancer recurrence.
16. 16. The method of claim 15, wherein the machine learning model comprises one or more of a random forest, a support vector machine (SVM), a neural network, a generalized linear model (GLM), a gradient boosting model (GBM), an extreme gradient boosting (XGB), and a deep learning algorithm.
17. 2. The method of claim 1, wherein the cancer sample and the non-cancerous sample comprise one or more of a breast cancer sample, a known metastatic breast cancer sample, a breast cancer recurrence sample, a sample from a subject who has completed a cancer treatment regimen, and a sample from a subject with no evidence of disease on standard of care.
18. 18. The method of any one of claims 1 to 17, further comprising treating the subject with the MRD, wherein the treatment comprises one or more of radiation therapy, surgery to remove the cancer, and administration of a therapeutic agent to the patient, thereby treating the MRD.