Method for Detection and Reduction of Methylation Artifacts Induced by Sample Preparation

JP2025523964A5Pending Publication Date: 2026-07-17GUARDANT HEALTH INC

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
GUARDANT HEALTH INC
Filing Date
2023-07-21
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing sequencing workflows inaccurately reflect the methylation states and variants of DNA due to synthesis during the end-repair process, leading to artifact information that misrepresents the original DNA fragment.

Method used

Incorporating modified deoxynucleotide triphosphates (dNTPs) such as dNTPs containing 5-methylcytosine or 5-hydroxymethyl-cytosine into the terminal repair reaction, followed by modified-sensitivity sequencing to identify and filter out synthesized regions, ensuring accurate sequencing data reflects the original DNA molecule.

Benefits of technology

The method provides accurate sequencing data by excluding artifact regions, allowing for precise determination of methylation states and variant detection in DNA samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present disclosure relates to improved sequencing reactions. Specifically, the present disclosure provides a method that enables the identification of regions of DNA molecules synthesized during end repair and / or A-tailing reactions. Sequence data derived from such synthesized regions may not represent the corresponding regions in the original DNA molecule; for example, it may contain the methylation state of artifacts. Thus, the identification of these synthesized regions allows for the identification of data of such potential artifacts and their corresponding filtering.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - References to Related Applications This application claims the benefit of priority based on U.S. Provisional Patent Application No. 63 / 391,213, filed on July 21, 2022, and U.S. Provisional Patent Application No. 63 / 513,250, filed on July 12, 2023, each of which is hereby incorporated by reference in its entirety for all purposes.

[0002] Field of the Invention The present disclosure relates to methods for identifying regions of sequenced DNA that are synthesized during the end - repair process, which are commonly present in sequencing workflows. Such methods are important for accurately detecting methylation states and variants present in DNA, which can in turn be important for inferring information about the cells and subjects from which the DNA sample was derived.

Background Art

[0003] Background DNA end - repair is a common step in sequencing workflows for preparing DNA for adapter ligation. Sample DNA molecules typically contain a mixture of blunt ends, 3' - overhangs, and 5' - overhangs. The end - repair step standardizes these ends by removing 3' - overhangs and filling in 5' - overhangs to generate blunt - end DNA, which can then be subjected to blunt - end adapter ligation or to A - tail addition prior to sticky - end ligation. Such end - repair steps are commonly used in sequencing workflows, such as those used for variant analysis and / or those used to determine the methylation state of nucleotides at single - base resolution.

[0004] For example, single nucleotide resolution assays for detecting nucleoside methylation often require the conversion of modified nucleosides or their corresponding unmodified nucleosides in order to alter their base pairing specificities. The conversion is then detected by sequencing. Examples of such methods include bisulfite and oxidative bisulfite as well as Tet-assisted bisulfite conversion, EM-seq, TAPS and TAPSβ conversion, ACE-seq and SEM-seq. See, for example, Moss et al., Nat Commun. 2018; 9: 5068; Booth et al., Science 2012; 336: 934-937; Yu et al., Cell 2012; 149: 1368-80; Liu et al., Nature Biotechnology 2019; 37:424-429; Schutsky, E.K. et al.; Vaisvila et al. Genome Research 2021 31(7): 1280-1289; and Vaisvila, et al. 2023. The discovery of novel DNA cytosine deaminase activity enables non-destructive single-enzyme methylation sequencing methods for base-resolution high-coverage methylome mapping of cell-free DNA and ultra-low input DNA. bioRxiv doi: 10.1101 / 2023.06.29.547047 (doi.org / 10.1101 / 2023.06.29.547047).

[0005] Bisulfite-based EM-seq and SEM-seq methylation assays convert unmethylated cytosine to uracil, which is PCR amplified and read as thymine in the sequencing reaction. In other methods, the epigenetic conversion is reversed, i.e., the modified nucleoside rather than the unmodified nucleoside is converted. For example, in the TAPS method, methylated cytosines (5mC and 5hmC) are converted to dihydrouridine (DHU), which is PCR amplified and read as thymine in the sequencing reaction. Alternatively, single molecule sequencing technologies such as nanopore-based sequencing and single molecule real-time (SMRT) sequencing can be used to detect the modification status of nucleosides. These single molecule technologies can directly detect the specific modification status of nucleosides and thus do not rely on converting the modified nucleoside or the corresponding unmodified nucleoside to change their base pairing specificities.

[0006] During the end repair step, repair and synthesis of regions of the DNA molecule can occur, which means that the modification status of nucleosides in these regions of the end-repaired DNA may not accurately reflect the modification status of the corresponding modification status in the original DNA fragment. Similarly, the templating nature of end repair can mean that it eradicates any mismatches between complementary strands (e.g., as a result of single-strand mutations). This can lead to the incorrect conclusion that the detected sequence has a double-stranded support (i.e., sequence data from both strands of the original DNA molecule that supports that sequence is present). As a result, using end repair can lead to artifact information that does not represent the original DNA fragment, which in turn can lead to inaccurate inferences regarding the DNA sample and the subject from which the DNA sample was obtained. Considering the problems surrounding artifact information in these assays, there is a need for an improved method that enables the identification of sequencing data that is known to accurately reflect the original DNA fragment.

Prior Art Documents

Non-Patent Documents

[0007] [Non-Patent Document 1] Moss et al., Nat Commun. 2018; 9: 5068 [Non-Patent Document 2] Booth et al., Science 2012; 336: 934-937 [Non-Patent Document 3] Yu et al., Cell 2012; 149: 1368-80 [Non-Patent Document 4] Liu et al., Nature Biotechnology 2019; 37:424-429 [Non-Patent Document 5] Vaisvila et al. Genome Research 2021 31(7): 1280-1289 [Non-Patent Document 6] bioRxiv doi: 10.1101 / 2023.06.29.547047 (doi.org / 10.1101 / 2023.06.29.547047). [Summary of the Invention] [Means for Solving the Problems]

[0008] Abstract The present disclosure provides embodiments that include methods that enable differentiating sequence data from an original DNA molecule from sequence data from regions of a DNA molecule synthesized during an end repair reaction (the "synthetic regions"). Thus, the method can be used to avoid using sequence data derived from these synthesized regions to determine the characteristics of the original DNA molecule. For example, if the method is used to determine the methylation state of cytosine in an original DNA molecule, the observed methylation state in the synthesized region may not accurately reflect the methylation state of the corresponding nucleoside in the original DNA molecule. Thus, the method can be used to determine the modified state of a nucleoside using data that is known to be derived from the original DNA molecule (i.e., not the synthesized regions) and thus accurately represents the methylation state of the nucleosides in the original DNA molecule. Similarly, identification of such synthesized regions can be used to interpret sequencing data regarding variant detection. For example, because the synthesized regions use the complementary strand as a template, any base mismatches that were present in this region of the original DNA molecule are eliminated. As such, any variants detected within the synthetic region cannot be reliably classified as having a duplex support (i.e., sequence data from both strands of the original DNA molecule that support that variant).

[0009] The disclosed method achieves this by using, in a terminal repair reaction, at least one type of dNTP that includes a modified base (such as a methylated deoxythymidine triphosphate, such as deoxycytidine triphosphate containing 5-methylcytosine (5mC) and / or 5-hydroxymethyl-cytosine (5hmC)). During the terminal repair reaction, the methylated deoxythymidine triphosphate is incorporated into the synthesized region regardless of the sequence context. This results in methylated cytosines at non-CpG positions, which are very rare in nature. Thus, these methylated non-CpG cytosines can be used as a label to identify the synthesized region in the terminally repaired DNA molecule. Similarly, other types of dNTPs containing modified bases that are not common or do not exist in nature can also be used. Identification of such modified bases can be performed using modification-sensitive sequencing, and regions containing these modifications can be interpreted as defining the regions synthesized in the terminal repair reaction.

[0010] Subsequent sequence analysis of the DNA sample can then focus on regions of the original DNA molecule that are known not to be synthesized during the terminal repair reaction and thus accurately reflect the characteristics (such as the modification state) of the original DNA molecule. Thus, this method provides sequencing data with improved accuracy because it allows subsequent sequence analysis to focus on regions that do not contain artifact information (such as methylation status) related to the regions of DNA synthesized during the terminal repair reaction. This method can also be used to quantify the level of DNA damage in the original DNA molecule through quantitative analysis of the synthesized regions.

[0011] In one aspect, the present disclosure provides a method comprising: (a) subjecting a DNA sample to end repair to generate end-repaired DNA, wherein the end repair is performed using deoxynucleotide triphosphates (dNTPs), and at least one type of dNTP contains a modified base; (b) performing a ligation reaction to ligate an adapter to the end-repaired DNA to generate adapted DNA, wherein the ligation reaction also seals nicks present in the end-repaired DNA; (c) subjecting the adapted DNA to modified-sensitivity sequencing to obtain sequencing data derived from the DNA sample, wherein the modified-sensitivity sequencing is capable of identifying base modifications in at least one type of dNTP; and (d) analyzing the sequence data obtained in step (c) to identify one or more regions of the end-repaired DNA synthesized during end repair based on the presence of base modifications in at least one type of dNTP used in step (a).

[0012] In some embodiments, the end repair is performed using a DNA polymerase that does not have 5'-3' exonuclease activity and / or is not a strand-displacing DNA polymerase. In some embodiments, the DNA polymerase is T4 DNA polymerase, T7 DNA polymerase, or the Klenow fragment. In some embodiments, the end repair is performed using a DNA polymerase that has 5'-3' exonuclease activity and / or is a strand-displacing DNA polymerase. In some embodiments, at least one type of dNTP containing a modified base includes a dNTP containing 4-methylcytosine (4mC), a dNTP containing 5-methylcytosine (5mC), a dNTP containing 5-hydroxymethyl-cytosine (5hmC), a dNTP containing N6-methyladenosine (6mA), a dNTP containing bromodeoxyuridine (BrdU), and / or a dNTP containing 8-oxoguanine (8oxoG). In some embodiments, the method further comprises performing an A-tailing reaction between steps (a) and (b).

[0013] In some embodiments, the end repair and A-tailing reactions are performed in a single tube. In some embodiments, A-tailing is performed using a DNA polymerase that does not have 5'-3' exonuclease activity and / or is not a strand-displacing DNA polymerase, and optionally, the DNA polymerase is HemoKlen Taq. In some embodiments, A-tailing is performed using Taq DNA polymerase, Tfl DNA polymerase, Bst DNA polymerase, the large fragment or Tth DNA polymerase.

[0014] In some embodiments, the end repair and A-tailing reactions are performed as separate reactions, and a reaction clean-up step is performed after the end repair and before the A-tailing reaction. In some embodiments, A-tailing is performed using a DNA polymerase that does not have 3'-5' exonuclease activity, such as the Klenow fragment lacking 3'-5' exonuclease activity. In some embodiments, A-tailing is performed using a DNA polymerase that has 5'-3' exonuclease activity and / or is a strand-displacing DNA polymerase. In some embodiments, the ligation reaction is a blunt-end ligation reaction. In some embodiments, the ligation reaction is a sticky-end ligation reaction.

[0015] In some embodiments, modification-sensitive sequencing includes a conversion procedure that changes or does not change the base pairing specificity of a base depending on the modified state of the base. In some embodiments, the modified base is methylated cytosine, and the conversion procedure converts methylated cytosine. For example, the conversion procedure is Tet-assisted conversion using a substituted borane reducing agent. Optionally, the substituted borane reducing agent is 2-picolyl borane, borane pyridine, tert-butylamine borane, or ammonia borane. In some embodiments, the modified base is methylated cytosine, and the conversion procedure converts non-methylated cytosine. For example, the conversion procedure is bisulfite conversion, oxidative bisulfite (Ox-BS) conversion, Tet-assisted bisulfite (TAB) conversion, APOBEC-coupled epigenetic (ACE) conversion, enzymatic methyl-seq (EM-seq), or single-enzyme 5-methylcytosine sequencing (SEM-seq). In some embodiments, modification-sensitive sequencing includes nanopore-based sequencing or single molecule real-time (SMRT) sequencing. In some embodiments, modification-sensitive sequencing includes nanopore-based sequencing, and at least one type of dNTP containing a modified base includes a dNTP containing 4mC (4-methyl deoxycytidine), a dNTP containing 5mC (5-methyl deoxycytidine), a dNTP containing 5hmC (5-hydroxymethyl deoxycytidine), a dNTP containing 6mA (6-methyl deoxyadenosine), a dNTP containing BrdU (5-bromo deoxyuridine), dUTP, a dNTP containing FldU (5-fluoro deoxyuridine), a dNTP containing IdU (5-iodo deoxyuridine), and / or a dNTP containing EdU (5-ethynyl deoxyuridine). In some embodiments, modification-sensitive sequencing includes single molecule real-time (SMRT) sequencing, and at least one type of dNTP containing a modified base includes a dNTP containing 4mC, a dNTP containing 5mC, a dNTP containing 5hmC, a dNTP containing 6mA, and / or a dNTP containing 8oxoG (8-oxo deoxyguanosine).In some embodiments, at least one type of dNTP containing a modified base includes a dNTP containing 4mC, a dNTP containing 5mC, a dNTP containing 5hmC, a dNTP containing 6mA, a dNTP containing BrdU, dUTP, a dNTP containing FldU, a dNTP containing IdU, a dNTP containing EdU, and / or a dNTP containing 8oxoG.

[0016] In some embodiments, the modified base is 5mC or 5hmC. The step of analyzing the sequence data obtained in step (c) includes identifying regions containing modified bases (e.g., 5mC or 5hmC) in a non-CpG sequence context and classifying these regions as regions synthesized during end repair. A modified base occurs in a non-CpG sequence context when a base other than G is immediately 3' to the modified base. In some embodiments, the modified base is other than 5mC or 5hmC, and one of the one or more regions is defined as: (i) a sequence between two unmodified bases spanning the modified base, where the bases are of the same entity as the modified base present in at least one type of dNTP; and / or (ii) a sequence between an unmodified base and the end of the sequence read, where there are no additional unmodified bases between the unmodified base and the end of the sequence read and the unmodified base is of the same entity as the modified base present in at least one type of dNTP.

[0017] In some embodiments, the modified base is a methylated cytosine such as 5mC or 5hmC, and a predetermined one of the one or more regions is defined as: (i) a sequence between two unmethylated cytosines spanning the methylated non-CpG cytosine; and / or (ii) a sequence between an unmethylated cytosine and the end of the sequence read, where there are no additional unmethylated cytosines between the unmethylated cytosine and the end of the sequence read.

[0018] In some embodiments, the method further includes: (i) filtering out and removing sequence data from one or more identified regions synthesized during end repair such that the sequence data is not used in subsequent analysis; or (ii) flagging sequence data from one or more identified regions synthesized during end repair as potentially containing artifact sequence data that may not represent the DNA sample. In some embodiments, the method further includes analyzing at least a portion of the sequence data corresponding to regions not identified as synthesized during end repair to detect the presence or absence of base modifications or mutations present in the DNA sample. In some embodiments, it is for detecting the methylation state of cytosine in a DNA sample, and the step of analyzing the sequence data includes filtering out and removing one or more regions of end-repaired DNA identified as synthesized during end repair such that the one or more regions are not used to determine the methylation state of cytosine in the DNA sample. In some embodiments, the method is for detecting single nucleotide variants (SNVs) in a DNA sample, and the step of analyzing the sequence data includes classifying all base calls within one or more regions as not having double-stranded support.

[0019] In some embodiments, the method further includes quantifying DNA damage in a DNA sample through identification of one or more regions of end-repaired DNA synthesized during end repair. In some embodiments, the level of DNA damage is used to predict whether a portion of the DNA in the DNA sample is derived from cancerous cells. In some embodiments, the DNA sample includes cell-free DNA (cfDNA), and the method further includes analyzing the sequence data obtained in step (c) to determine the level of measured artifacts in the cfDNA. In some embodiments, the method further compares the level of measured artifacts to one or more reference levels. In some embodiments, the method includes (i) predicting whether a portion of the cfDNA is derived from cancerous cells using the level of measured artifacts and / or the comparison of the level of measured artifacts to one or more reference levels, or (ii) determining the probability that a portion of the cfDNA is derived from cancerous cells using the level of measured artifacts and / or the comparison of the level of measured artifacts to one or more reference levels. In some embodiments, the DNA sample includes cell-free DNA or DNA from a formalin-fixed paraffin-embedded sample.

[0020] In another aspect, the present disclosure provides a method comprising: (a) subjecting a DNA sample to end repair to generate end-repaired DNA, wherein the end repair is performed using deoxynucleotide triphosphates (dNTPs) including 5mCTP; (b) performing a ligation reaction to ligate an adapter to the end-repaired DNA to generate adapted DNA, wherein the ligation reaction also seals nicks present in the end-repaired DNA; (c) subjecting the adapted DNA to bisulfite sequencing to obtain sequencing data derived from the DNA sample; (d) analyzing the sequence data obtained in step (c) to identify one or more regions of the end-repaired DNA synthesized during end repair by the presence of 5mC at non-CpG positions; and (e) optionally, further analyzing at least a portion of the sequence data corresponding to regions not identified as being synthesized during end repair to detect the presence or absence of cytosine methylation in the DNA sample.

[0021] In another aspect, the present disclosure provides a method comprising: (a) subjecting a DNA sample to end repair to generate end-repaired DNA, wherein the end repair is performed using deoxynucleotide triphosphates (dNTPs), and at least one type of dNTP comprises a modified base; (b) performing a ligation reaction to ligate an adapter to the end-repaired DNA to generate adapted DNA, wherein the ligation reaction also seals nicks present in the end-repaired DNA; (c) subjecting the adapted DNA to modified-sensitivity sequencing to obtain sequencing data derived from the DNA sample, wherein the modified-sensitivity sequencing is capable of identifying base modifications in at least one type of dNTP, and the modified-sensitivity sequencing is nanopore sequencing, single molecule real-time sequencing, or Tet-assisted pyridine borane sequencing (TAPS); (d) analyzing the sequence data obtained in step (c) to identify one or more regions of the end-repaired DNA synthesized during end repair based on the presence of base modifications in at least one type of dNTP used in step (a); and (e) further analyzing at least a portion of the sequence data corresponding to regions not identified as being synthesized during end repair to detect the presence or absence of mutations present in the DNA sample.

[0022] In another aspect, the present disclosure provides a method comprising: (a) subjecting a DNA sample to end repair to generate end-repaired DNA, wherein the end repair is performed using deoxynucleotide triphosphates (dNTPs), at least one type of dNTP contains a modified base, the end repair is performed using a DNA polymerase that does not have 5'-3' exonuclease activity and / or is not a strand-displacing DNA polymerase, such as T4 DNA polymerase, T7 DNA polymerase, or Klenow fragment; (b) performing an A-tailing reaction as a single-tube reaction using this end repair, wherein, optionally, the A-tailing is performed using a DNA polymerase that does not have 3'-5' exonuclease activity and / or is not a strand-displacing DNA polymerase; (c) performing a ligation reaction to ligate an adapter to the end-repaired A-tailed DNA to generate adapted DNA, wherein the ligation reaction also seals nicks present in the end-repaired DNA; (d) subjecting the adapted DNA to modified-sensitivity sequencing to obtain sequencing data derived from the DNA sample, wherein the modified-sensitivity sequencing is capable of identifying base modifications in at least one type of dNTP; (e) analyzing the sequence data obtained in step (d) to identify one or more regions of the end-repaired DNA synthesized during end repair based on the presence of base modifications in at least one type of dNTP used in step (a); and optionally, (f) further comprising analyzing at least a portion of the sequence data corresponding to regions not identified as being synthesized during end repair to detect the presence or absence of base modifications or mutations present in the DNA sample.

[0023] In another aspect, the present disclosure provides a method comprising: (a) subjecting a DNA sample to end repair to generate end-repaired DNA, wherein the end repair is performed using deoxynucleotide triphosphates (dNTPs), at least one type of dNTP contains a modified base, the end repair is performed using a DNA polymerase that does not have 5'-3' exonuclease activity and / or is not a strand-displacing DNA polymerase, such as T4 DNA polymerase, T7 DNA polymerase, or Klenow fragment; (b) performing a reaction cleanup step on the end-repaired DNA and then performing an A-tailing reaction to generate A-tailed DNA; (c) performing a ligation reaction to ligate an adapter to the A-tailed DNA to generate adapted DNA, wherein the ligation reaction seals nicks present in the end-repaired DNA; (d) subjecting the adapted DNA to modified-sensitivity sequencing to obtain sequencing data derived from the DNA sample, wherein the modified-sensitivity sequencing is capable of identifying base modifications in at least one type of dNTP; (e) analyzing the sequence data obtained in step (d) to identify one or more regions of the end-repaired DNA synthesized during end repair based on the presence of base modifications in at least one type of dNTP used in step (a); and optionally, (f) further analyzing at least a portion of the sequence data corresponding to regions not identified as being synthesized during end repair to detect the presence or absence of base modifications or mutations present in the DNA sample.

[0024] In another aspect, the present disclosure provides a method comprising: (a) subjecting a DNA sample to end repair to generate end-repaired DNA, wherein the end repair is performed using deoxynucleotide triphosphates (dNTPs), at least one type of the dNTPs contains a modified base, the end repair is performed using a DNA polymerase that does not have 5'-3' exonuclease activity and / or is not a strand-displacing DNA polymerase, such as T4 DNA polymerase, T7 DNA polymerase, or Klenow fragment; (b) performing a blunt-end ligation reaction to ligate an adapter to the end-repaired DNA to generate adapted DNA, wherein the ligation reaction also seals nicks present in the end-repaired DNA; (c) subjecting the adapted DNA to modified-sensitivity sequencing to obtain sequencing data derived from the DNA sample, wherein the modified-sensitivity sequencing is capable of identifying base modifications in at least one type of dNTP; (d) analyzing the sequence data obtained in step (c) to identify one or more regions of the end-repaired DNA synthesized during end repair based on the presence of base modifications in at least one type of dNTP used in step (a); and optionally, (e) further analyzing at least a portion of the sequence data corresponding to regions not identified as being synthesized during end repair to detect the presence or absence of base modifications or mutations present in the DNA sample.

[0025] In some embodiments, a kit is provided that includes: a) a first reagent for end repair to generate end-repaired DNA, the first reagent comprising at least one type of dNTP containing a modified base; b) a second reagent for ligating an adapter to the end-repaired DNA to generate adapted DNA, the second reagent also sealing nicks present in the end-repaired DNA; c) a reagent for modified-sensitivity sequencing capable of identifying base modifications in at least one type of dNTP and / or DNA polymerase for incorporation of the first reagent into DNA during end repair; and d) a library adapter having distinct molecular barcodes.

[0026] In some embodiments, the kit further comprises a plurality of oligonucleotide probes and / or primers for sequencing. In some embodiments, the first reagent of the kit comprises at least one type of dNTP containing a modified base selected from dNTP containing 4-methylcytosine (4mC), dNTP containing 5-methylcytosine (5mC), dNTP containing 5-hydroxymethyl-cytosine (5hmC), dNTP containing N6-methyladenosine (6mA), dNTP containing bromodeoxyuridine (BrdU), dNTP containing 8-oxoguanine (8oxoG), dUTP, dNTP containing fluorodeoxyuridine (FldU), dNTP containing iododeoxyuridine (IdU), and / or dNTP containing ethynyldeoxyuridine (EdU). In some embodiments, the DNA polymerase of the kit does not have 5'-3' exonuclease activity and / or is not a strand-displacing DNA polymerase. In some embodiments, the DNA polymerase of the kit is T4 DNA polymerase, T7 DNA polymerase or Klenow fragment. In some embodiments, the DNA polymerase of the kit has 5'-3' exonuclease activity and / or is a strand-displacing DNA polymerase. In some embodiments, the kit further comprises a reagent for performing an A-tailing reaction. In some embodiments, the reagent for performing the A-tailing reaction comprises a DNA polymerase that does not have 5'-3' exonuclease activity and / or is not a strand-displacing DNA polymerase, and optionally, the reagent for performing the A-tailing reaction is HemoKlen Taq. In some embodiments, the reagent for performing the A-tailing reaction comprises Taq DNA polymerase, Tfl DNA polymerase, Bst DNA polymerase, the large fragment or Tth DNA polymerase. In some embodiments, the reagent for performing the A-tailing reaction comprises a DNA polymerase that does not have 3'-5' exonuclease activity, and optionally, the reagent for performing the A-tailing reaction is a Klenow fragment lacking 3'-5' exonuclease activity.In some embodiments, the reagent for performing the A-tailing reaction comprises a DNA polymerase having 5'-3' exonuclease activity and / or is a strand-displacing DNA polymerase.

[0027] In some embodiments, the kit comprises: a) a plurality of oligonucleotide probes that selectively hybridize to at least 5, 6, 7, 8, 9, 10, 20, 30, 40 or all genes selected from ALK, APC, BRAF, CDKN2A, EGFR, ERBB2, FBXW7, KRAS, MYC, NOTCH1, NRAS, PIK3CA, PTEN, RBI, TP53, MET, AR, ABL1, AKT1, ATM, CDH1, CSF1R, CTNNB1, ERBB4, EZH2, FGFR1, FGFR2, FGFR3, FLT3, GNA11, GNAQ, GNAS, HNF1A, HRAS, IDH1, IDH2, JAK2, JAK3, KDR, KIT, MLH1, MPL, NPM1, PDGFRA, PROC, PTPN11, RET, SMAD4, SMARCB1, SMO, SRC, STK11, VHL, TERT, CCND1, CDK4, CDKN2B, RAF1, BRCA1, CCND2, CDK6, NF1, TP53, ARID1A, BRCA2, CCNE1, ESR1, RIT1, GATA3, MAP2K1, RHEB, ROS1, ARAF, MAP2K2, NFE2L2, RHOA and NTRK1; b) the library adapter does not contain a flow cell sequence nor a sequence that enables the formation of a hairpin loop for sequencing; c) the library adapter is blunt-ended and Y-shaped; and / or d) the library adapter is shorter than or equal to 40 nucleobases in length. In some embodiments, the kit further comprises instructions for performing any of the methods described herein.

[0028] The embodiments described herein include, but are not limited to, the following. Embodiment 1 is (a) A step of subjecting a DNA sample to end repair to generate end-repaired DNA, wherein the end repair is performed using deoxynucleotide triphosphates (dNTPs), and at least one type of dNTP contains a modified base; (b) A step of performing a ligation reaction to ligate an adapter to the end-repaired DNA to generate adapted DNA, wherein the ligation reaction also seals nicks present in the end-repaired DNA; (c) A step of subjecting the adapted DNA to modified-sensitivity sequencing to obtain sequencing data derived from the DNA sample, wherein the modified-sensitivity sequencing is capable of identifying base modifications in the at least one type of dNTP; and (d) A step of analyzing the sequence data obtained in step (c) to identify one or more regions of the end-repaired DNA synthesized during end repair based on the presence of base modifications in at least one type of dNTP used in step (a) A method comprising.

[0029] Embodiment 2 is the method according to Embodiment 1, wherein the end repair is performed using a DNA polymerase that does not have 5'-3' exonuclease activity and / or is not a strand-displacing DNA polymerase.

[0030] Embodiment 3 is the method according to Embodiment 2, wherein the DNA polymerase is T4 DNA polymerase, T7 DNA polymerase, or Klenow fragment.

[0031] Embodiment 4 is the method according to Embodiment 1, wherein the end repair is performed using a DNA polymerase that has 5'-3' exonuclease activity and / or is a strand-displacing DNA polymerase.

[0032] Embodiment 5 is the method according to any one of Embodiments 1 to 4, wherein the at least one type of dNTP containing a modified base includes a dNTP containing 4-methylcytosine (4mC), a dNTP containing 5-methylcytosine (5mC), a dNTP containing 5-hydroxymethyl-cytosine (5hmC), a dNTP containing N6-methyladenosine (6mA), a dNTP containing bromodeoxyuridine (BrdU), and / or a dNTP containing 8-oxoguanine (8oxoG).

[0033] Embodiment 6 is the method according to any one of Embodiments 1 to 5, further including a step of performing an A-tailing reaction between steps (a) and (b).

[0034] Embodiment 7 is the method according to Embodiment 6, wherein the end repair and the A-tailing reaction are performed in the same reaction mixture, and optionally, the end repair and the A-tailing reaction are performed in a single tube, and / or optionally, the end repair and the A-tailing reaction are performed without an intervening purification step.

[0035] Embodiment 8 is the method according to Embodiment 7, wherein the A-tailing is performed using a DNA polymerase that does not have 5'-3' exonuclease activity and / or is not a strand-displacing DNA polymerase, and optionally, the DNA polymerase is HemoKlen Taq.

[0036] Embodiment 9 is the method according to Embodiment 7, wherein the A-tailing is performed using a thermostable DNA polymerase.

[0037] Embodiment 10 is the method according to Embodiment 7, wherein the A-tailing is performed using Taq DNA polymerase, Tfl DNA polymerase, Bst DNA polymerase, large fragment or Tth DNA polymerase.

[0038] Embodiment 11 is the method according to Embodiment 6, wherein the end repair and the A-tailing reaction are carried out as separate reactions, and a reaction purification step is carried out after the end repair and before the A-tailing reaction.

[0039] Embodiment 12 is the method according to Embodiment 11, wherein the reaction purification step removes unincorporated dNTPs.

[0040] Embodiment 13 is the method according to Embodiment 11 or 12, wherein the A-tailing is carried out using a DNA polymerase that does not have 3'-5' exonuclease activity, and optionally, the DNA polymerase is a Klenow fragment lacking 3'-5' exonuclease activity.

[0041] Embodiment 14 is the method according to Embodiment 7 or Embodiment 11, wherein the A-tailing is carried out using a DNA polymerase that has 5'-3' exonuclease activity and / or is a strand-displacing DNA polymerase.

[0042] Embodiment 15 is the method according to any one of Embodiments 6 to 14, wherein the A-tailing reaction is carried out at a temperature higher than that of the end repair, and optionally, the end repair is carried out at about 15 to 35°C, and / or the A-tailing is carried out at a temperature exceeding about 60°C, and further optionally, the temperature exceeding 60°C is about 60°C to 75°C.

[0043] Embodiment 16 is the method according to any one of Embodiments 1 to 5, wherein the ligation reaction is a blunt-end ligation reaction.

[0044] Embodiment 17 is the method according to any one of Embodiments 6 to 15, wherein the ligation reaction is a sticky-end ligation reaction.

[0045] Embodiment 18 is the method according to any one of Embodiments 1 to 17, wherein the modified sensitivity sequencing includes a conversion procedure that changes or does not change the base pairing specificity of the base according to the modified state of the base.

[0046] Embodiment 19 is the method according to Embodiment 18, wherein the modified base is methylated cytosine, and the conversion procedure converts methylated cytosine. For example, the conversion procedure is Tet-assisted conversion using a substituted borane reducing agent. Optionally, the substituted borane reducing agent is 2-picolyl borane, borane pyridine, tert-butylamine borane, or ammonia borane.

[0047] Embodiment 20 is the method according to Embodiment 18, wherein the modified base is methylated cytosine, and the conversion procedure converts non-methylated cytosine. For example, the conversion procedure is bisulfite conversion, oxidative bisulfite (Ox-BS) conversion, Tet-assisted bisulfite (TAB) conversion, APOBEC-coupled epigenetic (ACE) conversion, enzymatic methyl-seq (EM-seq), or single-enzyme 5-methylcytosine sequencing (SEM-seq).

[0048] Embodiment 21 is the method according to any one of Embodiments 1 to 17, wherein the modified sensitivity sequencing includes nanopore-based sequencing or single molecule real-time (SMRT) sequencing.

[0049] Embodiment 22 is the method according to Embodiment 21, wherein the modified sensitivity sequencing includes nanopore-based sequencing, and the at least one type of dNTP containing a modified base includes a dNTP containing 4mC, a dNTP containing 5mC, a dNTP containing 5hmC, a dNTP containing 6mA, a dNTP containing BrdU, dUTP, a dNTP containing FldU, a dNTP containing IdU, and / or a dNTP containing EdU.

[0050] Embodiment 23 is the method according to embodiment 21, wherein the modified sensitivity sequencing includes single molecule real-time (SMRT) sequencing, and the at least one type of dNTP containing a modified base includes a dNTP containing 4mC, a dNTP containing 5mC, a dNTP containing 5hmC, a dNTP containing 6mA, and / or a dNTP containing 8oxoG.

[0051] Embodiment 24 is the method according to any one of embodiments 1 to 23, wherein the modified base is 5mC.

[0052] Embodiment 25 is the method according to any one of embodiments 1 to 23, wherein the modified base is 5hmC.

[0053] Embodiment 26 is the method according to any one of embodiments 1 to 25, wherein the step of analyzing the sequence data obtained in step (c) includes identifying regions containing the modified base in a non-CpG sequence context and classifying these regions as regions synthesized during the end repair.

[0054] Embodiment 27 is such that the modified base is other than 5mC or 5hmC, and one of the one or more regions is (i) a sequence between two unmodified bases spanning the modified base, wherein the bases are of the same entity as the modified base present in the at least one type of dNTP; and / or (ii) a sequence between an unmodified base and the end of the sequence read, wherein there is no further unmodified base between the unmodified base and the end of the sequence read, and the unmodified base is of the same entity as the modified base present in the at least one type of dNTP as defined in the method according to any one of embodiments 1 to 23.

[0055] Embodiment 28 is such that the modified base is methylated cytosine such as 5mC or 5hmC, and a predetermined region among the one or more regions is (i) a sequence between two unmethylated cytosines spanning a methylated non-CpG cytosine; and / or (ii) a sequence between an unmethylated cytosine and the end of the sequence read, defined as a sequence in which there is no additional unmethylated cytosine between the unmethylated cytosine and the end of the sequence read, the method according to any one of Embodiments 1 to 23.

[0056] Embodiment 29 further includes the step of (i) filtering and removing sequence data from one or more regions identified as synthesized during the end repair so that these sequence data are not used in subsequent analysis; or (ii) flagging sequence data from the one or more regions identified as synthesized during the end repair as potentially containing artifact sequence data that may not represent the DNA sample, the method according to any one of Embodiments 1 to 28.

[0057] Embodiment 30 further includes the step of analyzing at least a part of the sequence data corresponding to regions not identified as synthesized during the end repair to detect the presence or absence of base modifications or mutations present in the DNA sample, the method according to any one of Embodiments 1 to 29.

[0058] Embodiment 31 is for detecting the methylation state of cytosine in the DNA sample, and the step of analyzing the sequence data includes filtering and removing one or more regions of the end-repaired DNA identified as synthesized during the end repair so that the one or more regions are not used to determine the methylation state of cytosine in the DNA sample, the method according to any one of Embodiments 1 to 30.

[0059] Embodiment 32 is for detecting single nucleotide variants (SNVs) in the DNA sample, and the step of analyzing the sequence data includes classifying all base calls in the one or more regions as not having a double-stranded support, and is the method according to any one of Embodiments 1 to 31.

[0060] Embodiment 33 further includes the step of quantifying DNA damage in the DNA sample through identification of one or more regions of the end-repaired DNA synthesized during the end repair, and is the method according to any one of Embodiments 1 to 32.

[0061] Embodiment 34 is the method according to Embodiment 33, which uses the level of the DNA damage to predict whether a part of the DNA in the DNA sample is derived from cancerous cells.

[0062] Embodiment 35 is the method according to any one of Embodiments 1 to 34, wherein the DNA sample contains cell-free DNA (cfDNA), and the method further includes the step of analyzing the sequence data obtained in step (c) to determine the level of the measured artifacts in the cfDNA.

[0063] Embodiment 36 further includes the step of comparing the level of the measured artifacts with one or more reference levels, and is the method according to Embodiment 35.

[0064] Embodiment 37 further includes the step of (i) predicting whether a part of the cfDNA is derived from cancerous cells using the level of the measured artifacts and / or the comparison of the level of the measured artifacts with one or more reference levels, or (ii) determining the probability that a part of the cfDNA is derived from cancerous cells using the level of the measured artifacts and / or the comparison of the level of the measured artifacts with one or more reference levels, and is the method according to Embodiment 35 or 36.

[0065] Embodiment 38 is the method according to any one of Embodiments 1 to 34, wherein the DNA sample contains cell-free DNA or DNA from a formalin-fixed paraffin-embedded sample.

[0066] Embodiment 39 is the method according to the immediately preceding embodiment, wherein the DNA sample contains cell-free DNA.

[0067] Embodiment 40 is the method according to any one of the preceding embodiments, wherein the sample is derived from a subject, and the method further comprises determining the presence or absence of cancer in the subject based at least in part on the sequencing data.

[0068] Embodiment 41 is the method according to any one of the preceding embodiments, further comprising at least one DNA amplification step, and optionally, the DNA amplification step is performed after step (b) and before step (c).

[0069] Embodiment 42 is the method according to the immediately preceding embodiment, wherein the DNA amplification step comprises PCR.

[0070] Embodiment 43 is the method according to any one of the preceding embodiments, wherein the adapter contains a molecular barcode.

[0071] Embodiment 44 is the method according to any one of the preceding embodiments, further comprising enriching the DNA for a plurality of target regions before step (c).

[0072] Embodiment 45 is the method according to Embodiment 44, wherein the plurality of target regions contains epigenetic target regions.

[0073] Embodiment 46 is the method according to Embodiment 45, wherein the epigenetic target regions contain hypermethylation variable target regions.

[0074] Embodiment 47 is the method according to embodiment 45 or 46, wherein the epigenetic target region includes a hypomethylation variable target region.

[0075] Embodiment 48 is the method according to any one of embodiments 44 to 47, wherein the plurality of target regions includes an array variable target region.

[0076] Embodiment 49 is a) a first reagent for end repair for generating end-repaired DNA, the first reagent comprising at least one type of dNTP containing a modified base; b) a second reagent for ligating an adapter to the end-repaired DNA to generate adapted DNA, the second reagent also sealing nicks present in the end-repaired DNA; c) a reagent for modified-sensitivity sequencing capable of identifying a base modification in the at least one type of dNTP and / or DNA polymerase for incorporating the first reagent into DNA during end repair; and d) a library adapter having a distinct molecular barcode and is a kit.

[0077] Embodiment 50 is the kit according to embodiment 49, further comprising a plurality of oligonucleotide probes and / or primers for sequencing.

[0078] Embodiment 51 is the kit according to embodiment 49 or 50, wherein the first reagent contains at least one type of dNTP containing a modified base selected from dNTP containing 4-methylcytosine (4mC), dNTP containing 5-methylcytosine (5mC), dNTP containing 5-hydroxymethyl-cytosine (5hmC), dNTP containing N6-methyladenosine (6mA), dNTP containing bromodeoxyuridine (BrdU), dNTP containing 8-oxoguanine (8oxoG), dUTP, dNTP containing fluorodeoxyuridine (FldU), dNTP containing iododeoxyuridine (IdU), and / or dNTP containing ethynyl deoxyuridine (EdU).

[0079] Embodiment 52 is the kit according to any one of embodiments 49 to 51, wherein the DNA polymerase does not have 5'-3' exonuclease activity and / or is not a strand displacement DNA polymerase.

[0080] Embodiment 53 is the kit according to any one of embodiments 49 to 52, wherein the DNA polymerase is T4 DNA polymerase, T7 DNA polymerase or Klenow fragment.

[0081] Embodiment 54 is the kit according to embodiment 49, wherein the DNA polymerase has 5'-3' exonuclease activity and / or is a strand displacement DNA polymerase.

[0082] Embodiment 55 is the kit according to any one of embodiments 49 to 54, further comprising a reagent for performing an A-tailing reaction.

[0083] Embodiment 56 is the kit according to the immediately preceding embodiment, further comprising a reagent for a purification step after the A-tailing reaction, and optionally, the reagent for the purification step is for removing unincorporated dNTPs.

[0084] Embodiment 57 is the kit according to Embodiment 55 or 56, wherein the reagent for performing the A-tailing reaction contains a DNA polymerase that does not have 5'-3' exonuclease activity and / or is not a strand-displacing DNA polymerase, and optionally, the reagent for performing the A-tailing reaction is HemoKlen Taq.

[0085] Embodiment 58 is the kit according to any one of Embodiments 55 to 57, wherein the reagent for performing the A-tailing reaction contains Taq DNA polymerase, Tfl DNA polymerase, Bst DNA polymerase, large fragment or Tth DNA polymerase.

[0086] Embodiment 59 is the kit according to any one of Embodiments 55 to 58, wherein the reagent for performing the A-tailing reaction contains a thermostable DNA polymerase.

[0087] Embodiment 60 is the kit according to any one of Embodiments 55 to 56, wherein the reagent for performing the A-tailing reaction contains a DNA polymerase that does not have 3'-5' exonuclease activity, and optionally, the reagent for performing the A-tailing reaction is a Klenow fragment lacking 3'-5' exonuclease activity.

[0088] Embodiment 61 is the kit according to any one of Embodiments 55 to 56, wherein the reagent for performing the A-tailing reaction contains a DNA polymerase having 5'-3' exonuclease activity and / or is a strand-displacing DNA polymerase.

[0089] Embodiment 62 is a) The plurality of oligonucleotide probes selectively hybridize to at least 5, 6, 7, 8, 9, 10, 20, 30, 40 or all genes selected from ALK, APC, BRAF, CDKN2A, EGFR, ERBB2, FBXW7, KRAS, MYC, NOTCH1, NRAS, PIK3CA, PTEN, RBI, TP53, MET, AR, ABL1, AKT1, ATM, CDH1, CSF1R, CTNNB1, ERBB4, EZH2, FGFR1, FGFR2, FGFR3, FLT3, GNA11, GNAQ, GNAS, HNF1A, HRAS, IDH1, IDH2, JAK2, JAK3, KDR, KIT, MLH1, MPL, NPM1, PDGFRA, PROC, PTPN11, RET, SMAD4, SMARCB1, SMO, SRC, STK11, VHL, TERT, CCND1, CDK4, CDKN2B, RAF1, BRCA1, CCND2, CDK6, NF1, TP53, ARID1A, BRCA2, CCNE1, ESR1, RIT1, GATA3, MAP2K1, RHEB, ROS1, ARAF, MAP2K2, NFE2L2, RHOA and NTRK1; b) The library adapter does not include a flow cell sequence nor a sequence enabling the formation of a hairpin loop for sequencing; c) The library adapter is blunt-ended and Y-shaped; and / or d) The library adapter has a length shorter than or equal to 40 nucleobases, The kit according to any one of embodiments 49 to 61.

[0090] Embodiment 63 is the kit according to any one of embodiments 49 to 62, further comprising instructions for carrying out the method according to any one of embodiments 1 to 48.

[0091] In some embodiments, the results of the methods disclosed herein are used as input for generating a report. The report may be in paper or electronic form. For example, the true methylation status of cytosine or a variant, or information derived therefrom, as obtained by the methods disclosed herein, can be directly shown in such a report. Alternatively or additionally, diagnostic information or therapeutic recommendations that are at least partially based on the methods disclosed herein may be included in the report.

[0092] The various steps of the methods disclosed herein can be performed at the same or different times, at the same or different geographical locations, e.g., in different countries, and / or by the same or different people.

[0093] Further advantages will be described in part in the following description or can be learned by practice. The advantages are realized and achieved by the elements and combinations particularly pointed out in the appended claims.

[0094] In FIGS. 1-6, the solid black mushroom-shaped symbols represent methylated CpG sites, and the open mushroom-shaped symbols represent unmethylated CpG sites. The thick solid line represents the region of the DNA molecule that was present in the original DNA molecule (i.e., before end repair). The dashed line represents the synthetic region synthesized during end repair, including the region added to the 3' end of the DNA molecule with a 5' overhang during end repair, as well as the internal synthetic regions resulting from gap filling and nick translation. The thin solid line represents the synthetic region synthesized during A-tailing. The shaded circles represent the polymerases used for end repair. The open circles represent the polymerases used for A-tailing. The triangle represents the ligation site. The star represents the 5mCpH site. BRIEF DESCRIPTION OF THE DRAWINGS

[0095]

Figure 1

[0096]

Figure 2

[0097]

Figure 3

[0098]

Figure 4

[0099]

Figure 5

[0100]

Figure 6

[0101]

Figure 7A-B

[0102]

Figure 8

[0103]

Figure 9

[0104]

Figure 10

[0105]

Figure 11A-B

[0106]

Figure 12A-B

Figure 12C-D

[0107]

Figure 13

Mode for Carrying Out the Invention

[0108] Detailed Description Reference is made here in detail to certain embodiments of the present disclosure. Although the present disclosure is described in conjunction with such embodiments, it is understood that such embodiments are not intended to limit the present disclosure to those embodiments. In contrast, the present disclosure is intended to cover all alternatives, modifications, and equivalents that may be included within the scope of the disclosure as defined by the appended claims. Before describing the teachings of the present invention in detail, it should be understood that the present disclosure is not limited to any particular compositions or process steps and can itself vary. It should be noted that as used in this specification and the appended claims, the singular forms "a", "an", and "the" include plural references unless the context clearly dictates otherwise. Thus, for example, reference to "a nucleic acid" includes a plurality of nucleic acids.

[0109] Numeric ranges include the numbers defining the range. Measured and measurable values are understood to be approximate, taking into account significant figures and errors associated with the measurements.

[0110] Unless specifically noted otherwise in the above specification, embodiments in the specification listing various components as "comprising" can also be considered as "consisting of" or "essentially consisting of" the listed components.

[0111] The section headings used in this specification are for purposes of organization and should not be construed as in any way limiting the disclosed subject matter.

[0112] All patents, patent applications, websites, other publications or documents, etc., cited in this specification, whether supra or infra, are hereby incorporated by reference in their entirety for all purposes to the same extent as if each individual item were specifically and individually indicated to be incorporated by reference in such manner. If different versions of publications, websites, etc., are published at different times, unless otherwise indicated, the most recently published version as of the effective filing date of this application is meant.

[0113] Definitions "Reaction cleanup" refers to the removal of contaminants such as salts, enzymes, unincorporated dNTPs, primers, ethidium bromide, and other impurities that may interfere with downstream analysis. For example, if reaction cleanup is performed between end repair and A-tailing reaction, this removes unincorporated dNTPs so that the A-tailing reaction can be carried out alone in the presence of dATP (i.e., not dCTP, dGTP, and dCTP as used in the end-tailing reaction). Reaction cleanup can be performed using a commercially available kit such as the MinElute Reaction Cleanup Kit (Qiagen).

[0114] The "synthesized region" is also referred to as the "region of end-repaired DNA synthesized during end repair" and refers to a region of DNA that did not exist in the DNA prior to the end repair reaction and the A-tailing reaction. These are, if present, regions synthesized by the polymerase used in the end repair reaction and / or the A-tailing reaction. When A-tailing is performed in the same tube as the end repair reaction, all four types of dNTPs are present, and thus the polymerase used for A-tailing can generate the synthesized region, for example, via nick translation. When A-tailing is performed separately from the end repair reaction and these steps are separated by reaction cleanup, only dATP is present during the A-tailing reaction, and thus the polymerase used for A-tailing typically does not generate the synthesized region since not all dNTP components are present in the A-tailing reaction mixture.

[0115] As used herein, "base pairing specificity" refers to the standard DNA base (A, C, G or T) to which a given base most preferentially pairs. Thus, for example, unmodified cytosine and 5-methylcytosine have the same base pairing specificity (i.e., specificity for G), while uracil has a base pairing specificity for A and cytosine has a base pairing specificity for G, so uracil and cytosine have different base pairing specificities. The ability of uracil to form a wobbling pair with G is not important, for example, because uracil nonetheless most preferentially pairs with A among the four standard DNA bases.

[0116] "Type of dNTP" refers to a dNTP containing a specific base including A, T, G or C. Thus, when a terminal repair reaction is carried out using dNTPs and at least one type of dNTP contains a modified base, the terminal repair reaction can be carried out using dCTP containing 5mC, as well as dATP, dTTP and dGTP all of which contain unmodified bases.

[0117] "Capable of identifying a base modification in at least one type of dNTP" refers to the ability of a modification-sensitive sequencing method to detect the presence or absence of a base modification in at least one type of dNTP containing a modified base used for end repair. This detection of the base modification may be direct, such as in nanopore sequencing or single molecule real-time sequencing, where the sequencing data itself indicates the presence or absence of the base modification. Alternatively, the detection of the base modification may be indirect. For example, the method may include a conversion procedure that changes base pairing specificity depending on the base modification state. These changes in base pairing specificity can be detected by the sequencing method, for example, by comparing the sequencing data to a reference sequence. Furthermore, the modification-sensitive sequencing method can identify a base modification in at least one type of dNTP, regardless of whether it can distinguish one base modification from all other base modifications. For example, one form of modification-sensitive sequencing is sequencing after bisulfite conversion. This method can distinguish 5hmC and 5mC from unmethylated cytosine, but cannot distinguish 5hmC from 5mC.

[0118] Bases of the "same entity" refer to the same base, regardless of its modification state. For example, cytosine is considered to be of the "same entity" as 5-methylcytosine (5mC) and / or 5-hydroxymethyl-cytosine (5hmC), even though they have different modification states.

[0119] "Capturing" one or more target nucleic acids refers to preferentially isolating or separating one or more target nucleic acids from non-target nucleic acids.

[0120] "Cell-free DNA", "cfDNA molecule", or simply "cfDNA" includes DNA molecules that naturally exist extracellularly in a subject (e.g., in blood, serum, plasma, or other body fluids such as lymph, cerebrospinal fluid, urine, or sputum). CfDNA was originally present in cells (singular or plural) in a large complex organism, such as a mammal, but has undergone release from the cells into the fluid found in the organism and can be obtained from a sample of the fluid without the need to perform an in vitro cell lysis step.

[0121] As used herein, "cellular nucleic acid" means a nucleic acid that is located within one or more cells from which it originated, at least at the time the sample is obtained or collected from the subject, even if those nucleic acids are subsequently removed as part of a given analytical process (e.g., via cell lysis).

[0122] DNA is "derived from cancerous cells" if it has its origin in tumor cells. Cell-free DNA derived from cancerous cells includes ctDNA or circulating tumor DNA. Tumor cells are neoplastic cells that originate from a tumor, whether they remain in the tumor or are separated from the tumor (as is the case for metastatic cancer cells and circulating tumor cells).

[0123] The term "methylation" or "DNA methylation" refers to the addition of a methyl group to a nucleotide base in a nucleic acid molecule. In some embodiments, methylation refers to the addition of a methyl group to cytosine at a CpG site (cytosine-phosphate-guanine site, i.e., cytosine followed by guanine in the 5'→3' direction of a nucleic acid sequence). In some embodiments, DNA methylation is N 6Refers to the addition of a methyl group to adenine, such as in N6-methyladenine (6mA). In some embodiments, DNA methylation is 5-methylation (modification of the carbon at the 5th position of the cytosine ring). In some embodiments, 5-methylation refers to the addition of a methyl group to the 5C position of cytosine, which generates 5-methylcytosine (5mC). In some embodiments, methylation includes derivatives of 5mC. Examples of derivatives of 5mC include, but are not limited to, 5-hydroxymethylcytosine (5-hmC), 5-formylcytosine (5-fC), and 5-carboxylcytosine (5-caC). In some embodiments, DNA methylation is 3C methylation (modification of the carbon at the 3rd position of the cytosine ring). In some embodiments, 3C methylation includes the addition of a methyl group to the 3C position of cytosine, which generates 3-methylcytosine (3mC). Methylation can also occur at non-CpG sites; for example, methylation can occur at CpA, CpT, or CpC sites. DNA methylation can change the activity of the methylated DNA region. For example, if the DNA in the promoter region is methylated, gene transcription can be suppressed. DNA methylation is important for normal development, and abnormalities in methylation can disrupt epigenetic regulation. Disruptions in epigenetic regulation, such as suppression, can cause diseases such as cancer. Promoter methylation in DNA can indicate cancer.

[0124] "Modified nucleoside profile of DNA" means the position and identity of nucleosides within a DNA sequence, as well as the modified state of the nucleosides, such as methylation. As described above, various modification-sensitive sequencing methods can be used to detect such modifications. This includes methods that involve, after conversion, sequencing and detecting one or more different types of modified or unmodified nucleosides. For example, the TAPS method detects 5-methylcytosine (5mC) and 5-hydroxymethyl-cytosine (5hmC), but does not distinguish between them. Thus, a method for analyzing the modified nucleoside profile of DNA in a sample typically means identifying a specific modification or group of modifications, such as 5mC and / or 5hmC. Modified nucleosides are identified according to the specific method / conversion procedure used as described above. This generally involves comparing sequence data obtained from DNA that has been subjected to the conversion procedure to a reference sequence. Typically, this method includes (i) comparing the sequence data to (A) one or more predetermined reference sequences, or (B) sequence data obtained by sequencing a secondary sample of DNA that has not been subjected to the conversion procedure, such as a secondary sample separated prior to subjecting a separate secondary sample to the conversion procedure as described herein, and (ii) identifying the point differences between the converted DNA sequence and the reference sequence (A) or the unconverted DNA sequence (B) as nucleosides in the (original sample) having a modified state that allows for a change in base pairing specificity upon exposure to the conversion procedure.

[0125] As used herein, when the fraction of nucleotides having a modification or other feature is higher in a first sample or population than in a second population, the modification or other feature is present at a "higher rate" in the first sample or population of nucleic acids than in the second sample or population. For example, if in a first sample, one tenth of the nucleotides are mC and in a second sample, one twentieth of the nucleotides are mC, the first sample contains the 5-methylated cytosine modification at a higher rate than the second sample.

[0126] As used herein, "without substantially altering the base pairing specificity of a given nucleobase" means that the majority of molecules containing that nucleobase that can be sequenced have no alteration in the base pairing specificity of the second nucleobase as compared to the base pairing specificity when it was originally isolated in the sample. In some embodiments, 75%, 90%, 95% or 99% of the molecules containing that nucleobase that can be sequenced have no alteration in the base pairing specificity of the second nucleobase as compared to the base pairing specificity when it was originally isolated in the sample.

[0127] As used herein, "base pairing specificity" refers to the standard DNA base (A, C, G or T) to which a given base most preferentially pairs. Thus, for example, unmodified cytosine and 5-methylcytosine have the same base pairing specificity (i.e., specificity for G), while uracil has a base pairing specificity for A and cytosine has a base pairing specificity for G, so uracil and cytosine have different base pairing specificities. The ability of uracil to form a wobble pair with G is not important since uracil most preferentially pairs with A among the four standard DNA bases nonetheless.

[0128] As used herein, "modified cytosine" refers to cytosine in which at least one position of the cytosine is substituted with a chemical moiety different from the substituent at that position in unmodified cytosine, e.g., methyl or hydroxymethyl. To avoid ambiguity, "modified cytosine" does not include unmodified cytosine.

[0129] As used herein, a "combination" containing a plurality of members refers to, for example, a single composition containing these members in separate containers, or compartments within a larger container, e.g., multiwell plates, tube racks, refrigerators, freezers, incubators, water baths, ice buckets, machines, or other storage forms, or a set of adjacent compositions.

[0130] The "capture yield" of the collection of probes for a given target set refers to the amount of nucleic acid corresponding to the target set that the probe collection captures under typical conditions (e.g., the amount compared to another target set, or the absolute amount). Exemplary typical capture conditions are incubation of the sample nucleic acid and the probe at 65°C for 10 - 18 hours in a small reaction volume (about 20 μL) containing a stringent hybridization buffer. The capture yield can be expressed in absolute terms or, for multiple collections of probes, in relative terms. When the capture yields for multiple sets of target regions are compared, they are normalized with respect to the footprint size of the target region set (e.g., based on per kilobase). Thus, for example, if the footprint sizes of the first and second target regions are 50 kb and 500 kb respectively (giving a normalization factor of 0.1), the DNA corresponding to the first target region set is captured with a higher yield than the DNA corresponding to the second target region set when the mass concentration per volume of the captured DNA corresponding to the first target region set is higher than 0.1 times the mass concentration per volume of the captured DNA corresponding to the second target region set. As a further example, using the same footprint size, if the captured DNA corresponding to the first target region set has a mass concentration per volume that is 0.2 times the mass concentration per volume of the captured DNA corresponding to the second target region set, the DNA corresponding to the first target region set has been captured with a capture yield that is 2 times higher than the DNA corresponding to the second target region set.

[0131] "Capturing" one or more target nucleic acids refers to preferentially isolating or separating one or more target nucleic acids from non - target nucleic acids.

[0132] A "captured set" of nucleic acids refers to the nucleic acids that have undergone capture.

[0133] "Target region set" or "set of target regions" refers to a plurality of genomic loci that are targeted for capture and / or targeted by a set of probes (e.g., via sequence complementarity).

[0134] "Corresponding to a target region set" means that a nucleic acid, e.g., cfDNA, originated from a locus in the target region set or specifically binds to one or more probes for the target region set.

[0135] As used herein, a "differentially methylated region" (DMR) has a detectably different level of methylation in at least one cell or tissue type compared to the level of methylation in the same region of DNA from at least one other cell or tissue type; or has a detectably different level of methylation in at least one cell or tissue type obtained from a subject having a disease or disorder compared to the level of methylation in the same region of DNA in the same cell or tissue type obtained from a healthy subject. In some embodiments, the DMR has a detectably higher level of methylation (e.g., a hypermethylated region) in at least one cell or tissue type compared to the level of methylation in the same region of DNA from at least one other cell or tissue type or from the same cell or tissue type from a healthy subject. In some embodiments, the DMR has a detectably lower level of methylation (e.g., a hypomethylated region) in at least one cell or tissue type compared to the level of methylation in the same region of DNA from at least one other cell or tissue type or from the same cell or tissue type from a healthy subject.

[0136] "Specifically binds" in the context of a probe or other oligonucleotide and a target sequence means that, under appropriate hybridization conditions, the oligonucleotide or probe hybridizes to the target sequence or its replica to form a stable probe:target hybrid, while at the same time minimizing the formation of stable probe:non-target hybrids. Thus, the probe hybridizes to the target sequence or its replica to a sufficiently higher degree than to non-target sequences, enabling capture or detection of the target sequence. Appropriate hybridization conditions are well known in the art, can be predicted based on sequence composition, or can be determined using conventional testing methods (see, e.g., §§1.90-1.91, 7.37-7.57, 9.47-9.51, and 11.47-11.57 of Sambrook et al., Molecular Cloning, A Laboratory Manual, 2nd ed. (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989), which is hereby incorporated by reference herein, particularly see §§9.50-9.51, 11.12-11.13, 11.45-11.47, and 11.55-11.57).

[0137] "Set of sequence-variable target regions" refers to a set of target regions that can exhibit changes in sequence, such as nucleotide substitutions (i.e., single nucleotide variations), insertions, deletions, or gene fusions or translocations, in neoplastic cells (e.g., tumor cells and cancer cells).

[0138] The term "epigenetic target region set" refers to a set of target regions that can exhibit sequence-independent changes in neoplastic cells (e.g., tumor cells and cancer cells), or in cfDNA from a subject having cancer compared to cfDNA from a healthy subject. Examples of sequence-independent changes include, but are not limited to, changes in methylation (increase or decrease), nucleosome distribution, CTCF binding, transcription start sites, and regulatory protein binding regions. For the purposes of the present invention, loci that are prone to neoplastic-related, tumor-related, or cancer-related focal amplifications and / or gene fusions can also be included in the epigenetic target region set, provided that detection of changes in copy number by sequencing, or by a fused sequence mapped to more than one locus in the reference genome, is more similar to the detection of the exemplary epigenetic changes discussed above than to the detection of nucleotide substitutions, insertions, or deletions in that such detections do not rely on the accuracy of base calls at one or a few individual positions, and thus focal amplifications and / or gene fusions can be detected at a relatively shallow depth of sequencing.

[0139] The term "hypermethylated" refers to an increased level or degree of methylation of a nucleic acid molecule compared to other nucleic acid molecules within a population of nucleic acid molecules (e.g., a sample). In some embodiments, hypermethylated DNA can include DNA molecules that contain at least one methylated residue, at least two methylated residues, at least three methylated residues, at least five methylated residues, or at least ten methylated residues.

[0140] The term "hypomethylation" refers to a reduced level or degree of methylation of a nucleic acid molecule compared to other nucleic acid molecules within a population of nucleic acid molecules (e.g., a sample). In some embodiments, hypomethylated DNA includes DNA molecules that are not methylated. In some embodiments, hypomethylated DNA can include DNA molecules having zero methylated residues, at most one methylated residue, at most two methylated residues, at most three methylated residues, at most four methylated residues or at most five methylated residues.

[0141] As used herein, "methylation status" can refer to the presence or absence of a methyl group on a DNA base (e.g., cytosine) at a particular genomic position within a nucleic acid molecule. Methylation status can also refer to the degree of methylation within a nucleic acid sequence (e.g., a nucleic acid molecule that is highly methylated, hypomethylated, methylated intermediate, or not methylated). Methylation status can also refer to the number of methylated nucleotides within a particular nucleic acid molecule.

[0142] As used herein, "mutation" refers to a variation from a known reference sequence, and includes mutations such as single nucleotide variants (SNVs), and insertions or deletions (indels). Mutations can be germline mutations or somatic mutations. In some embodiments, the reference sequence for comparison purposes is the wild-type genomic sequence of the species of the subject providing the test sample, typically the human genome.

[0143] As used herein, the terms "neoplasm" and "tumor" are used interchangeably. These refer to abnormal growth of cells in a subject. A neoplasm or tumor can be benign, potentially malignant, or malignant. Malignant tumors are also referred to as cancers or cancerous tumors.

[0144] As used herein, "next-generation sequencing" or "NGS" refers to sequencing technologies that have increased throughput compared to conventional Sanger-based and capillary electrophoresis-based approaches, e.g., the ability to generate hundreds of thousands of relatively small sequence reads at once. Some examples of next-generation sequencing techniques include, but are not limited to, sequencing by synthesis, sequencing by ligation, and sequencing by hybridization. In some embodiments, next-generation sequencing includes the use of instruments capable of sequencing a single molecule. Examples of commercially available instruments for performing next-generation sequencing include, but are not limited to, NextSeq, HiSeq, NovaSeq, MiSeq, Ion PGM, and Ion GeneStudio S5.

[0145] As used herein, "nucleic acid tag" refers to different types of, or differently processed, short nucleic acids (e.g., having a length of about 500 nucleotides, about 100 nucleotides, about 50 nucleotides or less than about 10 nucleotides) used to distinguish nucleic acids from different samples (e.g., indicating a sample index), to distinguish nucleic acids from different fractions (e.g., indicating a fraction tag), or to distinguish different nucleic acid molecules within the same sample (e.g., indicating a molecular barcode). Nucleic acid tags include a predetermined, fixed, non-random, random or semi-random oligonucleotide sequence. Such nucleic acid tags can be used to label different nucleic acid molecules or different nucleic acid samples or subsamples. Nucleic acid tags can be single-stranded, double-stranded, or at least partially double-stranded. Nucleic acid tags can be of the same length or of varying lengths as needed. Nucleic acid tags can also include double-stranded molecules having one or more blunt ends, can include 5' or 3' single-stranded regions (e.g., overhangs), and / or can include one or more other single-stranded regions at other positions within a given molecule. Nucleic acid tags can be attached to one or both ends of other nucleic acids (e.g., sample nucleic acids to be amplified and / or sequenced). Nucleic acid tags can be decoded to reveal information such as the origin, form or processing of a given nucleic acid sample. For example, nucleic acid tags can also be used to enable pooling and / or parallel processing of multiple samples containing nucleic acids carrying different molecular barcodes and / or sample indices, where the nucleic acids are subsequently deconvolved by detecting (e.g., reading) the nucleic acid tags. Nucleic acid tags can also be referred to as identifiers (e.g., molecular identifiers, sample identifiers). Additionally or alternatively, nucleic acid tags can be used as molecular identifiers (e.g., to distinguish between different molecules or between amplicons of different parental molecules within the same sample or subsample). This includes, for example, uniquely tagging different nucleic acid molecules in a given sample or non-uniquely tagging such molecules.In the case of non-unique tagging applications, a limited number of tags (i.e., molecular barcodes) can be used to tag each nucleic acid molecule such that different molecules can be identified based on their endogenous sequence information (e.g., the start and / or stop positions at which they map to a selected reference genome, partial sequences at one or both ends of the sequence, and / or the length of the sequence) in combination with at least one molecular barcode. Typically, a sufficient number of different molecular barcodes are used such that there is a low probability (e.g., a probability of less than about 10%, less than about 5%, less than about 1% or less than about 0.1%) that any two molecules have the same endogenous sequence information (e.g., start and / or stop positions, partial sequences at one or both ends of the sequence, and / or length) and also have the same molecular barcode. Terms such as "library adapter having distinct molecular barcodes" encompass library adapters for uniquely or non-uniquely tagging molecules in that distinct barcodes will be present in a population of adapters regardless of whether the adapter is for unique or non-unique tagging.

[0146] As used herein, "unfixed" or "free in solution" DNA refers to DNA that is not covalently or non-covalently bound to a solid support, e.g., beads. Such DNA can be free in solution during any step (e.g., all steps) of the disclosed methods.

[0147] As used herein, "partitioning" refers to physically separating or fractionating a mixture of nucleic acid molecules in a sample based on the characteristics of the nucleic acid molecules. Partitioning can be a physical partitioning of the molecules. Partitioning can include separating nucleic acid molecules into groups or sets based on levels of epigenetic features (e.g., methylation). For example, nucleic acid molecules can be partitioned based on the level of methylation of the nucleic acid molecules. In some embodiments, the methods and systems used for partitioning can be found in PCT Patent Application No. PCT / US2017 / 068329, which is hereby incorporated by reference in its entirety.

[0148] As used herein, "fractionated set" or "partition" refers to a set of nucleic acid molecules that have been fractionated into sets or groups based on their differential binding affinity for a binder, either nucleic acid molecules or proteins associated with nucleic acid molecules. A fractionated set may also be referred to as a secondary sample. The binder preferentially binds to nucleic acid molecules containing nucleotides with epigenetic modifications. For example, if the epigenetic modification is methylation, the binder can be a methyl-binding domain (MBD) protein. In some embodiments, a fractionated set may include nucleic acid molecules belonging to a particular level or degree of epigenetic feature (e.g., methylation). For example, nucleic acid molecules can be fractionated into three sets for highly methylated nucleic acid molecules - one set (the first secondary sample, overfractionated, overfractionated set or overmethylated fractionated set), a second set for lowly methylated nucleic acid molecules (the second secondary sample, underfractionated, underfractionated set or undermethylated fractionated set), and a third set for moderately methylated nucleic acid molecules (the third secondary sample, intermediate fractionated set, intermediate methylated fractionated set, remaining fractionated set or remaining partition). In another example, nucleic acid molecules can be fractionated based on the number of methylated nucleotides - one fractionated set can have nucleic acid molecules with 9 methylated nucleotides, and another fractionated set can have nucleic acid molecules that are not methylated (zero methylated nucleotides).

[0149] As used herein, "polynucleotide", "nucleic acid", "nucleic acid molecule" or "oligonucleotide" refers to a linear polymer of nucleosides (including deoxyribonucleosides, ribonucleosides, or analogs thereof) connected by internucleoside linkages. Typically, a polynucleotide contains at least 3 nucleosides. Oligonucleotides often range in size from a few monomer units, e.g., 3 to 4, up to several hundred monomer units. Whenever a polynucleotide is represented by a sequence of letters, e.g., "ATGCCTG", the nucleotides are in the 5'→3' order from left to right, and in the case of DNA, unless otherwise specified, "A" represents deoxyadenosine, "C" represents deoxycytidine, "G" represents deoxyguanosine, and "T" represents deoxythymidine. The letters A, C, G, and T can be used to refer to the base itself, the nucleoside, or the nucleotide containing the base.

[0150] As used herein, "processing" refers to a set of steps used to generate a library of nucleic acids suitable for sequencing. The set of steps can include, but is not limited to, fractionating, end-repairing, adding sequencing adapters, tagging, and / or PCR amplification of the nucleic acids.

[0151] As used herein, "quantitative measure" refers to an absolute or relative measure. A quantitative measure can be, without limitation, a number, a statistical measurement (e.g., frequency, mean, median, standard deviation or quantile), or a degree or relative amount (e.g., high, medium and low). A quantitative measure can be a ratio of two quantitative measures. A quantitative measure can be a linear combination of quantitative measures. A quantitative measure can be a standardized measure.

[0152] As used herein, "reference sequence" refers to a known sequence used for the purpose of comparison with an experimentally determined sequence. For example, the known sequence can be an entire genome, a chromosome, or any segment thereof. The reference sequence can align with a single continuous sequence of a genome or chromosome or chromosomal arm, or can include non-continuous segments that align with different regions of a genome or chromosome. Examples of reference sequences include, for example, the human genome, such as hg19 and hg38.

[0153] As used herein, "sample" means anything that can be analyzed by the methods and / or systems disclosed herein.

[0154] As used herein, "determining a sequence" refers to any of several techniques used to determine the sequence (e.g., the identity and order of monomer units) of a biomolecule, such as a nucleic acid, such as DNA or RNA. Examples of sequencing methods include, but are not limited to, targeted sequencing, single molecule real-time sequencing, exon or exome sequencing, intron sequencing, electron microscopy-based sequencing, panel sequencing, transistor-mediated sequencing, direct sequencing, random shotgun sequencing, Sanger dideoxy termination sequencing, whole genome sequencing, sequencing by hybridization, pyrosequencing, duplex sequencing, cycle sequencing, single base extension sequencing, solid-phase sequencing, high-throughput sequencing, ultra-parallel signature sequencing, emulsion PCR, co-amplification-PCR at lower denaturation temperature (COLD-PCR), multiplex PCR, sequencing by reversible terminators, paired-end sequencing, short-read sequencing, single molecule sequencing, sequencing by synthesis, real-time sequencing, reverse-terminator sequencing, nanopore sequencing, 454 sequencing, Solexa Genome Analyzer sequencing, SOLiD™ sequencing, MS-PET sequencing, and combinations thereof. In some embodiments, the sequencing can be performed by a genetic analyzer, such as, among many others, a genetic analyzer commercially available from Illumina, Inc., Pacific Biosciences, Inc., or Applied Biosystems / Thermo Fisher Scientific.

[0155] As used herein, in the context of a nucleic acid polymer, "sequence information" means the order and identity of monomer units (e.g., nucleotides, etc.) in that polymer.

[0156] As used herein, the term "set of array-variable target regions" refers to a set of target regions that can exhibit changes in sequence, such as nucleotide substitutions, insertions, deletions, or gene fusions or translocations, in neoplastic cells (e.g., tumor cells and cancer cells).

[0157] As used herein, the terms "somatic mutation" or "somatic variation" are used interchangeably. These refer to mutations in the genome that occur after conception. Somatic mutations can occur in any cell of the body except germ cells and thus are not passed on to offspring.

[0158] As used herein, "subject" refers to an animal, e.g., a mammalian species (e.g., human) or an avian (e.g., bird) species, or other organisms, e.g., plants. More specifically, the subject can be a vertebrate, e.g., a mammal, e.g., a mouse, a primate, a monkey, or a human. Animals include farm animals (e.g., production cattle, dairy cows, poultry, horses, pigs, etc.), sports animals, and companion animals (e.g., pets or support animals). The subject can be a healthy individual, an individual having or suspected of having a disease or a predisposition to a disease, or an individual in need of or suspected of being in need of treatment. The terms "individual" or "patient" are intended to be interchangeable with "subject". For example, the subject can be an individual diagnosed with cancer, an individual who is going to receive cancer treatment, and / or an individual who has received at least one cancer treatment. The subject may be in remission from cancer. As another example, the subject can be an individual diagnosed with an autoimmune disease. As another example, the subject can be a female individual who is pregnant or planning to become pregnant, a female individual who may be diagnosed with or suspected of having a disease, e.g., cancer, an autoimmune disease.

[0159] As used herein, "target region set" or "set of target regions" or "target region" or "target target region" or "target region" or "target genomic region" refers to a plurality of genomic loci or a plurality of genomic regions that are targeted for capture and / or targeted by a set of probes (e.g., via sequence complementarity).

[0160] As used herein, "tumor fraction" refers to the proportion of cfDNA molecules originating from tumor cells for a given sample or sample-region pair.

[0161] As used herein, "asymmetric adapter" is a double-stranded adapter in which the two strands are not fully complementary or are otherwise distinguishable such that synthesis of the complementary sequence of one strand of the adapter yields a sequence distinguishable from the sequence of the other strand of the adapter. Examples of asymmetric adapters are Y-shaped adapters and bubble adapters.

[0162] As used herein, "Y-shaped adapter" refers to an adapter comprising two DNA strands that include a complementary portion and a non-complementary portion, the non-complementary portion forming a single-stranded arm. This adapter can be ligated to a sample or inserted DNA molecule, for example, such that the complementary (double-stranded) portion of the adapter is proximal to the sample or inserted DNA molecule. Prior to ligation, the double-stranded portion of the Y-shaped adapter can have blunt ends or, for example, an overhang of 1 to 3 nucleotides. The single-stranded arms can be of the same or different lengths.

[0163] As used herein, "bubble adapter" refers to an adapter comprising two DNA strands that contain a non-complementary portion sandwiched between complementary portions such that the adapter has a single-stranded region located between double-stranded regions. This adapter can be bound to a sample or an inserted DNA molecule, for example, by ligation, such that one of the complementary (double-stranded) portions of the adapter is proximal to the sample or the inserted DNA molecule. Prior to binding, the double-stranded portion of the Y-shaped adapter that binds to the inserted or sample molecule can have blunt ends or, for example, overhangs of 1 to 3 nucleotides. The single-stranded portions of the two strands may or may not be of the same length.

[0164] The terms "or combinations thereof (singular)" and "or combinations thereof (plural)", as used herein, refer to any and all permutations and combinations of the listed terms preceding this term. For example, "A, B, C, or combinations thereof" is intended to include at least one of the following: A, B, C, AB, AC, BC, or ABC, and, where order is important in a particular context, BA, CA, CB, ACB, CBA, BCA, BAC, or CAB. By way of extension of this example, combinations containing repetitions of one or more items or terms, such as BB, AAA, AAB, BBC, AAABCCCC, CBBAAA, CABABB, etc., are explicitly included. One of ordinary skill in the art will understand that there is typically no limitation with respect to the number of items or terms in any combination, unless otherwise apparent from the context.

[0165] "Or" is used in an inclusive sense, i.e., equivalent to "and / or", unless the context requires otherwise.

[0166] Samples and Subjects The present disclosure relates to a method for distinguishing artifact information from true information in a workflow including sequencing of a DNA sample. In some cases, the DNA sample is obtained from or has been obtained from a subject. In some embodiments, the DNA sample may comprise or consist of DNA from a biological sample obtained from the subject. The subject may be a human, mammal, animal, primate, rodent (including mice and rats), or other common laboratory, household, companion, business, or agricultural animal, such as a rabbit, dog, cat, horse, cow, sheep, goat, or pig. Preferably, the DNA sample is of human origin. The subject may, in some cases, have or be suspected of having cancer, a tumor, or a neoplasm. In other cases, the subject may not have cancer or detectable cancer symptoms. The subject may be treated with one or more cancer treatments, such as any one or more of chemotherapy, an antibody, a vaccine, or a biologic. The subject may be in remission from a tumor, cancer, or neoplasm (e.g., after treatment such as chemotherapy, surgical resection, radiation, or a combination thereof). The subject may or may not be diagnosed as being susceptible to cancer or any cancer-related genetic mutation / disorder. In some embodiments, the sample is a DNA sample obtained from a tumor tissue biopsy. The cancer, tumor, or neoplasm may generally be of any type, such as a cancer tumor or neoplasm of the lung, colon, rectum (or colorectal), kidney, breast, prostate, or liver, or other types of cancer described herein. In some embodiments, the sample is obtained from a subject in remission from a tumor, cancer, or neoplasm (e.g., after chemotherapy, surgical resection, radiation, or a combination thereof). In any of the above embodiments, the pre-cancer, cancer, tumor, or neoplasm or suspected pre-cancer, cancer, tumor, or neoplasm may be of the bladder, head and neck, lung, colon, rectum, kidney, breast, prostate, skin, or liver. In some embodiments, the pre-cancer, cancer, tumor, or neoplasm or suspected pre-cancer, cancer, tumor, or neoplasm is of the lung. In some embodiments, the pre-cancer, cancer, tumor, or neoplasm or suspected pre-cancer, cancer, tumor, or neoplasm is of the colon or rectum.In some embodiments, the pre - cancer, cancer, tumor or neoplasm or suspected pre - cancer, cancer, tumor or neoplasm is of the breast. In some embodiments, the pre - cancer, cancer, tumor or neoplasm or suspected pre - cancer, cancer, tumor or neoplasm is of the prostate. In any of the above - described embodiments, the subject may be a human subject. In some embodiments, the sample is obtained from a subject having stage I cancer, stage II cancer, stage III cancer or stage IV cancer.

[0167] In some embodiments, the subject may have an infectious disease, graft rejection, or other disease or disorder related to a change in the immune system. The subject may have no cancer or detectable cancer symptoms. The subject may be treated with one or more cancer treatments, such as any one or more of chemotherapy, antibodies, vaccines or biologics. The subject may be in a remission state. The subject may or may not be diagnosed as being susceptible to cancer or any cancer - related genetic mutations / disorders.

[0168] The biological sample can be any biological sample isolated from a subject. Biological samples can include body tissues, such as known or suspected solid tumors, whole blood, platelets, serum, plasma, feces, red blood cells, white blood cells (leucocytes), endothelial cells, tissue biopsies, cerebrospinal fluid, synovial fluid, lymphatic fluid, ascites, interstitial or extracellular fluid, gingival crevicular fluid (the fluid in the inter - cellular space containing cells), bone marrow, pleural effusion, cerebrospinal fluid, saliva, mucus, sputum, semen, sweat, urine. In some embodiments, the biological sample is a body fluid, particularly blood and its fractions, or urine. The sample can be in the form originally isolated from the subject or can be subjected to further processing to remove or add components, such as cells, or to enrich one component relative to another. The sample can be isolated or obtained from the subject and transported to the location of sample analysis. The sample can be stored and shipped at a desired temperature, such as room temperature, 4°C, - 20°C and / or - 80°C. The sample can be isolated or obtained from the subject at the location of sample analysis.

[0169] In a preferred embodiment, the DNA sample contains cell-free DNA. In another preferred embodiment, the DNA sample is a DNA sample from a formalin-fixed paraffin-embedded (FFPE) sample.

[0170] In some embodiments, the nucleic acid population is obtained from a serum, plasma, or blood sample of a subject suspected of having a neoplasm, tumor, pre-cancer, or cancer, or a subject previously diagnosed with a neoplasm, tumor, pre-cancer, or cancer. The population includes nucleic acids having various levels of sequence variation, epigenetic changes, and / or post-replicative or post-transcriptional modifications. Post-replicative modifications include, in particular, modifications of cytosine at the 5-position of the nucleobase, such as 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine, and 5-carboxylcytosine.

[0171] In some embodiments, the DNA sample is plasma-derived. The volume of plasma used to obtain the DNA sample may depend on the desired read depth for the region to be sequenced. Exemplary volumes are 0.4 - 40 ml, 5 - 20 ml, 10 - 20 ml. For example, the volume can be 0.5 mL, 1 mL, 5 mL, 10 mL, 20 mL, 30 mL, or 40 mL. The volume of plasma sampled can be 5 - 20 mL.

[0172] The sample can contain various amounts of DNA containing genomic equivalents. For example, a sample of about 30 ng of DNA contains about 10,000 (10 4 ) diploid human genomic equivalents, and in the case of cell-free DNA (cfDNA), can contain about 200 billion (2×10 11 ) individual polynucleotide molecules. Similarly, a sample of about 100 ng of DNA can contain about 30,000 diploid human genomic equivalents, and in the case of cfDNA, about 600 billion individual molecules.

[0173] The sample can include nucleic acids from different sources, such as nucleic acids and cell-free nucleic acids from cells of the same subject, as well as nucleic acids and cell-free nucleic acids from cells of different subjects. In some embodiments, the nucleic acid can be DNA. The sample can include DNA carrying a mutation. For example, the sample can include DNA carrying a germline mutation and / or a somatic mutation. A germline mutation refers to a mutation present in the germline DNA of the subject. A somatic mutation refers to a mutation originating from somatic cells of the subject, such as cancer cells. The sample can include DNA carrying a cancer-related mutation (e.g., a cancer-related somatic mutation). The sample can include an epigenetic variant, which is related to the presence of a genetic variant, such as a cancer-related mutation. In some embodiments, the sample includes an epigenetic variant related to the presence of a genetic variant, and the sample does not include that genetic variant.

[0174] The DNA sample may be, or may contain, cell-free nucleic acid or cfDNA. CfDNA can be obtained from a subject to be tested, for example, as described above. For example, a sample for analysis can be plasma or serum containing cell-free nucleic acid. "Cell-free DNA", "cfDNA molecule" or "cfDNA" includes, for example, DNA molecules that naturally exist in a subject in an extracellular form (e.g., in blood, serum, plasma, or other body fluids such as lymph, cerebrospinal fluid, urine or sputum). CfDNA was previously present in cells (singular or plural) in large complex organisms, such as mammals, but has undergone release from the cells into the fluid found in the organism in vivo and can be obtained by obtaining a sample of the fluid without the need to perform an in vitro cell lysis step. In other words, cell-free nucleic acid or cfDNA is nucleic acid or DNA that is not contained within a cell or not bound to a cell in another manner, or nucleic acid or DNA that remains in the sample after removal of intact cells. Cell-free nucleic acid includes DNA, RNA, and hybrids thereof, including genomic DNA, mitochondrial DNA, siRNA, miRNA, circular RNA (cRNA), tRNA, rRNA, small nucleolar RNA (snoRNA), Piwi-interacting RNA (piRNA), long non-coding RNA (long ncRNA), or fragments of any of these. Cell-free nucleic acid can be double-stranded, single-stranded, or hybrids thereof. Cell-free nucleic acid can be released into body fluids via secretion or cell death processes such as necrosis and apoptosis of cells. Some cell-free nucleic acids, such as circulating tumor DNA (ctDNA), are released from cancer cells into body fluids. Others are released from healthy cells. In some embodiments, cfDNA is cell-free fetal DNA (cffDNA). In some embodiments, cell-free nucleic acid is produced by tumor cells. In some embodiments, cell-free nucleic acid is produced by a mixture of tumor cells and non-tumor cells.

[0175] Exemplary amounts of cell-free nucleic acids (e.g., cfDNA) in a sample prior to amplification are in the range of about 1 fg to about 1 μg, such as 1 pg to 200 ng, 1 ng to 100 ng, 10 ng to 1000 ng. For example, the amount can be up to about 600 ng, up to about 500 ng, up to about 400 ng, up to about 300 ng, up to about 200 ng, up to about 100 ng, up to about 50 ng or up to about 20 ng of cell-free nucleic acid molecules. The amount can be at least 1 fg, at least 10 fg, at least 100 fg, at least 1 pg, at least 10 pg, at least 100 pg, at least 1 ng, at least 10 ng, at least 100 ng, at least 150 ng or at least 200 ng of cell-free nucleic acid molecules. The amount can be up to 1 femtogram (fg), 10 fg, 100 fg, 1 picogram (pg), 10 pg, 100 pg, 1 ng, 10 ng, 100 ng, 150 ng or 200 ng of cell-free nucleic acid molecules. The method can include obtaining from the sample from 1 femtogram (fg) to 200 ng of cell-free nucleic acid molecules.

[0176] Cell-free nucleic acids have an exemplary size distribution of about 100 to 500 nucleotides, with molecules of 110 to about 230 nucleotides accounting for about 90% of the molecules, the mode being about 168 nucleotides, and the second smallest peak being in the range between 240 and 440 nucleotides.

[0177] Cell-free nucleic acids can be isolated from a body fluid via a fractionation or fractionation step in which the cell-free nucleic acids found in solution are separated from intact cells and other insoluble components of the body fluid. The fractionation can include techniques such as centrifugation or filtration. Alternatively, the cells in the body fluid can be lysed and the cell-free and cellular nucleic acids can be processed together. Generally, after addition of buffer and washing steps, the nucleic acids can be precipitated using alcohol. Further purification steps such as silica-based columns to remove contaminants or salts can be used. Non-specific bulk carrier nucleic acids, DNA or proteins for sequencing (e.g., bisulfite sequencing), hybridization and / or ligation can be added through the reaction in certain embodiments of the procedure, e.g., to optimize yield.

[0178] After such processing, the sample can contain nucleic acids in various forms including double-stranded DNA, single-stranded DNA, and single-stranded RNA. In some embodiments, the single-stranded DNA and RNA can be converted to double-stranded form, and thus they can be included in subsequent processing and analysis steps.

[0179] The methods disclosed herein are also particularly suitable for the analysis of DNA from formalin-fixed paraffin-embedded (FFPE) tissue samples. The formalin fixation process preserves the ultrastructure of the tissue but causes various types of damage to the DNA within the tissue, such as nicks in the DNA. As described elsewhere herein, these nicks can lead to the synthesis of regions of the DNA molecule in the end repair process. The methods disclosed herein enable these regions to be identified and the sequence data to be interpreted accordingly.

[0180] End Repair and A-Tailing End repair refers to a method for repairing DNA by converting non-blunt-ended DNA to blunt-ended DNA. Sequencing workflows typically use end repair to make the ends of DNA molecules compatible with adapters, and then ligate the adapters to the DNA. Fragmented and / or damaged DNA (e.g., cfDNA or DNA from FFPE samples) often contains non-blunt ends that contain 3' overhangs and / or 5' overhangs. A 3' overhang refers to the 3' end of a DNA strand that extends beyond the 5' end of the complementary strand, resulting in one or more unpaired nucleotides at the 3' end of the DNA strand. Conversely, a 5' overhang refers to the 5' end of a DNA strand that extends beyond the 3' end of the complementary strand, resulting in one or more unpaired nucleotides at the 5' end of the DNA strand.

[0181] The process of end repair involves the conversion of double-stranded DNA with 3' overhangs and / or 5' overhangs to double-stranded DNA without overhangs. This can be done using enzymes such as T4 DNA polymerase and / or the Klenow fragment. The 3'→5' exonuclease activity of these enzymes removes the 3' end at the 3' overhang, and the 5'→3' polymerase activity of these enzymes extends the 3' end at the 5' overhang to remove the 5' overhang, thereby generating blunt-ended DNA molecules. To fill in these 5' overhangs, end repair is performed in the presence of dATP, dCTP, dGTP, and dTTP. End repair may also include a second step involving the addition of a phosphate group to the 5' end of the DNA by an enzyme such as polynucleotide kinase. This makes the 5' end of the end-repaired DNA molecule compatible with the subsequent action of DNA polymerase and DNA ligase.

[0182] As used herein, the term "A-tailing" refers to the addition of a single deoxyadenosine residue to the end of a blunt-ended double-stranded DNA fragment to form a 3'-deoxyadenosine single-base overhang. Such A-tailing reactions are performed using a polymerase that has the ability to add a non-template A to the 3' end of a blunt double-stranded DNA molecule. A polymerase capable of A-tailing typically does not have 3'-5' exonuclease activity. When A-tailing is performed as a separate reaction from end repair, it is typically done in the presence of dATP and in the absence of dCTP, dTTP, and dGTP. A-tailed fragments are not compatible with self-ligation (i.e., self-cyclization and ligation of DNA), but they are compatible with the 3'T overhangs that can be used on adapters. Methods that include end repair, A-tailing, and ligation to an adapter with 3'T-overhangs can result in higher efficiency ligation compared to blunt-end ligation, because blunt ligation can lead to self-ligation of the adapter and / or the DNA molecule.

[0183] In some cases, the methods disclosed herein include end repair of DNA molecules and subsequent blunt-end ligation of adapters. In other cases, the methods disclosed herein include end repair of DNA molecules and subsequent A-tailing, as well as sticky-end ligation of T-tailed adapters. When the methods disclosed herein include an A-tailing step, it may be performed separately from end repair using an intervening reaction cleanup step, or it may be performed in the same reaction as end repair (e.g., using the NEBNext® Ultra™ II End Repair / dA-Tailing Module (E7546)). When the A-tailing reaction is performed in the same reaction as end repair, sticky-end ligation may be performed using a mixture of T-tailed adapters and C-tailed adapters.

[0184] The inventors investigated the effects of the end repair and A-tailing steps on DNA molecules and their impact on the corresponding sequence data. As shown in FIGS. 1-10 (discussed in detail below), the end repair and A-tailing reactions can have various effects on the composition of DNA molecules depending on the exact workflow and reaction components used. These reactions can lead to the synthesis not only of regions at the 3' end of the DNA strand, but also of internal regions by nick translation, gap filling, and subsequent ligation. The effects of such synthesis are shown in FIGS. 1-10 in relation to its effect on cytosine methylation.

[0185] Figure 1 shows a schematic of the effects of combinations of end repair and A-tailing steps on intact double-stranded DNA, nicked double-stranded DNA, and gapped double-stranded DNA. In all types of DNA molecules shown, end repair can lead to 3’ filling by unmethylated cytosine, which may not reflect the true methylation state of that position in the DNA molecule prior to the generation of the 5’ overhang. Inferring the methylation state from this end-repaired DNA molecule can lead to inaccurate methylation data at this position. In nicked DNA, polymerases containing 5’→3’ exonuclease activity and / or strand displacement activity can lead to synthesis of the internal region of the DNA molecule via nick translation. If the end repair reaction is carried out using unmethylated deoxycytidine triphosphate (dCTP), the synthesized region may incorporate unmethylated dCTP at positions that originally contained methylated cytosine. Thus, using sequence data from these synthesized regions to infer the methylation state of the original DNA molecule can lead to inaccurate data. In gapped DNA, both DNA polymerases used for end repair and A-tailing can lead to the generation of synthesized regions. The gap can be filled with the DNA polymerase used in the end repair reaction, whether or not it has 5’→3’ exonuclease activity or strand displacement activity. After this gap filling, a nick still exists between the synthesized region and the region of the original DNA molecule on the 3’ side of the gap. The A-tailing enzyme can then introduce additional synthesized regions via nick translation, as described for nicked DNA. This synthesized region can extend to the 3’ end of the DNA molecule. As described above, using these synthesized regions to infer the methylation state of the original DNA molecule can lead to inaccurate results.

[0186] In some embodiments, the end repair and A-tailing reactions are performed within a single tube. In such cases, the A-tailing reaction can be performed at a higher temperature than the end repair. Optionally, the end repair is performed at ambient temperature (e.g., 15-35 °C), and the A-tailing is performed at a temperature above 60 °C. The A-tailing reaction can be performed using a thermostable polymerase (e.g., Taq DNA polymerase, Tfl DNA polymerase, Bst DNA polymerase, large fragment or Tth DNA polymerase), and the method further includes the step of raising the temperature of the sample after end repair to inactivate the polymerase (e.g., T4 DNA polymerase or Klenow fragment) used for end repair. In some embodiments, the A-tailing is performed using a DNA polymerase that (i) does not have 5'-3' exonuclease activity and / or (ii) is not a strand-displacing DNA polymerase. These properties reduce the ability of the DNA polymerase to extend from nicks. This reduces the level of synthesis that can occur during the end repair and A-tailing reactions, and thus reduces the proportion of sequencing data that can potentially contain artifact data and be filtered out as such. Thus, in some embodiments, the A-tailing is performed using a DNA polymerase that cannot extend from nicks in the DNA, such as HemoKlen Taq. In other embodiments, the A-tailing is performed using Taq DNA polymerase. In other embodiments, the A-tailing is performed using Tfl polymerase, Bst DNA polymerase, large fragment or Tth polymerase.

[0187] Figure 2 shows a schematic of the potential effects of separate end repair and A-tailing steps on intact double-stranded DNA, nicked double-stranded DNA, and gapped double-stranded DNA. To reduce the level of the synthesized region, the end repair reaction can be carried out using a DNA polymerase lacking 5'→3' exonuclease activity and / or strand displacement activity (e.g., T4 DNA polymerase or Klenow fragment). This approach means that nick translation is reduced so that a larger portion of the sequencing data can be used in the disclosed methods for methylation analysis, since less sequencing data is derived from the synthesized region.

[0188] In all types of DNA molecules shown in FIG. 2, end repair can lead to 3' filling by unmethylated cytosine, which may not reflect the true methylation state of that position in the DNA molecule before the generation of the 5' overhang. Inferring the methylation state from this end-repaired DNA molecule can lead to inaccurate data at this position. In nicked DNA and gapped DNA, nick translation in end repair is reduced by using polymerases lacking 5'→3' exonuclease activity and / or strand displacement activity. Separation of the end repair from the A-tailing reaction by reaction cleanup means that only dATP (not dCTP, dTTP, or dGTP) is present during the A-tailing reaction. This means that, since three of the four nucleotide components are not present in the reaction mixture, efficient nick translation cannot occur in the A-tailing reaction. In gapped DNA, the gaps can be filled with the DNA polymerase used in the end repair reaction, regardless of whether they have 5'→3' exonuclease activity or strand displacement activity. These filled gaps thereby generate the synthesized regions. As described above, inferring the methylation state of the original DNA molecule using these synthesized regions can lead to inaccurate results. However, the extent of the synthesized regions in such a workflow is significantly reduced compared to a workflow in which the end repair reaction and the A-tailing reaction are performed in a single reaction, as shown in FIG. 1.

[0189] Thus, in a preferred embodiment, in the methods disclosed herein, end repair is performed using a polymerase lacking 5’→3’ exonuclease activity and / or strand displacement activity. In some cases, the polymerase used in the end repair reaction may be Q5® High-Fidelity DNA Polymerase, Q5U® Hot Start High-Fidelity DNA Polymerase, Phusion® High-Fidelity DNA Polymerase, Hemo KlenTaq, phi29 DNA Polymerase, T7 DNA Polymerase, DNA Polymerase I (E. coli), DNA Polymerase I, the large (Klenow) fragment (“Klenow fragment”), or T4 DNA Polymerase. In some embodiments, the polymerase used for end repair is T4 DNA Polymerase or the Klenow fragment.

[0190] In some embodiments, the methods disclosed herein include an A-tailing reaction after end repair and before the ligation reaction, and the end repair and A-tailing reactions are separated by reaction cleanup. The A-tailing reaction is typically performed in the presence of dATP and in the absence of dCTP, dTTP, and dGTP. In some embodiments, the A-tailing reaction is performed using the Klenow fragment lacking 3’-5’ exonuclease activity.

[0191] Figure 3 shows a schematic of the potential effects of end repair and blunt-end ligation on intact double-stranded DNA, nicked double-stranded DNA, and gapped double-stranded DNA. To reduce the level of synthesized regions, the end repair reaction can be performed using a DNA polymerase lacking 5’→3’ exonuclease activity and / or strand displacement activity (e.g., T4 DNA Polymerase or the Klenow fragment). This approach means that nick translation is reduced so that a larger portion of the sequencing data can be used in the disclosed methods for methylation analysis, since less of the sequencing data is derived from the synthesized regions.

[0192] In all types of DNA molecules shown in FIG. 3, end repair can lead to 3′ filling by unmethylated cytosine, which may not reflect the true methylation state at that position in the DNA molecule prior to the generation of the 5′ overhang. Inferring the methylation state from this end-repaired DNA molecule can lead to inaccurate data at this position. In nicked DNA and gapped DNA, nick translation in end repair is reduced by using a polymerase lacking 5′→3′ exonuclease activity and / or strand displacement activity. In gapped DNA, the gaps can be filled with a DNA polymerase used in the end repair reaction, regardless of whether they have 5′→3′ exonuclease activity or strand displacement activity. These filled gaps thereby generate the synthesized regions. As described above, inferring the methylation state of the original DNA molecule using these synthesized regions can lead to inaccurate results. However, the extent of the synthesized regions in such workflows is significantly reduced compared to workflows in which end repair and A-tailing reactions are performed in a single reaction, as shown in FIG. 1.

[0193] Figure 4 is a repetition of the schematic diagram shown in Figure 1, but includes the use of dNTPs containing modified bases in the combination of end repair and A-tailing reaction, which in this example is methylated cytosine (5mC). This indicates that methylated cytosine is incorporated into the synthesized region at both CpG sites and CpH sites (i.e., CpA, CpC, and CpT sites). Although methylation of cytosine in non-CpG contexts has been described, this is thought to constitute 0.02% of the total methyl-cytosine in differentiated somatic cells (Jang et al. Genes (Basel). 2017 Jun; 8(6): 148). Therefore, the presence of methylated cytosine in non-CpG contexts in end-repaired DNA can be interpreted as being introduced during end repair and / or the A-tailing reaction. Thus, the use of dNTPs containing modified bases can be used to effectively label the synthesized region of end-repaired DNA.

[0194] Figure 5 is a repetition of the schematic diagram shown in Figure 2, but includes the use of dNTPs containing modified bases in end repair, which in this example is methylated cytosine (5mC). Similar to Figure 4, this indicates that methylated cytosine is incorporated into the synthesized region at both CpG sites and CpH sites (i.e., CpA, CpC, and CpT sites). Thus, the use of dNTPs containing modified bases can be used to effectively label the synthesized region of end-repaired DNA.

[0195] Figure 6 is a repetition of the schematic diagram shown in Figure 3, but includes the use of dNTPs containing modified bases in end repair, which in this example is methylated cytosine (5mC). Similar to Figure 4, this indicates that methylated cytosine is incorporated into the synthesized region at both CpG sites and CpH sites (i.e., CpA, CpC, and CpT sites). Thus, the use of dNTPs containing modified bases can be used to effectively label the synthesized region of end-repaired DNA.

[0196] dNTPs containing modified bases can contain any modified base, and the presence or absence of the modification can be detected depending on the type of modification-sensitive sequencing. Modified bases can be 4-methylcytosine (4mC), 5-methylcytosine (5mC), 5-hydroxymethyl-cytosine (5hmC), N6-methyladenosine (6mA), bromodeoxyuridine (BrdU), 5-fluorodeoxyuridine (FldU), 5-iododeoxyuridine (IdU), 5-ethynyl-deoxyuridine (EdU) and / or 8-oxoguanine (8oxoG).

[0197] When dNTPs containing modified bases are used, they can be used in place of equivalent unmodified bases in the end repair reaction. For example, when dCTP containing 5mC is used in the end repair reaction, dCTP containing unmodified cytosine may not be present. This ensures that the dCTP incorporated into the DNA molecule during the end repair reaction contains 5mC. In some embodiments, multiple types of dNTPs containing modified bases are used for end repair. For example, instead of dATP containing unmodified adenine and dCTP containing unmodified cytosine, dATP containing 6mA and dCTP containing 5mC can be used in the end repair reaction. The use of multiple types of dNTPs containing modified bases is advantageous because it provides an improvement in resolution when defining regions of the end-repaired DNA molecule synthesized during the end repair reaction. This is because in this example, the ends of the synthesized region can be defined as the first unmodified adenine or unmodified cytosine after the strand containing 6mA and / or 5mC, rather than relying solely on the detection of only unmodified adenine or only unmodified cytosine.

[0198] The modified sensitivity sequencing method used depends on the type of modified base used in the end repair reaction so as to be able to detect a specific modification. Exemplary conversion-based methods are described above along with the base modifications they can detect. Further, nanopore-based sequencing can be used to detect 4mC, 5mC, 5hmC, 6mA, BrdU, FdU, IdU, and EdU, and single molecule real-time (SMRT) sequencing from Pacific Biosciences can be used to detect 4mC, 5mC, 5hmC, 6mA, and 8oxoG.

[0199] Ligation to an adapter Once the DNA has been end-repaired, it can be subjected to blunt-end ligation with a blunt-end adapter if A-tailing is not performed, or sticky-end ligation with a T-tailed adapter if A-tailing is performed. The DNA molecule can be ligated to an adapter at either one or both ends. The DNA molecule can be ligated to an at least partially double-stranded adapter (e.g., a Y-shaped or bell-shaped adapter). In embodiments where the modified sensitivity sequencing includes a conversion procedure, the ligation step can be performed before or after the conversion step. In some embodiments, the ligation step is performed after the conversion step.

[0200] DNA ligase and an adapter are added to ligate the DNA molecules in the sample with the adapter on one or both ends, i.e., to form the adapted DNA. As used herein, an "adapter" is typically at least partially double-stranded and is a short nucleic acid (e.g., having less than about 500, less than about 100, or less than about 50 nucleotides in length, or 20-30, 20-40, 30-50, 30-60, 40-60, 40-70, 50-60, 50-70, 20-500, or 30-100 bases from end to end) that can be ligated to the ends of a given sample DNA molecule. In some cases, two adapters can be ligated to a single sample DNA molecule, with one adapter ligated to each end of the sample nucleic acid molecule.

[0201] In some embodiments, the ligase used in the ligation reaction can act on both single-stranded DNA nicks and double-stranded DNA ends. In some cases, the ligase is T4 DNA ligase or T3 DNA ligase. The adapter can include nucleic acid primer binding sites that enable amplification of sample DNA molecules flanked by the adapter at both ends, and / or sequencing primer binding sites for sequencing applications such as various next-generation sequencing (NGS) applications. The adapter can include a sequence for hybridizing to a solid support, such as a flow cell sequence. The adapter can also include a binding site for a capture probe, such as an oligonucleotide conjugated to a flow cell support or the like. The adapter can also include a sample index and / or a molecular barcode. These are typically arranged relative to the amplification primers and the sequencing primer binding sites such that the sample index and / or the molecular barcode are included in the amplicon and the sequencing reads of a given DNA molecule. Adapters of the same or different sequences can be ligated to each end of the sample DNA molecule. In some cases, adapters of the same or different sequences are ligated to each end of the DNA molecule, except that the sample index and / or the molecular barcode differ in their sequences. In some embodiments, the adapter is a Y-shaped adapter with one end blunt-ended or tailed as described herein for connecting to a nucleic acid molecule, and this nucleic acid molecule is also blunt-ended or tailed with one or more complementary nucleotides relative to the nucleotides in the tail of the adapter. In another exemplary embodiment, the adapter is a bell-shaped adapter including a blunt or tailed end for connecting to the DNA molecule to be analyzed. Other exemplary adapters include T-tail, C-tail, or hairpin-shaped adapters. For example, a hairpin-shaped adapter can include a complementary double-stranded portion and a loop portion, and the double-stranded portion can bind (e.g., ligate) to a double-stranded polynucleotide.Hairpin-shaped array determination adapters can be ligated to both ends of a polynucleotide fragment to generate a circular molecule, which can be sequenced multiple times. The adapters used in the methods of the present disclosure include one or more known modified nucleosides, such as methylated nucleosides. When two adapters are ligated to a sample nucleic acid (one at each end), either or both of the adapters may contain one or more known modified nucleosides. Typically, primer binding sites, sequencing primer binding sites, sample indices, and / or molecular barcodes, if present, do not contain known modified nucleosides that change base pairing specificity as a result of the conversion procedure.

[0202] In some embodiments, the adapter may be added to DNA or a secondary sample thereof. The adapter can be ligated to the DNA at any point in the methods herein. In some embodiments, the adapter is ligated to the DNA of the sample or a secondary sample thereof before annealing a primer to the DNA for capture probe generation. In some such embodiments, the DNA ligated to the adapter is amplified before annealing a primer to the DNA for capture probe generation. In some embodiments, the adapter is ligated to the DNA of the sample or a secondary sample thereof before contacting the DNA with a capture probe. In some embodiments, the DNA to which the adapter is ligated is in the same sample or secondary sample as the DNA used as a template for generating the capture probe. In some embodiments, the DNA to which the adapter is ligated is in a different sample or secondary sample than the DNA used as a template for generating the capture probe, such as a second sample or a second secondary sample of the first sample. In some embodiments, the adapter is ligated to the DNA captured by the capture probe.

[0203] In some embodiments, the primers used to generate the capture probes are not complementary to the adapters, and thus the resulting capture probes do not contain the adapters. Thus, DNA ligated to the adapter can be selectively amplified in the presence of capture probes that do not contain the adapter. Similarly, DNA ligated to the adapter can be separated from DNA that does not contain the adapter.

[0204] In some embodiments, the disclosed method includes analyzing DNA in a sample. In such methods, adapters may be added to the DNA. This can be done, for example, by providing an adapter to the 5' portion of a primer before or after an amplification step (when PCR is used, this can be referred to as library prep-PCR or LP-PCR) and can be done concurrently with the amplification procedure. In some embodiments, the adapter is added by other approaches such as ligation. In some such methods, a first adapter is added to the 3' end of a nucleic acid by ligation, which may include ligation to single-stranded DNA. In some embodiments, the first adapter is added to the nucleic acid by ligation, which may include ligation (e.g., to its 3' end) to single-stranded DNA, prior to any fractionation or capture step. In some embodiments, the capture probe can be isolated after fractionation and ligation. For example, a hypomethylated fraction can be ligated to an adapter, and then a portion of the ligated hypomethylated fraction can be used to generate a capture probe for rearrangement. The adapter can be used, for example, as a priming site for second-strand synthesis using a universal primer and DNA polymerase. A second adapter can then be ligated to at least the 3' end of the second strand of the double-stranded molecule. In some embodiments, the first adapter includes an affinity tag such as biotin, and the nucleic acid ligated to the first adapter is bound to a solid support (e.g., beads) that can include a binding partner for the affinity tag such as streptavidin. For further consideration of related procedures, see Gansauge et al., Nature Protocols 8:737-748 (2013)). Commercially available kits for preparation of sequencing libraries compatible with single-stranded nucleic acids, such as the Swift Biosciences Accel-NGS® Methyl-Seq DNA Library Kit, are available. In some embodiments, the nucleic acid is amplified after adapter ligation.

[0205] In some embodiments, the adapter contains a sufficient number of different tags to result in a low probability of a combination of tags, e.g., 95%, 99%, or 99.9% of two nucleic acids having the same start and stop points receive the same combination of tags. The adapter can contain the same or different primer binding sites, whether carrying the same or different tags, but preferably the adapter contains the same primer binding site.

[0206] In some embodiments, after binding of the adapter, the nucleic acid is subjected to amplification. For amplification, for example, universal primers that recognize primer binding sites in the adapter can be used.

[0207] In some embodiments, after binding of the adapter, the DNA or a secondary sample or portion of the DNA is fractionated, which includes contacting the DNA with an agent that preferentially binds to nucleic acids carrying epigenetic modifications. The nucleic acid is fractionated into at least two fractionated secondary samples that differ in the degree to which the nucleic acid carries modifications from binding to the agent. For example, if the agent has an affinity for nucleic acids carrying modifications, nucleic acids with overrepresented modifications (compared to the median presentation in the population) preferentially bind to the agent, while nucleic acids with underrepresented modifications do not bind to the agent or elute more readily from the agent. The nucleic acid can then be amplified from primers that bind to primer binding sites within the adapter. Fractionation can alternatively be performed prior to adapter binding, in which case the adapter can contain differential tags that include components that identify in which fraction the molecule originated.

[0208] In some embodiments, the nucleic acid is ligated at both ends to a Y-shaped adapter that includes primer binding sites and tags. The molecule is amplified.

[0209] Molecular tagging In some embodiments, the DNA molecules of the sample can be tagged with sample indices and / or molecular barcodes (commonly referred to as "tags").

[0210] A tag can be a molecule such as a nucleic acid that contains information indicating the characteristics of the molecule to which the tag is associated. For example, a DNA molecule can carry a sample tag or sample index (which discriminates molecules in one sample from those in different samples), a fraction tag (which discriminates molecules in one fraction from those in different fractions), and / or a molecule tag / molecule barcode (which discriminates different molecules from each other (in both unique tagging scenarios and non-unique tagging scenarios)).

[0211] Tagging strategies can be divided into unique tagging strategies and non-unique tagging strategies. In unique tagging, all or substantially all of the molecules in a sample carry different tags, such that reads can be assigned to the original molecules based on the tag information alone. Tags used in such methods are sometimes referred to as "unique tags". In non-unique tagging, different molecules in the same sample can carry the same tag, such that additional information in addition to the tag information is used to assign sequence reads to the original molecules. Such information can include start and stop coordinates, coordinates to which the molecule is mapped, start or stop coordinates alone, etc. Tags used in such methods are sometimes referred to as "non-unique tags". Thus, it is not necessary to uniquely tag all molecules in a sample. This is sufficient to uniquely tag molecules that fall within an identifiable class within the sample. Thus, molecules in different identifiable families can carry the same tag without loss of information about the identity of the tagged molecules.

[0212] In certain embodiments, the tag may include one or a combination of barcodes. As used herein, the term "barcode" refers, depending on the context, to a nucleic acid molecule having a specific nucleotide sequence, or to the nucleotide sequence itself. The barcode may have, for example, between 10 and 100 nucleotides. The collection of barcodes may have degenerate sequences or sequences having a particular Hamming distance if desired for a particular purpose. Thus, for example, a molecular barcode may be composed of one barcode or a combination of two barcodes each attached to a different end of the molecule. Additionally or alternatively, for different fractions and / or samples, different sets of molecular barcodes, molecular tags or molecular indexes may be used such that the barcodes function as molecular tags through their individual sequences and also function to identify the corresponding fractions and / or samples based on the set of which they are members.

[0213] Tags can be used to label individual polynucleotide population fractions so as to correlate the tag(s) with a particular fraction. Alternatively, tags can be used in embodiments of the present disclosure that do not use the step of fractionating. In some embodiments, a single tag can be used to label a particular fraction. In some embodiments, multiple different tags can be used to label a particular fraction. In embodiments where multiple different tags are used to label a particular fraction, the set of tags used to label one fraction can be readily differentiated from the set of tags used to label other fractions. In some embodiments, tags can have additional functionality; for example, tags can be used to index a sample source, or can be used as unique molecular identifiers (which can be used to improve the quality of sequencing data by differentiating sequencing errors from mutations, as, for example, in Kinde et al., Proc Nat'l Acad Sci USA 108: 9530-9535 (2011), Kou et al., PLoS ONE, 11: e0146638 (2016)), or can be used as non-unique molecular identifiers, as, for example, described in U.S. Patent No. 9,598,731. Similarly, in some embodiments, tags can have additional functionality; for example, tags can be used to index a sample source, or can be used as non-unique molecular identifiers (which can be used to improve the quality of sequencing data by differentiating sequencing errors from mutations).

[0214] Tags can be incorporated into adapters or otherwise attached by other means, including, among others, chemical synthesis, ligation (e.g., as described above, e.g., by blunt-end ligation or sticky-end ligation), or overlap extension polymerase chain reaction (PCR). Such adapters are ultimately attached to the sample DNA molecules. In other embodiments, one or more amplification cycles (e.g., PCR amplification) can be applied to introduce sample indices into nucleic acid molecules using conventional nucleic acid amplification methods. Amplification can be performed in one or more reaction mixtures (e.g., multiple microwells in an array). Molecular barcodes and / or sample indices can be introduced simultaneously or in any sequential order. In some embodiments, molecular barcodes and / or sample indices are introduced before and / or after any conversion procedure. When molecular barcodes and / or sample indices are introduced via an amplification process, the conversion step is performed before the molecular barcodes and / or sample indices are introduced. In some embodiments, molecular barcodes and / or sample indices are introduced before and / or after a sequence capture step, if any. In some embodiments, only the molecular barcode is introduced before probe capture, and the sample index is introduced after the sequence capture step is performed. In some embodiments, both the molecular barcode and the sample index are introduced before performing a probe-based capture step, if any. In some embodiments, the sample index is introduced after the sequence capture step is performed, if any. In some embodiments, the sample index is incorporated through overlap extension polymerase chain reaction (PCR).

[0215] In some embodiments, the tag can be located at one or both ends of the sample DNA molecule. In some embodiments, the tag is an oligonucleotide of a predetermined or random or semi-random sequence. In some embodiments, the tags can together be less than about 500, 200, 100, 50, 20, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 nucleotide in length. Typically, the tag is about 5-20 or 6-15 nucleotides in length. The tag can be ligated to the sample DNA molecule randomly or non-randomly.

[0216] In some embodiments, each sample or fraction (discussed below) is uniquely tagged by a sample index or combination of sample indices. In some embodiments, each nucleic acid molecule of a sample or subsample is uniquely tagged by a molecular barcode or combination of molecular barcodes. In other embodiments, multiple molecular barcodes can be used, such that the molecular barcodes are not necessarily unique to each other among the plurality (e.g., non-unique molecular barcodes). In these embodiments, the molecular barcodes are generally attached to individual molecules (e.g., by ligation as part of an adapter), such that the combination of the molecular barcode and the sequence to which it can be attached creates a unique sequence that can be individually tracked. Detection of non-unique molecular barcodes in combination with endogenous sequence information (e.g., the first (start) and / or last (stop) genomic location / position corresponding to the sequence of the original DNA molecule in the sample, the start and stop genomic positions corresponding to the sequence of the original DNA molecule in the sample, the first (start) and / or last (stop) genomic location / position of a sequence read mapped to a reference sequence, the start and stop genomic positions of a sequence read mapped to a reference sequence, a subsequence of a sequence read at one or both ends, the length of the sequence read, and / or the length of the original DNA molecule in the sample) typically enables assignment of a unique identity to a particular molecule. In some embodiments, the start region includes the first 1, first 2, first 5, first 10, first 15, first 20, first 25, first 30, or at least first 30 base positions at the 5' end of a sequencing read that aligns to a reference sequence. In some embodiments, the end region includes the last 1, last 2, last 5, last 10, last 15, last 20, last 25, last 30, or at least last 30 base positions at the 3' end of a sequencing read that aligns to a reference sequence. The length of an individual sequence read, or number of base pairs, can also be used, if desired, to assign a unique identity to a given molecule. As described herein, a fragment from a single strand of a nucleic acid to which a unique identity has been assigned can thereby enable subsequent identification of fragments from the parental strand and / or complementary strand.

[0217] In certain embodiments with non-unique tagging, the number of different tags used can be sufficient such that there is a very high probability (e.g., at least 99%, at least 99.9%, at least 99.99% or at least 99.999%) that all DNA molecules in a particular group carry different tags. When barcodes are used as tags, and the barcodes are attached, for example randomly, to both ends of the molecule, it should be noted that combinations of barcodes can together constitute a tag. This number in the term is a function of the number of molecules entering the call. For example, a class can be all molecules that map to the same start-stop position on a reference genome. A class can be all molecules that map across a particular locus, e.g., a particular base or a particular region (e.g., up to 100 bases, or a gene, or an exon of a gene).

[0218] In certain embodiments, the number z of different tags used to uniquely identify some of the molecules in a class is 2 * z, 3 * z, 4 * z, 5 * z, 6 * z, 7 * z, 8 * z, 9 * z, 10 * z, 11 * z, 12 * z, 13 * z, 14 * z, 15 * z, 16 * z, 17 * z, 18 * z, 19 * z, 20 * z or 100 * either z (e.g., a lower limit) and 100,000 * z, 10,000 * z, 1000 * z or 100 *It can be between any of z (e.g., the upper limit). In some embodiments, the molecular barcodes are introduced into the molecules in the sample at an expected ratio of a set of identifiers (e.g., a combination of unique or non-unique molecular barcodes). One example format uses from about 2 to about 1,000,000 different molecular barcode sequences, or from about 5 to about 150 different molecular barcode sequences, or from about 20 to about 50 different molecular barcode sequences, ligated to both ends of the target molecule. Alternatively, from about 25 to about 1,000,000 different molecular barcode sequences can be used. For example, 20 - 50 kinds × 20 - 50 kinds of molecular barcode sequences (i.e., one of 20 - 50 different molecular barcode sequences can be attached to each end of the target molecule) can be used. Such a number of identifiers is typically sufficient to increase the probability that different molecules with the same start and end points receive different combinations of identifiers (e.g., at least 94%, 99.5%, 99.99%, or 99.999%). In some embodiments, about 80%, about 90%, about 95%, or about 99% of the molecules have the same combination of molecular barcodes. For example, in a sample of about 5 ng to 30 ng of cell-free DNA, approximately 3000 molecules are mapped to specific nucleotide coordinates and are expected to have any starting coordinate such that between about 3 and 10 molecules share the same stop coordinate. Thus, from about 50 to about 50,000 different tags (e.g., barcode combinations between about 6 and 220) may be sufficient to uniquely tag all such molecules. To uniquely tag all 3000 molecules mapped across nucleotide coordinates, from about one million to about 20 million different tags are required.

[0219] In some embodiments, the assignment of unique or non-unique molecular barcodes in a reaction is performed using, for example, the methods and systems described in U.S. Patent Application Nos. 20010053519, 20030152490, and 20110160078, and U.S. Patents Nos. 6,582,908, 7,537,898, 9,598,731, and 9,902,992, each of which is hereby incorporated by reference in its entirety. Alternatively, in some embodiments, different nucleic acid molecules of a sample can be identified using only endogenous sequence information (e.g., start and / or stop positions, partial sequences at one or both ends of the sequence, and / or length). Tags can be linked to the sample nucleic acids either randomly or non-randomly.

[0220] In some embodiments, the tagged nucleic acids are sequenced after being loaded into a microwell plate. The microwell plate can have 96, 384, or 1536 microwells. In some cases, these are introduced in an expected ratio to the microwells of the unique tag. For example, the unique tags can be loaded such that more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000, or 1,000,000,000 unique tags are loaded per genomic sample. In some cases, the unique tags can be loaded such that less than about 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000, or 1,000,000,000 unique tags are loaded per genomic sample. In some cases, the average number of unique tags loaded per sample genome is less than, more than, or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000, or 1,000,000,000 unique tags per genomic sample.

[0221] In some embodiments, the format uses 20 to 50 different tags (e.g., barcodes) ligated to both ends of the target nucleic acid. For example, 35 different tags (e.g., barcodes) ligated to both ends of the target molecule create a 35×35 permutation, which equals 1225 for 35 tags. Such a number of tags is sufficient for different molecules having the same starting and stopping points to have a high probability (e.g., at least 94%, 99.5%, 99.99%, 99.999%) of receiving different combinations of tags. Other barcode combinations include any number between 10 and 500, such as approximately 15×15, approximately 35×35, approximately 75×75, approximately 100×100, approximately 250×250, approximately 500×500.

[0222] In some cases, the unique tags can be predetermined or random or semi-random sequence oligonucleotides. In other cases, multiple barcodes can be used where the barcodes do not necessarily have to be unique from each other among the plurality. In this example, the barcode can be ligated to an individual molecule such that the combination of the barcode and the sequence to which it can be ligated creates a unique sequence that can be individually tracked. As described herein, the detection of non-unique barcodes in combination with the sequence data at the beginning (start) and end (stop) positions of the sequence read can enable the assignment of a unique identity to a particular molecule. The length or number of base pairs of an individual sequence read can also be used to assign a unique identity to such a molecule. Thereby, as described herein, a fragment derived from a single strand of a nucleic acid to which a unique identity has been assigned can enable the subsequent identification of fragments derived from the parental strand.

[0223] In some embodiments, the method includes adding one or more internal control DNAs, as well as a forward primer and a reverse primer for amplifying the internal control DNA. The internal control DNA can be added prior to amplification using primers that anneal upstream and downstream of the rearrangement breakpoint. The forward and reverse primers for amplifying the internal control DNA may be included with, or added simultaneously with, primers that anneal upstream and downstream of the rearrangement breakpoint. The internal control DNA may comprise, or consist of, a sequence that is not present in the genome of the subject, or in the genome of the species of which the subject is a member (e.g., the human genome). The forward and / or reverse primers for amplifying the internal control DNA may comprise a sequence that is not complementary to any sequence in the genome of the subject, e.g., the human genome. The internal control DNA can be used to ensure that the amplification process proceeded as designed. Thus, the method can include detecting (e.g., sequencing) molecules amplified from and / or captured by one or more internal control DNAs. The method can include comparing the amount of the internal control DNA (e.g., the number of detected molecules or reads corresponding to the internal control DNA sequence) to a predetermined threshold and either rejecting the sequencing results if the predetermined threshold is not met or accepting the sequencing results if the predetermined threshold is met. The predetermined threshold can be established, for example, based on past data or by testing the method on DNA samples from test subjects such as healthy volunteers. For example, amplification and detection of one or more internal control DNAs provides confirmation that the amplification process proceeded properly and thus reduces the likelihood of false negatives.

[0224] Modified Sensitivity Sequencing The methods disclosed herein use modified sensitivity sequencing to detect the modified state of at least one type of dNTP containing a modified base used in the end repair reaction. "Modified sensitivity sequencing" refers to any sequencing workflow that can distinguish at least two modified states of nucleotide bases. These two states can be (i) whether the base is modified (e.g., is 5mC and / or 5hmC, or is an unmethylated cytosine), or (ii) the type of modification the base exhibits (e.g., is 5mC or 5hmC). Modified sensitivity sequencing does not necessarily need to identify whether a particular type of modification is present or absent at a particular position, regardless of whether one or more types of modifications (e.g., 5mC and 5hmC) are present or absent. For example, in some embodiments, modified sensitivity sequencing includes sequencing that includes a bisulfite conversion step that can distinguish 5mC and 5hmC from unmethylated C, but cannot distinguish 5mC from 5hmC.

[0225] The type of modified sensitivity sequencing used will, of course, depend on the type of modified base used in the end repair, and as a result, the type of modified sensitivity sequencing can at least detect the presence or absence of at least that modified base. Table 1 summarizes exemplary forms of modified sensitivity sequencing along with the types of modified bases detectable using these methods. These are described in more detail below.

Table 1-1

[0226] As outlined below, there are various methods for detecting and / or identifying modified nucleosides that rely on conversion procedures that change the base pairing specificity of the nucleoside based on its modified state. These changes in base pairing specificity can then be detected, and thus the modified state of the nucleoside can be inferred by sequencing. At the same time, the conversion procedure and the sequencing itself constitute one form of modified recognition sequencing, as referred to herein.

[0227] In some cases, the conversion procedures used in the methods of the present disclosure change the base pairing specificity of a modified nucleoside (e.g., methylated cytosine), but do not change or do not change the base pairing specificity of the corresponding unmodified nucleoside (e.g., cytosine), or of any unmodified nucleosides (e.g., cytosine, adenosine, guanosine, and thymidine (or uracil)). Advantages of methods that do not convert the base pairing specificity of unmodified nucleosides include reduced loss of sequence complexity, improved sequencing efficiency, and reduced alignment loss. In addition, methods such as TAPS are in some cases less destructive (which is particularly important for low-yield samples such as cfDNA or FFPE samples) and do not require denaturation, and may therefore be preferred over methods such as bisulfite sequencing and EM-seq, which means that non-conversion errors are more likely to be theoretically random. In methods that require denaturation for conversion, failure to denature the DNA molecule results in non-conversion of all bases in the DNA molecule. Since biological changes in methylation occur primarily coordinately with respect to the localization region of interest, these non-random (localized) non-conversion events can appear as false negatives (non-methylated regions). Random non-conversion methods can have a maximum impact on a low percentage of bases within a region, and thus the specificity of methylation change detection (reducing false positives) can be maximized by setting a threshold for the percentage of bases within the methylation / non-methylation region. Therefore, in some cases, conversion procedures without denaturation are preferred.

[0228] In other cases, the conversion procedures used in the methods of the present disclosure change the base pairing specificity of unmodified nucleosides (e.g., cytosine), but do not change the base pairing specificity of the corresponding modified nucleosides (e.g., methylated cytosines such as 5hmC and / or 5mC). Such methods include, for example, bisulfite sequencing.

[0229] One of ordinary skill in the art can select an appropriate method as needed, including which nucleoside modifications should be detected and / or identified, and which types of modified bases are used in the end repair reaction.

[0230] In some embodiments, the conversion procedure converts modified nucleosides. In some embodiments, the conversion procedure for converting modified nucleosides includes Tet-assisted conversion with a substituted borane reducing agent, and optionally, the substituted borane reducing agent is 2-picoline borane, borane pyridine, tert-butylamine borane, ammonia borane, or pyridine borane. In Tet-assisted pic-borane conversion using a substituted borane reducing agent conversion, the TET protein is used to convert 5mC and 5hmC to 5caC without affecting unmodified C. Then, 5caC, and 5fC if present, are converted to dihydrouracil (DHU) by treatment with 2-picoline borane (pic-borane) or another substituted borane reducing agent, such as borane pyridine, tert-butylamine borane, or ammonia borane, also without affecting unmodified C. See, for example, Liu et al., Nature Biotechnology 2019; 37:424-429 (e.g., Supplementary Figure 1 and Supplementary Note 7). DHU is read as T in sequencing. Sequencing of the converted DNA identifies positions read as cytosine as unmodified C positions. On the other hand, positions read as T are identified as T, 5mC, 5fC, 5caC, or 5hmC. Thus, performing TAP conversion facilitates identifying positions containing unmodified C using the resulting sequence reads.

[0231] Thus, in these embodiments, the end repair reaction can be carried out using dNTPs, and at least one type of dNTP contains 5mC or 5hmC, and the region synthesized during the end repair reaction can be identified as a region containing 5mC or 5hmC at non-CpG positions (through T called at positions that are C in the reference). This procedure encompasses Tet-assisted pyridine borane sequencing (TAPS), which is described in more detail above in Liu et al. 2019. In this method, Tet enzymes are used to gradually oxidize 5mC and 5hmC to 5fC or 5caC, and then pyridine borane deaminates 5fC, 5caC to DHU, which is amplified as T.

[0232] Alternatively, protection of 5hmC (e.g., using βGT) can be combined with Tet-assisted conversion using a borane reducing agent substituted as described above. In this method (TAPS-β), 5hmC can be protected from conversion, for example, via glucosylation using β-glucosyltransferase (βGT) to form 5ghmC (to form 5-glucosylhydroxymethylcytosine). This is described in Yu et al., Cell 2012; 149: 1368-80. Treatment with a TET protein, e.g., mTet1, then converts 5mC to 5caC, but does not convert C or 5ghmC. 5caC is then converted to DHU by treatment with pic-borane or another substituted borane reducing agent, e.g., borane pyridine, tert-butylamine borane, or ammonia borane, without affecting the unmodified C or 5ghmC. Sequencing of the converted DNA identifies positions read as cytosine as either 5hmC or unmodified C positions. Positions read as T, on the other hand, are identified as T, 5fC, 5caC, or 5mC. Thus, performing TAPSβ conversion on a sample as described herein facilitates using the resulting sequence reads to distinguish positions containing unmodified C or 5hmC on the one hand from positions containing 5mC on the other. Thus, in these embodiments, the end repair reaction can be performed using dNTPs, at least one type of dNTP contains 5mC, and the region synthesized during the end repair reaction can be identified as a region containing 5mC at non-CpG positions (via T called as C at the position in the reference). For an exemplary illustration of this type of conversion, see, for example, Liu et al., Nature Biotechnology 2019; 37:424-429.

[0233] In some embodiments, the conversion procedure converts a modified nucleoside. In some embodiments, the conversion procedure that converts a modified nucleoside includes chemical-assisted conversion using a substituted borane reducing agent. Optionally, the substituted borane reducing agent is 2-picolyl borane, borane pyridine, tert-butylamine borane, borane pyridine, or ammonia borane. In the chemical-assisted conversion using a substituted borane reducing agent, an oxidizing agent such as potassium perruthenate (KRuO4, also suitable for use in ox-BS conversion) is used to specifically oxidize 5hmC to 5fC. Treatment with pic-borane or another substituted borane reducing agent, such as borane pyridine, tert-butylamine borane, or ammonia borane, converts 5fC and 5caC to DHU, but does not affect 5mC or unmodified C. Sequencing of the converted DNA identifies positions read as cytosine as either 5mC or unmodified C positions. On the other hand, positions read as T are identified as T, 5fC, 5caC, or 5hmC. Thus, performing this type of conversion described herein facilitates using the resulting sequence reads to distinguish, on the one hand, positions containing unmodified C or 5mC from positions containing 5hmC. Thus, in these embodiments, the end repair reaction can be performed using dNTPs, and at least one type of dNTP contains 5hmC, and the region synthesized during the end repair reaction can be identified as a region containing 5hmC at non-CpG positions (via T called as C at positions in the reference). For an exemplary description of this type of conversion, see, for example, Liu et al., Nature Biotechnology 2019; 37:424-429.

[0234] Exemplary conversion procedures are described that alter the base pairing specificity of modified cytosine. However, the methods described herein can, in principle, be used with any modified nucleoside and appropriate conversion procedure (i.e., single base epigenetic conversion assay) that alters the base pairing specificity of the modified nucleoside such that, when sequenced, the modified base can be distinguished from the corresponding unmodified nucleoside and / or other types of modifications. For example, any conversion procedure can be used to distinguish N 6 -methyladenine (6mA), N 6 -hydroxymethyladenine (6hmA), or N 6 -formyladenine (6fA) from unmodified adenosine.

[0235] In some embodiments, the conversion procedure converts an unmodified nucleoside. In some embodiments, the conversion procedure that converts an unmodified nucleoside includes bisulfite conversion. Treatment with bisulfite converts unmodified cytosine and certain modified cytosine nucleotides (e.g., 5-formylcytosine (5fC) or 5-carboxylcytosine (5caC)) to uracil, but other modified cytosines (e.g., 5mC and 5hmC) are not converted. Thus, when bisulfite conversion is used, the converted nucleobases are presumed to include one or more of unmodified cytosine, 5fC, 5caC, or other cytosine forms affected by bisulfite. The unconverted nucleobases are presumed to include one or more of 5mC and 5hmC. Sequencing of bisulfite-treated DNA identifies positions read as cytosine as positions of 5mC or 5hmC. On the other hand, positions read as T are identified as being in a bisulfite-sensitive form of T or C (e.g., unmodified cytosine, 5fC or 5caC). Thus, performing bisulfite conversion as described herein facilitates identifying positions containing 5mC or 5hmC. Thus, in these embodiments, the end repair reaction can be performed using dNTPs, at least one type of dNTP includes 5mC and / or 5hmC, and the region synthesized during the end repair reaction can be identified as a region containing 5mC or 5hmC at non-CpG positions (through the C called at these positions). For an exemplary description of bisulfite conversion, see, e.g., Moss et al., Nat Commun. 2018; 9: 5068.

[0236] In some embodiments, the procedure for converting an unmodified nucleoside includes oxidative bisulfite (Ox-BS) conversion. This procedure first converts 5hmC to 5-formylcytosine (5fC), which is bisulfite-sensitive, and then performs bisulfite conversion. Thus, when oxidative bisulfite conversion is used, the nucleobases to be converted are presumed to include one or more of unmodified cytosine, 5fC, 5-carboxylcytosine (5caC), 5hmC, or other cytosine forms affected by bisulfite. The unconverted nucleobases are presumed to include 5mC. Sequencing of Ox-BS-converted DNA identifies positions read as cytosine as 5mC positions. On the other hand, positions read as thymine are identified as thymine or bisulfite-sensitive forms of cytosine, such as unmodified cytosine, 5fC, or 5hmC. Thus, in these embodiments, the end repair reaction can be performed using dNTPs, at least one type of dNTP includes 5mC, and regions synthesized during the end repair reaction can be identified as regions containing 5mC at non-CpG positions (through the C called at these positions). Thus, performing Ox-BS conversion facilitates identifying positions containing mC. For an exemplary description of oxidative bisulfite conversion, see, for example, Booth et al., Science 2012; 336: 934-937.

[0237] In some embodiments, the procedure for converting an unmodified nucleoside includes Tet-assisted bisulfite (TAB) conversion. In TAB conversion, 5hmC is protected from conversion, and 5mC is oxidized prior to bisulfite treatment, such that the positions originally occupied by 5mC are converted to U, while the positions originally occupied by 5hmC remain in the protected form of cytosine. For example, as described in Yu et al., Cell 2012; 149: 1368-80, β-glucosyltransferase can be used to protect 5hmC (forming 5-glucosylhydroxymethylcytosine (5ghmC)), and then a TET protein, e.g., mTet1, can be used to convert 5mC to 5caC, and then bisulfite treatment can be used to convert C and 5caC to U, while 5ghmC remains unaffected. Thus, when TAB conversion is used, the nucleobases to be converted are presumed to include one or more of unmodified cytosine, 5fC, 5caC, 5mC, or other cytosine forms affected by bisulfite. The unconverted nucleobases are presumed to include 5hmC. Sequencing of TAB-converted DNA identifies positions read as cytosine as 5hmC positions. On the other hand, positions read as T are identified as T, or a bisulfite-sensitive form of C, e.g., unmodified cytosine, 5mC, 5fC, or 5caC. Thus, performing TAB conversion on a first secondary sample as described herein facilitates identifying positions containing 5hmC. Thus, in these embodiments, the end repair reaction can be performed using dNTPs, where at least one type of dNTP contains 5hmC, and the region synthesized during the end repair reaction can be identified as a region containing 5hmC at non-CpG positions (through the C called at these positions).

[0238] In some embodiments, the conversion procedure for converting an unmodified nucleoside includes APOBEC-coupled epigenetic (ACE) conversion. In ACE conversion, an AID / APOBEC family DNA deaminase enzyme, such as APOBEC3A (A3A), is used to deaminate unmodified cytosine and 5mC without deaminating 5hmC, 5fC, or 5caC. Sequencing of ACE-converted DNA identifies positions read as cytosine as being 5hmC, 5fC, or 5caC positions. Meanwhile, positions read as T are identified as being T, unmodified C, or 5mC. Thus, performing ACE conversion as described herein facilitates identifying positions containing 5hmC from positions containing 5mC or unmodified C using sequence reads obtained from a first secondary sample. Thus, in these embodiments, the end repair reaction can be performed using dNTPs, at least one type of dNTP includes 5hmC, and the region synthesized during the end repair reaction can be identified as a region containing 5hmC at non-CpG positions (through the C called at these positions). For an exemplary description of ACE conversion, see, for example, Schutsky et al., Nature Biotechnology 2018; 36: 1083-1090.

[0239] Modification-sensitive sequencing also includes sequencing methods that do not rely on a conversion step, where the base pairing specificity of a base changes depending on its modification state. For example, single molecule techniques such as nanopore-based sequencing and single molecule real-time sequencing can be used to directly detect modified bases.

[0240] For example, some sequencing reactions involve the use of enzymes to control the passage of nucleic acids through a nanopore, and in such cases, the reaction data can include both the kinetics and other behaviors of the enzyme, as well as fluctuations in the current passing through the nanopore. For example, ratchet proteins, helicases, or motor proteins can be used to push or pull nucleic acid molecules through holes in biological or synthetic membranes. The kinetics of these proteins can vary depending on the sequence context of the nucleic acid on which they act. For example, they may slow down or stop at sites of modified bases, and this behavior, captured as part of the reaction data, indicates the presence of modified bases even if the modified bases are not inside the sensing portion of the nanopore. An example of a nanopore sequencing system is one commercially available from Oxford Nanopore Technologies (ONT) (see, e.g., Weirather et al., F1000Research, 6:100, 2017). ONT sequencing directly sequences native single-stranded DNA (ssDNA) molecules by measuring characteristic current changes as bases pass through the nanopore by a molecular motor protein. ONT sequencing uses a hairpin library structure similar to that of PacBio circular DNA templates:DNA templates, to which hairpin adapters are ligated on their complements. Thus, the DNA template passes through the nanopore, then the hairpin, and finally the complement. Raw reads can be split into two "1D" reads ("template" and "complement") by removing the adapters. The consensus sequence of the two "1D" reads is a "2D" read with higher accuracy.

[0241] Base modifications including 5mC, 5hmC, 6mA, BrdU, FdU, IdU, and EdU can be detected using nanopore sequencing (see, e.g., Gouil & Keniry Essays in Biochemistry (2019) 63 639-648; Kutyavin, Biochemistry (2008), 47, 51, 13666-1367; Mueller et al., Nature Methods (2019), volume 16, pages 429-436; Hennion et al., Genome Biology (2020), volume 21, Article number: 125). Thus, in some embodiments, modification-sensitive sequencing includes nanopore sequencing. In such embodiments, end repair can be performed using dNTPs including 4mC, 5mC, 5hmC, 6mA, BrdU, FdU, IdU, and / or EdU.

[0242] Another modification-sensitive single molecule sequencing technique is single molecule real-time (SMRT) sequencing, commercialized by Pacific Biosciences. SMRT sequencing relies on synthetic sequencing in which the sequence of a circular DNA template is determined from the succession of fluorescence pulses resulting from the addition of one labeled nucleotide by a polymerase fixed to the bottom of a well. Base modifications do not affect the sequence called by the base, but do affect the dynamics of the polymerase. By considering the interpulse duration (IPD), base modifications can be inferred from a comparison of the modified template with an in silico model or an unmodified template. Thus, such methods can use the pulse width of the signal from the sequencing of the base, the interpulse duration (IPD) of the base, and the identity of the base to detect modifications in the base or adjacent bases. (See, e.g., Weirather et al., F1000Research, 6:100, 2017).

[0243] Base modifications such as 4mC, 5mC, 5hmC, 6mA, and 8oxoG can be detected using single molecule real-time sequencing (Gouil & Keniry Essays in Biochemistry (2019) 63 639-648). Thus, in some embodiments, modification-sensitive sequencing includes single molecule real-time sequencing. In such embodiments, end repair can be performed using dNTPs that include 4mC, 5mC, 5hmC, 6mA, and / or 8oxoG.

[0244] Fractionation In some examples, a heterogeneous nucleic acid sample is fractionated into two or more fractions (sub-samples). In some embodiments, each fraction is differentially tagged. The tagged fractions can then be pooled together for pooled sample preparation and / or sequencing. The steps of fractionating - tagging - pooling can be performed more than once, and each round of fractionation is performed based on different characteristics and tagged using differential tags that distinguish it from other fractions and fractionation means.

[0245] Examples of features that can be used for fractionation include array length, methylation level, nucleosome binding, array mismatch, immunoprecipitation, and / or proteins that bind to DNA. The resulting fractions can contain one or more of the following nucleic acid forms: single-stranded DNA (ssDNA), double-stranded DNA (dsDNA), shorter DNA fragments, and longer DNA fragments. In some embodiments, the step of fractionating based on cytosine modification (e.g., cytosine methylation) or methylation is generally performed and, optionally, combined with at least one additional fractionation step that can be based on any of the aforementioned features or forms of DNA. In some embodiments, a heterogeneous population of nucleic acids is fractionated into nucleic acids having one or more base modifications and nucleic acids having no base modifications. Examples of base modifications are described elsewhere in this specification. Alternatively or additionally, a heterogeneous population of nucleic acids can be fractionated into nucleic acid molecules associated with nucleosomes and nucleic acid molecules without nucleosomes. Alternatively or additionally, a heterogeneous population of nucleic acids can be fractionated into single-stranded DNA (ssDNA) and double-stranded DNA (dsDNA). Alternatively or additionally, a heterogeneous population of nucleic acids can be fractionated based on nucleic acid length (e.g., molecules up to 160 bp in length and molecules having a length greater than 160 bp).

[0246] In some cases, different procedures are applied to different fractions to determine different characteristics of the initial sample. The DNA of at least one fraction is subjected to end repair and modified sensitivity sequencing procedures according to the methods of the present disclosure described herein. In some embodiments, at least one fraction is not subjected to end repair and modified sensitivity sequencing procedures according to the methods of the present disclosure described herein. If the modified sensitivity sequencing procedure includes a conversion procedure, the corresponding sequences from the converted and unconverted fractions can be compared to identify the single nucleotide that has undergone conversion and, therefore, the corresponding modified nucleoside in the initial sample can be identified.

[0247] In some embodiments, tagging of fractions involves tagging the molecules in each fraction with a fraction tag. After the fractions are recombined (e.g., to reduce the number of sequencing runs required and avoid unnecessary costs) and the molecules are sequenced, the fraction tag identifies the source fraction. In another embodiment, different fractions are tagged with different sets of molecule tags, e.g., consisting of pairs of barcodes. In this way, each molecule barcode is useful not only for indicating the source fraction but also for identifying the molecules within the fraction. For example, a first set of 35 barcodes can be used to tag the molecules in a first fraction, and a second set of 35 barcodes can be used to tag the molecules in a second fraction.

[0248] In some embodiments, after fractionation and tagging with a fraction tag, the molecules can be pooled for sequencing in a single run. In some embodiments, a sample tag is added to the molecules, e.g., in a step subsequent to the addition and pooling of the fraction tags. The sample tag can facilitate pooling materials generated from multiple samples for sequencing in a single sequencing run.

[0249] Alternatively, in some embodiments, the fraction tag can be correlated with both the sample and the fraction. As a simple example, a first tag can indicate the first fraction of a first sample; a second tag can indicate the second fraction of the first sample; a third tag can indicate the first fraction of a second sample; a fourth tag can indicate the second fraction of the second sample.

[0250] Tags can be attached to molecules that have already been fractionated based on one or more characteristics, although the final tagged molecules in the library may no longer have those characteristics. For example, single-stranded DNA molecules can be fractionated and tagged, but the final tagged molecules in the library are likely to be double-stranded. Similarly, DNA can be subjected to fractionation based on different levels of methylation, but in the final library, the tagged molecules derived from these molecules are likely to be unmethylated. Thus, tags attached to molecules in the library typically indicate the characteristics of the "parent molecule" from which the final tagged molecule is derived, and do not necessarily indicate the characteristics of the tagged molecule itself.

[0251] As an example, barcodes 1, 2, 3, 4, etc. are used to tag and label molecules in the first fraction; barcodes A, B, C, D, etc. are used to tag and label molecules in the second fraction; barcodes a, b, c, d, etc. are used to tag and label molecules in the third fraction. The differentially tagged fractions can be pooled prior to sequencing. The differentially tagged fractions can be sequenced separately or, for example, sequenced together in parallel in the same flow cell of an Illumina sequencer.

[0252] After sequencing, the analysis of the reads can be performed at the level of each fraction, as well as at the level of the total DNA population. Tags are used to sort reads from different fractions. The analysis can include in silico analysis to determine genetic and epigenetic variations (one or more of methylation, chromatin structure, etc.) using sequence information, genomic locus length, coverage, and / or copy number. In some embodiments, higher coverage may correlate with higher nucleosome occupancy in the genomic region, while lower coverage may correlate with lower nucleosome occupancy or nucleosome-depleted regions (NDRs).

[0253] The methods disclosed herein include the step of analyzing DNA in a sample. In some embodiments described herein, the disclosed methods include the step of fractionating DNA. In such methods, different forms of DNA (e.g., hypermethylated and hypomethylated DNA) can be physically fractionated based on one or more characteristics of the DNA. This approach can be used, for example, to determine whether a particular sequence is hypermethylated or hypomethylated. In some embodiments, a first subsample or aliquot of the sample is subjected to the steps for making capture probes as described elsewhere herein, and a second subsample or aliquot of the sample is subjected to fractionation. In some embodiments, the sample or its subsample or aliquot is subjected to fractionation and differential tagging, and then to a capture step using capture probes for the rearranged sequences and, optionally, additional capture probes, e.g., capture probes for sequence variable regions and / or epigenetic target regions.

[0254] Methylation profiling can include determining methylation patterns across different regions of the genome. For example, after fractionating and sequencing molecules based on the degree of methylation (e.g., the relative number of methylated nucleobases per molecule), the sequences of the molecules in the different fractions can be mapped to a reference genome. This can indicate regions of the genome that are more highly methylated or less highly methylated compared to other regions. In this method, genomic regions can differ in the degree of methylation, as opposed to individual molecules.

[0255] Fractionating nucleic acid molecules in a sample can increase rare signals, for example, by enriching rare nucleic acid molecules that are more common in one fraction of the sample. For example, gene variations that are present in highly methylated DNA but less (or not) present in lowly methylated DNA can be more easily detected by fractionating the sample into highly methylated nucleic acid molecules and lowly methylated nucleic acid molecules. By analyzing multiple fractions of a sample, multi-dimensional analysis of single molecules can be performed, and thus higher sensitivity can be achieved. Fractionation can include physically fractionating nucleic acid molecules into fractions or secondary samples based on the presence or absence of one or more methylated nucleobases. A sample can be fractionated into fractions or secondary samples based on features indicative of differential gene expression or disease state. A sample can be fractionated during the analysis of nucleic acids, such as cell-free DNA (cfDNA), non-cfDNA, tumor DNA, circulating tumor DNA (ctDNA), and cell-free nucleic acid (cfNA), based on features or combinations thereof that provide a signal difference between a normal state and an affected state.

[0256] In some embodiments, hypermethylated and / or hypomethylated variable epigenetic target regions are analyzed to determine whether they exhibit differential methylation characteristic of types of cells (such as cfDNA) that do not normally contribute to tumor cells or the DNA sample being analyzed, and / or specific immune cell types.

[0257] In some cases, the heterogeneous DNA in the sample is fractionated into two or more fractions (e.g., at least 3, 4, 5, 6, or 7 fractions). In some embodiments, each fraction is separately tagged. The tagged fractions can then be pooled together for collective sample preparation and / or sequencing. The fractionation-tagging-pooling steps can be performed more than once, and each round of fractionation is performed based on different characteristics (examples provided herein) and tagged using differential tags that distinguish it from other fractions and fractionation means. In other examples, the separately tagged fractions are sequenced separately.

[0258] Agents used to fractionate a population of nucleic acids in a sample include affinity agents such as antibodies with desired specificities, natural binding partners or their variants (Bock et al., Nat Biotech 28: 1106-1114 (2010); Song et al., Nat Biotech 29: 68-72 (2011)), or artificial peptides selected, for example, by phage display to have specificity for a given target. In some embodiments, the agent used for fractionation is an agent that recognizes a modified nucleobase. In some embodiments, the modified nucleobase recognized by the agent is a modified cytosine, such as methylcytosine (e.g., 5-methylcytosine). In some embodiments, the modified nucleobase recognized by the agent is the product of a procedure that affects the first nucleobase in the DNA differently from the second nucleobase in the sample's DNA. In some embodiments, the modified nucleobase can be a "converted nucleobase," which means that its base pairing specificity has been changed by the procedure. For example, a particular procedure converts unmethylated or unmodified cytosine to dihydrouracil, or more generally, at least one modified or unmodified form of cytosine undergoes deamination, resulting in uracil (considered a modified nucleobase in the context of DNA) or a further modified form of uracil. Examples of fractionating agents include antibodies that recognize modified nucleobases, which can be modified cytosines such as methylcytosine (e.g., 5-methylcytosine). In some embodiments, the fractionating agent is an antibody that recognizes a modified cytosine other than 5-methylcytosine, such as 5-carboxylcytosine (5caC). Alternative fractionating agents include the methyl-binding domain (MBD) and methyl-binding proteins (MBP) described herein, including proteins such as MeCP2.

[0259] A further non-limiting example of a fractionating agent is a histone-binding protein that can separate nucleic acids bound to histones from free or unbound nucleic acids. Examples of histone-binding proteins that can be used in the methods disclosed herein include RBBP4, RbAp48, and SANT domain peptides.

[0260] In some embodiments, fractionation can include both binary fractionation and fractionation based on degree / level of modification. For example, methylated fragments can be fractionated by methylated DNA immunoprecipitation (MeDIP), or all methylated fragments can be fractionated from unmethylated fragments using a methyl-binding domain protein (e.g., MethylMinder Methylated DNA Enrichment Kit (ThermoFisher Scientific)). Subsequently, additional fractionation can include eluting fragments with different levels of methylation by adjusting the salt concentration in a solution containing the methyl-binding domain and the bound fragments. As the salt concentration increases, fragments with higher levels of methylation are eluted.

[0261] Analyzing DNA can include detecting or quantifying the DNA of interest. Analyzing DNA can include detecting genetic variants and / or epigenetic features (e.g., DNA methylation and / or DNA fragmentation).

[0262] In some embodiments, the methylation level can be determined using fractionation, modification-sensitive conversion, such as bisulfite conversion, direct detection during sequencing, methylation-sensitive restriction enzyme digestion, methylation-dependent restriction enzyme digestion, or any other suitable approach. For example, different forms of DNA (e.g., hypermethylated DNA and hypomethylated DNA) can be physically fractionated based on one or more characteristics of the DNA. For example, methylated DNA-binding proteins (e.g., MBDs such as MBD2, MBD4, or MeCP2) or antibodies specific for 5-methylcytosine (such as in MeDIP) can be used to fractionate the DNA. This approach can be used, for example, to determine whether a particular sequence is hypermethylated or hypomethylated. In some embodiments, the DNA fragmentation pattern can be determined based on the endpoints and / or midpoints of DNA molecules such as cfDNA molecules.

[0263] In some examples, the final fractions are enriched in nucleic acids having different degrees of modification (presenting an excess or a deficiency of the modification). Excess presentation and deficiency presentation can be defined by the number of modifications produced by the nucleic acid compared to the median number of modifications per strand in the population. For example, if the median number of 5-methylcytosine residues in the nucleic acids in a sample is 2, nucleic acids containing more than 2 5-methylcytosine residues have this modification presented in excess, and nucleic acids having 1 or zero 5-methylcytosine residues have it presented deficiently. The effect of affinity separation is to enrich nucleic acids with an excess presentation of the modification in the bound phase and nucleic acids with a deficiency presentation of the modification in the unbound phase (i.e., in solution). The nucleic acids in the bound phase can be eluted prior to further processing.

[0264] When using MeDIP or the MethylMiner® Methylated DNA Enrichment Kit (ThermoFisher Scientific), various levels of methylation can be fractionated using sequential elution. For example, a hypomethylated fraction (no methylation) can be separated from a methylated fraction by contacting a nucleic acid population with MBD from this kit that is bound to magnetic beads. The beads are used to separate methylated nucleic acids from unmethylated nucleic acids. Subsequently, one or more elution steps are sequentially performed to elute nucleic acids with different levels of methylation. For example, a first set of methylated nucleic acids can be eluted at a salt concentration of 160 mM or higher, such as at least 150 mM, at least 200 mM, 300 mM, 400 mM, 500 mM, 600 mM, 700 mM, 800 mM, 900 mM, 1000 mM, or 2000 mM. After such methylated nucleic acids are eluted, magnetic separation is used again to separate more highly methylated nucleic acids from nucleic acids having a lower level of methylation. The elution step and the magnetic separation step can be repeated multiple times to generate various fractions, such as a hypomethylated fraction (enriched in nucleic acids without methylation), a methylated fraction (enriched in nucleic acids with low levels of methylation), and a hypermethylated fraction (enriched in nucleic acids with high levels of methylation).

[0265] In some methods, nucleic acids bound to the agent used for fractionation based on affinity separation are subjected to a washing step. The washing step washes away nucleic acids that weakly bind to the affinity agent. Such nucleic acids can be enriched in nucleic acids having modifications to an extent close to the average or median (i.e., the intermediate between nucleic acids remaining bound to the solid phase and nucleic acids not bound to the solid phase at the first contact of the sample with the agent).

[0266] Affinity separation yields at least two, sometimes three or more fractions of nucleic acids having different degrees of modification. These fractions are still separate, but the nucleic acids of at least one fraction, usually the nucleic acids of two or three (or more) fractions, are linked to nucleic acid tags that are usually provided as components of the adapters, and the nucleic acids in different fractions receive different tags that identify a member of one fraction from another. The tags linked to the nucleic acid molecules of the same fraction can be the same or different from each other. However, if different from each other, these tags can have a common part of their codes in order to identify the molecules to which they are attached as being those of a particular fraction.

[0267] For further details regarding fractionating nucleic acid samples based on features such as methylation, see WO2018 / 119452, which is hereby incorporated by reference herein.

[0268] In some embodiments, nucleic acid molecules can be fractionated into different fractions based on nucleic acid molecules that bind to a specific protein or fragment thereof and nucleic acid molecules that do not bind to the specific protein or fragment thereof.

[0269] Nucleic acid molecules can be fractionated based on DNA-protein binding. Protein-DNA complexes can be fractionated based on specific properties of the protein. Examples of such properties include various epitopes, modifications (e.g., histone methylation or acetylation), or enzyme activity. Examples of proteins that can bind to DNA and serve as a basis for fractionation include, but are not limited to, Protein A and Protein G. Any suitable method can be used to fractionate nucleic acid molecules based on the protein-bound region. Examples of methods used to fractionate nucleic acid molecules based on the protein-bound region include, but are not limited to, SDS-PAGE, chromatin-immunoprecipitation (ChIP), heparin chromatography, and asymmetrical field flow fractionation (AF4).

[0270] In some embodiments, fractionation includes contacting the DNA with a methylation-sensitive restriction enzyme (MSRE) and / or a methylation-dependent restriction enzyme (MDRE). After treating the DNA with an MSRE or MDRE, the DNA is fractionated based on size to generate hypomethylated (the longest DNA molecules after MSRE treatment and the shortest DNA fragments after MDRE treatment), intermediate methylation (DNA molecules of intermediate length after MSRE or MDRE treatment), and hypermethylated (the shortest DNA molecules after MSRE treatment and the longest DNA fragments after MDRE treatment) sub-samples.

[0271] In some embodiments, fractionation is performed by contacting the nucleic acid with the methyl-binding domain (MBD) of a methyl-binding protein (“MBP”). In some such embodiments, the nucleic acid is contacted with the entire MBP. In some embodiments, the MBD binds to 5-methylcytosine (5mC), the MBP contains the MBD, and is herein referred to interchangeably as a methyl-binding protein or methyl-binding domain protein. In some embodiments, the MBD is coupled via a biotin linker to paramagnetic beads, such as Dynabeads® M-280 streptavidin. Fractionation into fractions having different degrees of methylation can be performed by eluting the fractions by increasing the NaCl concentration.

[0272] In some embodiments, the bound DNA is eluted by contacting the antibody or MBD with a protease, such as proteinase K. This can be done instead of, or in addition to, the elution step using NaCl discussed above.

[0273] Examples of agents that recognize modified nucleobases contemplated herein include, but are not limited to: (a) MeCP2 is a protein that preferentially binds 5-methyl-cytosine over unmodified cytosine. (b) RPL26, PRP8, and DNA mismatch repair protein MHS6 preferentially bind 5-hydroxymethyl-cytosine over unmodified cytosine. (c) FOXK1, FOXK2, FOXP1, FOXP4, and FOXI3 preferably bind 5-formyl-cytosine over unmodified cytosine (Iurlaro et al., Genome Biol. 14: R119 (2013)). (d) One or more antibodies specific for methylated or modified nucleobases or their conversion products, such as 5mC, 5caC, or DHU.

[0274] Generally, elution is a function of the number of modifications, e.g., the number of methylated sites per molecule, and molecules with more methylation elute under increased salt concentrations. To elute DNA into separate populations based on the degree of methylation, a series of elution buffers with increasing NaCl concentrations can be used. The salt concentration can range from about 100 nM to about 2500 mM NaCl. In one embodiment, the process yields three fractions. The molecules are contacted with a solution containing an agent that recognizes the modified nucleobase at a first salt concentration, and this molecule can be bound to a capture moiety, e.g., streptavidin. At the first salt concentration, a population of molecules binds to the agent and a population remains unbound. The unbound population can be separated as the "low-methylated" population. For example, the first fraction enriched in the low-methylated form of DNA is the fraction that remains unbound at a low salt concentration, e.g., 100 mM or 160 mM. The second fraction enriched in intermediate-methylated DNA is eluted using an intermediate salt concentration, e.g., a concentration from 100 mM to 2000 mM. This is also separated from the sample. The third fraction enriched in the hypermethylated form of DNA is eluted using a high salt concentration, e.g., at least about 2000 mM.

[0275] In some embodiments, a monoclonal antibody produced against 5-methylcytidine (5mC) is used to purify methylated DNA. The DNA is denatured, e.g., at 95 °C, to obtain single-stranded DNA fragments. Protein G coupled to standard beads or magnetic beads, along with washing after incubation with the anti-5mC antibody, is used to immunoprecipitate the DNA bound to the antibody. The such DNA can then be eluted. The fractions can include the DNA that did not precipitate and one or more fractions eluted from the beads.

[0276] In some embodiments, the DNA fractions are desalted and enriched in preparation for the enzymatic steps of library preparation.

[0277] Arrays containing abnormally high copy numbers may tend to be hypermethylated. Thus, in some embodiments, DNA contacted with capture probes specific for members of an epigenetic target region set comprising a plurality of target regions that are both type-specific differentially methylated regions and copy number variants contains at least a portion of the hypermethylated fraction. DNA from or containing at least a portion of the hypermethylated fraction may or may not be combined with DNA from or containing at least a portion of one or more other fractions, such as intermediate or hypomethylated fractions.

[0278] Amplification Prior to, or as part of, modified sensitivity sequencing, the adapted DNA can be amplified (e.g., by PCR). For example, in a modified sensitivity sequencing procedure that includes a conversion step, the adapted DNA can be amplified after the conversion step. In a modified sensitivity sequencing procedure that includes single molecule sequencing (such as nanopore-based sequencing or SMRT sequencing), an amplification step may not be required.

[0279] Amplification is typically primed by a primer that binds to a primer binding site in an adapter adjacent to the DNA molecule to be amplified. The amplification method can include cycles of denaturation, annealing, and extension resulting from thermocycling, or can be isothermal in the case of transcription-mediated amplification. Other amplification methods include ligase chain reaction, strand displacement amplification, nucleic acid sequence-based amplification, and self-sustained sequence-based replication.

[0280] In some embodiments, the method performs dsDNA ligation using T-tailed and C-tailed adapters. The addition of the C-tailed adapter can increase ligation efficiency, which is because when A-tailing is performed in the presence of dGTP, for example, when A-tailing is performed in the same reaction as end repair, the A-tailing reaction can also add a G tail to a small portion of the DNA molecule. The use of T-tailed and C-tailed adapters can result in amplification of at least 50, 60, 70, or 80% of the double-stranded nucleic acid. Preferably, the method of the present invention increases the amount or number of amplified molecules by at least 10, 15, or 20% compared to a control method performed with the T-tailed adapter alone.

[0281] In some embodiments, the adapted DNA is amplified prior to sequencing. Amplification can, in some cases, be prior to one or more capture steps. In some embodiments, the ligation step is performed after the conversion step. In some embodiments, ligation occurs before or simultaneously with amplification.

[0282] Enrichment, capture, and use of capture probes The DNA molecules in the sample can be subjected to a capture step in which molecules having the target sequence are captured for subsequent analysis. In some embodiments, the methods disclosed herein include the step of capturing one or more sets of target regions of DNA, such as cfDNA. The capture can be performed using any suitable approach known in the art. Target capture can involve the use of a bait set that includes a capture moiety, such as an oligonucleotide bait labeled with biotin or other examples noted below. The probes can have sequences selected to tile across a region, such as a panel of genes. Such a bait set is combined with the sample under conditions that allow hybridization of the target molecules to the baits. The captured molecules are then isolated using the capture moiety. For example, biotin capture moiety with bead-based streptavidin. Such methods are further described, for example, in U.S. Patent No. 9,850,523, issued December 26, 2017, which is hereby incorporated by reference herein.

[0283] Capture moieties include, but are not limited to, biotin, avidin, streptavidin, nucleic acids containing specific nucleotide sequences, haptens recognized by antibodies, and magnetically attractable particles. The extraction moiety can be a member of a binding pair, such as biotin / streptavidin or hapten / antibody. In some embodiments, the capture moiety bound to the analyte is captured by an isolatable moiety, such as a magnetically attractable particle, or its binding pair bound to a large particle that can be sedimented via centrifugation. The capture moiety can be any type of molecule that allows for affinity separation of the nucleic acid bearing the capture moiety from the nucleic acid lacking the capture moiety. Exemplary capture moieties are biotin, which allows for affinity separation by binding to streptavidin that is linked or linkable to a solid phase, or an oligonucleotide that allows for affinity separation via binding to a complementary oligonucleotide that is linked or linkable to a solid phase.

[0284] The panel of regions targeted for enrichment can be selected to exclude regions known to contain base modifications used in the end repair reaction. If end repair is performed using dNTPs containing 5mC or 5hmC, the panel of regions targeted for enrichment can be selected to exclude CpH dinucleotides known to be naturally methylated in the subject (e.g., human). Such CpH dinucleotides can be identified through the use of publicly available resources (e.g., MethBank 3.0: a database of DNA methylomes across a variety of species Nucleic Acids Res 2018). Such an approach has the advantage that any detected methylated CpH dinucleotides can clearly arise from regions synthesized during end repair.

[0285] In some embodiments, the capturing step comprises contacting the DNA to be captured with a set of target-specific probes. The set of target-specific probes can have any of the features described herein for sets of target-specific probes, including but not limited to those shown in the embodiments above and in the section on probes below. The capturing step can be performed on one or more sub-samples prepared during the methods disclosed herein. In some embodiments, the DNA is captured from at least a first sub-sample or a second sub-sample, e.g., from at least a first sub-sample and a second sub-sample. In some embodiments, the sub-samples are differentially tagged (e.g., as described herein) and then pooled prior to undergoing capture.

[0286] The capturing step can be performed using conditions appropriate for specific nucleic acid hybridization, which generally depend to some extent on features of the probes, such as length, base composition, etc. Those skilled in the art are familiar with appropriate conditions, taking into account general knowledge in the art regarding nucleic acid hybridization. In some embodiments, a complex is formed between the target-specific probe and the DNA.

[0287] In some embodiments, the methods described herein include capturing cfDNA obtained from a subject for a plurality of sets of target regions. The target regions include epigenetic target regions that may exhibit differences in methylation levels and / or fragmentation patterns depending on whether they originate from tumors or healthy cells. The target regions also include sequence-variable target regions that may exhibit differences in sequence depending on whether they originate from tumors or healthy cells. The capturing step generates a captured set of cfDNA molecules, and cfDNA molecules corresponding to the set of sequence-variable target regions are captured with a higher capture yield in the captured set of cfDNA molecules than cfDNA molecules corresponding to the set of epigenetic target regions. For additional discussion of the capturing step, capture yield, and related aspects, see WO2020 / 160414, which is hereby incorporated by reference herein for all purposes.

[0288] In some embodiments, the methods described herein include contacting cfDNA obtained from a subject with a set of target-specific probes, the set of target-specific probes being configured to capture cfDNA corresponding to the set of sequence-variable target regions with a higher capture yield than cfDNA corresponding to the set of epigenetic target regions.

[0289] A greater depth of sequencing may be required to analyze the array-variable target regions with sufficient confidence or accuracy than may be required when analyzing epigenetic target regions. Thus, it may be beneficial to capture cfDNA corresponding to a set of array-variable target regions with a higher capture yield than cfDNA corresponding to a set of epigenetic target regions. The volume of data required to determine fragmentation patterns (e.g., to test for disruption of transcription start sites or CTCF binding sites) or fragment abundance (e.g., in hypermethylated and hypomethylated fractions) is generally smaller than the volume of data required to determine the presence or absence of cancer-related sequence variations. Capturing sets of target regions with different yields can facilitate sequencing the target regions to different depths of sequencing in the same sequencing run (e.g., using pooled mixtures and / or in the same sequencing cell).

[0290] In various embodiments, these methods further include, consistent with the discussion herein, sequencing the captured cfDNA, for example, to different extents of sequencing depth for a set of epigenetic target regions and a set of array-variable target regions.

[0291] In some embodiments, the complex of the target-specific probe and DNA is separated from DNA not bound to the target-specific probe. For example, if the target-specific probe is covalently or non-covalently bound to a solid support, a washing or aspiration step can be used to separate the unbound material. Alternatively, if the complex has chromatographic properties distinct from the unbound material (e.g., if the probe contains a ligand that binds to a chromatographic resin), chromatography can be used.

[0292] As discussed in detail elsewhere in this specification, a set of target-specific probes can include multiple sets, e.g., probes for a set of array variable target regions and probes for a set of epigenetic target regions. In some such embodiments, the capturing step is performed simultaneously in the same container using probes for a set of array variable target regions and probes for a set of epigenetic target regions, e.g., the probes for a set of array variable target regions and the probes for a set of epigenetic target regions are in the same composition. This approach provides a relatively streamlined workflow. In some embodiments, the concentration of the probes for the set of array variable target regions is higher than the concentration of the probes for the set of epigenetic target regions.

[0293] Alternatively, the capturing step is performed using a set of array variable target region probes in a first container and using a set of epigenetic target region probes in a second container, or the contacting step is performed using a set of array variable target region probes at a first time point and in a first container and using a set of epigenetic target region probes at a second time point before or after the first time point. This approach enables the preparation of separate first and second compositions containing captured DNA corresponding to a set of array variable target regions and captured DNA corresponding to a set of epigenetic target regions. These compositions can be processed separately, if desired (e.g., for fractionation based on methylation as described elsewhere in this specification), and recombined at appropriate ratios to provide material for further processing and analysis, e.g., for sequencing.

[0294] In some embodiments, a captured set of DNA (e.g., cfDNA) is provided. With respect to the disclosed methods, the captured set of DNA can be provided, for example, by performing a capturing step prior to the sequencing step described herein. The captured set can include DNA corresponding to a set of sequence-variable target regions, a set of epigenetic target regions, or a combination thereof. In some embodiments, the capturing step is performed before or after the conversion step.

[0295] In some embodiments, a first set of target regions is captured (e.g., from a sample or a first subsample) and includes at least epigenetic target regions. The epigenetic target regions captured from the first subsample can include hypermethylation-variable target regions. In some embodiments, the hypermethylation-variable target regions are CpG-containing regions that are unmethylated or hypomethylated (e.g., have methylation below the average compared to bulk cfDNA) in cfDNA from healthy subjects. In some embodiments, the hypermethylation-variable target regions are regions in healthy cfDNA that exhibit lower methylation than in at least one other tissue type. Without wishing to be bound by any particular theory, cancer cells may shed more DNA into the bloodstream than healthy cells of the same tissue type. Thus, the distribution of the tissue from which cfDNA originates can change during carcinogenesis. Thus, an increase in the level of hypermethylation-variable target regions in the first subsample can be an indicator of the presence of cancer (or recurrence, depending on the subject's medical history).

[0296] In some embodiments, the second set of target regions is captured from a second sub-sample that includes at least epigenetic target regions. The epigenetic target regions can include hypomethylated variable target regions. In some embodiments, the hypomethylated variable target regions are CpG-containing regions in cfDNA from a healthy subject that have methylation or high methylation (e.g., methylation above the average compared to bulk cfDNA). In some embodiments, the hypomethylated variable target regions are regions in healthy cfDNA that exhibit higher methylation than in at least one other tissue type. Without wishing to be bound by any particular theory, cancer cells may shed more DNA into the bloodstream than healthy cells of the same tissue type. Thus, the distribution of the tissue from which the cfDNA originated can change during carcinogenesis. Thus, an increase in the level of hypomethylated variable target regions in the second sub-sample can be an indicator of the presence of cancer (or recurrence, depending on the subject's medical history).

[0297] In some embodiments, the amount of captured sequence variable target region DNA is greater than the amount of captured epigenetic target region DNA when normalized for differences in the size of the targeted regions (footprint size).

[0298] Alternatively, first and second captured sets can be provided, each including DNA corresponding to a set of sequence variable target regions and DNA corresponding to a set of epigenetic target regions, respectively. The first and second captured sets can be combined to provide a combined captured set.

[0299] In some embodiments, the captured set containing DNA corresponding to the array-variable target region set and the epigenetic target region set includes the combined captured set discussed above. The DNA corresponding to the array-variable target region set may be present at a higher concentration than the DNA corresponding to the epigenetic target region set, for example, 1.1 to 1.2 times higher, 1.2 to 1.4 times higher, 1.4 to 1.6 times higher, 1.6 to 1.8 times higher, 1.8 to 2.0 times higher, 2.0 to 2.2 times higher, 2.2 to 2.4 times higher, 2.4 to 2.6 times higher, 2.6 to 2.8 times higher, 2.8 to 3.0 times higher, 3.0 to 3.5 times higher, 3.5 to 4.0 times higher, 4.0 to 4.5 times higher, 4.5 to 5.0 times higher, 5.0 to 5.5 times higher, 5.5 to 6.0 times higher, 6.0 to 6.5 times higher, 6.5 to 7.0 times higher, 7.0 to 7.5 times higher, 7.5 to 8.0 times higher, 8.0 to 8.5 times higher, 8.5 to 9.0 times higher, 9.0 to 9.5 times higher, 9.5 to 10.0 times higher, 10 to 11 times higher, 11 to 12 times higher, 12 to 13 times higher, 13 to 14 times higher, 14 to 15 times higher, 15 to 16 times higher, 16 to 17 times higher, 17 to 18 times higher, 18 to 19 times higher, 19 to 20 times higher, 20 to 30 times higher, 30 to 40 times higher, 40 to 50 times higher, 50 to 60 times higher, 60 to 70 times higher, 70 to 80 times higher, 80 to 90 times higher, or 90 to 100 times higher. The degree of difference in concentration is the main cause of the normalization with respect to the footprint size of the target region, as discussed in the Definitions section.

[0300] In some embodiments, the DNA to be captured includes intron regions. In some embodiments, the intron regions include one or more introns that are likely to differentiate DNA from neoplastic (e.g., tumor or cancer) cells from DNA of healthy cells, such as non-neoplastic circulating cells. For example, introns containing rearrangements that are known to be present in some neoplastic cells and not in healthy cells can be used to differentiate DNA from neoplastic (e.g., tumor or cancer) cells from DNA of healthy cells. In some embodiments, the rearrangement is a translocation.

[0301] In some embodiments, the captured intron region has a footprint of at least 30 bp, such as at least 100 bp, at least 200 bp, at least 500 bp, at least 1 kb, at least 2 kb, at least 5 kb, at least 10 kb, at least 20 kb, at least 50 kb, at least 200 kb, at least 300 kb, or at least 400 kb. In some embodiments, the set of intron target regions has a footprint in the range of 30 bp to 1000 kb, such as 30 bp to 100 bp, 100 bp to 200 bp, 200 bp to 500 bp, 500 bp to 1 kb, 1 kb to 2 kb, 2 kb to 5 kb, 5 kb to 10 kb, 10 kb to 20 kb, 20 kb to 50 kb, 50 kb to 100 kb, 100 to 200 kb, 200 to 300 kb, 300 to 400 kb, 400 to 500 kb, 500 to 600 kb, 600 to 700 kb, 700 to 800 kb, 800 to 900 kb, and 900 to 1,000 kb.

[0302] Exemplary rearrangements that can be detected using the methods described herein, such as intron translocations, include, but are not limited to, translocations in which at least one of the two genes involved in the translocation is a receptor tyrosine kinase. Exemplary translocation products are the BCR-ABL fusion and fusions containing any of ALK, FGFR2, FGFR3, NTRK1, RET, or ROS1.

[0303] In some embodiments, the DNA to be captured includes target regions having type-specific epigenetic variations. In some embodiments, the set of epigenetic target regions consists of target regions having type-specific epigenetic variations. In some embodiments, type-specific epigenetic variations, such as patterns of differential methylation or type-specific fragmentation, are likely to differentiate DNA from one or more associated cell types or tissue types present in the sample or subject from DNA from other cell types or tissue types.

[0304] In some embodiments, the nucleic acids captured or enriched using the methods described herein include the captured DNA, e.g., one or more captured sets of DNA. In some embodiments, the captured DNA includes target regions that are differentially methylated in different immune cell types. In some embodiments, the immune cell types include rare or closely related immune cell types, such as activated and naive lymphocytes or myeloid cells at different stages of differentiation.

[0305] In some embodiments, the set of captured epigenetic target regions captured from a sample or a first secondary sample includes hypermethylated variable target regions. In some embodiments, the hypermethylated variable target regions are differentially or exclusively hypermethylated in one or more related cell or tissue types. In some embodiments, the hypermethylated variable target regions are differentially or exclusively hypermethylated in one cell type, or in one immune cell type, or in one immune cell type within a cluster. In some embodiments, the hypermethylated variable target regions are hypermethylated to a degree that is discriminably higher or exclusive such that they are present in one cell type, or in one immune cell type, or in one immune cell type within a cluster. Such hypermethylated variable target regions may be hypermethylated in other cell or tissue types, but not to the extent observed in one or more related cell or tissue types. In some embodiments, the hypermethylated variable target regions exhibit lower methylation in healthy cfDNA than at least one other tissue type. In some embodiments, the hypermethylated variable target regions exhibit even higher methylation in cfDNA derived from diseased cells of one or more related cell or tissue types. In some embodiments, the target regions include hypermethylated regions having an abnormally high copy number. In some such embodiments, the target regions are hypermethylated in healthy and diseased colon tissue and have an abnormally high copy number in pre-cancerous or cancerous colon tissue. Examples of such target regions are shown in Table 1 below.

Table 1-2

Table 2

[0306] In some embodiments, the set of captured epigenetic target regions captured from a sample or a secondary sample includes hypomethylated variable target regions. In some embodiments, the hypomethylated variable target regions are hypomethylated exclusively in one or more associated cell or tissue types. In some embodiments, the hypomethylated variable target regions are hypomethylated exclusively in one cell type, or in one immune cell type, or in one immune cell type within a cluster. In some embodiments, the hypomethylated variable target regions are hypomethylated to the extent of being present exclusively in one cell type, or in one immune cell type, or in one immune cell type within a cluster. Such hypomethylated variable target regions may be hypomethylated in other cell or tissue types, but cannot be hypomethylated to the extent observed in one or more cell or tissue types. In some embodiments, the hypomethylated variable target regions exhibit higher methylation in healthy cfDNA than at least one other tissue type.

[0307] Without wishing to be bound by any particular theory, in an individual having cancer, proliferating or activated immune cells and / or dying cancer cells may each shed more DNA into the bloodstream than immune cells and / or healthy cells of the same tissue type in a healthy individual. Thus, the distribution of cell types and / or tissues from which cfDNA originates can change during carcinogenesis. Thus, the presence and / or level of cfDNA derived from a particular cell or tissue type can be an indicator of disease. Variations in hypermethylation and / or hypomethylation can be indicators of disease. For example, an increase in the level of hypermethylated variable target regions and / or hypomethylated variable target regions in a secondary sample after a fractionation step can be an indicator of the presence of cancer (or recurrence, depending on the subject's medical history).

[0308] For example, as described in Scott, C.A., Duryea, J.D., MacKay, H. et al., "Identification of cell type-specific methylation signals in bulk whole genome bisulfite sequencing data," Genome Biol 21, 156 (2020) (doi.org / 10.1186 / s13059-020-02065-5), exemplary hypermethylated variable target regions and hypomethylated variable target regions useful for identifying various cell types, including but not limited to immune cell types, have been identified by analyzing DNA obtained from various cell types via whole genome bisulfite sequencing. Whole genome bisulfite sequencing data is available from the Blueprint Consortium and is available on the internet at dcc.blueprint-epigenome.eu.

[0309] In some embodiments, the first and second captured target region sets each include, for example, as described in WO2020 / 160414, DNA corresponding to a set of variable sequence target regions and DNA corresponding to a set of epigenetic target regions. The first and second captured sets can be combined to provide a combined captured set. The set of variable sequence target regions and the set of epigenetic target regions can have any of the features described for such sets in WO2020 / 160414, which is hereby incorporated by reference in its entirety. In some embodiments, the set of epigenetic target regions includes a set of hypermethylated variable target regions. In some embodiments, the set of epigenetic target regions includes a set of hypomethylated variable target regions. In some embodiments, the set of epigenetic target regions includes CTCF binding regions. In some embodiments, the set of epigenetic target regions includes fragmented variable target regions. In some embodiments, the set of epigenetic target regions includes transcription start sites. In some embodiments, the set of epigenetic target regions includes regions that can exhibit focal amplification in cancer, such as one or more of AR, BRAF, CCND1, CCND2, CCNE1, CDK4, CDK6, EGFR, ERBB2, FGFR1, FGFR2, KIT, KRAS, MET, MYC, PDGFRA, PIK3CA, and RAF1. For example, in some embodiments, the set of epigenetic target regions includes at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or 18 of the aforementioned targets.

[0310] In some embodiments, the set of variable target regions includes a plurality of regions known to undergo somatic mutations in cancer. In some aspects, the set of variable target regions targets a plurality of different genes or genomic regions (a "panel") selected such that a determined percentage of subjects with cancer exhibit genetic variants or tumor markers in one or more different genes or genomic regions in the panel. The panel can be selected to limit the regions for sequencing up to a fixed number of base pairs. The panel can be selected to sequence a desired amount of DNA, for example, by adjusting the affinity and / or amount of the probes as described elsewhere herein. The panel can be further selected to achieve a desired sequence read depth. The panel can be selected to achieve a desired sequence read depth or sequence read coverage for a certain amount of base pairs to be sequenced. The panel can be selected to achieve a theoretical sensitivity, theoretical specificity, and / or theoretical accuracy for detecting one or more genetic variants in a sample.

[0311] Probes for detecting a panel of regions can include probes for detecting the genomic regions of interest (hotspot regions). Information regarding chromatin structure can be taken into account when designing the probes, and / or the probes can be designed to maximize the likelihood that a particular site (e.g., KRAS codons 12 and 13) can be captured, and can be designed to optimize capture based on analysis of cfDNA coverage and fragment size variations affected by nucleosome binding patterns and GC sequence composition. The regions used herein can also include non-hotspot regions optimized based on nucleosome positions and GC models.

[0312] Probes for detecting panels of regions may include probes for detecting a genomic region of interest (hotspot region). Information regarding chromatin structure can be taken into account when designing the probes and / or the probes can be designed to maximize the likelihood that a particular site (e.g., KRAS codons 12 and 13) can be captured and can be designed to optimize capture based on analysis of cfDNA coverage and fragment size variations influenced by nucleosome binding patterns and GC sequence composition. Regions used herein may also include non-hotspot regions optimized based on nucleosome positions and GC models.

[0313] Examples of the listing of genomic positions of interest can be found in Tables 3 and 4 of WO2020 / 160414. In some embodiments, the set of array variable target regions used in the methods of the present disclosure comprises at least a portion of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, or 70 genes of Table 3 of WO2020 / 160414. In some embodiments, the set of array variable target regions used in the methods of the present disclosure comprises at least a portion of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or 73 genes of Table 4 of WO2020 / 160414. Additionally or alternatively, suitable target region sets are available from the literature. For example, Gale et al., PLoS One 13: e0194630 (2018), which is hereby incorporated by reference herein, describes a panel of 35 cancer-related gene targets that can be used as part or all of a set of array variable target regions. These 35 targets are AKT1, ALK, BRAF, CCND1, CDK2A, CTNNB1, EGFR, ERBB2, ESR1, FGFR1, FGFR2, FGFR3, FOXL2, GATA3, GNA11, GNAQ, GNAS, HRAS, IDH1, IDH2, KIT, KRAS, MED12, MET, MYC, NFE2L2, NRAS, PDGFRA, PIK3CA, PPP2R1A, PTEN, RET, STK11, TP53, and U2AF1.

[0314] In some embodiments, the set of array variable target regions comprises target regions from at least 10, 20, 30, or 35 cancer-related genes, such as the cancer-related genes listed above and in WO2020 / 160414.

[0315] In some embodiments, for example, a collection of capture probes, including capture probes prepared by any of the methods disclosed herein for doing so, is used in the methods described herein. In some embodiments, the collection of capture probes further comprises target binding probes specific for a set of array-variable target regions and / or target binding probes specific for a set of epigenetic target regions. In some embodiments, the capture yield of target binding probes specific for a set of array-variable target regions is higher (e.g., at least 2-fold higher) than the capture yield of target binding probes specific for a set of epigenetic target regions. In some embodiments, the collection of capture probes is configured to have a capture yield specific for a set of array-variable target regions that is higher (e.g., at least 2-fold higher) than its capture yield specific for a set of epigenetic target regions.

[0316] In some embodiments, the capture yield of target binding probes specific for a set of array-variable target regions is at least 1.25-fold, 1.5-fold, 1.75-fold, 2-fold, 2.25-fold, 2.5-fold, 2.75-fold, 3-fold, 3.5-fold, 4-fold, 4.5-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 11-fold, 12-fold, 13-fold, 14-fold or 15-fold higher than the capture yield of target binding probes specific for a set of epigenetic target regions. In some embodiments, the capture yield of target binding probes specific for a set of array-variable target regions is 1.25-fold to 1.5-fold, 1.5-fold to 1.75-fold, 1.75-fold to 2-fold, 2-fold to 2.25-fold, 2.25-fold to 2.5-fold, 2.5-fold to 2.75-fold, 2.75-fold to 3-fold, 3-fold to 3.5-fold, 3.5-fold to 4-fold, 4-fold to 4.5-fold, 4.5-fold to 5-fold, 5-fold to 5.5-fold, 5.5-fold to 6-fold, 6-fold to 7-fold, 7-fold to 8-fold, 8-fold to 9-fold, 9-fold to 10-fold, 10-fold to 11-fold, 11-fold to 12-fold, 13-fold to 14-fold or 14-fold to 15-fold higher than the capture yield of target binding probes specific for a set of epigenetic target regions.

[0317] In some embodiments, the collection of capture probes is configured to have a capture yield specific to the set of array-variable target regions that is at least 1.25-fold, 1.5-fold, 1.75-fold, 2-fold, 2.25-fold, 2.5-fold, 2.75-fold, 3-fold, 3.5-fold, 4-fold, 4.5-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 11-fold, 12-fold, 13-fold, 14-fold, or 15-fold higher than its capture yield for the set of epigenetic target regions. In some embodiments, the collection of capture probes is 1.25-fold to 1.5-fold, 1.5-fold to 1.75-fold, 1.75-fold to 2-fold, 2-fold to 2.25-fold, 2.25-fold to 2.5-fold, 2.5-fold to 2.75-fold, 2.75-fold to 3-fold, 3-fold to 3.5-fold, 3.5-fold to 4-fold, 4-fold to 4.5-fold, 4.5-fold to 5-fold, 5-fold to 5.5-fold, 5.5-fold to 6-fold, 6-fold to 7-fold, 7-fold to 8-fold, 8-fold to 9-fold, 9-fold to 10-fold, 10-fold to 11-fold, 11-fold to 12-fold, 13-fold to 14-fold, or 14-fold to 15-fold higher than its capture yield specific to the set of epigenetic target regions and is configured to have a capture yield specific to the set of array-variable target regions.

[0318] The collection of probes can be configured to provide a higher capture yield for the set of array-variable target regions in a variety of ways, including concentration, different lengths, and / or chemistry (e.g., affecting affinity), and combinations thereof. Affinity can be modulated by adjusting probe length and / or including nucleotide modifications, as discussed below.

[0319] In some embodiments, capture probes specific for a set of array-variable target regions are present at a higher concentration than capture probes specific for a set of epigenetic target regions. In some embodiments, the concentration of target-binding probes specific for a set of array-variable target regions is at least 1.25-fold, 1.5-fold, 1.75-fold, 2-fold, 2.25-fold, 2.5-fold, 2.75-fold, 3-fold, 3.5-fold, 4-fold, 4.5-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 11-fold, 12-fold, 13-fold, 14-fold, or 15-fold higher than the concentration of target-binding probes specific for a set of epigenetic target regions. In some embodiments, the concentration of target-binding probes specific for a set of array-variable target regions is 1.25-fold to 1.5-fold, 1.5-fold to 1.75-fold, 1.75-fold to 2-fold, 2-fold to 2.25-fold, 2.25-fold to 2.5-fold, 2.5-fold to 2.75-fold, 2.75-fold to 3-fold, 3-fold to 3.5-fold, 3.5-fold to 4-fold, 4-fold to 4.5-fold, 4.5-fold to 5-fold, 5-fold to 5.5-fold, 5.5-fold to 6-fold, 6-fold to 7-fold, 7-fold to 8-fold, 8-fold to 9-fold, 9-fold to 10-fold, 10-fold to 11-fold, 11-fold to 12-fold, 13-fold to 14-fold, or 14-fold to 15-fold higher than the concentration of target-binding probes specific for a set of epigenetic target regions. In such embodiments, the concentration can refer to the average mass concentration per volume of the individual probes in each set.

[0320] In some embodiments, the capture probes specific for the set of variable target regions have a higher affinity for those targets than the capture probes specific for the set of epigenetic target regions. The affinity can be modulated by any method known to those of skill in the art, including by using different probe chemistries. For example, certain nucleotide modifications, such as cytosine 5-methylation (in certain sequence contexts), modifications that provide a heteroatom at the 2'-sugar position, and LNA nucleotides can increase the stability of double-stranded nucleic acids, indicating that oligonucleotides with such modifications have a relatively high affinity for their complementary sequences. See, for example, Severin et al., Nucleic Acids Res. 39: 8740-8751 (2011); Freier et al., Nucleic Acids Res. 25: 4429-4443 (1997); U.S. Patent No. 9,738,894. Also, longer sequence lengths generally provide increased affinity. Other nucleotide modifications, such as substitution of guanine with the nucleobase hypoxanthine, reduce affinity by reducing the amount of hydrogen bonding between the oligonucleotide and its complementary sequence. In some embodiments, the capture probes specific for the set of variable target regions have modifications that increase their affinity for those targets. In some embodiments, alternatively or additionally, the capture probes specific for the set of epigenetic target regions have modifications that decrease their affinity for those targets. In some embodiments, the capture probes specific for the set of variable target regions have a longer average length and / or a higher average melting temperature than the capture probes specific for the set of epigenetic target regions. These embodiments can be combined with each other and / or with the differences in concentrations discussed above to achieve the desired differences in capture yields, such as any of the differences in yields or ranges thereof described above.

[0321] In some embodiments, the capture probe includes a capture moiety. The capture moiety can be any of the capture moieties described herein, for example, biotin. In some embodiments, the capture probe is linked to a solid support, for example, covalently or non-covalently, via interaction of a binding pair of the capture moiety. In some embodiments, the solid support is a bead, for example, a magnetic bead.

[0322] In some embodiments, a capture probe specific for a set of variable target regions and / or a capture probe specific for a set of epigenetic target regions is a probe that includes a capture moiety and a sequence selected to tile across a panel of regions, for example, genes, such as the capture probe set discussed above.

[0323] In some embodiments, the capture probe is provided in a single composition. The single composition can be a solution (liquid or frozen). Alternatively, the single composition can be a lyophilized product.

[0324] Alternatively, the capture probe can be provided as a plurality of compositions, for example, a first composition that includes probes specific for a set of epigenetic target regions and a second composition that includes probes specific for a set of variable target regions. These probes can be mixed in a ratio appropriate to provide either a difference in magnification in concentration and / or capture yield in the combined probe composition. Alternatively, they can be used in separate capture procedures (e.g., using aliquots of the sample or sequentially using the same sample) to provide first and second compositions each including captured epigenetic target regions and variable target regions.

[0325] Probes specific for epigenetic target regions Probes for an epigenetic target region set may include probes specific for one or more types of target regions that are likely to differentiate DNA from neoplastic (e.g., tumor or cancer) cells from healthy cells, such as non-neoplastic circulating cells. Exemplary types of such regions are discussed in detail herein, for example, in the above section regarding the captured set. Probes for an epigenetic target region set may also include probes for one or more control regions, such as those described herein.

[0326] In some embodiments, probes for an epigenetic target region set have a footprint of at least 100 kbp, such as at least 200 kbp, at least 300 kbp, or at least 400 kbp. In some embodiments, the epigenetic target region set has a footprint in the range of 100 to 20 Mbp, such as 100 to 200 kbp, 200 to 300 kbp, 300 to 400 kbp, 400 to 500 kbp, 500 to 600 kbp, 600 to 700 kbp, 700 to 800 kbp, 800 to 900 kbp, 900 to 1,000 kbp, 1 to 1.5 Mbp, 1.5 to 2 Mbp, 2 to 3 Mbp, 3 to 4 Mbp, 4 to 5 Mbp, 5 to 6 Mbp, 6 to 7 Mbp, 7 to 8 Mbp, 8 to 9 Mbp, 9 to 10 Mbp, or 10 to 20 Mbp. In some embodiments, the epigenetic target region set has a footprint of at least 20 Mbp.

[0327] Hypermethylation variable target region In some embodiments, the probes for the set of epigenetic target regions include probes specific to one or more hypermethylated variable target regions. Hypermethylated variable target regions may also be referred to herein as DMRs (differentially methylated regions) that are hypermethylated. The hypermethylated variable target regions can be any of those shown above. For example, in some embodiments, the probes specific to the hypermethylated variable target regions include a plurality of loci listed in Table 1, for example, at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or 100% of the loci listed in Table 1. In some embodiments, the probes specific to the hypermethylated variable target regions include a plurality of loci listed in Table 2, for example, at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or 100% of the loci listed in Table 2. In some embodiments, the probes specific to the hypermethylated variable target regions include a plurality of loci listed in Table 1 or Table 2, for example, at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or 100% of the loci listed in Table 1 or Table 2. In some embodiments, for each locus included as a target region, there may be one or more probes having a hybridization site that binds between the transcription start site of the gene and the stop codon (for alternatively spliced genes, the last stop codon). In some embodiments, the one or more probes bind within 300 bp of the listed position, for example, within 200 or 100 bp. In some embodiments, the probe has a hybridization site that overlaps the position listed above. In some embodiments, the probes specific to the hypermethylated target regions include one, two, three, four or five probes specific to one, two, three, four or five subsets of the hypermethylated target regions that collectively show hypermethylation in one, two, three, four or five of breast, colon, kidney, liver and lung cancers.

[0328] Hypomethylated variable target region In some embodiments, probes for an epigenetic target region set include probes specific to one or more hypomethylated variable target regions. Hypomethylated variable target regions may also be referred to herein as hypomethylated DMRs (differentially methylated regions). Hypomethylated variable target regions can be any of those shown above. For example, probes specific to one or more hypomethylated variable target regions can include regions that may show reduced methylation in tumor cells, such as repetitive elements, such as LINE1 elements, Alu elements, centromeric tandem repeats, pericentromeric tandem repeats, and satellite DNA, and probes for intergenic regions that are normally methylated in healthy cells.

[0329] In some embodiments, probes specific to hypomethylated variable target regions include probes specific to repetitive elements and / or intergenic regions. In some embodiments, probes specific to repetitive elements include probes specific to 1, 2, 3, 4, or 5 of LINE1 elements, Alu elements, centromeric tandem repeats, pericentromeric tandem repeats, and / or satellite DNA.

[0330] Exemplary probes specific to genomic regions showing cancer-related hypomethylation include probes specific to nucleotides 8403565 - 8953708 and / or 151104701 - 151106035 of human chromosome 1. In some embodiments, probes specific to hypomethylated variable target regions include probes specific to regions overlapping with or including nucleotides 8403565 - 8953708 and / or 151104701 - 151106035 of human chromosome 1.

[0331] CTCF binding region In some embodiments, the probes for the set of epigenetic target regions include probes specific to CTCF binding regions. In some embodiments, the probes specific to CTCF binding regions are specific to at least 10, 20, 50, 100, 200, or 500 CTCF binding regions, or 10 - 20, 20 - 50, 50 - 100, 100 - 200, 200 - 500, or 500 - 1000 CTCF binding regions, such as, for example, the CTCF binding regions described in one or more of the above, or the CTCFBSDB or the papers of Cuddapah et al., Martin et al., or Rhee et al. cited above. The probes include. In some embodiments, the probes for the set of epigenetic target regions include regions at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 750 bp, or at least 1000 bp upstream and downstream of the CTCF binding sites.

[0332] Transcription start site In some embodiments, the probes for the set of epigenetic target regions include probes specific to transcription start sites. In some embodiments, the probes specific to transcription start sites are specific to at least 10, 20, 50, 100, 200, or 500 transcription start sites, or 10 - 20, 20 - 50, 50 - 100, 100 - 200, 200 - 500, or 500 - 1000 transcription start sites, such as, for example, the transcription start sites listed in DBTSS. In some embodiments, the probes for the set of epigenetic target regions include probes for sequences at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 750 bp, or at least 1000 bp upstream and downstream of the transcription start sites.

[0333] Limited amplification As noted above, focal amplifications are somatic mutations, but these can be detected by sequencing based on read frequency in a manner similar to approaches for detecting certain epigenetic changes, such as changes in methylation. Thus, regions that may exhibit focal amplifications in cancer can be included in a set of epigenetic target regions as discussed above. In some embodiments, the probes specific to the set of epigenetic target regions include probes specific to focal amplifications. In some embodiments, the probes specific to focal amplifications include probes specific to one or more of AR, BRAF, CCND1, CCND2, CCNE1, CDK4, CDK6, EGFR, ERBB2, FGFR1, FGFR2, KIT, KRAS, MET, MYC, PDGFRA, PIK3CA, and RAF1. For example, in some embodiments, the probes specific to focal amplifications include probes specific to one or more of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or 18 of the above targets.

[0334] Control region To facilitate data validation, it may be useful to include control regions. In some embodiments, the probes specific to the set of epigenetic target regions include probes specific to methylated control regions that are expected to be methylated in essentially all samples. In some embodiments, the probes specific to the set of epigenetic target regions include probes specific to hypomethylated control regions that are expected to be hypomethylated in essentially all samples.

[0335] Probes specific to sequence variable target regions Probes for the array-variable target region set may include probes specific to a plurality of regions known to undergo somatic mutations in cancer. The probes can be specific to any array-variable target region set described herein. Exemplary array-variable target region sets are discussed in detail herein, for example, in the above section regarding the captured set.

[0336] In some embodiments, the array-variable target region probe set has a footprint of at least 0.5 kb, such as at least 1 kb, at least 2 kb, at least 5 kb, at least 10 kb, at least 20 kb, at least 30 kb, or at least 40 kb. In some embodiments, the epigenetic target region probe set has a footprint in the range of 0.5 - 100 kb, such as 0.5 - 2 kb, 2 - 10 kb, 10 - 20 kb, 20 - 30 kb, 30 - 40 kb, 40 - 50 kb, 50 - 60 kb, 60 - 70 kb, 70 - 80 kb, 80 - 90 kb, and 90 - 100 kb. In some embodiments, the array-variable target region probe set has a footprint of at least 50 kbp, such as at least 100 kbp, at least 200 kbp, at least 300 kbp, or at least 400 kbp. In some embodiments, the array-variable target region probe set has a footprint in the range of 100 - 2000 kbp, such as 100 - 200 kbp, 200 - 300 kbp, 300 - 400 kbp, 400 - 500 kbp, 500 - 600 kbp, 600 - 700 kbp, 700 - 800 kbp, 800 - 900 kbp, 900 - 1,000 kbp, 1 - 1.5 Mbp, or 1.5 - 2 Mbp. In some embodiments, the array-variable target region set has a footprint of at least 2 Mbp.

[0337] In some embodiments, the probes specific to the set of array-variable target regions include probes specific to at least a portion of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, or 70 of the genes in Table 3. In some embodiments, the probes specific to the set of array-variable target regions include probes specific to at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, or 70 of the SNVs in Table 3. In some embodiments, the probes specific to the set of array-variable target regions include probes specific to at least 1, at least 2, at least 3, at least 4, at least 5, or 6 of the fusions in Table 3. In some embodiments, the probes specific to the set of array-variable target regions include probes specific to at least a portion of at least 1, at least 2, or 3 of the indels in Table 3. In some embodiments, the probes specific to the set of array-variable target regions include probes specific to at least a portion of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or 73 of the genes in Table 4. In some embodiments, the probes specific to the set of array-variable target regions include probes specific to at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or 73 of the SNVs in Table 4. In some embodiments, the probes specific to the set of array-variable target regions include probes specific to at least 1, at least 2, at least 3, at least 4, at least 5, or 6 of the fusions in Table 4.In some embodiments, probes specific to the set of array-variable target regions include probes specific to at least a portion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, or 18 of the indels in Table 4. In some embodiments, probes specific to the set of array-variable target regions include probes specific to at least a portion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, or at least 20 of the genes in Table 5.

[0338]

Table 3

[0339]

Table 4

[0340]

Table 5-1

Table 5-2

Table 5-3

[0341] In some embodiments, probes specific to the set of array-variable target regions include probes specific to target regions from at least 10, 20, 30, or 35 cancer-related genes, such as AKT1, ALK, BRAF, CCND1, CDK2A, CTNNB1, EGFR, ERBB2, ESR1, FGFR1, FGFR2, FGFR3, FOXL2, GATA3, GNA11, GNAQ, GNAS, HRAS, IDH1, IDH2, KIT, KRAS, MED12, MET, MYC, NFE2L2, NRAS, PDGFRA, PIK3CA, PPP2R1A, PTEN, RET, STK11, TP53, and U2AF1.

[0342] Sequencing Generally, sample nucleic acids adjacent to the adapter can be subjected to sequencing, with or without prior amplification. Sequencing methods include, for example, Sanger sequencing, high-throughput sequencing, pyrosequencing, sequencing by synthesis, single molecule sequencing, nanopore sequencing, semiconductor sequencing, sequencing by ligation, sequencing by hybridization, digital gene expression (Helicos), next-generation sequencing (NGS), single molecule sequencing by synthesis (SMSS) (Helicos), ultra-parallel sequencing, clonal single molecule array (Solexa), shotgun sequencing, Ion Torrent, Oxford Nanopore, Roche Genia, Maxim-Gilbert sequencing, primer walking, and sequencing using PacBio, SOLiD, Ion Torrent, or Nanopore platforms. The sequencing reaction can be performed in various sample processing units, including those that can substantially simultaneously process multiple lanes, multiple channels, multiple wells, or other means for processing multiple sample sets. Sample processing units can also include multiple sample chambers capable of simultaneously processing multiple runs.

[0343] In some embodiments, the genomic sequence coverage can be, for example, less than 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9% or 100%. In some embodiments, the array reaction can provide an array coverage of, for example, at least 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70% or 80% of the genome. The array coverage can be performed on, for example, at least 5, 10, 20, 70, 100, 200 or 500 different genes, or at most, for example, 5000, 2500, 1000, 500 or 100 different genes.

[0344] The simultaneous sequencing reaction can be performed using multiplex sequencing. In some cases, the cell-free nucleic acid can be sequenced in at least, for example, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. In other cases, the cell-free nucleic acid can be sequenced in, for example, less than 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. The sequencing reactions can be performed sequentially or simultaneously. Subsequent data analysis can be performed on all or part of the sequencing reactions. In some cases, the data analysis can be performed on at least, for example, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. In other cases, the data analysis can be performed on, for example, less than 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. An exemplary read depth is 1000 - 50000 or 1000 - 10000 or 1000 - 20000 reads per locus (base).

[0345] Generally, for example, sequencing of epigenetic target regions for analyzing modified nucleoside profiles of DNA requires, for example, a shallower depth of sequencing than sequencing of variable target regions for mutation analysis. Thus, as described herein, a shallower sequencing depth may, in some cases, be appropriate for the methods described herein.

[0346] a. Differential depth of sequencing In some embodiments, the nucleic acids corresponding to the set of array-variable target regions are sequenced to a deeper depth of sequencing than the nucleic acids corresponding to the set of epigenetic target regions. In some embodiments, the nucleic acids corresponding to the set of hydroxymethylation-variable target regions are sequenced to a deeper depth of sequencing than the nucleic acids corresponding to at least one other set of target regions. For example, the depth of sequencing for the nucleic acids corresponding to the set of array-variable and / or hydroxymethylation-variable target regions may be at least 1.25-fold, 1.5-fold, 1.75-fold, 2-fold, 2.25-fold, 2.5-fold, 2.75-fold, 3-fold, 3.5-fold, 4-fold, 4.5-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 11-fold, 12-fold, 13-fold, 14-fold or 15-fold deeper than the depth of sequencing for the nucleic acids corresponding to the epigenetic target region set or the nucleic acids corresponding to at least one other set of target regions, or 1.25-fold to 1.5-fold, 1.5-fold to 1.75-fold, 1.75-fold to 2-fold, 2-fold to 2.25-fold, 2.25-fold to 2.5-fold, 2.5-fold to 2.75-fold, 2.75-fold to 3-fold, 3-fold to 3.5-fold, 3.5-fold to 4-fold, 4-fold to 4.5-fold, 4.5-fold to 5-fold, 5-fold to 5.5-fold, 5.5-fold to 6-fold, 6-fold to 7-fold, 7-fold to 8-fold, 8-fold to 9-fold, 9-fold to 10-fold, 10-fold to 11-fold, 11-fold to 12-fold, 13-fold to 14-fold, 14-fold to 15-fold or 15-fold to 100-fold deeper. In some embodiments, the depth of sequencing is at least 2-fold deeper. In some embodiments, the depth of sequencing is at least 5-fold deeper. In some embodiments, the depth of sequencing is at least 10-fold deeper. In some embodiments, the depth of sequencing is 4-fold to 10-fold deeper. In some embodiments, the depth of sequencing is 4-fold to 100-fold deeper. Each of these embodiments refers to the extent to which the nucleic acids corresponding to the set of array-variable target regions are sequenced to a deeper depth of sequencing than the nucleic acids corresponding to the set of epigenetic target regions.

[0347] In some embodiments, the captured cfDNA corresponding to the set of variable target regions and the captured cfDNA corresponding to the set of epigenetic target regions are sequenced in parallel, e.g., in the same sequencing cell (e.g., the flow cell of an Illumina sequencer), and / or in the same composition resulting from recombining separately captured sets, or in the same composition that can be the composition obtained by capturing the cfDNA corresponding to the set of variable target regions and the captured cfDNA corresponding to the set of epigenetic target regions in the same container.

[0348] In some embodiments, the captured cfDNA corresponding to the set of hydroxymethylation variable target regions and the captured cfDNA corresponding to at least one other set of target regions are sequenced in parallel, e.g., in the same sequencing cell (e.g., the flow cell of an Illumina sequencer), and / or in the same composition resulting from rebinding separately captured sets or in the same composition that can be the composition obtained by capturing the cfDNA corresponding to the set of hydroxymethylation variable target regions and the captured cfDNA corresponding to at least one other set of target regions in the same container.

[0349] Analysis Defining the region of end-repaired DNA synthesized during end repair Generally, the methods described herein rely on the use of at least one dNTP containing a modified base in a terminal repair reaction, in combination with the use of modified-sensitivity sequencing that can detect the modified base. This enables the regions synthesized in the terminal repair to be identified in the sequencing data. This is important because these synthesized regions can lead to artifact data in typical sequencing reactions that do not control these synthesized regions. For example, if the terminal repair is performed using unmodified dNTPs prior to methylation-sensitive sequencing (e.g., bisulfite sequencing), the terminal repair can lead to 5' overhang filling, nick translation, and gap filling with dCTP containing unmodified cytosine. These unmodified cytosines may not reflect the original methylation state at these positions in the original DNA molecule (i.e., before the formation of overhangs, nicks, and gaps), and thus, the terminal repair can lead to artifact methylation information. The methods disclosed herein avoid such artifact information by identifying the sequencing data corresponding to these synthesized regions. Such regions can be filtered out and removed, for example, so as not to be used for classifying the methylation state of a DNA molecule.

[0350] The regions synthesized in the terminal repair can be classified in various ways, and the exact approach depends on the identity of the modified base used in the terminal repair reaction, as well as the modified-sensitivity sequencing method used. Further, the exact endpoints of the regions classified as being synthesized during the terminal repair can be determined by the user. The basic step in the identification of the regions synthesized during the terminal repair reaction is the identification of the presence of a base modification in at least one type of dNTP used in the terminal repair reaction.

[0351] In some embodiments, end repair is performed using dNTPs containing 5mC or 5hmC. Both 5mC and 5hmC are naturally occurring base modifications, and thus the identification of these modified bases in sequencing data can be attributed to either (i) modified bases present in the original DNA molecule or (ii) modified bases introduced during the end repair reaction. However, 5mC and 5hmC can be classified as being introduced during the end repair reaction when they occur in non-CpG sequence contexts. CpH (i.e., CpA, CpT, CpC) methylation has been described in humans and is thought to constitute 0.02% of total methyl-cytosine in differentiated somatic cells (Jang et al. Genes (Basel). 2017 Jun; 8(6): 148). Thus, methylated cytosines in the context of CpH sequences can reliably result from regions synthesized during end repair. This is particularly true when the disclosed methods involve enrichment of sequence panels that do not include regions known to contain methylated CpH sites. Alternatively, classification of whether methylated CpH is part of the synthesized region can be performed by considering (i) the position of specific CpH sites in the reference sequence and / or (ii) the methylation status of surrounding CpH sites. For example, if a CpH site is known to be naturally methylated (e.g., by comparison with reference data), the methylation detected at that CpH site can be ignored when defining the region synthesized during end repair. Similarly, the methylation detected at such a CpH site can be referred to as the true methylation state in the DNA sample. If a CpH site known to be naturally methylated is detected as being methylated in the sequencing data but is included within a stretch of other methylated CpH sites (some of which are not known to be naturally methylated (e.g., by comparison with reference data)), that region can still be classified as being synthesized during end repair.

[0352] In embodiments where end repair is performed using dNTPs that include 5mC and / or 5hmC, one region of one or more regions of the end-repaired DNA synthesized during end repair is (i) a sequence between two unmethylated cytosines that spans a methylated non-CpG cytosine, and / or (ii) a sequence between an unmethylated cytosine and the end of the sequence read, defined as a sequence in which there is no additional unmethylated cytosine between the unmethylated cytosine and the end of the sequence read. Alternatively, one or more regions of the end-repaired DNA synthesized during end repair can be defined as (i) the sequence from the first methylated non-CpG cytosine to the last methylated non-CpG cytosine in one or more consecutive methylated non-CpG cytosines, and / or (ii) the sequence from a methylated cytosine (5mC or 5hmC) in a non-CpG context to the end of the sequence read, defined as a sequence in which there is no unmethylated cytosine between the methylated cytosine in the non-CpG context and the end of the sequence read. "The end of the sequence read" corresponds to the end-repaired DNA molecule and refers to, for example, the portion of the sequence read that does not include the adapter sequence.

[0353] In some embodiments, end repair is performed using dNTPs that include base modifications that are not found naturally in the subject from which the DNA sample is derived, or that occur only at very low frequencies. For example, 4mC does not occur in mammals (e.g., humans), and 6mA occurs only at very low frequencies (Xiao et al. Molecular Cell Volume 71, Issue 2, 19 July 2018, Pages 306-318.e7). In these embodiments, the region of end-repaired DNA synthesized during end repair can simply be classified as any region where the modified base is detected. Such an approach can result in misclassifying naturally occurring low-frequency base modifications as being the result of end repair, which is rare and can simply result in corresponding sequence data that is not used for further analysis. This is preferable to using sequence data from regions synthesized during end repair, which can include artifact data that can lead to incorrect inferences about the corresponding DNA sample and subject.

[0354] Thus, in some embodiments where the modified base is other than 5mC or 5hmC, one region of the one or more regions can be defined as: (i) a sequence between two unmodified bases spanning the modified base, where the bases have the same identity as the bases present in at least one type of dNTP containing the modified base, and / or (ii) a sequence between an unmodified base and the end of the sequence read, where there are no additional unmodified bases between the unmodified base and the end of the sequence read, and the unmodified base has the same identity as the modified base present in at least one type of dNTP containing the modified base. Alternatively, one or more regions of the end-repaired DNA synthesized during end repair can be defined as: (i) a sequence from the first modified base to the last modified base in one or more consecutive modified bases, where the bases are the same entity as the bases present in at least one type of dNTP containing the modified base, and / or (ii) a sequence from the modified base to the end of the sequence read, where there are no unmodified bases between the modified base and the end of the sequence read, and the modified and unmodified bases are the same entity as at least one type of dNTP containing the modified base.

[0355] Once identified, regions of end-repaired DNA classified as being synthesized during end repair can be filtered out and removed from the sequence data so that they are not used for further analysis, such as for variant calling, or to determine the modification status of bases in the original DNA molecule (i.e., prior to end repair). Thus, in some embodiments, the methods disclosed herein further include analyzing at least a portion of the sequence data corresponding to regions not identified as being synthesized during end repair to detect the presence or absence of base modifications or mutations present in the DNA sample. The disclosed methods for identifying regions synthesized during end repair are advantageous over prior art methods that use "end clipping" without information, as these prior art methods do not synthesize during the end repair reaction and thus potentially remove regions representing the original DNA molecule.

[0356] Exemplary uses The methods disclosed herein enable the detection of regions of end-repaired DNA synthesized during end repair. This information is useful in a wide range of situations, including determination of the methylation state of DNA in a DNA sample (i.e., prior to end repair) and detection of mutations in the DNA.

[0357] The methods presented herein can be used as part of any method that benefits from obtaining an accurate modified nucleoside profile of DNA in any DNA sample and / or an accurate variant calling of DNA in any DNA sample. This is because the methods disclosed herein correspond to regions of end-repaired DNA molecules synthesized during end repair and thus enable the identification of sequencing data that may not represent the original DNA molecule. Identification of these regions avoids reliance on such potentially artifact data for subsequent analysis, such as variant calling and / or subsequent methylation analysis.

[0358] For example, the classification of whether a variant is present or absent in a DNA molecule depends on whether a double-stranded support for the variant is present. The double-stranded support refers to the presence of sequencing data derived from both DNA strands that support the presence of the variant. However, in synthesized regions, a double-stranded support can be artificially introduced by end repair and / or A-tailing reactions that use a complementary strand as a template to synthesize the region. This can occur when there is a mismatch at the equivalent position in the original DNA molecule and thus no variant was present in both strands of the original DNA molecule prior to end repair and / or A-tailing. Accordingly, the methods disclosed herein can be used to identify synthesized regions and filter out and remove variants within regions that would otherwise be misclassified as having a double-stranded support. The use of the disclosed methods in variant calling is particularly advantageous when the modified sensitivity sequencing method used does not require the conversion of unmethylated cytosine and is thus compatible with high-sensitivity variant calling. Accordingly, in some embodiments, the methods disclosed herein are used to detect SNVs and the modified sensitivity sequencing is nanopore-based sequencing, single molecule real-time (SMRT) sequencing, or Tet-assisted pyridine borane sequencing (TAPS).

[0359] One important exemplary application of the methods of the present disclosure is the use of the obtained sequencing data in the diagnosis and prognosis of cancer or other genetic diseases or conditions.

[0360] Accordingly, in some embodiments, the methods described herein include the step of identifying or predicting the presence or absence of DNA produced by a tumor (or neoplastic cells, or cancer cells), the step of determining the probability that a subject under test has a tumor or cancer, and / or the step of characterizing a tumor, neoplastic cells, or cancer as described herein.

[0361] The methods of the present invention can be used to diagnose a condition in a subject, particularly the presence or absence of cancer, to characterize a condition (e.g., to stage cancer or to determine cancer heterogeneity), to monitor the response of a condition to treatment, or to provide a prognostic risk of developing a condition or a subsequent course of a condition. The present disclosure can also be useful in determining the effectiveness of a particular treatment option. If a treatment is successful, more cancer may die and shed DNA, so a successful treatment option can increase the amount of rare mutations detected in the subject's blood. In other instances, this may not occur. In another instance, perhaps a particular treatment option can be correlated with the genetic profile of cancer over time. This correlation can be useful in selecting a treatment. In some embodiments, target regions are analyzed to determine their methylation characteristics in tumor cells or cells that do not significantly contribute to cfDNA normally, and / or to determine whether target regions exhibit methylation characteristics in tumor cells or cells that do not significantly contribute to cfDNA normally.

[0362] In some embodiments, the methods of the present invention are used in a method for screening for cancer, e.g., metastasis, or for screening for cancer, e.g., a method for detecting the presence or absence of metastasis. For example, the sample can be a sample from a subject previously diagnosed or not diagnosed with cancer. In some embodiments, one or more, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more samples are collected from a subject as described herein, e.g., before and / or after the subject is diagnosed with cancer. In some embodiments, the subject may or may not have cancer. In some embodiments, the subject may or may not have early stage cancer. In some embodiments, the subject has one or more risk factors for cancer, e.g., tobacco use (e.g., smoking), being overweight or obese, having a high body mass index (BMI), being elderly, malnutrition, high alcohol intake, or having a family history of cancer.

[0363] In some embodiments, the subject has been using tobacco, for example, for at least 1, 5, 10, or 15 years. In some embodiments, the subject has a high BMI, such as a BMI of 25 or higher, 26 or higher, 27 or higher, 28 or higher, 29 or higher, or 30 or higher. In some embodiments, the subject is at least 40, 45, 50, 55, 60, 65, 70, 75, or 80 years old. In some embodiments, the subject has malnutrition, such as a high intake of one or more of red and / or processed meat, trans fat, saturated fat, and refined sugar, and / or a low intake of fruits and vegetables, complex carbohydrates, and / or unsaturated fat. High intake and low intake can be defined as exceeding or falling below the recommendations in the Dietary Guidelines for Americans 2020 - 2025, available, for example, at dietaryguidelines.gov / sites / default / files / 2021-03 / Dietary_Guidelines_for_Americans-2020-2025.pdf. In some embodiments, the subject has a high alcohol intake, such as having an average of at least 3, 4, or 5 drinks per day (where a drink is about 1 ounce or 30 mL of 80-proof hard liquor or equivalent). In some embodiments, the subject has a family history of cancer, for example, at least 1, 2, or 3 blood relatives have been previously diagnosed with cancer. In some embodiments, the blood relatives are at least third-degree blood relatives (e.g., great-grandparents, great-aunts or uncles, cousins), at least second-degree blood relatives (e.g., grandparents, aunts or uncles, or siblings with different parents), or first-degree blood relatives (e.g., parents or siblings with the same parents).

[0364] Furthermore, if it is observed that the cancer is in a remission state after treatment, the methods of the invention can be used to monitor for residual disease or disease recurrence.

[0365] Typically, the disease assumed is a type of cancer such as any of those mentioned herein. The types and numbers of cancers that can be detected can include blood cancers, brain cancers, lung cancers, skin cancers, nasal cancers, throat cancers, liver cancers, bone cancers, lymphomas, pancreatic cancers, skin cancers, intestinal cancers, rectal cancers, thyroid cancers, bladder cancers, kidney cancers, oral cancers, stomach cancers, solid tumors, heterogeneous tumors, homogeneous tumors, and the like. Examples of such cancers include biliary tract cancer, bladder cancer, transitional cell carcinoma, urothelial carcinoma, brain cancer, glioma, astrocytoma, breast cancer, metaplastic carcinoma, cervical cancer, cervical squamous cell carcinoma, rectal cancer, colorectal cancer, colon cancer, hereditary non-polyposis colorectal cancer, colorectal adenocarcinoma, gastrointestinal stromal tumor (GIST), endometrial cancer, endometrial stromal sarcoma, esophageal cancer, esophageal squamous cell carcinoma, esophageal adenocarcinoma, ocular melanoma, uveal melanoma, gallbladder cancer, gallbladder adenocarcinoma, renal cell carcinoma, clear cell renal cell carcinoma, transitional cell carcinoma, urothelial carcinoma, Wilms tumor, leukemia, acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), chronic lymphocytic leukemia (CLL), chronic myeloid leukemia (CML), chronic myelomonocytic leukemia (CMML), liver cancer, hepatocarcinoma, hepatoma, hepatocellular carcinoma, cholangiocarcinoma, hepatoblastoma, lung cancer, non-small cell lung cancer (NSCLC), mesothelioma, B-cell lymphoma, non-Hodgkin lymphoma, diffuse large B-cell lymphoma, mantle cell lymphoma, T-cell lymphoma, non-Hodgkin lymphoma, precursor T-lymphoblastic lymphoma / leukemia, peripheral T-cell lymphoma, multiple myeloma, nasopharyngeal carcinoma (NPC), neuroblastoma, oropharyngeal cancer, oral squamous cell carcinoma, osteosarcoma, ovarian cancer, pancreatic cancer, pancreatic ductal adenocarcinoma, pseudopapillary neoplasm, acinar cell carcinoma, prostate cancer, prostatic adenocarcinoma, skin cancer, melanoma, malignant melanoma, cutaneous melanoma, small intestine cancer, stomach cancer, gastric cancer, gastrointestinal stromal tumor (GIST), uterine cancer or uterine sarcoma.

[0366] Cancer type and / or stage can be detected from genetic variations such as mutations, rare mutations, indels, copy number variations, transversions, translocations, inversions, deletions, aneuploidy, segmental aneuploidy, ploidy, chromosomal instability, chromosomal structural alterations, gene fusions, chromosomal fusions, gene truncations, gene amplifications, gene duplications, chromosomal damage, DNA damage, abnormal changes in chemical modifications of nucleic acids, and abnormal changes in epigenetic patterns, e.g., 5mC and 5mC profiles. Thus, the method can, in some cases, be used in combination with methods used to detect other genetic / epigenetic variations, e.g., in methods for detecting or characterizing cancer or other methods described herein.

[0367] In some embodiments, the methods described herein include identifying the presence of a target region and / or DNA produced by a tumor (or neoplastic cell, or cancer cell) or a pre-cancerous cell. In some embodiments, the methods described herein include determining the level of the target region and / or identifying the presence of DNA produced by a tumor (or neoplastic cell, or cancer cell) or a pre-cancerous cell. In some embodiments, determining the level of the target region includes determining either an increase or a decrease in the level of the target region, and the increase or decrease in the level of the target region is determined by comparing the level of the target region to a threshold level / value.

[0368] Genetic and / or epigenetic data can also be used to characterize specific forms of cancer. Cancer is often heterogeneous both in composition and staging. Genetic and / or epigenetic profile data can enable the characterization of specific subtypes of cancer, which may be important in the diagnosis or treatment of that specific subtype. This information can also provide clues to the prognosis of a specific type of cancer to the subject or practitioner, and may enable either the subject or the practitioner to adapt treatment options as the disease progresses. Some cancers can progress to become more malignant and genetically unstable. Other cancers may remain benign, inactive, or dormant. The systems and methods of the present disclosure may be useful in determining disease progression.

[0369] Furthermore, the methods of the present disclosure can be used to characterize the heterogeneity of an abnormal condition in a subject. Such methods can include, for example, generating a genetic and / or epigenetic profile of cfDNA derived from the subject, where the genetic and / or epigenetic profile includes a plurality of data resulting from the analysis of copy number variations and rare mutations. In some embodiments, the abnormal condition is cancer, such as as described herein. In some embodiments, the abnormal condition can result in a heterogeneous genomic population. In the example of cancer, it is known that some tumors contain tumor cells at different stages of the cancer. In other examples, the heterogeneity can include multiple lesions of the disease, for example, where one or more lesions (e.g., one or more tumor lesions) are the result of metastases spreading from the primary site of the cancer. The tissue(s) of origin can be useful in identifying the organ(s) affected by the cancer that includes the primary cancer and / or metastatic tumors.

[0370] The method of the present invention can also be used to quantify the levels of different cell types, such as rare immune cell types, such as activated lymphocytes and myeloid cells, at specific stages of differentiation. Such quantification can be based on the number of molecules corresponding to a given cell type in the sample. The sequence information obtained in the method of the present invention can include sequence reads of nucleic acids generated by a nucleic acid sequencer. In some embodiments, the nucleic acid sequencer performs pyrosequencing, single molecule sequencing, nanopore sequencing, semiconductor sequencing, sequencing by synthesis, five-letter sequencing, six-letter sequencing, sequencing by ligation, or sequencing by hybridization on the nucleic acid to generate sequencing reads. In some embodiments, the method further includes the step of grouping the sequence reads into families of sequence reads, each family containing sequence reads generated from the nucleic acids in the sample. In some embodiments, these methods include the step of determining the probability that the subject from whom the sample was taken has cancer or pre-cancer, or has metastases, which is related to changes in the proportions of immune cell types.

[0371] The method of the present invention can be used to generate or profile a set of fingerprints or data that is the sum of genetic and / or epigenetic information from different cells in a heterogeneous disease. This set of data can include, alone or in combination, analysis of copy number variations, epigenetic variations, and mutations.

[0372] The method of the present invention can be used to diagnose, prognose, monitor, or observe cancer or other diseases. In some embodiments, the methods herein are not involved in diagnosing, prognosing, or monitoring a fetus, and thus are not related to non-invasive prenatal testing. In other embodiments, these methodologies can be used in pregnant subjects to diagnose, prognose, monitor, or observe cancer or other diseases in a prenatal subject whose DNA and other polynucleotides can co-circulate with maternal molecules.

[0373] Non-limiting examples of other gene-based diseases, disorders or conditions that may be evaluated as needed using the methods and systems disclosed herein include achondroplasia, alpha1-antitrypsin deficiency, antiphospholipid antibody syndrome, autism, autosomal dominant polycystic kidney disease, Charcot-Marie-Tooth (CMT), meowing cat, Crohn's disease, cystic fibrosis, Dercum's disease, Down syndrome, Duane syndrome, Duchenne muscular dystrophy, factor V Leiden thrombophilia, familial hypercholesterolemia, familial Mediterranean fever, fragile X syndrome, Gaucher's disease, hemochromatosis, hemophilia, holoprosencephaly, Huntington's disease, Klinefelter syndrome, Marfan syndrome, myotonic dystrophy, neurofibromatosis, Noonan syndrome, osteogenesis imperfecta, Parkinson's disease, phenylketonuria, Poland anomaly, porphyria, progeria, retinitis pigmentosa, severe combined immunodeficiency (SCID), sickle cell disease, spinal muscular atrophy, Tay-Sachs, thalassemia, trimethylaminuria, Turner syndrome, velocardiofacial syndrome, WAGR syndrome, Wilson's disease, and the like.

[0374] In some embodiments, the method can provide a measure of the extent of DNA damage through quantification of regions synthesized during end repair, and the methods disclosed herein can also be used to quantify the level of DNA damage present in the original DNA sample. This is because the level of end repair depends in part on the amount of DNA damage (e.g., gaps, nicks, and overhangs) present in the DNA, and this damage can act as a priming site for synthesis in end repair (see FIGS. 1-6 and the corresponding description).

[0375] In some embodiments, the method further includes calculating a synthesis index, which is a quantitative measure of the synthesized region in end repair. The synthesis index may be at the molecular level and / or at the sample level. The synthesis index can be the proportion of the sequencing data corresponding to the synthesized region. In some embodiments, the method further includes classifying the DNA sample by comparing the synthesis index with one or more reference values. The classification can be whether the DNA sample is from a subject with cancer or from a subject without cancer. The reference values can be obtained from one or more control DNA samples known to have specific characteristics, for example, from subjects known to have cancer, such as a specific type of cancer. The reference values can be obtained by performing the method used to obtain the synthesis index for the control samples (i.e., using the same end repair, ligation, and sequencing methods).

[0376] In some embodiments, the method described herein includes detecting the presence or absence of DNA originating from or derived from tumor cells at a preselected time point after a previous cancer treatment of a subject previously diagnosed with cancer, using a set of sequence information obtained as described herein. The method may further include determining, for the subject, a cancer recurrence score indicative of the presence or absence of DNA originating from or derived from tumor cells.

[0377] When a cancer recurrence score is determined, it can be further used to determine the cancer recurrence status. The cancer recurrence status can be, for example, at risk of cancer recurrence if the cancer recurrence score exceeds a predetermined threshold. The cancer recurrence status can be, for example, at a low or lower risk of cancer recurrence if the cancer recurrence score exceeds a predetermined threshold. In certain embodiments, a cancer recurrence score equal to the predetermined threshold can result in a cancer recurrence status of either at risk of cancer recurrence or at a low or lower risk of cancer recurrence.

[0378] In some embodiments, the cancer recurrence score is compared to a predetermined cancer recurrence threshold, and the subject is classified as a candidate for subsequent cancer treatment if the cancer recurrence score exceeds the cancer recurrence threshold, or not classified as a candidate for treatment if the cancer recurrence score is below the cancer recurrence threshold. In certain embodiments, a cancer recurrence score equal to the cancer recurrence threshold can result in either classification as a candidate for subsequent cancer treatment or not as a candidate for treatment.

[0379] The methods discussed above may further include any singular or plural compliance features shown elsewhere in this specification that are included in the section related to methods for determining the risk of cancer recurrence in a subject and / or methods for classifying a subject as a candidate for subsequent cancer treatment.

[0380] Methods for determining the risk of cancer recurrence in a subject and / or methods for classifying a subject as a candidate for subsequent cancer treatment In some embodiments, the methods provided herein are, or include, methods for determining the risk of cancer recurrence in a subject. In some embodiments, the methods provided herein are, or include, methods for detecting the presence or absence of metastases in a subject. In some embodiments, the methods provided herein are, or include, methods for classifying a subject as a candidate for subsequent cancer treatment.

[0381] Any such method can include collecting a sample (e.g., DNA, e.g., DNA originating from or derived from tumor cells) from a subject diagnosed with cancer at one or more preselected time points after one or more previous cancer treatments administered to the subject. The subject can be any of the subjects described herein. The sample can include chromatin, cfDNA or other cellular material. The sample, e.g., the DNA sample, can be a tissue sample.

[0382] Any such method may include capturing multiple sets of target regions from DNA derived from a subject, the multiple sets of target regions including a set of sequence-variable target regions and a set of epigenetic target regions, whereby a captured set of DNA molecules is produced. The capturing step may be performed according to any of the embodiments described elsewhere in this specification.

[0383] In any such method, previous cancer treatments may include surgery, administration of therapeutic compositions, and / or chemotherapy.

[0384] Any such method includes sequencing the captured DNA molecules, whereby a set of sequence information is produced. The captured DNA molecules of the set of sequence-variable target regions may be sequenced to a greater depth of sequencing than the captured DNA molecules of the set of epigenetic target regions.

[0385] Any such method may include using the set of sequence information to detect the presence or absence of DNA originating from or derived from tumor cells at a preselected time point. Detection of the presence or absence of DNA originating from or derived from tumor cells may be performed according to any of its embodiments described elsewhere in this specification.

[0386] A method for determining the risk of cancer recurrence in a subject may include, for the subject, determining a cancer recurrence score indicative of the presence, absence, or amount of DNA originating from or derived from tumor cells, e.g., of a target genomic region and a target area. The cancer recurrence score may be further used to determine a cancer recurrence status. The cancer recurrence status may be, for example, in a risk state of cancer recurrence if the cancer recurrence score exceeds a predetermined threshold. The cancer recurrence status may be, for example, in a low or lower risk state of cancer recurrence if the cancer recurrence score exceeds a predetermined threshold. In certain embodiments, a cancer recurrence score equal to the predetermined threshold may result in a cancer recurrence status of either being in a risk state of cancer recurrence or in a low or lower risk state of cancer recurrence.

[0387] A method for detecting the presence or absence of metastasis in a subject may include comparing the presence or level of tissue-specific cell material to the presence or level of tissue-specific cell material obtained from the subject at different time points, a reference level of tissue-specific cell material, or a comparative cell material for comparison. The methods herein may include additional steps for determining whether metastasis is present.

[0388] A method for classifying a subject as a candidate for subsequent cancer treatment may include comparing the subject's cancer recurrence score to a predetermined cancer recurrence threshold, whereby the subject is classified as a candidate for subsequent cancer treatment if the cancer recurrence score exceeds the cancer recurrence threshold, or is not classified as a candidate for treatment if the cancer recurrence score is below the cancer recurrence threshold. In certain embodiments, a cancer recurrence score equal to the cancer recurrence threshold may result in either a classification as a candidate for subsequent cancer treatment or a classification as not being a candidate for treatment. In some embodiments, subsequent cancer treatment includes chemotherapy or administration of a therapeutic composition.

[0389] Any such method may include determining a disease-free survival (DFS) period for a subject based on a cancer recurrence score; for example, the DFS period may be 1 year, 2 years, 3 years, 4 years, 5 years, or 10 years.

[0390] In some embodiments, an array variable target region array is obtained, and determining the cancer recurrence score may include determining at least a first subscore indicative of the amount of SNVs, insertions / deletions, CNVs, and / or fusions present in the array variable target region array.

[0391] In some embodiments, the number of mutations in an array variable target region selected from 1, 2, 3, 4, or 5 is sufficient for the first subscore to result in a cancer recurrence score classified as positive for cancer recurrence. In some embodiments, the number of mutations is selected from 1, 2, or 3.

[0392] In some embodiments, an epigenetic target region array is obtained, and determining the cancer recurrence score includes determining a second subscore indicative of the amount of molecules (obtained from the epigenetic target region array) that exhibit an epigenetic state different from that of DNA found in a corresponding sample from a healthy subject (e.g., cfDNA found in a blood sample from a healthy subject, or DNA found in a tissue sample from a healthy subject if the tissue sample is of the same type as that obtained from the subject). These abnormal molecules (i.e., molecules having an epigenetic state different from that of DNA found in a corresponding sample from a healthy subject) may correspond to epigenetic changes associated with cancer (such as those with metastases), for example, methylation of hypermethylated variable target regions and / or disrupted fragmentation of fragmented variable target regions, where "disrupted" means different from that of DNA found in a corresponding sample from a healthy subject.

[0393] In some embodiments, the proportion of molecules corresponding to a set of hypermethylated variable target regions showing hypermethylation and / or abnormal fragmentation in a set of fragmented variable target regions greater than or equal to a value in the range of 0.001% to 10% is sufficient for the subscore to be classified as positive for cancer recurrence. This range can be 0.001% to 1%, 0.005% to 1%, 0.01% to 5%, 0.01% to 2% or 0.01% to 1%.

[0394] In some embodiments, any of such methods may include determining the fraction of tumor DNA from the fraction of molecules in a set of sequence information showing one or more characteristics indicative of a tumor cell origin. This can be carried out, for example, for molecules corresponding to some or all of the target regions including one or more of hypermethylated variable target regions, hypomethylated variable target regions and fragmented variable target regions (hypermethylation of hypermethylated variable target regions and / or abnormal fragmentation of fragmented variable target regions can be considered indicative of a tumor cell origin). This can be carried out for molecules corresponding to sequence variable target regions, for example, molecules containing changes consistent with cancer, such as SNV, indel, CNV and / or fusion. The fraction of tumor DNA can be determined based on a combination of molecules corresponding to epigenetic target regions and molecules corresponding to sequence variable target regions.

[0395] The determination of the cancer recurrence score can be based at least in part on the fraction of tumor DNA, and a fraction of tumor DNA greater than a threshold in the range of 10 -11 ~1 or 10 -10 ~1 is sufficient for the cancer recurrence score to be classified as positive for cancer recurrence. In some embodiments, 10 -10 ~10 -9 、10 -9 ~10 -8 、10 -8 ~10 -7 、10 -7 ~10 -6 、10 -6 ~10 -5 、10 -5 ~10-4 , 10 -4 ~10 -3 , 10 -3 ~10 -2 or 10 -2 ~10 -1 In some embodiments, a fraction of tumor DNA greater than or equal to a threshold in the range of at least 10 is sufficient for the Cancer Recurrence Score to be classified as positive for cancer recurrence. -7 The fraction of tumor DNA greater than the threshold is sufficient for the cancer recurrence score to be classified as positive for cancer recurrence. The determination that the fraction of tumor DNA is greater than a threshold, for example, a threshold corresponding to any of the above-mentioned embodiments, can be based on cumulative probability. For example, a sample is considered positive if the cumulative probability that the tumor fraction is greater than a threshold in any of the above-mentioned ranges exceeds a probability threshold of at least 0.5, 0.75, 0.9, 0.95, 0.98, 0.99, 0.995 or 0.999. In some embodiments, the probability threshold is at least 0.95, for example 0.99.

[0396] In some embodiments, the set of sequence information includes sequence variable target region sequences and epigenetic target region sequences, and determining the cancer recurrence score includes determining subscores indicative of the amount of SNVs, insertions / deletions, CNVs and / or fusions present in the sequence variable target region sequences, and subscores indicative of the amount of abnormal molecules in the epigenetic target region sequences, and combining the subscores to provide a cancer recurrence score. When subscores are combined, they can be combined by applying a threshold value (e.g., for sequence variable target regions, greater than a predetermined number of mutations (e.g., >1), and for epigenetic target regions, greater than a predetermined fraction of abnormal molecules (i.e., molecules with an epigenetic state different from that found in the corresponding sample from a healthy subject; e.g., tumor)) to each subscore independently, or by training a machine learning classifier to determine the status based on multiple positive and negative training samples.

[0397] In some embodiments, a value for a combined score in the range of -4 to 2 or -3 to 1 is sufficient for the cancer recurrence score to be classified as positive for cancer recurrence.

[0398] In any embodiment where the cancer recurrence score is classified as positive for cancer recurrence, the subject's cancer recurrence status may be at risk of cancer recurrence and / or the subject may be classified as a candidate for subsequent cancer treatment.

[0399] In some embodiments, the cancer is any one of the types of cancer described elsewhere herein, for example, colorectal cancer.

[0400] Method for monitoring cancer in a subject over time; sample collection at two or more time points In some embodiments, the method may be used to monitor one or more aspects of a subject's condition over time, such as the subject's response to treatment for the condition (e.g., response to a chemotherapeutic or immunotherapeutic agent), the severity of the condition in the subject (e.g., cancer stage), recurrence of the condition (e.g., cancer), and / or the risk of the subject developing the condition (e.g., cancer), and / or to monitor the subject's health as part of a preventive health monitoring program (e.g., to determine whether and / or when the subject requires further diagnostic screening). In some embodiments, the monitoring includes the analysis of at least two samples collected from the subject at at least two different time points described herein.

[0401] The methods according to the present disclosure can be useful, for example, in predicting a subject's response to certain treatment options over a period of time. As described elsewhere herein, for example, if the treatment is successful such that more cancer dies and can shed DNA, a successful treatment option may increase the amount of cancer-related DNA sequences detected in the subject's blood. In such an example, a particular treatment option may correlate with the genetic profile of the cancer over time. This correlation can be useful in selecting a treatment.

[0402] As disclosed herein, in some embodiments, the amount of each of a plurality of cell types, such as immune cell types, is determined based on sequencing and analysis (e.g., determination of epigenetic and / or genomic signatures) of DNA isolated from at least one sample (e.g., a tissue sample or a blood sample, such as a whole blood sample, a buffy coat sample, a leukapheresis sample, or a PBMC sample) containing cells from the subject. In some embodiments, differences in the levels and / or presence of particular gene signatures and / or epigenetic signatures in DNA isolated from a blood sample from the subject can be used to quantify cell types, such as immune cell types, within the sample. Thus, comparison of the disclosed gene signatures and / or epigenetic signatures in DNA isolated from blood samples taken from the subject at two or more time points can be used to monitor changes in the amount of cell types in the subject under different conditions (e.g., before and after treatment), or over time (e.g., as part of a preventive health monitoring program).

[0403] The disclosed method may include the step of evaluating (e.g., quantifying) and / or interpreting the cell types (e.g., immune cell types) present in one or more samples (e.g., tissue samples or blood samples, such as whole blood samples, buffy coat samples, leukapheresis samples, or PBMC samples) collected from a subject at one or more time points, compared to a selected baseline value or reference standard (or a selected set of baseline values or reference standards). The baseline value or reference standard may be the amount of cell types (e.g., the average amount or range of amounts of cell types present in at least two samples) measured in one or more samples collected from a subject at one or more time points, e.g., before receiving treatment, before diagnosis of a condition (e.g., cancer), or as part of a preventive health monitoring program. The baseline value or reference standard may be the amount of cell types (e.g., the average amount or range of amounts of cell types present in at least two samples) measured in one or more samples collected from one or more subjects without the condition (e.g., healthy subjects without cancer), one or more subjects who responded favorably to a treatment, or one or more subjects not undergoing treatment, at one or more time points. In certain embodiments, the baseline value or reference standard utilized is a standard or profile derived from a single reference subject. In other embodiments, the baseline value or reference standard utilized is a standard or profile derived from data averaged from multiple reference subjects. The reference standard may, in various embodiments, be a single value, mean, average, numerical average or range of numerical averages, numerical pattern, or graphical pattern generated from cell type amount data derived from a single reference subject or multiple reference subjects. The selection of a particular baseline value or reference standard, or the selection of one or more reference subjects, depends, for example, on how the methods described herein are to be used by a research scientist or clinician (e.g., a physician).

[0404] In some embodiments, one or more samples (e.g., tissue samples or blood samples, such as whole blood samples, buffy coat samples, leukapheresis samples, or PBMC samples) can be collected from a subject at two or more time points to evaluate changes in cell types (e.g., changes in the amount of cell types) between two or more time points. In some embodiments, the sample collected at the first time point is a tissue sample or a blood sample, and the sample collected at a subsequent time point (e.g., the second time point) is a blood sample. In some embodiments, the sample collected at the first time point is a tissue sample, and the sample collected at a subsequent time point (e.g., the second time point) is a blood sample. By monitoring cell types in samples collected from a subject at two or more time points and identifying differences between cell types, the method can be used, for example, to determine the presence or absence of a condition (e.g., cancer), the subject's response to treatment, one or more characteristics of a condition (e.g., cancer stage) in the subject, recurrence of a condition (e.g., cancer), and / or the risk of a subject developing a condition (e.g., cancer). Thus, in some embodiments, there is provided a method in which the amount of cell types present in at least one sample (e.g., at least one tissue sample and / or at least one blood sample, such as whole blood sample, buffy coat sample, leukapheresis sample, or PBMC sample) collected from a subject at one or more time points (e.g., before receiving treatment) is compared to the amount of cell types present in at least one sample collected from the subject at one or more different time points (e.g., after receiving treatment, etc.). The disclosed method can enable patient-specific monitoring such that differences in the amount of cell types between samples collected from a subject at different time points can indicate changes (e.g., presence or absence of a condition, response to treatment, prognosis, etc.) that are significant for the subject but still within the normal range for a general healthy population.

[0405] As disclosed herein, methods are provided for monitoring one or more aspects of the state of a subject over time, such as, but not limited to, the subject's response to treatment of a condition (e.g., response to a chemotherapeutic or immunotherapeutic agent). In certain embodiments, one or more samples are collected from the subject at least 1-10, at least 1-5, at least 2-5, or at least 1, at least 2, at least 3 (least 3), at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15 or at least 20 time points before the subject undergoes treatment. In certain embodiments, one or more samples are collected from the subject at least 1-10, at least 1-5, at least 2-5, or at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15 or at least 20 time points after the subject undergoes treatment. Sample collection from the subject can continue during and / or after treatment to monitor the subject's response to the treatment.

[0406] In some embodiments, samples are not collected from the subject before diagnosis of a condition (e.g., cancer) or before treatment. In such embodiments where the subject's response to treatment or the course or stage of a condition (e.g., cancer) in the subject is being monitored over time, the cell types are compared between samples obtained at least 2-10, at least 2-5, at least 3-6, or at least 2, e.g., at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15 or at least 20 time points after the subject is diagnosed and / or after the subject undergoes treatment. Sample collection from the subject can continue during and / or after treatment to monitor the subject's response to the treatment.

[0407] In some embodiments of the disclosed methods, one or more samples (e.g., one or more tissue, whole blood, buffy coat, leukapheresis, or PBMC samples) are collected from the subject at least once per year, e.g., about 1-12 times or about 2-6 times per year, e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 times. In other embodiments, one or more samples are collected from the subject less than once per year, e.g., about once every 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 months. In some embodiments, one or more samples are collected from the subject about once every 1-5 years or about once every 1-2 years, e.g., about every 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, or 5 years.

[0408] In other embodiments of the disclosed methods, one or more samples (e.g., one or more tissue samples or blood samples, e.g., or one or more buffy coat samples, leukocyte samples, leukapheresis samples, or PBMC samples) are collected from the subject at least once per week, e.g., 1-4 days, 1-2 days, or 1, 2, 3, 4, 5, 6, or 7 days per week. In certain embodiments, one or more samples are collected from the subject at least once per month, e.g., 1-15 times, 1-10 times, 2-5 times, or 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 times per month. In other embodiments, one or more samples are collected from the subject every month, every two months, every three months, every four months, every five months, every six months, every seven months, every eight months, every nine months, every ten months, every eleven months, or every twelve months. In some embodiments, one or more samples are collected from a subject at least once per day, e.g., 1, 2, 3, 4, 5, or 6 times per day. The selection of one or more sample collection time points (e.g., frequency of sample collection), or the number of samples collected at each time point, depends on how the methods described herein are used, e.g., by a research scientist or clinician (e.g., physician).

[0409] Treatment and Related Administration In certain embodiments, the methods disclosed herein relate to customized treatment, e.g., identifying a customized treatment and administering it to a patient. In some embodiments, the patient or subject has a given disease, disorder, or condition, e.g., any of the cancers or other conditions described elsewhere herein. Essentially any cancer treatment (e.g., surgical treatment, radiation treatment, chemotherapy, immunotherapy, etc.) can be included as part of these methods. In certain embodiments, the treatment administered to the subject includes at least one chemotherapeutic drug. In some embodiments, chemotherapeutic drugs include alkylating agents (e.g., without limitation, chlorambucil, cyclophosphamide, cisplatin, and carboplatin), nitrosoureas (e.g., without limitation, carmustine and lomustine), antimetabolites (e.g., without limitation, Fluorauracil, methotrexate, and fludarabine), plant alkaloids and natural products (e.g., without limitation, vincristine, paclitaxel, and topotecan), antitumor antibiotics (e.g., without limitation, bleomycin, doxorubicin, and mitoxantrone), hormonal agents (e.g., without limitation, prednisone, dexamethasone, tamoxifen, and leuprolide), and biological response modifiers (e.g., without limitation, Herceptin and Avastin, Erbitux and Rituxan). In some embodiments, the chemotherapy administered to the subject can include FOLFOX or FOLFIRI. In certain embodiments, a treatment comprising at least one PARP inhibitor can be administered to the subject. In certain embodiments, PARP inhibitors can include, inter alia, OLAPARIB, TALAZOPARIB, RUCAPARIB, NIRAPARIB (trade name ZEJULA). Typically, the treatment includes at least one immunotherapy (or immunotherapeutic agent). Immunotherapy generally refers to methods of enhancing the immune response to a given cancer type. In certain embodiments, immunotherapy refers to methods of enhancing the T cell response to a tumor or cancer.

[0410] In some embodiments, the immunotherapy or immunotherapeutic agent targets immune checkpoint molecules. Certain tumors can evade the immune system by exploiting the immune checkpoint pathway. Thus, targeting immune checkpoints has emerged as an effective approach to disable a tumor's ability to evade the immune system and activate antitumor immunity against certain cancers. Pardoll, Nature Reviews Cancer, 2012, 12:252-264.

[0411] In certain embodiments, the immune checkpoint molecule is an inhibitory molecule ...