Helicase-cytidine deaminase complexes and methods of use thereof

By performing one-step enzymatic mapping using altered cytidine deaminase, the problems of DNA degradation, complexity loss and detection resolution limitation in the prior art are solved, and efficient and simplified detection of DNA methylation status is achieved.

CN120153070APending Publication Date: 2025-06-13ILLUMINA INC
View PDF 102 Cites 0 Cited by

Patent Information

Application Number
CN202380076909.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-14
Filing Date
2023-09-29
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art has limitations in mapping modified DNA cytosines, including DNA degradation, loss of complexity, the need for enzyme and chemical treatment of multiple transformations, and detection resolution limitations.

Method used

Using a one-step, completely enzymatic approach, using altered cytidine deaminase selectively acts on certain modified cytosines of the target nucleic acid, converting them into thymidine or modified thymidine analogs, avoiding chemical treatment at neutral pH and high temperatures and maintaining the complexity of sample DNA.

Benefits of technology

The detection of methylated C at single base resolution is achieved, maintaining DNA complexity, simplifying next-generation sequencing analysis, and avoiding the limitations of DNA loss and detection resolution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005385134840000261
    Figure BDA0005385134840000261
  • Figure BDA0005385134840000271
    Figure BDA0005385134840000271
  • Figure HDA0005385134850000011
    Figure HDA0005385134850000011
Patent Text Reader

Abstract

The protein complex comprises cytidine deaminase and helicase. In some embodiments, the cytidine deaminase is an altered cytidine deaminase. In some embodiments, the protein complexes convert 5 methylcytosine to thymine. Kits, compositions, and methods of use comprising the protein complexes of cytidine deaminase and helicase are also described.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims the benefit of U.S. Provisional Application Serial No. 63 / 412,221, filed Sep. 30, 2022, and U.S. Provisional Application Serial No. 63 / 416,143, filed Oct. 14, 2022, each of which is hereby incorporated by reference in its entirety.

[0003] Sequence Listing

[0004] This application contains a Sequence Listing that has been electronically submitted to the United States Patent and Trademark Office via EFS - Web and is hereby incorporated by reference in its entirety. The Sequence Listing is an XML file named "0531_002462WO01_SL.xml", size 429 kilobytes, and created on Sep. 25, 2023. The information contained in the Sequence Listing is hereby incorporated by reference. Technical Field

[0005] Embodiments of the present disclosure relate to the preparation of nucleic acids for sequencing or other applications. Specifically, embodiments of the proteins, methods, compositions, and kits provided herein relate to mapping methylation status by using sequencing libraries and other methods. Background Art

[0006] Modified DNA cytosines, including 5 - methylcytosine (5mC) and 5 - hydroxymethylcytosine (5hmC), are well - studied epigenetic modifications that play fundamental roles in human development and disease. Their genome - wide distributions are different between tissue types and between healthy and diseased states. In recent years, 5mC has also received attention as a clinical diagnostic tool: its distribution in cell - free DNA (cfDNA) obtained from liquid biopsy samples can be used for tissue - specific prediction of early cancer or monitoring of cancer recurrence or remission after treatment. Thus, the development of methods for mapping modified DNA cytosines at single - base resolution with minimal loss of sample DNA quantity, quality, and complexity has been closely followed. However, current methods for mapping modified DNA cytosines show limitations, including (i) sample DNA degradation due to long - term chemical treatment at non - neutral pH and high temperature, (ii) loss of sample DNA complexity due to conversion of unmethylated DNA bases to uracil, resulting in low - complexity genomic mapping, (iii) multi - step conversions that require both enzymatic and chemical treatment, and (iv) for antibody - based 5mC detection, the detection resolution is limited to about 150 bp, which hinders the identification of its exact location in the genome.

[0007] 5-Hydroxymethylcytosine (5hmC) is an oxidized derivative of the widely studied epigenetic modification 5-methylcytosine (5mC). Increasing evidence supports the biological importance of 5hmC in multiple developmental processes in mammals, such as neurogenesis. Accordingly, determining the localization of 5hmC in DNA from healthy and diseased patients has received wide attention. Most methods for mapping 5-hydroxymethylcytosine (5hmC) require bisulfite treatment, which results in significant DNA loss and damage. Methods for mapping 5hmC have recently been developed, such as oxBS-seq, TAB-seq, and ACE-seq, but some methods include bisulfite treatment and all involve multiple steps using different enzymes. Summary of the Invention

[0008] The present disclosure provides proteins, methods, compositions, and kits for determining the methylation status of DNA and RNA using physically proximate cytidine deaminase and helicase. Different from current methods for mapping the methylation status of cytosine (C) nucleotides, the present disclosure provides a one-step, fully enzymatic method using an altered cytidine deaminase that selectively acts on certain modified cytosines of a target nucleic acid and converts them to thymidine (T) or a modified thymidine analog. The altered cytidine deaminases described herein circumvent the limitations of currently available methods for mapping methylated cytosine nucleotides because (i) they are active at near-neutral pH and physiological temperature, (ii) unmethylated cytosines react at a reduced rate, thus preserving sample DNA complexity, (iii) the conversion of methylated C to T is a one-step enzymatic reaction, (iv) the method enables the detection of methylated C at single-base resolution, and (v) the method maintains DNA complexity, thereby simplifying analysis by next-generation sequencing. The altered and wild-type (WT) cytidine deaminases described herein generally have higher catalytic activity towards single-stranded nucleic acids. During nucleic acid processing, the use of physically proximate helicase and cytidine deaminase is expected to demonstrate improved catalytic activity towards double-stranded nucleic acids.

[0009] The present disclosure also provides proteins, methods, compositions, and kits for mapping 5-hydroxymethylcytosine (5hmC) nucleotides present in DNA and RNA. Current methods for mapping the methylation status of 5hmC nucleotides include steps of modifying or blocking 5hmC nucleotides. For example, the ACE-seq method (Schutsky et al., Nature biotechnology, 10.1038 / nbt.4204. October 8, 2018, doi: 10.1038 / nbt.4204) blocks 5hmC by conversion to 5ghmC using β-glucosyltransferase (βGT). Different from current methods for mapping the methylation status of 5hmC nucleotides, the methods provided herein do not require modifying or blocking 5hmC nucleotides. Instead, the present disclosure provides a one-step, fully enzymatic method using an altered cytidine deaminase that selectively acts on certain modified cytosines of a target nucleic acid and converts them to uracil (U) or thymidine (T), but does not act on 5hmC, 5-formylcytosine (5fC), or 5-carboxylcytosine (5-caC). The altered cytidine deaminase described herein circumvents the limitations of currently available methods for mapping 5hmC nucleotides because (i) it does not require harsh chemical treatments that result in substantial loss of DNA and RNA, (ii) the conversion is a one-step enzymatic reaction, and (iii) the method enables detection of 5hmC at single-base resolution.

[0010] The present disclosure includes altered cytidine deaminases, protein complexes including an altered cytidine deaminase and a second protein such as a helicase. In one embodiment, the altered cytidine deaminase comprises amino acid substitution mutations at positions that are functionally equivalent to (Tyr / Phe)130 and Tyr132 in the wild-type APOBEC3A protein. In another embodiment, the altered cytidine deaminase comprises an amino acid substitution mutation at a position that is functionally equivalent to (Tyr / Phe)130 in the wild-type APOBEC3A protein, wherein the substitution mutation is (Tyr / Phe)130Trp. The (Tyr / Phe)130 of the altered cytidine deaminase can be Tyr130, and the wild-type APOBEC3A protein is SEQ ID NO: 3. The present disclosure also includes polynucleotides encoding the altered cytidine deaminases.

[0011] The present disclosure also provides compositions comprising the altered cytidine deaminases described herein, including protein complexes of a cytidine deaminase and a second protein (such as a helicase). In one embodiment, the composition may additionally comprise at least one of the following: (i) a sample containing DNA, the DNA comprising at least one modified cytosine, wherein the modified cytosine is 5-methylcytosine (5mC), 5-hydroxymethylcytosine (5hmC), 5-formylcytosine (5fC), 5-carboxylcytosine (5caC), or a combination thereof; or (ii) a buffer, such as a buffer having a pH below 7; or (iii) a combination thereof.

[0012] The present disclosure also provides methods of using the cytidine deaminases described herein, the cytidine deaminases including protein complexes of an altered cytidine deaminase and a second protein (such as a helicase). In one embodiment, a method includes providing a DNA sample suspected of containing single-stranded DNA, the single-stranded DNA comprising at least one 5-methylcytosine (5mC), at least one 5-hydroxymethylcytosine (5hmC), at least one 5-formylcytosine (5fC), at least one 5-carboxylcytosine (5CaC), or a combination thereof; contacting the single-stranded DNA with the altered cytidine deaminase under conditions suitable for (i) converting 5-methylcytosine (5mC) to thymidine (T) by deamination at a rate greater than the rate of converting cytosine (C) to uracil (U) by deamination to produce a converted single-stranded DNA, or under conditions suitable for (ii) converting C to U and converting 5mC to T by deamination at a rate greater than the rate of converting 5-hydroxymethylcytosine (5hmC) to 5-hydroxymethyluracil (5hmU) by deamination; and processing the converted single-stranded DNA to produce a sequencing library.

[0013] In another embodiment, a method includes providing a DNA sample suspected of containing double-stranded DNA that includes at least one 5-methylcytosine (5mC), at least one 5-hydroxymethylcytosine (5hmC), at least one 5-formylcytosine (5fC), at least one 5-carboxylcytosine (5caC), or a combination thereof; processing the double-stranded DNA to produce a sequencing library; denaturing the sequencing library to produce single-stranded DNA; contacting the single-stranded DNA with an altered cytidine deaminase under conditions suitable for (i) converting 5-methylcytosine (5mC) to thymidine (T) by deamination at a rate greater than the rate of converting cytosine (C) to uracil (U) by deamination, or under conditions suitable for (ii) converting C to U and converting 5mC to T by deamination at a rate greater than the rate of converting 5-hydroxymethylcytosine (5hmC) to 5-hydroxymethyluracil (5hmU) by deamination, to produce a converted single-stranded DNA; and converting the converted single-stranded DNA to a converted double-stranded DNA sequencing library.

[0014] In one embodiment, a method can include detecting the position of a modified cytosine in a target nucleic acid. The method can include (a) contacting a target nucleic acid suspected of containing at least one modified cytosine with the altered cytidine deaminase according to claim 1 or claim 2 to produce a converted nucleic acid that includes at least one converted cytosine; and (b) detecting the at least one converted cytosine in the converted nucleic acid of (a). The detecting can include sequencing the converted nucleic acid or hybridizing a nucleic acid probe to the converted nucleic acid.

[0015] In an embodiment where the detecting includes sequencing the converted nucleic acid, the method can further include (c) comparing the sequence of the converted nucleic acid to an untreated reference sequence to determine which cytosines in the target nucleic acid are modified.

[0016] In embodiments where detection involves hybridizing the converted nucleic acid to a nucleic acid probe, the method may additionally include that the nucleic acid probe may be present on an analyte array, and the method may additionally include sequencing the hybridized converted nucleic acid. In another embodiment, the method may additionally include amplifying the converted nucleic acid, where the nucleic acid probe comprises two primers for amplifying a predetermined sequence, where the primers have a greater affinity for annealing to a region of the converted nucleic acid comprising at least one converted cytosine than to a region of the converted nucleic acid where at least one cytosine is not a converted cytosine, and where the presence of the amplified product indicates the presence of a modified cytosine in the target nucleic acid. In another embodiment, the method may additionally include cleaving a single-stranded DNA (ssDNA) reporter substrate by a CRISPR-based system, where the ssDNA reporter substrate comprises a fluorophore and a quencher, and where the presence of fluorescence indicates the presence of a modified cytosine in the target nucleic acid. In another embodiment, the converted nucleic acid may be present in fixed cells, where the nucleic acid probe comprises a fluorescence-labeled probe, and where the nucleic acid probe has a greater affinity for annealing to a predetermined sequence of the converted nucleic acid comprising at least one converted cytosine than to a region of the converted nucleic acid where at least one cytosine is not a converted cytosine, and where the presence of cell-associated fluorescence indicates the presence of a modified cytosine in the target nucleic acid.

[0017] In one embodiment of detecting the position of modified cytosine in a target nucleic acid, the target nucleic acid may be obtained from a subject, and the detection may include obtaining the pattern of cytosine modification in the converted nucleic acid. In some embodiments, the method may additionally include comparing the pattern of cytosine modification in the converted nucleic acid with the pattern of cytosine modification in a reference nucleic acid. For example, the subject may be a subject having a disease or disorder or at risk of having a disease or disorder, and the reference nucleic acid may be from a normal subject. In one embodiment, the pattern of cytosine modification is cis-linked to a coding region associated with the disease or disorder. For example, the pattern of cytosine modification is cis-linked to a coding region, where the coding region in the reference nucleic acid is transcriptionally active or transcriptionally inactive. The comparison may additionally include determining whether the pattern of cytosine modification of the converted nucleic acid indicates that the coding region in the subject is transcriptionally active or transcriptionally inactive. Transcription of the coding region may be associated with the disease or disorder.

[0018] In one aspect, the present disclosure describes a protein complex comprising a cytidine deaminase and a helicase. The cytidine deaminase may be a wild-type cytidine deaminase, or it may be an altered cytidine deaminase. In some embodiments, the altered cytidine deaminase may have an amino acid substitution mutation at a position functionally equivalent to (Tyr / Phe)130 in the wild-type APOBEC3A protein.

[0019] In another aspect, the present disclosure describes a composition comprising a protein complex and a DNA sample, the protein complex comprising a cytidine deaminase and a helicase, and the DNA sample comprising double-stranded DNA containing at least one 5mC.

[0020] In another aspect, the present disclosure describes a method comprising providing a DNA sample suspected of containing double-stranded DNA (dsDNA) containing at least one 5mC, contacting the sample with the protein complex described herein under conditions suitable for converting 5mC to T to produce converted dsDNA, wherein 5mC is converted to T, and processing the converted dsDNA.

[0021] In another aspect, the present disclosure describes a protein complex comprising a pathway for converting 5mC to T and a pathway for directional separation of double-stranded nucleic acids.

[0022] Unless otherwise specified, the terms used herein should be understood to have their ordinary meaning in the relevant art. Several terms used herein and their meanings are listed below.

[0023] As used herein, the terms "organism" and "subject" are used interchangeably and refer to microorganisms (e.g., prokaryotic or eukaryotic), animals, and plants. Examples of animals are mammals, such as humans.

[0024] As used herein, the term "target nucleic acid" is intended as a semantic identifier for a nucleic acid in the context of the methods, compositions, or kits shown herein and does not necessarily limit the structure or function of the nucleic acid, unless otherwise explicitly stated. Unless otherwise specified, references to nucleic acids such as target nucleic acids include both single-stranded nucleic acids and double-stranded nucleic acids, and both DNA and RNA. The term library refers to a collection of target nucleic acids containing known common sequences (such as universal sequences or adaptors) at the 3' and 5' ends.

[0025] As used herein, the term "adaptor" and its derivatives (e.g., universal adaptor) generally refer to any linear oligonucleotide that can be attached to a target nucleic acid. The adaptor can be single-stranded or double-stranded DNA, or can include both double-stranded regions and single-stranded regions. The adaptor can contain: a universal sequence that is substantially identical or substantially complementary to at least a portion of a primer (e.g., a universal primer); an index (also referred to herein as a barcode or tag) that is used to assist in downstream error correction, identification, or sequencing; and / or a unique molecular identifier. In some embodiments, the adaptor is substantially non-complementary to the 3' or 5' end of any target sequence present in the sample. In some embodiments, a suitable adaptor length ranges from about 6 to 100 nucleotides, from about 12 to 60 nucleotides, or from about 15 to 50 nucleotides. For example, the terms "linker" and "adaptor" can be used interchangeably.

[0026] As used herein, when used to describe nucleotide sequences, the term "common" refers to a sequence region shared by two or more nucleic acid molecules, where these molecules also have sequence regions that are different from each other. The common sequence present in different members of a nucleic acid collection can be used, for example, as a "landing zone" in subsequent steps to anneal a nucleotide sequence that can be used as a primer to add another nucleotide sequence (such as an index) to a target nucleic acid. The common sequence present in different members of a nucleic acid collection can allow for the capture of a variety of different nucleic acids using a common capture nucleic acid population (e.g., capture oligonucleotides complementary to a portion of the common sequence (e.g., common capture sequence)). Non-limiting examples of common capture sequences include sequences identical or complementary to the P5 and P7 primers. Similarly, the common sequence present in different members of a molecule collection can allow for the replication (e.g., sequencing) or amplification of a variety of different nucleic acids using a population of common primers complementary to a portion of the common sequence (e.g., common anchor sequence). In one embodiment, the common anchor sequence serves as the site to which a common primer (e.g., a sequencing primer for read 1 or read 2) anneals for sequencing. Thus, the capture oligonucleotide or common primer contains a sequence that can specifically hybridize to the common sequence.

[0027] When referring to common capture sequences or capture oligonucleotides, the terms "P5" and "P7" can be used. The terms "P5′" (P5 prime) and "P7′" (P7 prime) refer to the complementary sequences of P5 and P7, respectively. It should be understood that any suitable common capture sequence or capture oligonucleotide can be used in the methods presented herein, and the use of P5 and P7 is only an exemplary embodiment. The use of capture oligonucleotides such as P5 and P7 or their complementary sequences on a flow cell is known in the art, as illustrated by the disclosures of WO 2007 / 010251, WO 2006 / 064199, WO 2005 / 065814, WO 2015 / 106941, WO 1998 / 044151, and WO 2000 / 018957, the contents of which regarding P5 and P7 and their use are incorporated by reference. For example, any suitable forward amplification primer, whether immobilized or in solution, can be used in the methods presented herein for hybridizing to a complementary sequence and amplifying a sequence. Similarly, any suitable reverse amplification primer, whether immobilized or in solution, can be used in the methods presented herein for hybridizing to a complementary sequence and amplifying a sequence. Those skilled in the art will understand how to design and use primer sequences suitable for capturing and / or amplifying the nucleic acids presented herein.

[0028] As used herein, the term "primer" and its derivatives generally refer to any nucleic acid capable of hybridizing to a target sequence of interest. Generally, a primer serves as a substrate to which nucleotides can be polymerized by a polymerase or to which a polynucleotide can be ligated; however, in some embodiments, a primer can be incorporated into a synthetic nucleic acid strand and provide a site to which another primer can hybridize to initiate synthesis of a new strand complementary to the synthetic nucleic acid molecule. In some embodiments, a primer can be used to hybridize to a predetermined sequence, such as a predetermined sequence of one or more nucleotides that includes identification of the position of a modified cytosine. In one embodiment, a "primer" comprises a sequence present in a guide RNA used in a CRISPR-based system to hybridize to a predetermined sequence. A primer can include any combination of nucleotides or their analogs. In some embodiments, a primer is a single-stranded oligonucleotide or polynucleotide.

[0029] The terms "polynucleotide", "oligonucleotide", and "nucleic acid" are used interchangeably herein and refer to the polymeric form of nucleotides of any length and can include ribonucleotides, deoxyribonucleotides, their analogs, or mixtures thereof. These terms are to be understood to include analogs of DNA, RNA, cDNA, or antibody-oligonucleotide conjugates made from nucleotide analogs as equivalents, and apply to single-stranded (such as sense or antisense) and double-stranded polynucleotides. As used herein, the term also encompasses cDNA, i.e., complementary DNA or copy DNA produced from an RNA template, for example, by the action of reverse transcriptase.

[0030] As used herein, an "index" (also referred to as an "index region", "index adapter", "tag", or "barcode") refers to a unique nucleic acid tag that can be used to identify a sample or source of nucleic acid material, or a compartment in which a target nucleic acid is present. An index can be present in solution, on a solid support, or attached to or associated with a solid support and released into solution or the compartment. When nucleic acid samples are derived from multiple sources, the nucleic acids in each nucleic acid sample can be tagged with a different nucleic acid tag such that the source of the sample can be identified. Any suitable index or set of indexes can be used, as known in the art and exemplified by the disclosures of U.S. Patent 8,053,192, PCT Publication WO 05 / 068656, and U.S. Patent Publication 2013 / 0274117. In some embodiments, an index can comprise a six-base index 1 (i7) sequence, an eight-base index 1 (i7) sequence, an eight-base index 2 (i5e) sequence, a ten-base index 1 (i7) sequence, or a ten-base index 2 (i5) sequence obtained from Illumina, Inc. (San Diego, CA).

[0031] As used herein, the term "amplicon," when used in reference to a nucleic acid, means a product of replicating the nucleic acid, wherein the product has a nucleotide sequence that is the same as or complementary to at least a portion of the nucleotide sequence of the nucleic acid. An amplicon can be generated by any of a variety of amplification methods that use the nucleic acid or an amplicon thereof as a template, including, for example, polymerase extension, polymerase chain reaction (PCR), rolling circle amplification (RCA), ligation extension, or ligase chain reaction. An amplicon can be a nucleic acid molecule that is a single copy (e.g., a PCR product) or multiple copies of a particular nucleotide sequence (e.g., a tandem product of RCA) of that nucleotide sequence. The first amplicon of a target nucleic acid is typically a complementary copy. Subsequent amplicons are copies formed from the target nucleic acid or from the first amplicon after the first amplicon is generated. Subsequent amplicons can have a sequence that is substantially complementary to or substantially the same as the target nucleic acid.

[0032] As used herein, "amplify" or "amplification reaction" and their derivatives generally refer to any action or process in which at least a portion of a nucleic acid molecule is replicated or copied into at least one additional nucleic acid molecule. The additional nucleic acid molecule optionally contains a sequence that is substantially the same as or substantially complementary to at least some portions of the template nucleic acid molecule. The template nucleic acid molecule can be single-stranded or double-stranded, and the additional nucleic acid molecule can independently be single-stranded or double-stranded. Amplification is generally an exponential replication of a nucleic acid molecule. In some embodiments, such amplification can be carried out using isothermal conditions; in other embodiments, such amplification can include thermal cycling. In some embodiments, the amplification is multiplex amplification, which includes simultaneously amplifying multiple target sequences in a single amplification reaction. In some embodiments, "amplify" includes amplifying at least some portions of DNA- and RNA-based nucleic acids, either alone or in combination. An amplification reaction can include any amplification process known to those of ordinary skill in the art. In some embodiments, the amplification reaction includes polymerase chain reaction (PCR).

[0033] As used herein, the term "polymerase chain reaction" ("PCR") refers to the methods of U.S. Pat. Nos. 4,683,195 and 4,683,202 to Mullis, which describe methods for increasing the concentration of a segment of a polynucleotide of interest in a mixture of genomic DNA without cloning or purification. The method for amplifying the polynucleotide of interest involves introducing a large excess of two oligonucleotide primers into a DNA mixture containing the desired polynucleotide of interest, followed by a series of thermal cycles in the presence of a DNA polymerase. The two primers are complementary to the strands of their respective double-stranded polynucleotide of interest. The mixture is first denatured at a higher temperature, and then the primers are annealed to the complementary sequences within the polynucleotide molecule of interest. After annealing, the primers are extended with a polymerase to form a new pair of complementary strands. The steps of denaturation, primer annealing, and polymerase extension can be repeated multiple times (referred to as thermal cycling) to obtain a high concentration of the amplified fragment of the desired polynucleotide of interest. The length of the amplified fragment of the desired polynucleotide of interest (amplicon) is determined by the relative positions of the primers relative to each other, and thus, this length is a controllable parameter. Because this process is repeated, the method is called PCR. Since the desired amplified fragments of the polynucleotide of interest become the major nucleic acid sequences in the mixture (in terms of concentration), they are considered to be "PCR amplified". In a modified form of the above method, multiple different primer pairs (in some cases, one or more primer pairs for each target nucleic acid molecule of interest) can be used to PCR amplify the target nucleic acid molecules, thereby forming a multiplex PCR reaction.

[0034] As used herein, "amplification conditions" and its derivatives generally refer to conditions suitable for amplifying one or more nucleic acid sequences. In some embodiments, the amplification conditions can include isothermal conditions, or can include thermal cycling conditions, or a combination of isothermal conditions and thermal cycling conditions. In some embodiments, the conditions suitable for amplifying one or more nucleic acid sequences include polymerase chain reaction (PCR) conditions. Generally, amplification conditions refer to a reaction mixture sufficient to amplify a nucleic acid (such as one or more target sequences flanked by universal sequences or target-specific primers) or to amplify an amplified target sequence flanked by one or more adaptors. Generally speaking, amplification conditions include a catalyst for amplification or for nucleic acid synthesis, such as a polymerase; primers having a degree of complementarity to the nucleic acid to be amplified; and nucleotides, such as deoxyribonucleotide triphosphates (dNTPs), which promote primer extension once hybridized to the nucleic acid. Amplification conditions may require hybridization or annealing of the primers to the nucleic acid, extension of the primers, and a denaturation step in which the extended primers are separated from the nucleic acid sequence undergoing amplification. Generally, but not necessarily, amplification conditions can include thermal cycling; in some embodiments, the amplification conditions include multiple cycles in which the steps of annealing, extension, and separation are repeated. Generally, amplification conditions include cations such as Mg 2+ or Mn 2+and may also include various ionic strength modifiers.

[0035] As defined herein, "multiplex amplification" refers to the selective and non-random amplification of two or more target sequences within a sample using at least one target-specific primer pair. In some embodiments, multiplex amplification is performed such that some or all of the target sequences are amplified within a single reaction vessel. The "multiplicity" or "plexity" of a given multiplex amplification generally refers to the number of different target-specific sequences amplified during a single multiplex amplification. In some embodiments, the multiplicity can be about 12-plex, 24-plex, 48-plex, 96-plex, 192-plex, 384-plex, 768-plex, 1536-plex, 3072-plex, 6144-plex or higher. The amplified target sequences can also be detected by several different methods (e.g., gel electrophoresis followed by densitometry, quantification using a bioanalyzer or quantitative PCR, hybridization with a labeled probe; incorporation of biotinylated primers followed by avidin-enzyme conjugate detection; incorporation of 32 p-labeled deoxynucleotide triphosphates into the amplified target sequences).

[0036] As used herein, the term "amplification site" refers to a site within or on an array at which one or more amplicons can be generated. An amplification site can also be configured to contain, hold, or attach at least one amplicon generated at that site.

[0037] As used herein, the terms "array", "analyte array", and "microarray" are used interchangeably and refer to a set of sites that can be distinguished from one another based on their relative positions. Different molecules located at different sites of the array can be distinguished from one another based on the position of the site within the array. A single site of the array can contain one or more specific types of molecules. For example, a site can contain a single target nucleic acid molecule having a specific sequence, or a site can contain several nucleic acid molecules having the same sequence (and / or its complementary sequence). The sites of the array can be different features located on the same substrate. Exemplary features include, but are not limited to, droplets, wells in a substrate, beads (or other particles) in or on a substrate, protrusions of a substrate, ridges on a substrate, or channels in a substrate. The sites of the array can be separate substrates each bearing a different molecule. The different molecules attached to the separate substrates can be identified based on the position of the substrate on a surface associated with the substrate, or based on the position of the substrate in a liquid or gel. Exemplary arrays in which the separate substrates are located on a surface include, but are not limited to, those arrays having beads in wells.

[0038] As used herein, the term "compartment" is intended to denote an area or volume that separates or isolates something from other things. Exemplary compartments include, but are not limited to, vials, tubes, wells, droplets, agglomerates, beads, containers, surface features, flow cells, or areas or volumes separated by physical forces such as fluid flow, magnetic force, electric current, etc. In one embodiment, the compartment is a well of a microtiter plate (such as a 96-well plate or a 384-well plate). As used herein, a droplet can include a hydrogel bead, which is a bead for encapsulating one or more cell nuclei or cells and contains a hydrogel composition. In some embodiments, the droplet is a homogeneous droplet of a hydrogel material or a hollow droplet having a polymeric hydrogel shell. Whether homogeneous or hollow, the droplet is capable of encapsulating one or more cell nuclei or cells. In some embodiments, the droplet is a surfactant-stabilized droplet. In some embodiments, there is a single cell or cell nucleus in each compartment. In some embodiments, there are two or more cells or cell nuclei in each compartment. In some embodiments, each compartment contains a compartment-specific index. In some embodiments, the index is in solution or attached or associated with a solid phase in each compartment.

[0039] As used herein, the term "flow cell" refers to a chamber that includes a solid surface over which one or more fluid reagents can flow. Examples of flow cells and associated fluid systems and detection platforms that can be readily used in the methods of the present disclosure are described, for example, in the following documents: Bentley et al., Nature 456: 53-59 (2008); WO 04 / 018497; US 7,057,026; WO 91 / 06678; WO 07 / 123744; US 7,329,492; US 7,211,414; US 7,315,019; US 7,405,281 and US2008 / 0108082.

[0040] As used herein, the term "clonal population" refers to a population of nucleic acids that are homologous with respect to a particular nucleotide sequence. The homologous sequences are typically at least 10 nucleotides in length, but can even be longer, including, for example, at least 50, 100, 250, 500, or 1000 nucleotides in length. The clonal population can be derived from a single target nucleic acid or template nucleic acid. Generally, all of the nucleic acids in the clonal population will have the same nucleotide sequence. It should be understood that a small number of mutations (e.g., due to amplification artifacts) can occur in the clonal population without departing from clonality.

[0041] As used herein, "pattern of cytosine modification" (also referred to as "methylation profile") refers to a pattern in which both methylated and unmethylated cytosines are distributed in the genome of a cell or organism. The "pattern" includes both modified and unmodified cytosines. The pattern can be defined in several distribution dimensions: by organ, by tissue, by disease state or pathological condition (e.g., cancer, neurophysiology), by genomic segment (e.g., chromosome or genetic coordinates on a chromosome), by gene, by CpG island, a group of cytosines, or by the site of a modified cytosine. The pattern of cytosine modification may have a known correlation with a disease or pathological condition, or the correlation of the pattern of cytosine modification with a disease or pathological condition can be identified using the methods described herein. The pattern of cytosine modification may be present at a specific locus (e.g., position) in the genome, and the specific position can be a single modified cytosine or a group of modified cytosines, such as a CpG island. The pattern of cytosine modification can be identified by using a predefined sequence. For example, a method using an altered cytidine deaminase can be designed and implemented to determine the pattern of cytosine modification, e.g., the methylation status of one or more specific cytosines, the methylation status of one or more specific cytosines present at a specific position in the genome, or a combination thereof.

[0042] As used herein, the term "each" when used in reference to a collection of items is intended to identify an individual item in the collection, but not necessarily every item in the collection, unless expressly specified otherwise in the context.

[0043] As used in this specification and the appended claims, unless expressly specified otherwise in the context, the term "or" is generally employed in the sense including "and / or". The term "and / or" means one or all of the listed elements, or a combination of any two or more of the listed elements. In some cases, the use of "and / or" does not imply that in other cases the use of "or" may not mean "and / or".

[0044] Unless otherwise specified, "a", "an", "the", and "at least one" are used interchangeably and mean one or more than one.

[0045] As used in this specification and the appended claims, unless expressly specified otherwise in the context, the term "or" is generally employed in the sense including "and / or". The term "and / or" means one or all of the listed elements, or a combination of any two or more of the listed elements. In some cases, the use of "and / or" does not imply that in other cases the use of "or" may not mean "and / or".

[0046] The terms "preferred" and "preferably" refer to embodiments of the present disclosure that may provide certain benefits in certain circumstances. However, in the same or other circumstances, other embodiments may also be preferred. Additionally, the recitation of one or more preferred embodiments does not imply that other embodiments are unavailable and is not intended to exclude other embodiments from the scope of the present disclosure.

[0047] As used herein, "having", "including", "comprising", etc. are used in their open, inclusive sense and generally mean "including but not limited to".

[0048] It should be understood that wherever an embodiment is described herein in language such as "having", "including", "comprising", etc., other similar embodiments described in terms of "consisting of" and / or "consisting essentially of" are also provided. The term "consisting of" means including and limited to whatever follows the phrase "consisting of". That is, "consisting of" indicates that the listed elements are required or mandatory and that no other elements are present. The term "consisting essentially of" indicates including any elements listed after the phrase and may include other elements in addition to those listed, provided that those elements do not interfere with or contribute to the activity or action specified in the disclosure of the listed elements.

[0049] "Suitable" conditions or "appropriate" conditions for an event to occur, such as the conversion of 5-methylcytosine to thymidine by deamination, are conditions that do not prevent such an event from occurring. Thus, these conditions permit, enhance, facilitate, and / or favor the event.

[0050] As used herein, in the context of a sample of a protein, DNA, or RNA, or a composition, "providing" means preparing a sample of a protein, DNA, or RNA, or a composition, purchasing a sample of a protein, DNA, or RNA, or a composition, or obtaining a sample of a protein, DNA, or RNA, or a composition.

[0051] References throughout this specification to "one embodiment", "an embodiment", "certain embodiments", or "some embodiments", etc., mean that a particular feature, configuration, composition, or property described in connection with the embodiment is included in at least one embodiment of the present disclosure. Thus, the appearances of such phrases throughout this specification are not necessarily referring to the same embodiment of the present disclosure. Additionally, in one or more embodiments, a particular feature, configuration, composition, or property may be combined in any suitable manner.

[0052] Although polynucleotide sequences encoding altered cytidine deaminases, helicases, or fusion proteins are described herein as DNA sequences, it should be understood that the complementary, reverse, and reverse-complementary sequences of a DNA sequence can be readily determined by one of ordinary skill in the art. It should also be understood that sequences described herein as DNA sequences can be converted from a DNA sequence to an RNA sequence by replacing each thymidine nucleotide with a uracil nucleotide.

[0053] Throughout this disclosure, various aspects of the disclosure may be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the disclosure. Accordingly, the description of a range should be considered to have specifically disclosed all the possible subranges as well as the individual numerical values within that range. For example, a description of a range such as 1 to 6 should be considered to have specifically disclosed subranges (such as 1 to 3, 1 to 4, 1 to 5, 2 to 4, 2 to 6, 3 to 6, etc.) as well as the individual numbers within that range, for example 1, 2, 2.7, 3, 4, 4.5, 5, 5.3, and 6. This applies regardless of the breadth of the range.

[0054] In the description herein, specific embodiments may be described in isolation for clarity. Unless specifically stated otherwise, certain embodiments may include combinations of compatible features described herein in connection with one or more embodiments.

[0055] For any method disclosed herein that includes discrete steps, those steps may be performed in any feasible order. And, where appropriate, any combination of two or more steps may be performed simultaneously.

[0056] The foregoing summary of the disclosure is not intended to describe every disclosed embodiment or every implementation of the disclosure. The following description more specifically illustrates exemplary embodiments. Throughout the application, guidance is provided by way of lists of examples, which may be used in various combinations. In each case, the recited lists are only used as representative groups and should not be construed as exclusive lists. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] The following detailed description of the exemplary embodiments of the disclosure is best understood when read in conjunction with the following drawings.

[0058] Figures 1A to 1C show the deamination schemes of APOBEC3A. The cytosine (C), 5-methylcytosine (5mC), and 5-hydroxymethylcytosine (5hmC) nucleobases in single-stranded DNA are well-characterized substrates of APOBEC3A. Figure 1A shows the conversion of C to uracil (U) by APOBEC3A. Figure 1B shows the conversion of 5mC to thymidine (T) by APOBEC3A. Figure 1C shows the conversion of 5hmC to 5-hydroxymethyluracil (5hmU) by APOBEC3. Represents the linkage of the nucleobase to the DNA molecule.

[0059] Figures 1D to 1F Shows the results of treating a DNA sample with wild-type APOBEC3A enzyme ( Figure 1D ) or an example of one-step detection of 5mC using the altered cytidine deaminase described herein ( Figure 1E ). Figures 1D to 1F The upper strand of shows the C, 5mC, and / or 5hmC bases, and the altered bases in the lower strand of Figures 1C to Figure 1D are underlined. In Figure 1D , the 5mC nucleobase is labeled with CH 3 , the 5hmC nucleobase is labeled with CH2-OH, the 5-hydroxymethyluracil nucleobase is denoted by the lowercase letter "u", and the uracil nucleobase is denoted by the uppercase letter "U".

[0060] Figure 2 Shows examples of cytosine and modified cytosine nucleobases in DNA. Represents the linkage of the nucleobase to the DNA molecule.

[0061] Figure 3It is a schematic diagram showing the alignment of cytidine deaminase amino acid sequences using the Clustal O algorithm. The "*" (asterisk) indicates positions with a single fully conserved residue among all cytidine deaminases. The ":" (colon) indicates conservation between groups with the following strong similarity properties - roughly equivalent to a score > 0.5 in the Gonnet PAM 250 matrix. The "." (period) indicates conservation between groups with the following weak similarity properties - roughly equivalent to a score =< 0.5 and > 0 in the Gonnet PAM 250 matrix. Amino acids marked with "^" show the ZDD motif SEQ ID NO: 12 (e.g., amino acids 70 to 106 of sp|P31941|1-199 and above). Amino acids marked with "^" and "#" show the ZDD motif SEQ ID NO: 13 (e.g., amino acids 70 to 153 of sp|P31941|1-199 and above). sp|P31941|1-199 is human APOBEC3A, SEQ ID NO: 3; XP_045219544.1 is APOBEC3A from Macaca fascicularis, SEQ ID NO: 19; AER45717.1 is APOBEC3A from Pongo pygmaeus, SEQ ID NO: 20; XP003264816.1 is APOBEC3A from Nomascus leucogenys, SEQ ID NO: 21; PNI48846.1 is APOBEC3A from Pan troglodytes, SEQ ID NO: 22; and ADO85886.1 is APOBEC3A from Gorilla gorilla, SEQ ID NO: 23.

[0062] Figure 4 It shows a schematic diagram of a restriction enzyme (SwaI)-based assay for deamination by cytidine deaminase. In the case of C, X is H, and in the case of 5mC, X is methyl, and in the case of 5hmC, X is hydroxymethyl. "Matched oligonucleotides" refer to perfect complementarity between two oligonucleotides; "mismatched oligonucleotides" refer to imperfect complementarity between two oligonucleotides. The mismatch causes the double-stranded oligonucleotide to be cleaved by SwaI.

[0063] Figures 5A to 5B It shows a positive control experiment of a SwaI-based assay using synthetic oligonucleotides. Figure 5AShows the sequences of the synthesized oligonucleotides. oLB1609, SEQ ID NO: 24; oLB1610, SEQ ID NO: 25; oLB1611, SEQ ID NO: 26; oLB1612, SEQ ID NO: 27; oLB1679, SEQ ID NO: 28; oJT1910, and oJT1911. Figure 5B Shows the visualization of the SwaI digestion results. "Match" refers to the perfect complementarity between two oligonucleotides; "mismatch" refers to the imperfect complementarity between two oligonucleotides. Mismatches result in the cleavage of double-stranded oligonucleotides by SwaI.

[0064] Figure 6 Shows the positive control deamination experiment using commercially obtained APOBEC3A enzyme.

[0065] Figure 7 A to Figure 7 B shows the SDS-PAGE panel of APOBEC3A(Y130X) protein. Figure 7 A shows the SDS-PAGE analysis of purified APOBEC3A(Y130X) mutant protein. Figure 7 B shows the SDS-PAGE analysis of purified APOBEC3A(Y130AY132H) mutant protein.

[0066] Figure 8 Shows the results of the APOBEC3A(Y130X) deamination endpoint assay panel using the SwaI assay readout.

[0067] Figure 9 Shows the bar graph illustration of the APOBEC3A(Y130X) deaminase activity.

[0068] Figure 10 Shows the deaminase activities of wild-type (NEB APOBEC) and mutant APOBEC variants on C, 5mC, and 5hmC substrates. The deamination percentage values were determined by SwaI restriction enzyme assay and quantified as Figure 4 shown. The C deamination activity was measured in two independent experiments corresponding to the left and right panels.

[0069] Figures 11A to 11F show the time course of the deamination reactions of C and 5mC by APOBEC3A (Y130A). Figures 11A to 11C show the reactions at 37 °C, and Figures 11D to 11F show the reactions at 22 °C. Figures 11A and 11D are SwaI-based assays of Y130A deamination performed at 37 °C and 22 °C, respectively, as shown by 15% urea-PAGE and FAM filter. Figures 11B and 11E are graphical representations of the time course of Y130A deamination reactions of C and 5mC performed at 37 °C and 22 °C, respectively. Figures 11C and 11F are tables depicting the percentage of deamination at different time points at 37 °C and 22 °C, respectively.

[0070] Figure 12 Showed the initial Michaelis-Menten kinetics of Y130A against C and 5mC oligonucleotide substrates.

[0071] Figure 13 Showed the DNA oligonucleotide substrates used to evaluate the deaminase activity of double mutant cytidine deaminases. Group (A), SEQ ID NO: 29; Group (B), SEQ ID NO: 30, SEQ ID NO: 31, SEQ ID NO: 32, and SEQ ID NO: 36, respectively.

[0072] Figure 14 Showed the percentage of deamination at each NCN motif in DNA oligonucleotide substrate (A) after incubation with APOBEC3A mutants (37 °C, 6-hour reaction). This metric was calculated as the percentage of C>T (cytosine to thymidine) mutations at each position determined by DNA sequencing.

[0073] Figure 15 Showed the percentage of deamination at each NCN motif in DNA oligonucleotide substrate (A) after incubation with APOBEC3A mutants (37 °C, 1-hour reaction). This metric was calculated as the percentage of C>T mutations at each position.

[0074] Figure 16 Showed the percentage of deamination at each NCpGN motif in DNA oligonucleotide substrate (B) after incubation with APOBEC3A mutants. Four different DNA oligonucleotides were mixed together as substrates for APOBEC3A deamination. The percentage of deamination was calculated as the percentage of C>T mutations in each NCpGN motif. The methylated and unmethylated forms of each NCpGN motif (32 sites in total) were assayed.

[0075] Figure 17 A to Figure 17Panel B shows the time course of deamination of APOBEC3A (Y130W) against substrates containing C, 5mC, and 5hmC. A SwaI restriction enzyme assay was performed to measure the deaminase activity of APOBEC3A (Y130W). The percentage of deamination was calculated as the ratio of the intensity of the cut band to the intensity of the (uncut + cut) bands. Figure 17 Panel A is the gel image, and Figure 17 Panel B is the quantification of the band intensity.

[0076] Figure 18 Panels A through Figure 18 B show a comparison of the deaminase activities of APOBEC3A against C, 5mC, and oxidized derivatives. A SwaI restriction enzyme assay was performed for a reaction time of 90 minutes to measure the deaminase activities of the wild-type APOBEC3A enzyme and the Y130W mutant enzyme. A protein-free control was included to account for potential degradation of the oligonucleotide substrate and to account for non-specific activity of SwaI during the assay. The percentage of deamination was calculated as the ratio of the intensity of the cut band to the intensity of the (uncut + cut) bands. Figure 18 Panel A is the gel image, and Figure 18 Panel B is the quantification of the band intensity, with subtraction from the corresponding protein-free control lane.

[0077] Figure 19 A method for 5hmC detection using deaminase-based sequencing is shown.

[0078] Figure 20 Panels A through Figure 20 B show a comparison of different reaction conditions and the resulting methylation reports for pUC19 (CG methylated) and λ (fully unmethylated), where the altered cytidine deaminases have Y130A and Y132H. Figure 20 Panel A is the methylation levels of methylated pUC19 and unmethylated λ DNA obtained using the altered cytidine deaminases at different concentrations. Figure 20 Panel B is the methylation levels of methylated pUC19 and unmethylated λ DNA obtained using the altered cytidine deaminases in different buffers.

[0079] Figure 21 The effect of ribonuclease A on the deamination of methylated pUC19 and unmethylated λ as determined by sequencing is shown.

[0080] Figures 22A to 22B The activity of APOBEC Y130A - Y132H against 5hmC is shown. ( Figure 22A ) Construction strategy for the control oligonucleotides; ( Figure 22B ) Observed methylation levels for the control oligonucleotide substrates. mC, 5mC; hmC, 5hmC.

[0081] Figure 23 Shows analysis of regional methylation in CpG islands, which indicates that the deaminase-based assay generates the expected methylation profiles. mC-deaminase-seq, APOBEC3A Y130A-Y132H.

[0082] Figure 24 Shows the performance of SNV and indel calling with and without methylation conversion using APOBEC3A Y130A-Y132H. mC-deaminase-seq, APOBEC3A Y130A-Y132H.

[0083] Figure 25 A to Figure 25 F show visualization of DMRs identified by EM-Seq TM and the modified cytidine deaminase assay in the ZNF154 gene. Representation of methylation levels across the region is shown in: ( Figure 25 A) HCC2218-normal, EM-Seq TM conversion, ( Figure 25 B) HCC2218-tumor, EM-Seq TM conversion, ( Figure 25 C) HCC2218-normal, modified cytidine deaminase assay, ( Figure 25 D) HCC2218-tumor, modified cytidine deaminase assay, ( Figure 25 E) Differentially methylated regions (DMRs) called between tumor / normal samples using EM-Seq TM data, and ( Figure 25 F) Differentially methylated regions (DMRs) called between tumor / normal samples using modified cytidine deaminase data.

[0084] Figure 26 Shows re-call and precision of DMRs in HCC1187 tumor / normal and HCC2218 tumor / normal paired genomes. DMRs from each workflow were compared to DMRs called by EM-Seq TM and used as the truth set. Bisulfite identified most of the DMRs identified by EM-Seq TM . Additionally, the mC-selective deamination protocols described in Methods A, B, and C herein were able to identify most of the DMRs identified by EM-Seq TM . mC-deaminase-seq, APOBEC3AY130A-Y132H.

[0085] Figure 27 A to Figure 27 B show the use of Method A and EM-SeqTM Methylation levels detected in the promoter region.( Figure 27 A) Methylation levels at the H3K36me3 region, which is expected to be hypermethylated,( Figure 27 B) Methylation levels at the H3K27ac region, which is expected to be hypomethylated. Dotted line, methylation levels detected in the promoter region using method A; solid line, using EM-Seq TM Methylation levels detected in the promoter region.

[0086] Figure 28 Tumor signals for 0% spike-in of tumor DNA added to normal DNA and 10% spike-in of tumor DNA added to normal DNA are shown. The methylation levels of HCC2218 normal DNA at individual CpG sites within the PanSeer cancer panel were evaluated to generate a baseline. Then, the methylation levels at individual CpG sites were evaluated in separate replicates of HCC2218 normal DNA and 10% spike-in of HCC2218 tumor DNA added to HCC2218 normal DNA. Tumor signals indicate the fraction of CpGs with methylation levels significantly different from the background.

[0087] Figure 29 A schematic diagram depicting the beneficial effect of 5mC>T conversion for enrichment is shown.

[0088] Figure 30 A to Figure 30 C show different workflows for enrichment of methylated-converted libraries.( Figure 30 A) Hybridization after conversion and amplification, where specialized probe designs as shown in Figure 29 are typically employed.( Figure 30 B) Illumina Methyl Capture EPIC workflow, where standard probe designs are employed before bisulfite conversion, requiring higher DNA input.( Figure 30 C) Workflow providing data on the altered cytidine deaminase (mC-deaminase) for enrichment.

[0089] Figure 31 A to Figure 31 B show the enrichment performance of the altered cytidine deaminase (mC-deaminase-seq) libraries.( Figure 31 A) Read enrichment performance of the altered cytidine deaminase libraries compared to libraries without methylation conversion.( Figure 31 B) Correlation of regional methylation levels in CpG islands measured in unenriched (WGS) samples compared to enriched samples.

[0090] Figure 32Shows a schematic diagram of the qPCR-based detection of 5mC in the genomic locus of interest. Selective 5mC deamination by APOBEC3A (Y130A / Y132H) generates a DNA template that is fully complementary to the qPCR primer and can thus be amplified. Deamination by APOBEC3A (Y130A / Y132H) is not observed in unmethylated substrates, resulting in a mismatch between the qPCR primer and the target site and thus disrupting amplification. "Selective 5mC->T deamination" refers to the result of incubating an ssDNA substrate with the altered cytidine deaminase APOBEC3A Y130A / Y132H, and the thymidine nucleotide indicated by the underline is the result of selective deamination of 5mC. "Non-selective 5mC->T and C->U deamination" refers to the result of incubating an ssDNA substrate with wild-type APOBEC3A, and the uracil and the thymidine nucleotide indicated by the underline are the results of non-selective deamination by wild-type APOBEC3A.

[0091] Figure 33 Shows a qPCR assay using purified APOBEC proteins, which shows a decrease in the Cq value after treatment of methylated ssDNA substrates with Y130A_Y132H. NEB APOBEC, APOBEC proteins from New England Biolabs; Y130A, Y130A / Y132H, Y130A / E72A, and Y130A / Y132 / E72A are substitution mutations present in four APOBEC3A proteins. The E72A mutation abolishes APOBEC activity and is used as a negative control to show that the observed difference in Cq values is due to the mutant APOBEC enzyme.

[0092] Figure 34 Shows the detection of 5mC mediated by CRISPR-Cas12. The conversion of 5mC to T mediated by the altered cytidine deaminase restores the perfect complementarity of the Cas12 guide RNA to its DNA target, allowing the Cas12-guide RNA protein complex to bind to the converted substrate. This activates the collateral cleavage activity of Cas12, resulting in the cleavage of a reporter ssDNA containing a fluorophore and a quencher. The release of the fluorophore increases the fluorescence, which is measured in a standard fluorometer. F, fluorophore; q, quencher.

[0093] Figures 35A to 35C show that incubation of substrate DNA with APOBEC3A (Y130A / Y132H) results in high 5mC deamination and low but detectable C deamination. Figure 32The ssDNA oligonucleotide substrates described in detail were treated with enzymes and then analyzed by Illumina sequencing. The % methylation at each C or 5mC site in the oligonucleotide was calculated as the % conversion to T. In this experiment, different concentrations of APOBEC3A (Y130A / Y132H) were tested at different reaction temperatures and times. As the enzyme concentration and reaction time increased, an elevated level of C deamination was observed at 25 °C. 1 h, 3 h, and 6 h refer to the hours of incubation; 0.75 μM, 1.5 μM, and 4 μM refer to the micromolar per liter amounts of enzyme used in each reaction; and 25 °C, 30 °C, and 37 °C refer to the reaction temperatures. In each histogram, the 17 bars on the left are C deamination and the 16 bars on the right are 5mC deamination. For the micromolar per liter amounts of enzyme, the nucleotide triplets on the X-axis of the histogram are as follows: ACT, ACT, CCT, TCT, ACA, GCA, CCA, GCT, GCC, ACG, ACG, ACT, TCG, CCT, TCG, GCA, ACG, TCA, TCC, ACG, ACA, ACG, GCC, ACG, TCG, CCA, GCA, GCG, TCG, GCC, ACA, GCG, TCA.

[0094] Figure 36 Detection of 5mC in fixed cell or tissue preparations using FISH is shown. The fixed biological sample is permeabilized and denatured to make the DNA accessible to the modified cytidine deaminase. Enzymatic deamination selectively converts 5mC to T. Methylation events are detected using fluorescent probes specific for the converted DNA sequences.

[0095] Figures 37A to 37O The amino acid sequences of SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NOs: 37 - 67, SEQ ID NOs: 92 - 101, and SEQ ID NOs: 105 - 108 are shown.

[0096] Figure 38 A schematic diagram of one embodiment consistent with the present disclosure is shown. The dsDNA is unwound by a helicase fused to a cytidine deaminase that acts on 5mC. After treatment with the helicase fused to the cytidine deaminase, 5mC is converted to thymidine. This results in a T - G mismatch.

[0097] The schematic diagrams are not necessarily drawn to scale. Like reference numerals used in the figures refer to like components, steps, etc. However, it should be understood that using a number to refer to a component in a given figure is not intended to limit the component labeled with the same number in another figure. Additionally, using different numbers to refer to components is not intended to indicate that components with different numbers cannot be the same or similar to other numbered components. Detailed Description

[0098] This document describes a one-step enzymatic method for mapping modified cytosines, such as 5mC and 5hmC, at single-base resolution using a cytidine deaminase in combination with a helicase. The working examples provided herein describe APOBEC3A-based cytidine deaminases, including altered cytidine deaminases. Other APOBEC proteins are expected to be used, either as unmodified, wild-type proteins or as modified proteins as described herein. As used herein, the term "cytidine deaminase" refers to a wild-type cytidine deaminase or a cytidine deaminase comprising any of the mutations described herein. The predictive examples provided herein describe helicases used in combination with altered APOBEC proteins and protein complexes, such as fusion proteins, including helicases and altered APOBEC proteins. As used herein, the term "protein complex" refers to a grouping of more than one protein. The proteins of the protein complex may be attached, or they may be unattached. The protein complex may include a fusion protein composed of more than one protein fused together using a peptide linker. The protein complex may include a protein covalently attached to another protein using a linkage other than a peptide linker. The protein complex may include two or more unattached proteins, but they are combined for common use.

[0099] This document describes a one-step enzymatic method for mapping modified cytosines, such as 5mC and 5hmC, at single-base resolution using an altered cytidine deaminase. The working examples provided herein describe altered APOBEC3A-based cytidine deaminases, and other APOBEC proteins modified as described herein are expected to be used.

[0100] Wild-type APOBEC3A efficiently deaminates cytosine (C), 5-methylcytosine (5mC), and 4-hydroxymethylcytosine (5hmC) in single-stranded DNA (Figures 1A to 1C). Treating DNA, such as genomic DNA, with wild-type APOBEC3A results in the conversion of C to uracil (U), 5mC to thymidine (T), and 5hmC to 5-hydroxymethyluracil. These conversions reduce the complexity of the DNA to three bases for sequencing ( Figure 1D ). Point mutations were generated in the human APOBEC3A protein in a previous analysis, and the ability of the mutant APOBEC3A protein to convert cytosine to uracil was determined. Modifying the tyrosine residue at position 130 to alanine (Y130A) consistently produced an inactive APOBEC protein (see Bulliard et al., 2011, JViro1., 85(4):1765-1776 for Figure 6FIG. 5a) of C and Shi et al., 2017, Nat Struct Mol Biol., 24(2):131-139. Contrary to Bulliard and Shi, the inventors made a surprising and unexpected discovery: certain mutations at position 130 of APOBEC3A altered the deamination rate of the enzyme towards 5mC compared to the C substrate.

[0101] As described herein, the homologous tyrosine (Y) at position 130 was individually mutated to all possible canonical amino acid substitutions, including A, C, D, E, F, G, H, I, K, L, M, N, P, Q, R, S, T, V, and W, and the activities towards C, 5mC, and 5hmC substrates were evaluated. For example, it was found that the APOBEC3A mutant containing a tyrosine-to-alanine point mutation (Y130A) at position 130 preferentially deaminated 5mC rather than C (the rate of conversion of 5mC to T was greater than the rate of conversion of C to U), and it was found that the APOBEC3A mutant containing a tyrosine-to-leucine point mutation (Y130L) at position 130 preferentially deaminated C rather than 5mC (the rate of conversion of C to U was greater than the rate of conversion of 5mC to T). The deamination of 5mC to T converts C to T, which can be identified by standard sequencing methods. Thus, in one embodiment, treating DNA with the protein complex of the present disclosure preferentially converts 5mC to thymidine ( Figure 1E ). After treatment with the altered cytidine deaminase described herein, analysis of the sample DNA, e.g., by sequencing the sample DNA and optionally comparing it to a reference (e.g., a reference sequence), allows for easy identification of C-to-T point mutations, and these point mutations are inferred to be 5mC positions. Wild-type APOBEC proteins and other cytidine deaminases typically act on single-stranded (ss) DNA. For some workflows, the treatment of double-stranded (ds) DNA is preferred. Thus, to modify the cytidine deaminase to act on dsDNA, the inventors explored coupling the cytidine deaminase to a protein capable of separating the two strands of a double-stranded nucleic acid such as dsDNA. Helicases are a class of such proteins. The inventors explored different methods of attaching or complexing these two enzymes and studied many different helicases to determine their compatibility with recombinant expression. Additionally, the compatibility of the protein complex with existing sequencing workflows was determined.

[0102] Sequencing of the sample DNA after treatment with the protein complex described herein and optional comparison to a reference sequence allows for easy identification of C-to-T point mutations, and these point mutations are inferred to be 5mC positions. The protein complex described herein enables one-step modification of dsDNA samples for easy identification of 5mC.

[0103] In another example, an APOBEC3A mutant containing a tyrosine-to-tryptophan point mutation (Y130W) at position 130 maintains the ability to deaminate C and 5mC to U and T, respectively, but loses the ability to deaminate 5hmC, 5fC, and 5caC. When determining the sequence of a nucleic acid exposed to this APOBEC3A mutant, C and 5mC are deaminated to U and T, respectively, and read as T by a sequencer. On the other hand, 5hmC is not deaminated by the APOBEC3A mutant and is read as C.( Figure 1F )。5fC and 5caC are also not deaminated by this APOBEC3A mutant and thus cannot be distinguished from 5hmC. However, the abundances of 5fC and 5caC in human genomic DNA are several orders of magnitude lower than that of 5hmC and are close to the detection limit of the mass spectrometers used for such measurements (Ito et al., 2011, Science 333, 1300-1303 (2011); Wagner et al., 2015, Angew. Chem. Int. Edn Engl. 54, 12511-12514 (2015); Bachman et al., 2015, Nat. Chem. Biol. 11, 555-557 (2015)). Therefore, the signals from 5fC and 5caC should be insignificant compared to 5hmC.

[0104] Cytidine deaminase

[0105] The present disclosure provides protein complexes comprising a cytidine deaminase, fusion proteins comprising a cytidine deaminase, compositions comprising a cytidine deaminase, methods of using a cytidine deaminase, and kits comprising a cytidine deaminase. Each of these protein complexes is compatible with a wild-type cytidine deaminase or with an altered cytidine deaminase. Cytidine deaminases include apolipoprotein B mRNA editing enzyme, catalytic polypeptide-like (APOBEC), and activation-induced cytidine deaminase (AID). Wild-type APOBEC and AID cytidine deaminases have the activity of deaminating cytidine (C) in DNA and / or RNA to form uridine (U). One type of altered cytidine deaminase preferentially deaminates 5mC rather than C (i.e., converts 5mC to T at a higher rate than it converts C to U), and is referred to herein as having "cytosine-deficient deaminase activity" or "5mC-enhanced" or "5mC-selective" deaminase activity. A second type of altered cytidine deaminase preferentially deaminates C rather than 5mC (i.e., converts C to U at a higher rate than it converts 5mC to T), and is referred to herein as having "5mC-deficient deaminase activity". A third type of altered cytidine deaminase preferentially deaminates C and 5mC to U and T, respectively, and has a significantly reduced degree of deamination of 5hmC, 5fC, and 5caC. The third type is referred to herein as having "5hmC-deficient deaminase activity". Unless the context otherwise indicates, references to an altered cytidine deaminase include altered cytidine deaminases having cytosine-deficient deaminase activity, altered cytidine deaminases having 5mC-deficient deaminase activity, and altered cytidine deaminases having 5hmC-deficient deaminase activity. Unless the context otherwise indicates, references to "cytidine deaminase" include both wild-type cytidine deaminases and altered cytidine deaminases, and wild-type cytidine deaminases and altered cytidine deaminases are structurally similar to a reference cytidine deaminase. Structural similarities are described herein.

[0106] "Altered cytidine deaminases", "mutant cytidine deaminases", "modified cytidine deaminases", or "recombinant cytidine deaminases" of the present disclosure are those that comprise one or more of the substitution mutations described herein. Typically, one or more substitution mutations provide unexpected properties of an altered deamination profile, e.g., altering its ability to preferentially deaminate one form of cytosine over another.

[0107] Whether a protein has cytidine deaminase activity can be determined by an in vitro assay. One example of an in vitro assay is based on digestion with the restriction enzyme SwaI (see Example 1). A protein that can deaminate 5mC to thymidine has cytidine deaminase activity.

[0108] An altered cytidine deaminase that preferentially deaminates 5mC rather than C (i.e., has cytosine-deficient deaminase activity) can have a catalytic efficiency for a 5mC substrate that is at least 10-fold, at least 50-fold, or at least 100-fold that of a C substrate. In one embodiment, the altered cytidine deaminase that preferentially deaminates 5mC rather than C has a catalytic efficiency for a 5mC substrate that is no more than 1500-fold higher than that for a C substrate. An altered cytidine deaminase that preferentially deaminates C rather than 5mC (i.e., has 5mC-deficient deaminase activity) can have a catalytic efficiency for a C substrate that is at least 10-fold, at least 50-fold, or at least 100-fold that of a 5mC substrate. In one embodiment, the altered cytidine deaminase that preferentially deaminates C rather than 5mC has a catalytic efficiency for a C substrate that is no more than 1500-fold higher than that for a 5mC substrate.

[0109] When compared to a wild-type cytidine deaminase (an altered cytidine deaminase that deaminates C and 5mC to U and T, respectively, and has a significantly reduced degree of deamination of 5hmC (i.e., has 5hmC-deficient deaminase activity)), the degree of deamination of 5hmC by the altered cytidine deaminases disclosed herein is reduced by at least 80%, at least 90%, or at least 99% compared to the wild-type cytidine deaminase. In one embodiment, using a SwaI-based assay such as described herein, the deamination of 5hmC by the altered cytidine deaminases disclosed herein is undetectable.

[0110] In certain embodiments, the altered cytidine deaminases of the present disclosure are based on wild-type cytidine deaminases, such as members of the APOBEC protein family. An altered cytidine deaminase of the present disclosure “based on” a member of the APOBEC protein family means that the altered cytidine deaminase is an APOBEC protein that contains one or more of the substitution mutations described herein compared to a reference APOBEC sequence. An altered cytidine deaminase of the present disclosure “based on” a member of the APOBEC protein family can also contain conservative and / or non-conservative mutations as described herein.

[0111] The APOBEC protein family includes, but is not limited to, the subfamilies AID, APOBEC1, APOBEC2, APOBEC3 (including 3A, 3B, 3C, 3D, 3F, 3G, 3H), and APOBEC4. A wild-type cytidine deaminase can be a member of the APOBEC protein family or have structural similarity to a member of the APOBEC protein family. The altered cytidine deaminases of the present disclosure can be based on members of the AID subfamily, APOBEC1 subfamily, APOBEC2 subfamily, APOBEC3 subfamily (e.g., 3A subfamily, 3B subfamily, 3C subfamily, 3D subfamily, 3F subfamily, 3G subfamily, or 3H subfamily), or APOBEC4 subfamily. The wild-type cytidine deaminase can be a member of the APOBEC protein family from a vertebrate such as a mammal. Examples of mammals include, but are not limited to, rodents, primates, rabbits, bovines (e.g., cows), suids (e.g., pigs), and equines (e.g., horses). An example of a primate is a human or a chimpanzee.

[0112] The APOBEC protein family is a member of the large cytidine deaminase superfamily that contains the canonical zinc-dependent deaminase (ZDD) signature motif embedded within the core cytidine deaminase fold. The fold consists of a five-stranded mixed β(b)-sheet surrounded by six α(a) helices in the order a1-b1-b2-a2-b3-a3-b4-a4-b5-a5-a6 (Salter et al., Trends Biochem Sci. 2016 41(7):578-594. doi:10.1016 / j.tibs.2016.05.001; Salter et al., Trends Biochem. Sci. 2018, 43(8):606-622 doi.org / 10.1016 / j.tibs.2018.04.013). Each cytidine deaminase domain core structure of an APOBEC protein contains the highly conserved zinc-binding motif H-[P / A / V]-E-X [23-28] -P-C-X [2-4] -C spatial arrangement of catalytic center residues (SEQ ID NO: 12) (referred to herein as the ZDD motif, where X is any amino acid and the subscript range of the number after X refers to the number of amino acids) (Salter et al., Trends Biochem Sci. 2016 41(7):578-594. doi:10.1016 / j.tibs.2016.05.001). Without being bound by theory, the H and two C residues coordinate with the Zn atom, and the E residue polarizes a water molecule near the Zn atom for catalysis (Chen et al., 2021, Viruses, 13:497, doi.org / 10.3390 / v13030497).

[0113] Some members of the APOBEC protein family (e.g., the AID subfamily, the APOBEC1 subfamily, the APOBEC2 subfamily, the APOBEC3A subfamily, the APOBEC3C subfamily, the APOBEC3H subfamily, and the APOBEC4 subfamily) contain one copy of the ZDD motif. Other members of the APOBEC protein family (e.g., the APOBEC3B subfamily, the APOBEC3D subfamily, the APOBEC3F subfamily, and the APOBEC3G subfamily) contain two copies of the ZDD motif, but usually only the C-terminal copy is active (Salter et al., Trends Biochem Sci. 2016 41(7):578-594. doi:10.1016 / j.tibs.2016.05.001). Thus, the cytidine deaminases disclosed herein contain one or two ZDD motifs. In one embodiment, a wild-type cytidine deaminase that is a member of the APOBEC3A subfamily, or an altered cytidine deaminase based on a member of the APOBEC3A subfamily, contains the following ZDD motif: HXEX2 4 SW(S / T)PCX [2-4] CX 6 FX 8 LX 5 R(L / I)YX [8-11] LX 2 LX

[10] M(SEQ ID NO:13) (where X is any amino acid, and the subscript number or number range after X refers to the number of amino acids) (Salter et al., Trends Biochem Sci. 2016 41(7):578-594. doi:10.1016 / j.tibs.2016.05.001).

[0114] In one embodiment, a wild-type cytidine deaminase that is a member of the following subfamilies, or an altered cytidine deaminase based on a member of the following subfamilies: APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D, APOBEC3F, and APOBEC3G, may contain one or more highly conserved sites that are part of the active site and within the ZDD motif SEQ ID NO:12. These sites include tryptophan at position 98 and serine or threonine at position 99 (Kouno et al., 2017, Nat.Comm, 8:15024, DOI:10.1038 / ncomms15024).

[0115] In addition to the ZDD motif, members of the APOBEC protein family also contain other highly conserved residues that are part of the active site but do not exist as part of the ZDD motif SEQ ID NO: 12. Members of the APOBEC3A subfamily, APOBEC3B subfamily, APOBEC3C subfamily, APOBEC3D subfamily, APOBEC3F subfamily, and APOBEC3G subfamily typically contain one or more of the following highly conserved sites, which are part of the following active sites: arginine at position 28; histidine, asparagine, or arginine at position 29; serine or threonine, preferably threonine, at position 31; asparagine or aspartic acid at position 57; tyrosine or phenylalanine at position 130; asparagine or tyrosine at position 131; asparagine, tyrosine, or phenylalanine, preferably tyrosine, at position 132; and arginine or lysine at position 189 (Kouno et al., 2017, Nat.Comm, 8:15024, DOI: 10.1038 / ncomms15024).

[0116] When compared to wild-type cytidine deaminase, the altered cytidine deaminase of the present disclosure contains substitution mutations at one or more residues. The substitution mutations can be at the same position or at a functionally equivalent position compared to wild-type cytidine deaminase. Wild-type cytidine deaminase and functionally equivalent positions are described in detail herein. Those skilled in the art will readily understand that the altered cytidine deaminase described herein is not naturally occurring.

[0117] Wild-type cytidine deaminase can be a member of the APOBEC protein family. Essentially any known member of the APOBEC protein family can be wild-type cytidine deaminase. By using publicly available databases, such as the Protein Data Bank available at the National Center for Biotechnology Information in the United States (ncbi.nlm.nih.gov / protein), and searching for APOBEC 1, APOBEC2, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D, APOBEC3F, APOBEC3G, APOBEC3H, APOBEC4, or when identifying members of the AID (activation-induced cytidine deaminase family), those skilled in the art can readily identify the members of each subfamily. Wild-type cytidine deaminase has the activity of binding single-stranded DNA (ssDNA) and deaminating cytosine present on ssDNA to convert it to uracil. In one embodiment, wild-type cytidine deaminase has the activity of binding single-stranded RNA (ssRNA) and deaminating cytosine present on ssRNA to convert it to uracil. Methods for determining whether a protein binds ssDNA or ssRNA and deaminates the cytosine present are known to those skilled in the art.

[0118] In one embodiment, the cytidine deaminase has an amino acid sequence based on a reference sequence that is a member of the APOBEC protein family, such as a wild-type APOBEC protein. The altered cytidine deaminase may comprise the ZDD motif H-[P / A / V]-E-X [23-28] -P-C-X [2-4] -C (SEQ ID No: 12) and at least one substitution mutation disclosed herein. Optionally, the altered cytidine deaminase comprises other active site residues disclosed herein. Non-limiting examples of reference cytidine deaminase proteins are shown in the table below.

[0119] Table 1. Examples of members of the APOBEC protein subfamily.

[0120]

[0121]

[0122] UniProt, a database of protein sequences and functional information, is available at UniProt.org; GenBank, a collection of nucleotide sequences and their protein translations, is available at ncbi.nlm.nih.gov / protein / .

[0123] In one embodiment, the cytidine deaminase has an amino acid sequence based on a reference sequence that is a member of the APOBEC3A subfamily and includes the ZDD motif HXEX 24 SW(S / T)PCX [2-4] CX 6 FX 8 LX 5 R(L / I)YX [8-11] LX 2 LX

[10] M (SEQ ID NO: 13) (where X is any amino acid and the subscript number or range of numbers after X refers to the number of amino acids) and at least one substitution mutation disclosed herein. In one embodiment, the substitution mutation is a substitution mutation at the underlined tyrosine, such as a substitution mutation to alanine (A). Optionally, the altered cytidine deaminase comprises other active site residues disclosed herein. In one embodiment, the substitution mutation is a substitution mutation at the underlined tyrosine (Y), such as a substitution mutation to alanine (A) or tryptophan (W).

[0124] In one embodiment, the amino acid sequence of the cytidine deaminase comprises the amino acids of a member of the APOBEC3A subfamily: X [16-26] -GRXXTXLCYXV-X 15 -GXXXN-X12 -HAEXXF-X 14 -YXXTWXXSWSPC-X [2-4] -CA-X 5 -FL-X 7 -LXIXXXR(L / I)Y-X 8 -GLXXLXXXG-X 5 -M-X 4 -FXXCWXXFV-X 6 -FXPW-X 13 -LXXI-X [2-6] (SEQ ID NO: 14) (wherein X is any amino acid, and the subscript number or range of numbers after X refers to the number of amino acids present) and the altered cytidine deaminase further comprises at least one substitution mutation disclosed herein. In one embodiment, the substitution mutation is a substitution mutation at the underlined tyrosine (Y), such as a substitution mutation to alanine (A) or tryptophan (W).

[0125] In one embodiment, the amino acid sequence of the cytidine deaminase comprises a subset of the amino acids of a member of the APOBEC3A subfamily: X 26 -GRXXTXLCYXV-X 15 -G-X 16 -HAEXXF-X 14 -YXXTWXXSWSPC-X 4 -CA-X 5 -FL-X 7 -LXIFXXR(L / I)Y-X 8 -GLXXLXXXG-X 5 -M-X 4 -FXXCWXXFV-X 6 -FXPW-X 13 -LXXI-X 6 (SEQ ID NO: 15) (wherein X is any amino acid, and the subscript number or range of numbers after X refers to the number of amino acids), or a subset thereof, and at least one substitution mutation disclosed herein. In one embodiment, the substitution mutation is a substitution mutation at the underlined tyrosine (Y), such as a substitution mutation to alanine (A) or tryptophan (W).

[0126] Compared to the reference cytidine deaminase, the substitution mutation can be at the same position or a functionally equivalent position. "Functionally equivalent" means that the altered cytidine deaminase has an amino acid substitution at the amino acid position in the reference cytidine deaminase that has the same functional role in the reference cytidine deaminase and the altered cytidine deaminase.

[0127] Generally, functionally equivalent substitution mutations in two or more different cytidine deaminases occur at homologous amino acid positions in the amino acid sequences of these cytidine deaminases. Thus, the term "functionally equivalent" as used herein also encompasses mutations that are "positionally equivalent" or "homologous" to a given mutation, regardless of whether the specific function of the mutated amino acid is known. The positions of functionally equivalent and positionally equivalent amino acid residues in the amino acid sequences of two or more different cytidine deaminases can be identified based on sequence alignment and / or molecular modeling. An example of a sequence alignment for identifying positionally equivalent and / or functionally equivalent residues is shown in Figure 3 . For example, Figure 3 the residues aligned vertically among members of the APOBEC3A subfamily in Figure 3 are considered to be positionally equivalent and functionally equivalent to the corresponding residues in the human APOBEC3A amino acid sequence. Thus, for example, as

[0128] shown in

[0129] the tyrosine at residue 130 of the APOBEC3A protein of Homo sapiens, Pongo pygmaeus, Nomascus leucogenys, Pan troglodytes, and Gorilla gorilla is functionally and positionally equivalent to the tyrosine at residue 133 of the APOBEC3A protein from Macaca fascicularis. Those skilled in the art can readily identify functionally equivalent residues in cytidine deaminases.

[0130] The structural similarity of two amino acid sequences can be determined by aligning the residues of the two sequences (e.g., a candidate cytidine deaminase and a reference or wild-type cytidine deaminase as described herein) to optimize the number of identical amino acids along their sequence lengths; to optimize the number of identical amino acids, gaps are allowed in either or both sequences when performing the alignment, although the amino acids in each sequence must maintain their correct order. A candidate cytidine deaminase is a cytidine deaminase compared to a reference cytidine deaminase. A candidate cytidine deaminase that has structural similarity to a reference cytidine deaminase and cytidine deaminase activity is an altered cytidine deaminase.

[0131] Unless otherwise indicated herein that modifications have been made, pairwise comparison analysis of amino acid sequences can be performed, for example, by the local homology algorithm of Smith & Waterman, Adv. Appl. Math. 2:482 (1981); the homology alignment algorithm of Needleman & Wunsch, J. Mol. Biol. 48:443 (1970); by the method of searching for similarity by Pearson & Lipman, Proc. Nat'l. Acad. Sci. USA 85:2444 (1988); by computerized implementations of these algorithms (GAP, BESTFIT, FASTA, and TFASTA in the Wisconsin Genetics Software Package, Genetics Computer Group, 575 Science Dr., Madison, Wis.); or by visual inspection (see generally, Current Protocols in Molecular Biology, Ausubel et al., eds., Current Protocols, a joint venture between Greene Publishing Associates, Inc. and John Wiley & Sons, Inc., supplemented to 2004). An example of an algorithm suitable for determining structural similarity is the BLAST algorithm, which is described in Person, J. Mol. Biol. 215:403-410 (1990). The algorithms can be used to calculate the percent sequence identity and percent sequence similarity between two sequences. The software for performing BLAST analysis is publicly available through the National Center for Biotechnology Information.

[0132] In the comparison of two amino acid sequences, structural similarity can be expressed as a percentage of "identity" or can be expressed as a percentage of "similarity". "Identity" refers to the presence of the same amino acids. "Similarity" refers to the presence of not only the same amino acids but also conservative substitutions. Thus, in one embodiment, an amino acid sequence of a cytidine deaminase protein having sequence similarity to a reference sequence can include conservative substitutions of the amino acids present in the reference sequence.

[0133] Conservative substitutions of amino acids in a protein can be selected from other members of the class to which the amino acid belongs. For example, it is well known in the field of protein biochemistry that an amino acid belonging to a group of amino acids having a particular size or property (such as charge, hydrophobicity or hydrophilicity) can be replaced by another amino acid without altering the activity of the protein, particularly in regions of the protein not directly related to biological activity. For example, amino acids having nonpolar side chains include alanine, glycine, isoleucine, leucine, methionine, phenylalanine, proline, tryptophan and valine; amino acids having hydrophobic side chains include glycine, alanine, valine, leucine, isoleucine, proline, phenylalanine, methionine and tryptophan; amino acids having polar side chains include arginine, asparagine, aspartic acid, glutamine, glutamic acid, histidine, lysine, serine, cysteine, tyrosine and threonine; and amino acids having uncharged side chains include glycine, serine, cysteine, asparagine, glutamine, tyrosine and threonine.

[0134] Thus, as used herein, reference to a cytidine deaminase as described herein, such as reference to the amino acid sequence of one or more of the SEQ ID NOs described herein, can include proteins having at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% amino acid sequence similarity to a reference cytidine deaminase. Examples of altered cytidine deaminases having similarity to a reference amino acid sequence include those having, for example, at least 80%, at least 85%, at least 90% or at least 95% similarity to SEQ ID NO:16 and having alanine at amino acid 130. Other examples of altered cytidine deaminases having similarity to a reference amino acid sequence include those having, for example, at least 80%, at least 85%, at least 90% or at least 95% similarity or identity to SEQ ID NO:17 and having alanine at amino acid 130 and histidine at amino acid 132.

[0135] Alternatively, as used herein, reference to a cytidine deaminase as described herein, such as reference to the amino acid sequences of one or more SEQ ID NOs described herein, may include proteins having at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% amino acid sequence identity to a reference cytidine deaminase. Examples of altered cytidine deaminases having identity to a reference amino acid sequence include those having, for example, at least 80%, at least 85%, at least 90% or at least 95% similarity to SEQ ID NO: 16 and having alanine (A) at amino acid 130. Other examples of altered cytidine deaminases having identity to a reference amino acid sequence include those having, for example, at least 80%, at least 85%, at least 90% or at least 95% similarity or identity to SEQ ID NO: 17 and having alanine (A) at amino acid 130 and histidine (H) at amino acid 132.

[0136] Substitution mutation

[0137] The altered cytidine deaminases of the present disclosure contain substitution mutations at positions that are functionally equivalent to tyrosine (Y130) at position 130 in members of the APOBEC3A subfamily (e.g., SEQ ID NO: 3). Thus, members of the APOBEC3A subfamily (e.g., SEQ ID NO: 3) and another candidate cytidine deaminase from the APOBEC3A subfamily or a different APOBEC subfamily can be used to generate an alignment. In one embodiment, the candidate is selected from the APOPEC subfamily APOBEC1 or AID. An example of an algorithm that can be used to generate an alignment is Clustal O. In some APOBEC family proteins, the wild-type residue at the position functionally equivalent to Y130 is phenylalanine (F).

[0138] In another embodiment, the altered cytidine deaminases of the present disclosure contain a ZDD motif HXEX 24 SW(S / T)PCX [2-4] CX 6 FX 8 LX 5 R(L / I)YX [8-11] LX 2 LX

[10] Substitution mutations at the position of tyrosine (Y) of M (SEQ ID NO: 13). The underlined tyrosine (Y) of SEQ ID NO: 13 is at the position that is functionally equivalent to tyrosine amino acid 130 of APOBEC3A protein SEQ ID NO: 3.

[0139] In one embodiment, the substitution mutation at the position functionally equivalent to Y130 increases cytidine deaminase activity and preferentially acts on 5mC compared to cytosine (i.e., has cytosine-deficient deaminase activity). The substitution mutation can be a mutation to alanine (A), glycine, phenylalanine, histidine, glutamine, methionine, asparagine, lysine, valine, aspartic acid, glutamic acid, serine, cysteine, proline, or threonine. For example, the altered cytidine deaminase can comprise SEQ ID NO: 102, wherein X is selected from A, G, F, H, Q, M, N, K, V, D, E, S, C, P, or T (and is not Y), or can comprise SEQ ID NO: 103, wherein Z is selected from A, G, F, H, Q, M, N, K, V, D, E, S, C, P, or T (and is not Y). Preferably, in one embodiment, X or Z is A or L. In an exemplary aspect of this embodiment, the substitution mutation at the position functionally equivalent to Y130 is a mutation to alanine. Specific examples of the altered cytidine deaminase have increased activity and preferentially act on 5mC compared to cytosine (including SEQ ID NO: 16).

[0140] The modified cytidine deaminase of the present disclosure optionally comprises a position of two, three, four, or five amino acids on the C-terminal side of position Y130, or a second substitution mutation functionally equivalent to the position at Y130. In one embodiment, the second mutation is at a position of two, three, four, or five amino acids on the C-terminal side of position Y130, or a tyrosine, tryptophan, cysteine, histidine, or phenylalanine functionally equivalent to the position at Y130. In one embodiment, the second mutation is at a position functionally equivalent to the position of tyrosine (Y132) at position 132 in a member of the APOBEC3A subfamily (e.g., SEQ ID NO: 3). APOBEC proteins, such as the APOBEC3A protein, contain substitution mutations at both a first site (a position functionally equivalent to Y130) and a second site (a position of two, three, four, or five amino acids on the C-terminal side of the Y130 position), which increases the preferential activity against 5mC compared to the same APOBEC protein (such as an APOBEC3A protein containing a single substitution mutation at Y130). In one embodiment, the substitution mutation at the second position is an amino acid having a positively charged side chain and selected from arginine, histidine, lysine, or having a polar side chain and selected from glutamine. In one embodiment, the substitution mutation at the second position is histidine, such as Y132 to histidine. The double mutant containing both the first mutation and the second mutation can be any combination of any substitution mutation at a position functionally equivalent to the position of Y130 described herein and any second substitution mutation at a position of two, three, four, or five amino acids on the C-terminal side of the Y130 position described herein. For example, the modified cytidine deaminase can be SEQ ID NO: 3, SEQ ID NO: 15, or SEQ ID NO: 16, and has substitutions at Y130 and Y132, or at positions functionally equivalent to Y130 and Y132 as described herein. An example of a modified cytidine deaminase is SEQ ID NO: 104 comprising Y130X and Y132Z, where X is selected from (A), (L), or (W) (preferably (A)), and Z is selected from (R), (H), (L), or (0), preferably (H). This encompasses examples including but not limited to the following, such as Y130A and Y132R, Y130A and Y132H, Y130A and Y132L, Y130A and Y132Q, Y130L and Y132R, Y130L and Y132H, Y130L and Y132L, Y130L and Y132Q, Y130W and Y132R, Y130W and Y132H, Y130W and Y132L, Y130W and Y130Q, or any suitable combination thereof. In one embodiment, the double mutant comprises the substitution mutations Y130A and Y132H.Specific examples of altered cytidine deaminases having two substitution mutations and preferentially acting on 5mC, compared to APOBEC proteins having only a single substitution mutation at cytosine, include SEQ ID NO: 17 or sequences having at least 90%, at least 95%, at least 98%, at least 99% sequence identity to SEQ ID NO: 17 and comprising Y130A and Y132H. In another embodiment, the 5mC-selective deaminase further includes a single substitution (e.g., D133, preferably D133W) at a position functionally equivalent to position 133 in wild-type APOBEC3A (such as SEQ ID NO: 3).

[0141] One of ordinary skill in the art can confirm the 5mC or 5hmC preferential deaminase activity of the arginine, glutamine, histidine, and lysine substitution mutations at the second position in the double mutants described above. For example, double mutants are constructed to generate such an altered cytidine deaminase having a first substitution mutation (at a position functionally equivalent to Y130) and a second arginine, glutamine, histidine, or lysine substitution mutation (at the tyrosine position of two amino acids on the C-terminal side of the Y130 position), and then the deamination of C residues is evaluated in one assay and the deamination of 5mC residues is evaluated in a second assay. Using an assay such as the SwaI-based assay described herein, the rates of 5mC deamination and C deamination are compared to identify those double mutants that preferentially deaminate 5mC compared to C. One of ordinary skill in the art can similarly test double mutants having tyrosine at three, four, or five positions at the C-terminal of a position functionally equivalent to Y130 and confirm that the substitution mutation to arginine, glutamine, histidine, or lysine at that position, in combination with a mutation (such as Y130A) at a position functionally equivalent to Y130, is a double mutant that preferentially deaminates 5mC compared to C. Some embodiments provided herein relate to substitution mutations that produce 5mC-deficient deaminase activity (i.e., convert C to U at a higher rate than convert 5mC to T). In one embodiment, the substitution mutation at a position functionally equivalent to Y130 increases cytidine deaminase activity and preferentially acts on cytosine and is a mutation to an amino acid having a non-polar side chain or a hydrophobic side chain (such as leucine (L) or tryptophan (W)). In an exemplary aspect of this embodiment, the substitution mutation at a position functionally equivalent to Y130 is a mutation to leucine. Other examples of mutations that produce increased preferential deamination activity against cytosine compared to 5mC include single mutants having Y132P and double mutants having substitution mutations at Y130V and Y132H or Y130W and Y132H. Specific examples of altered cytidine deaminases having increased cytidine deaminase activity and preferentially acting on cytosine compared to 5mC include SEQ ID NO: 18 or sequences having at least 90%, at least 95%, at least 98%, at least 99% sequence identity to SEQ ID NO: 18 and containing Y130L.

[0142] In one embodiment, the substitution mutation is at a position functionally equivalent to Y130, and this substitution mutation results in 5hmC-deficient deaminase activity (i.e., preferentially deaminates C and 5mC to U and T, respectively, and has a significantly reduced degree of deamination of 5hmC). In an exemplary aspect of this embodiment, the substitution mutation at a position functionally equivalent to Y130 is a mutation to an amino acid with a non-polar side chain or a hydrophobic side chain, such as tryptophan (W). Specific examples of the altered cytidine deaminase that has the ability to deaminate C and 5mC to U and T, respectively, but has a reduced ability to deaminate 5hmC, preferably has no detectable ability to deaminate 5hmC, include SEQ ID NO: 59 or a sequence having at least 90%, at least 95%, at least 98%, at least 99% sequence identity to SEQ ID NO: 59 and containing Y130W.

[0143] The altered cytidine deaminase described herein may contain additional mutations. Generally, the additional mutations do not overly alter the activity of the altered cytidine deaminase. One or more additional mutations may be conservative mutations.

[0144] The altered cytidine deaminase described herein may also contain additional protein domains. For example, the altered cytidine deaminase may include a domain for detection, such as a fluorescent protein, an affinity tag such as a His tag, or a conjugation domain such as a SpyTag for attachment to another protein.

[0145] The altered cytidine deaminase described herein may be a truncated protein. The truncated protein is a fragment of the altered cytidine deaminase of the present disclosure that retains the ability to deaminate 5mC to thymidine. The truncated altered cytidine deaminase may contain a deletion of 1 amino acid to 13 amino acids at the N-terminus of the protein, a deletion of 1 amino acid to 3 amino acids at the C-terminus of the protein, or a combination thereof.

[0146] The altered cytidine deaminase of the present application may be the altered cytidine deaminase described in International Patent Application No. PCT / US2023 / 017846, titled "Altered Cytidine Deaminases and Methods of Use", filed on April 7, 2023, and the altered deaminase described therein is hereby incorporated by reference in its entirety. Other altered cytosine deaminases can be used in the practice of the present invention, including but not limited to, for example, hyperactive altered forms of APOBEC, such as those described in U.S. Patent No. 10,961,525, titled "Hyperactive AID / APOBEC and HMC Dominant TET enzymes", the content of the altered APOBEC is incorporated by reference herein in its entirety (e.g., but not limited to chimeric deaminases that comprise a first domain of APOBEC 4B (A3B) and a catalytic domain of APOBEC C3A (A3A), wherein the A3B protein has amino acid mutations at D196H, T197I, δ(206-210), Ins(206)GIG, R212H, Q213K, W228S, I230K, M235R, C239H, E241K, E342K, Y343H, Y350D, R351H, and E363D). Additionally, other altered APOBEC proteins are described in U.S. Patent No. 9,896,726, titled "Methods and compositions for discrimination between cytosine and modifications thereof, and for methylome analysis", the altered cytosine deaminase is incorporated by reference herein. The list of wild-type and altered cytosine deaminases that can be used in the present disclosure is not exhaustive, and it is contemplated that the present disclosure can use and contemplate other cytosine deaminases.

[0147] Helicase

[0148] The present disclosure provides compositions comprising a helicase and a cytidine deaminase (such as an altered cytidine deaminase as described herein), and methods of using a helicase and a cytidine deaminase (such as an altered cytidine deaminase as described herein) in combination. Wild-type helicases typically direct the separation of double-stranded nucleic acids during replication of the duplex. During DNA replication, each separated DNA strand is amplified using a polymerase to prepare two daughter strands from a single double-stranded parental template. There are also a variety of helicases that separate RNA and are involved in RNA metabolism.

[0149] Although there are some passive helicases, most typically require energy to disrupt the hydrogen bonds between base pairs. This energy is typically provided by ATP hydrolysis (Johnson DS, Bai L, Smith BY, Patel SS, Wang MD. Single-molecule studies reveal dynamics of DNA unwinding by the ring-shaped T7 helicase. Cell. Jun 29, 2007;129(7):1299 - 309. doi:10.1016 / j.cell.2007.04.038. PMID:17604719; PMCID:PMC2699903.). A helicase can move along a segment of double-stranded DNA, leaving a window of open single-stranded DNA. A helicase can move in the 3′ to 5′ direction or in the 5′ to 3′ direction. Helicase movement can be led by the N-terminus or the C-terminus of the protein. Some helicases localize to the negative strand of the nucleic acid, while other helicases localize to the positive strand. Each of these properties can be considered when selecting a helicase for a given function or for use in a given context.

[0150] Helicases are thought to be essential for all forms of life and have been identified in eukaryotes, prokaryotes, archaea, and viruses. There are six helicase superfamilies (SF1 - SF6), defined by gross structure and shared sequence motifs (Singleton, Martin R et al., "Structure and mechanism of helicases and nucleic acid translocases." Annual Review of Biochemistry Vol. 76 (2007): 23 - 50. doi:10.1146 / annurev.biochem.76.052305.115300). Helicases belonging to SF3 - SF6 form a loop around the nucleic acid as they process along the strand. Helicases typically form multimers, such as homohexamers. In other cases, helicases are active as monomers or dimers. Other numbers of subunits are possible and are considered, particularly for the modified proteins of the present disclosure. Although it is generally thought that, in nature, multimeric helicases exist as homotetramers or homodimers, the formation of heterotetramers and heterodimers is possible.

[0151] The helicases of the present disclosure can be mutated to alter their function. For example, the protein sequence can be mutated to produce a helicase with a faster or slower processive movement, depending on the desired activity of the resulting enzyme.

[0152] In embodiments using a helicase with more than one monomer, any suitable number of monomers can be altered, such as two, three, four, five, six, seven, or eight monomers. In an embodiment, the alteration can be present on each monomer. Alternatively, the modification can be present only on a subset of the monomers, resulting in a heteromultimeric helicase. In particular, when the helicase is modified to include large structural changes, such as in the attachment of a cytidine deaminase, the multimeric helicase assembly can have improved structural stability when only a portion of the monomers include the structural change. For example, it may be desirable to form a mutant DnaB helicase assembly in which two or three of the six monomers are fused to a cytidine deaminase. In embodiments using a dimeric helicase, one or both of the monomers can be altered.

[0153] The assembly of helicase monomers is typically random. Providing a predetermined ratio of modified monomers to unmodified monomers in a given helicase multimer can affect the number of modified monomers in each multimer. The ratio of modified monomers to unmodified monomers can be controlled by any suitable method. In embodiments where the helicase is expressed as a nucleic acid construct, different expression cassettes can be included for each unaltered and altered monomer. Expression cassette components such as promoters, polyadenylation signals, and post-transcriptional elements can be selected to achieve the desired expression levels. Each expression cassette can be provided on a single construct, such as a single plasmid. Alternatively, separate constructs with each nucleic acid expression cassette can be provided in a predetermined ratio to control the levels of each protein produced. The ratio of modified monomers to unmodified monomers can alternatively or additionally be controlled using protein modifications. Suitable protein modifications include, but are not limited to, degradation tags such as ubiquitin and localization tags such as nuclear localization signals.

[0154] In certain embodiments, the helicase is attached to the cytidine deaminase via an amino acid linker. Amino acid linkers are described herein. Nucleic acid constructs can be prepared for recombinant expression of the helicase as a fusion protein with the cytidine deaminase. From such constructs, the fusion protein can be recombinantly expressed in cells and purified from the culture medium or cell lysate. Standard methods for overexpressing recombinant proteins can be applied to prepare the fusion protein.

[0155] In an embodiment, it may be desirable to provide a cell with a nucleic acid construct encoding a fusion protein of cytidine deaminase and helicase and a nucleic acid construct encoding an unmodified helicase. The cell can then be incubated under conditions suitable for simultaneous recombinant expression of the two constructs, and the resulting proteins can be purified from the cell lysate or culture medium. In this way, a heteromeric helicase comprising a combination of modified and unmodified monomers can be obtained.

[0156] In one or more embodiments, the helicase comprises an ATP-dependent helicase. In one or more embodiments, the helicase does not comprise an ATP-dependent helicase.

[0157] In the following amino acid motifs, the character "X" is used to represent any amino acid. The character "h" is used to represent any hydrophobic amino acid, typically valine, leucine, isoleucine, phenylalanine, methionine, or tryptophan. The character "y" is used to represent any hydrophilic amino acid, typically serine, threonine, histidine, aspartic acid, glutamine, glutamic acid, lysine, arginine, or asparagine.

[0158] In one or more embodiments, the amino acid sequence of the helicase comprises an amino acid subset of a member of helicase superfamily 1, including motif I: hhXGXAGyGKS (SEQ ID NO: 70), motif Ia: XXhXXyy (SEQ ID NO: 71), motif II: hhhDEXy (SEQ ID NO: 72), motif III: hhhhGDXyQ (SEQ ID NO: 73), motif IV: xxhXyyXR (SEQ ID NO: 74), motif V: XXThXXXQGhyhyyV (SEQ ID NO: 75), motif VI: VAhTRXyy (SEQ ID NO: 76), or combinations thereof. (Singleton, Martin R et al., "Structure and mechanism of helicases and nucleic acid translocases." Annual Review of Biochemistry Vol. 76 (2007): 23-50., Hall, M C, and S W Matson. "Helicase motifs: the engine that powers DNA unwinding." Molecular microbiology Vol. 34, 5 (1999): 867-77).

[0159] In one or more embodiments, the amino acid sequence of the helicase comprises an amino acid subset of a member of helicase superfamily 2, including motif I: hhXXXyGXGKT (SEQ ID NO: 77), motif Ia: XhhhXPγy (SEQ ID NO: 78), motif II: hhhDEXH (SEQ ID NO: 79), motif III: hXhSAThhh (SEQ ID NO: 80), motif IV: hhFXXyXy (SEQ ID NO: 81), motif V: hXXTXXXXXGhyhXyh (SEQ ID NO: 82), motif VI: QXXGRXXR (SEQ ID NO: 83), or combinations thereof. (Singleton, Martin R et al "Structure and mechanism of helicases and nucleic acid translocases." Annual Review of Biochemistry Vol. 76 (2007): 23-50., Hall, M C, and S W Matson. "Helicase motifs: the engine that powers DNA unwinding." Molecular microbiology Vol. 34, 5 (1999): 867-77).

[0160] In one or more embodiments, the amino acid sequence of the helicase comprises an amino acid subset of a member of the helicase superfamily 3, including motif A: hhhXGPXGTGKS (SEQ ID NO: 84), motif B: hhXhDD (SEQ ID NO: 85), motif C: hhhTTN (SEQ ID NO: 86), or a combination thereof. (Singleton, Martin R et al. “Structure and mechanism of helicases and nucleic acid translocases.” Annual review of biochemistry Vol. 76 (2007): 23 - 50., Hall, M C, and S W Matson. “Helicase motifs: the engine that powers DNA unwinding.” Molecular microbiology Vol. 34, 5 (1999): 867 - 77).

[0161] In one or more embodiments, the amino acid sequence of the helicase comprises an amino acid subset of a member of the helicase family 4, including motif 1: XhXhhXARXXhGKT (SEQ ID NO: 87), motif 1a: VLXhSLEM (SEQ ID NO: 88), motif 2, hIhhDYL (SEQ ID NO: 89), motif 3, IXXIXXyLKAhAXyLXPhXXhXQ (SEQ ID NO: 90), motif 4, PXXXDLRXSGXIXQXADXIh (SEQ ID NO: 91), or a combination thereof. (Singleton, Martin R et al., “Structure and mechanism of helicases and nucleic acid translocases.” Annual review of biochemistry Vol. 76 (2007): 23 - 50., Hall, M C, and S W Matson. “Helicase motifs: the engine that powers DNA unwinding." Molecular microbiology Vol. 34, 5 (1999): 867 - 77).

[0162] In one or more embodiments, the amino acid sequence of the helicase comprises Walker motifs, including the Walker A motif: GXXXXGKTS (SEQ ID NO: 68) and / or the Walker B motif: RKXXXGXXXLhhhDE (SEQ ID NO: 69).

[0163] In one or more embodiments, the helicase comprises a RecQ family helicase. RecQ family helicases are a subfamily of the helicase superfamily 2 and are typically identified by conserved RecQ motifs. Examples of suitable RecQ family helicases include RecQ from Escherichia coli (E. coli) (UniProt: P15043, SEQ ID NO: 92), BLM helicase from Homo sapiens (H. sapiens) (UniProt: B7ZKN7, SEQ ID NO: 57), RecQ from Bacillus subtilis (B. subtilis) (UniProt: O34748, SEQ ID NO: 93), WRN helicase from H. sapiens (UniProt: W14191, SEQ ID NO: 94), Sgs1 helicase from Saccharomyces cerevisiae (S. cerevisiae) (UniProt: P35187, SEQ ID NO: 95), and DDM1 helicase from Arabidopsis thaliana (A. thaliana) (UniProt: Q9XFH4, SEQ ID NO: 96).

[0164] In one or more embodiments, the helicase can be of bacterial or archaeal origin. In one or more embodiments, the helicase is RadA from Streptococcus pneumoniae (S. pneumoniae) (UniProt: Q8DRP0, SEQ ID NO: 97), MCM from Methanobacterium thermoautotrophicum (UniprotO27798, SEQ ID NO: 98), or the MCM4,6,7 complex from Schizosaccharomyces pombe (S. pombe) (UniProt: P29458, P49731, O75001, SEQ ID NO: 99 - 101).

[0165] In one or more embodiments, the helicase is a mammalian helicase. Examples of suitable mammalian helicases include BLM DNA helicase from Homo sapiens (UniProt: B7ZKN7, SEQ ID NO: 57), DnaB helicase from Homo sapiens (UniProt: Q8NG08, SEQ ID NO: 58), DnaB helicase from bonobo (P. paniscus) (UniProt: A0A2R8ZJK1, SEQ ID NO: 59), DnaB helicase from Sumatran orangutan (P. abelii) (UniProt: A0A2R8ZJK1, SEQ ID NO: 60), DnaB helicase from chimpanzee (P. troglodytes) (UniProt: A0A2J8KTT1, SEQ ID NO: 105), DNA helicase B from western gorilla (G. gorilla) (UniProt: G3RJF6, SEQ ID NO: 62), DNA helicase B from cynomolgus macaque (M fascicularis) (UniProt: A1A7N9CCZ5, SEQ ID NO: 63), and DNA helicase B from red colobus monkey (P. tephrosceles) (UniProt: A0A8C9LS01, SEQ ID NO: 64).

[0166] In some embodiments, the helicase is a reference helicase or has structural similarity to a reference helicase. Examples of reference helicases include, but are not limited to, the helicase proteins of SEQ NOs: 57 - 62. As used herein, a helicase may be "structurally similar" to a reference helicase if the amino acid sequence of the helicase has a specific amount of sequence similarity and / or sequence identity compared to the reference helicase. The structural similarity of two amino acid sequences can be determined as described herein. A candidate helicase that has structural similarity to a reference helicase and has the helicase activity described herein is a helicase expected to be useful in the disclosed methods. In one embodiment, the amino acid sequence of a helicase protein having sequence similarity to a reference sequence may contain conservative substitutions of the amino acids present in that reference sequence. Conservative substitutions are described herein.

[0167] Thus, as used herein, reference to a helicase as described herein, such as reference to the amino acid sequences of one or more SEQ ID NOs of a cytidine deaminase as described herein such as SEQ ID NOs: 57 - 62, may include a protein having at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% amino acid sequence similarity to a reference UDG.

[0168] Alternatively, as used herein, reference to a helicase as described herein, such as reference to the amino acid sequences of one or more SEQ ID NOs as described herein such as SEQ ID NOs: 57 - 62, may include a protein having at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% amino acid sequence identity to a reference UDG.

[0169] Method for attaching helicase-cytidine deaminase

[0170] In certain embodiments, it may be desirable for the helicase and the cytidine deaminase to be physically close during nucleic acid processing. As described herein, the cytidine deaminases of the present disclosure (including altered cytidine deaminases) can deaminate 5mC. The catalytic activity of these proteins on single-stranded nucleic acids may be significantly higher than on double-stranded nucleic acids. In some cases, it may be desirable for the cytidine deaminases of the present disclosure to exhibit improved catalytic activity on double-stranded nucleic acids. To this end, a protein complex can be used that includes a cytidine deaminase (such as an altered cytidine deaminase) and an enzyme (such as a helicase) that can separate the two strands of a double-stranded nucleic acid. In some embodiments, the protein complex includes a cytidine deaminase attached to an enzyme (such as a helicase) that separates the two strands of a double-stranded nucleic acid. When the cytidine deaminase is spatially close to an enzyme such as a helicase, the cytidine deaminase can act on the strand separated by the helicase. In some embodiments, the region of the separated single strand may re-form a double-stranded structure after the protein complex such as a helicase-cytidine deaminase fusion has moved away.

[0171] In some embodiments, the cytidine deaminase can be attached to the helicase to form a helicase-cytidine deaminase complex. In some of these embodiments, treating a dsDNA substrate with the helicase-cytidine deaminase complex may exhibit a higher level of 5mC deamination than a comparable treatment where the helicase and the cytidine deaminase are not attached.

[0172] In some embodiments, the helicase and the cytidine deaminase can be recombinantly expressed. In some of these embodiments, the helicase and the cytidine deaminase are each expressed as a separate recombinant protein and prepared as a helicase-cytidine deaminase protein complex. In some other embodiments, the helicase and the cytidine deaminase can be expressed as a fusion protein.

[0173] In some embodiments, the helicase-cytidine deaminase complex can be a fusion protein. In some embodiments, the cytidine deaminase and the helicase can be attached N-terminus to N-terminus, N-terminus to C-terminus, or C-terminus to C-terminus. The helicase and the cytidine deaminase can be attached in any suitable order. For example, the helicase can be attached to the C-terminus of the cytidine deaminase.

[0174] In some embodiments, the helicase can be attached to a non-terminal region of the cytidine deaminase. As used herein, attachment of a protein to the "non-terminal region" of another protein describes attachment of the protein at any residue other than the N-terminal residue and the C-terminal residue. Alternatively, the cytidine deaminase can be attached to a non-terminal region of the helicase. Suitable loops can be identified by their high B-factors. In some embodiments, the helicase can be attached to a loop that is functionally equivalent to loops 60 to 67 in APOBEC3A. In some embodiments, the helicase can be attached to a loop that is functionally equivalent to loops 40 to 45 in APOBEC3A.

[0175] In embodiments where the helicase-cytidine deaminase complex is a fusion protein, the fusion protein can include any additional components. For example, the fusion protein can include a purification tag, such as a His tag, a FLAG tag, or a Myc tag. The fusion protein can be further modified to include a reactive handle, such as a free cysteine. In embodiments, the fusion protein can include a fluorescent protein or another detectable marker.

[0176] In embodiments where the helicase-cytidine deaminase complex is a fusion protein, they can be linked by any suitable linker. The physical proximity of the helicase and the cytidine deaminase may play a role in the efficiency of 5mC deamination. For example, when the cytidine deaminase and the helicase are far apart, the cytidine deaminase cannot access the ssDNA window opened by the helicase at a high frequency, resulting in a lower rate of 5mC deamination. Conversely, if the helicase and the cytidine deaminase are attached with no linker or a short linker, the two may not be able to act on the same dsDNA substrate spatially. Those of ordinary skill in the art should understand that the selection of the linker for the bifunctional fusion protein involves a considerable amount of creative effort. In some embodiments, a structured linker, such as an α-helical linker, may be preferred. In some other embodiments, an unstructured linker, such as a glycine-serine linker, is preferred. The linker can have any suitable length. For example, the linker can comprise at least two amino acids, at least four amino acids, at least six amino acids, at least ten amino acids, at least 12 amino acids, at least 16 amino acids, at least 20 amino acids, at least 30 amino acids, or at least 40 amino acids. The linker can be up to 100 amino acids, up to 80 amino acids, up to 60 amino acids, up to 40 amino acids, up to 20 amino acids, or up to ten amino acids.

[0177] In some embodiments, the helicase-cytidine deaminase complex is covalently attached. In some of these embodiments, each protein can be provided with a modification that enables attachment. For example, one protein can be attached to an avidin moiety (such as streptavidin or neutravidin), while the other protein can be attached to a biotin moiety.

[0178] The attachment moieties can be attached to each protein by any suitable method. The helicase and / or the cytidine deaminase can be recombinantly expressed as a fusion with a peptide attachment moiety. In some embodiments, the cytidine deaminase and the helicase can be obtained as separate proteins and further modified to include a peptide attachment moiety. For example, each protein can be recombinantly expressed individually, or they can be co-expressed and co-purified. Each protein can be purchased. Each protein can be purified from a natural source such as a tissue lysate. The peptide attachment moiety can include any additional protein or peptide that is generally capable of forming a covalent bond with a partner. Other attachment systems can be suitable, including systems that do not form covalent bonds but form strong bonds with a dissociation constant of less than, for example, 10 nM. Examples of suitable attachment pairs include SpyTag / SpyCatcher, SnoopTag / SnoopCatcher, biotin / avidin, and leucine zippers.

[0179] Alternatively or additionally, the helicase and / or the cytidine deaminase can be recombinantly expressed or provided as a fusion with a reactive handle. The reactive handle can be further reacted to enable the helicase to attach to the cytidine deaminase.

[0180] In embodiments where a reactive handle is used, the reactive handle can be attached to any suitable portion of a cytidine deaminase and / or a helicase. The reactive handle can be attached to the N-terminus, C-terminus, or non-terminal region of the cytidine deaminase. The reactive handle can be attached to the N-terminus, C-terminus, or non-terminal region of the helicase.

[0181] Additional considerations

[0182] In nature, when acting on dsDNA, helicases typically work in concert with additional enzymes (Seo, Yeon-Soo, and Young-Hoon Kang. “The Human Replicative Helicase, the CMG Complex, as a Target for Anti-cancer Therapy.” Frontiers in molecular biosciences Vol. 5 26. 29 Mar. 2018, doi:10.3389 / fmolb.2018.00026). In some embodiments, it may be advantageous to include additional enzymes in the compositions and methods described herein to improve helicase activity. Examples of enzymes that can be used include, but are not limited to, topoisomerases. Topoisomerases that can be used include DNA topoisomerase from Homo sapiens (UniProt: P11387, SEQ ID NO: 65) and DNA topoisomerase 2-α from Homo sapiens (UniProt: P11388, SEQ ID NO: 66).

[0183] In embodiments where an ATP-dependent helicase is used, it may be advantageous to include a high concentration of ATP. In some embodiments, the helicase may require the use of additional cofactors. Any suitable cofactor can be used to improve the activity of the helicase.

[0184] Polynucleotide encoding cytidine deaminase and helicase

[0185] The cytidine deaminases, helicases, and fusion proteins described herein can also be identified based on the polynucleotides encoding the proteins. Accordingly, the present disclosure provides polynucleotides encoding the cytidine deaminases, helicases, or fusion proteins described herein. Alternatively, the present disclosure provides polynucleotides that hybridize under standard hybridization conditions to the polynucleotides encoding the cytidine deaminases, helicases, or fusion proteins described herein, as well as the complements of such polynucleotide sequences.

[0186] The polynucleotides as described herein can include any polynucleotide encoding a cytidine deaminase (e.g., an altered cytidine deaminase), helicase, fusion protein, or combination thereof of the present disclosure. Thus, the nucleotide sequence of a polynucleotide can be inferred from the amino acid sequence encoded by the polynucleotide. The proteins of the present disclosure can be encoded by multiple codons, and certain translation systems (e.g., prokaryotic or eukaryotic cells) often exhibit codon bias. For example, different organisms often prefer one of several synonymous codons that encode the same amino acid. Thus, the polynucleotides provided herein are optionally "codon-optimized," meaning that the synthetic polynucleotide is made to include the codons preferred by a particular translation system for expressing the protein. For example, when it is desired to express a protein in a bacterial cell (or even a particular bacterial strain), the polynucleotide can be synthesized to include the codons most common in the genome of that bacterial cell for efficient expression of the altered cytidine deaminase. When it is desired to express the altered cytidine deaminase in a eukaryotic cell, a similar strategy can be employed. For example, the nucleic acid can include the codons preferred by that eukaryotic cell.

[0187] The polynucleotides described herein can also advantageously be included in a suitable expression vector for expressing the altered cytidine deaminase encoded thereby in a suitable host. Incorporating cloned DNA into a suitable expression vector for subsequent transformation of a host cell and subsequent selection of the transformed cells is well known to those skilled in the art, as provided in Sambrook et al. (1989), Molecular cloning: A Laboratory Manual, Cold Spring Harbor Laboratory. Suitable host cells include, but are not limited to, Escherichia coli (E. coli) and Saccharomyces cerevisiae.

[0188] Such expression vectors include vectors having the polynucleotides described herein, the nucleic acid being operably linked to a heterologous regulatory sequence capable of effecting the expression of the DNA fragment, such as a promoter region. The term "operably linked" refers to a juxtaposition in which the components described are in a relationship that permits them to function in their intended manner. Such vectors can be transformed into suitable host cells to provide expression of the altered cytidine deaminase.

[0189] The nucleic acid molecule can encode a mature protein or a protein having a presequence, including a nucleic acid encoding a leader sequence on a preprotein, which is then cleaved by the host cell to form the mature protein. The vector can be, for example, a plasmid, virus, or phage vector that is provided with an origin of replication, and optionally a promoter for expressing the nucleotide and optionally a regulator of the promoter. The vector can contain one or more selectable markers, such as, for example, antibiotic resistance genes.

[0190] The regulatory elements required for expression include a promoter sequence for binding RNA polymerase and directing an appropriate level of transcriptional initiation, and a translation initiation sequence for ribosome binding. For example, a bacterial expression vector can include a promoter such as the lac promoter, and the Shine - dalarno sequence and the start codon AUG for translation initiation. Similarly, a eukaryotic expression vector can include a heterologous or homologous promoter for RNA polymerase II, a downstream polyadenylation signal, the start codon AUG, and a stop codon for dissociating ribosomes. Such vectors can be obtained commercially or assembled from sequences described by methods well known in the art.

[0191] Transcription of the DNA encoding the altered cytidine deaminase can be optimized by including enhancer sequences in the vector. Enhancers are cis - acting elements of DNA that act on promoters to increase the level of transcription. Vectors generally also include an origin of replication in addition to a selectable marker.

[0192] Preparation and isolation of cytidine deaminase, helicase, and fusion protein

[0193] Generally, polynucleotides encoding cytidine deaminase, helicase, or fusion proteins as provided herein can be prepared by cloning, recombination, in vitro synthesis, in vitro amplification, and / or other available methods. A variety of recombinant methods can be used to express expression vectors encoding cytidine deaminase, helicase, or fusion proteins as provided herein. Methods for preparing recombinant polynucleotides, expressing, and isolating the expression products are well known in the art and have been described.

[0194] Polynucleotides encoding wild - type cytidine deaminase can be obtained from sources and mutagenized to introduce the substitution mutations described herein. Generally, any available mutagenesis procedure can be used to prepare the cytidine deaminases described herein, particularly the altered cytidine deaminases. Methods that can be used include, but are not limited to: site - directed mutagenesis, in vitro or in vivo homologous recombination, oligonucleotide - directed mutagenesis, mutagenesis by total gene synthesis, and many other methods known to those skilled in the art.

[0195] Polynucleotides encoding wild - type helicase can be obtained from sources and mutagenized to introduce the substitution mutations described herein. Generally, any mutagenesis procedure can be used to prepare the helicases described herein. Methods that can be used include, but are not limited to: site - directed mutagenesis, in vitro or in vivo homologous recombination, oligonucleotide - directed mutagenesis, mutagenesis by total gene synthesis, and many other methods known to those skilled in the art.

[0196] Polynucleotides encoding the fusion proteins described herein can be prepared using any available molecular cloning technique. Preferably, polynucleotides encoding fusions of helicases and cytidine deaminases can be prepared using scarless or semi-scarless gene assembly techniques. Examples of suitable techniques include, but are not limited to, Golden Gate assembly, Gibson assembly, and restriction enzyme ligation. Template polynucleotides encoding helicases and / or cytidine deaminases can be obtained from sources as described herein.

[0197] Polynucleotides encoding additional protein components required for the preparation of cytidine deaminases or helicases can be obtained from a source and mutagenized or cloned to prepare the altered proteins described herein. Examples of additional protein components include linkers, affinity tags, protein conjugation tags, and additional enzymes.

[0198] Additional useful references regarding mutagenesis, recombination, and in vitro nucleic acid manipulation methods (including cloning, expression, PCR, etc.) include Berger and Kimmel, Guide to Molecular Cloning Techniques, Methods in Enzymology Volume 152 Academic Press, Inc., San Diego, Calif. (Berger); Kaufman et al. (2003) Handbook of Molecular and Cellular Methods in Biology and Medicine Second Edition Ceske (ed.) CRC Press (Kaufman); The Nucleic Acid Protocols Handbook Ralph Rapley (ed.) (2000) Cold Spring Harbor, Humana Press Inc (Rapley); Chen et al. (eds.) PCR Cloning Protocols, Second Edition (Methods in Molecular Biology, Volume 192) Humana Press; and Viljoen et al. (2005) Molecular Diagnostic PCR Handbook Springer, ISBN 1402034032.

[0199] In addition, many commercially available kits are available for the purification of plasmids or other relevant nucleic acids from cells. The isolated polynucleotides can be further manipulated to generate other polynucleotides for transfection or transformation of cells, incorporation into relevant vectors and introduction into cells for expression, etc. Typical cloning vectors contain transcription and translation terminators, transcription and translation initiation sequences, and promoters that can be used to regulate the expression of specific target nucleic acids. The vector optionally includes universal expression cassettes that contain at least one independent terminator sequence, sequences that allow replication of the cassette in eukaryotes or prokaryotes or both (e.g., shuttle vectors), and selectable markers for both prokaryotic and eukaryotic systems. The vector is suitable for replication and integration in prokaryotes, eukaryotes, or both.

[0200] Other useful references, such as for cell isolation and culture (e.g., for subsequent nucleic acid isolation), include Freshney (1994) Culture of Animal Cells, a Manual of Basic Technique, Third Edition, Wiley-Liss, New York and references cited therein; Payne et al. (1992) Plant Cell and Tissue Culture in Liquid Systems John Wiley & Sons, Inc. New York, N.Y.; Gamborg and Phillips (eds.) (1995) Plant Cell, Tissue and Organ Culture; Fundamental Methods Springer Lab Manual, Springer-Verlag (Berlin Heidelberg New York); and Atlas and Parks (eds.) The Handbook of Microbiological Media (1993) CRC Press, Boca Raton, Fla. Standard ligation techniques known in the art are used to construct vectors containing nucleic acids encoding the cytidine deaminases described herein. See, for example, Sambrook et al., Molecular Cloning: A Laboratory Manual., Cold Spring Harbor Laboratory Press (1989) or Ausubel, R.M. ed. Current Protocols in Molecular Biology (1994).

[0201] A variety of protein separation and detection methods are known and can be used to isolate cytidine deaminase, helicase, or a fusion protein (e.g., from a recombinant culture of cells expressing the recombinant proteins provided herein). A variety of protein separation and detection methods are well known in the art and include, for example, those described in the following: R. Scopes, Protein Purification, Springer-Verlag, N.Y. (1982); Deutscher, Methods in Enzymology Volume 182: Guide to Protein Purification, Academic Press, Inc. N.Y. (1990); Sandana (1997) Bioseparation of Proteins, Academic Press, Inc.; Bollag et al. (1996) Protein Methods, 2nd Edition Wiley-Liss, NY; Walker (1996) The Protein Protocols Handbook Humana Press, NJ, Harris and Angal (1990) Protein Purification Applications: A Practical Approach IRL Press at Oxford, Oxford, England; Harris and Angal Protein Purification Methods: A Practical Approach IRL Press at Oxford, Oxford, England; Scopes (1993) Protein Purification: Principles and Practice 3rd Edition Springer Verlag, NY; Janson and Ryden (1998) Protein Purification: Principles, High Resolution Methods and Applications, Second Edition Wiley-VCH, NY; and Walker (1998) Protein Protocols on CD-ROM Humana Press, NJ; and the references cited therein. Additional details on protein purification and detection methods can be found in Handbook of Bioseparations, Academic Press (2000), edited by Satinder Ahuja.

[0202] A cytidine deaminase, helicase, or fusion protein can be prepared as an isolated protein or polynucleotide. An "isolated" protein or polynucleotide is a protein or polynucleotide that has been removed from a cell. For example, an isolated protein is a polypeptide that has been removed from the cytoplasm or from the cell membrane, and many other cellular materials of the protein, nucleic acid, and its natural environment are no longer present. A protein produced extracellularly by chemical or recombinant means is by definition considered isolated because it has never been present in a cell.

[0203] Method of use

[0204] The cytidine deaminases, helicases, and fusion proteins provided by the present disclosure can be readily incorporated into substantially any application for the identification of modified cytosines. For example, the altered cytidine deaminases, helicases, and fusion proteins can be incorporated into applications including sequencing library preparation. Examples of sequencing library preparation include, but are not limited to, whole genome, accessibility (e.g., ATAC), conformational state (e.g., HiC), and reduced representation bisulfite sequencing (RRBS). It can be particularly useful for substantially any application using low input DNA or RNA, such as, but not limited to, single cell combinatorial indexing (sci) methods such as sci-WGS-seq, sci-MET-seq, and sci-ATAC-seq, sci-RNA-seq, and cell-free DNA-based methods. Specific applications include, but are not limited to, identifying one or more patterns of cytosine modification, such as determining methylation on CpG islands (Example 15) and reduced representation bisulfite sequencing (RRBS); variant calling, including SNV / indel, copy number variation (CNV), short tandem repeat (STR), and structural variant (SV) (Example 16); detecting differentially methylated regions (DMRs) (Example 17); measuring methylation at promoters (Example 18); and detecting tumor DNA (Example 19).

[0205] The altered cytidine deaminases provided by the present disclosure can be readily incorporated into substantially any application including locus-specific methylation profiling. Typical locus-specific detection of epigenetically methylated cytosines, such as 5mC, requires the use of 5mC-specific antibodies or multi-step chemical or chemoenzymatic conversions that deaminate C or 5mC to U / T in order to be able to distinguish between the two C-isotypes. When combined with various in vitro detection methods, these methods can be powerful methods for detecting 5mC at defined loci. However, these methods can be confounded by antibody cross-reactivity and stability, or the toxicity and complex workflows required for chemical and chemoenzymatic methods. Using the altered cytidine deaminases described herein in a single enzymatic deamination protocol allows for the selective conversion of 5mC to T, which is compatible with many in vitro diagnostic modalities, enabling locus-specific detection of 5mC.

[0206] As an alternative to using destructive methods to identify methylated cytosines, the cytidine deaminases, helicases, and fusion proteins provided by this disclosure are incorporated into methods for identifying modified cytosines, such as sequencing library generation and locus-specific methylation profiling, such that more efficient enzymatic catalytic conversion of modified cytosines occurs during the production of the target nucleic acid, thereby allowing for more efficient detection of modified cytosines, particularly those from double-stranded nucleic acids. The altered cytidine deaminases, helicases, and fusion proteins of this disclosure also provide better sequencing data and better retention of genetic information, as demonstrated by high variant calling performance (see Example 16). Additionally, as an enzymatic method of conversion, the use of the altered cytidine deaminase enables high coverage uniformity and low sample damage, which can result in lower nucleic acid input requirements. A variety of sequencing library methods that can be used to construct whole genome or targeted libraries are known to those of skill in the art.

[0207] Generally, methods for using the cytidine deaminases of this disclosure include contacting a target nucleic acid (e.g., DNA or RNA) with the enzyme under conditions suitable for converting a modified cytidine, such as 5mC, to thymidine or suitable for converting an unmodified cytidine to uracil. Because the amplification of DNA does not maintain the modified state of cytidine (e.g., does not maintain the methylation state of 5mC and 5hmC), the use of cytidine deaminases typically occurs prior to the amplification of the target DNA. So long as the DNA is double-stranded, the target nucleic acid can be contacted with the cytidine deaminase at substantially any time in the pre-amplification method. For example, after isolating genomic DNA or mRNA or cell-free DNA or mRNA, before or after fragmentation, or before or after tagmentation, when the nucleic acid is within fixed or unfixed cells or nuclei, the target nucleic acid can be contacted with the cytidine deaminase. Those of skill in the art will recognize that the target nucleic acid can be contacted with the cytidine deaminase after the addition of universal sequences and / or adapters, provided that the universal sequences and / or adapters are not added by amplification.

[0208] Methods using the altered cytidine deaminases can include the optional step of comparing the treated target nucleic acid to an untreated nucleic acid or comparing the treated target nucleic acid to a nucleic acid treated with wild-type cytidine deaminase. For example, in an embodiment where the treated nucleic acid is sequenced, the sequence can be compared to a reference sequence, allowing for easy identification of point mutations and inference of modified cytosines. Thus, in an embodiment using an altered cytidine deaminase with cytosine-deficient deaminase activity (i.e., converting 5mC to T at a higher rate than converting C to U), C to T point mutations can be easily identified and these point mutations are inferred to be at 5mC positions. In an embodiment using an altered cytidine deaminase with 5hmC-deficient deaminase activity (i.e., preferentially deaminating C and 5mC to U and T, respectively, and having a significantly reduced degree of deamination of 5hmC), the absence of C to T point mutations can be easily identified and the absence of point mutations is inferred to be at 5hmC positions. In an embodiment where the treated nucleic acid is not sequenced, the nucleic acid can be treated with the altered cytidine deaminase and the nucleic acid can be compared to an untreated (i.e., not contacted with the altered cytidine deaminase) nucleic acid. Here, the readout generally depends on the assay method, e.g., when amplification is used (e.g., Example 21), the relative amount of amplification can be easily identified and the presence or absence of a pattern of 5mC or cytosine modification at a predetermined sequence can be inferred.

[0209] Reaction conditions suitable for the conversion of modified cytosines (such as converting 5mC to thymidine) by the cytidine deaminases described herein include, but are not limited to: a substrate of a double-stranded (ds) DNA or RNA target nucleic acid suspected of containing at least one modified cytosine, pH, temperature of the reaction, time of the reaction, and the concentration of the cytidine deaminase and / or the ds DNA or RNA substrate.

[0210] Target nucleic acids that can be used in the methods of the present disclosure are described herein. Modified cytosines present on the substrate double-stranded (ds) DNA or RNA include, but are not limited to, 5-methylcytosine (5mC), 5-hydroxymethylcytosine (5hmC), 5-formylcytosine (5fC), and 5-carboxylcytosine (5CaC)( Figure 2 ). In one embodiment, the modified cytosine is 5-methylcytosine. Methods for generating sequencing libraries using double-stranded target DNA can be modified to include treatment with the cytidine deaminases described herein. In some embodiments, dsDNA used for tag fragmentation reactions or for adapter attachment can be treated with cytidine deaminase.

[0211] In some embodiments, the target nucleic acid that can be used in the methods of the present disclosure may contain additional motifs. For some helicases of interest, certain motifs may preferably improve the recruitment of the helicase to the target nucleic acid. These types of motifs may be referred to herein as "helicase recruitment motifs". Examples of helicase recruitment motifs that may be included in the target dsDNA sequence include Y-shaped adaptor structures, partial duplexes, ssDNA bubbles, forks, flat duplexes, triple junctions, etc. In some embodiments, the adaptor may include more than one motif to recruit the helicase. The 3' end or the 5' end of the target nucleic acid may contain a motif. In some embodiments, both the 3' end and the 5' end of the target nucleic acid contain motifs.

[0212] In some embodiments, the cytidine deaminases provided herein can be used to distinguish 5-methylcytosine (5mC) and 5-hydroxymethylcytosine (5hmC). In such embodiments, a DNA sample suspected of containing single-stranded DNA comprising at least one 5-methylcytosine (5mC) or 5-hydroxymethylcytosine (5hmC) is modified to prevent the cytidine deaminase from converting 5hmC to thymidine. Methods of blocking deaminase activity are known in the art, and any of a variety of methods can be used to protect 5hmC from deaminase activity. As an example, the target DNA can be treated to modify 5hmC but not 5mC such that 5hmC is an unsuitable substrate for cytidine deaminase activity. In a specific example, a glucosyltransferase can be used to glucosylate 5hmC but not 5mC. Glucosyltransferases are known to those of skill in the art and include, for example, β-glucosyltransferase (βGT). For example, the enzyme T4 β-glucosyltransferase is commercially available (βGT, NEB) and can be used for the modification of 5hmC. Methods of using βGT to glucosylate 5hmC are known in the art and can be used in conjunction with the use of the cytidine deaminases provided herein. For example, a DNA sample can be treated with βGT prior to treatment with the cytidine deaminase to glucosylate 5hmC in the sample DNA. By treating the sample DNA with βGT, 5hmC can be protected from the deaminase activity of the cytidine deaminase. Thus, 5mC will be detected as thymidine in downstream readouts such as sequencing, PCR, arrays, etc. In contrast, any protected 5hmC sites will be detected as cytosine in the same readout. Enzymes, buffers, and conditions for performing the glucosylation of 5hmC are known in the art, such as Schutsky et al., Nature biotechnology, 10.1038 / nbt.4204. October 8, 2018, doi: 10.1038 / nbt.4204. Specific examples of modifying 5hmC to protect it from cytidine deaminase activity are provided in Example 9, Example 10, and Example 11. In some embodiments, the cytidine deaminases provided herein can be used in in vitro diagnostic (IVD) methods for profiling methylation in a locus-specific manner. Currently available methods for detecting methylation biomarkers generally involve digesting genomic DNA with methylation-sensitive enzymes followed by quantitative PCR (qPCR) at the locus of interest to quantify the extent of restriction enzyme digestion and thus the percentage of methylation at that locus. Then, mismatch-sensitive qPCR of bisulfite-treated DNA is performed, where 5mC is read out as a lack of 5mC>T conversion. However, these methods have drawbacks. The recognition site of the methylation-sensitive restriction enzyme must be present in the methylated region of the target locus.Bisulfite treatment requires large amounts of starting DNA and results in conversion to a low-complexity genome (unmethylated cytosines - which represent the majority of cytosines in the genome - are converted to U and read as T). This reduced complexity of the genomic template limits the design of qPCR primers for locus-specific hybridization to the gene of interest. During bisulfite conversion, DNA is inherently damaged or lost, which can impede downstream analysis. DNA damage reduces the coverage uniformity of the genome, which can lead to biased coverage. In addition, incomplete bisulfite conversion has the potential to adversely affect the results because it can inflate DNA methylation levels (Sam et al., PLoS One. 2018; 13(6); Ehrich et al., Nucleic Acids Res. Oxford University Press; 2007; 35: e29).

[0213] Conversion of 5mC to T by the altered cytidine deaminases described herein eliminates the need for restriction enzymes or bisulfite treatment and preserves DNA complexity. Modifications of one or more cytosines generated can be detected using established in vitro diagnostic (IVD) methods for profiling methylation in a locus-specific manner. Examples of methods include detecting 5mC loci by amplification, such as quantitative PCR (qPCR) (Example 21), using CRISPR-based systems to detect 5mC loci, such as CRISPR-Cas12 (Example 22), spatially detecting 5mC using molecular cytogenetic methods, such as fluorescence in situ hybridization (FISH) (Example 23), and array-based 5mC detection (Example 24). In one embodiment, an in vitro diagnostic (IVD) method for locus-specific methylation profiling uses one or more primers that anneal to a predetermined sequence (which may contain one or more modified cytosines). After treatment of the target nucleic acid with the altered cytidine deaminase, modified cytosines present in the target nucleic acid are converted as described herein (e.g., 5mC is converted to T), and when the predetermined sequence contains a nucleotide generated by deaminase treatment (e.g., a T nucleotide in the case where 5mC was present prior to treatment), primers can be readily designed to anneal to the predetermined sequence with higher affinity. For example, as described in Example 21, primers for amplification bind with greater affinity to a nucleic acid containing a T nucleotide, where a 5mC nucleotide was present prior to treatment. Annealing of the primer to the predetermined sequence (containing the expected 5mC to T conversion) allows inference of the position of the modified cytosine in the untreated target nucleic acid. In the case where a 5mC nucleotide was present prior to treatment, a primer that binds with higher affinity to a nucleic acid containing a T nucleotide can contain at least 1, at least 2, at least 3, at least 4, or at least 5 nucleotides that will base pair with the nucleotide generated by the conversion of 5mC to T (i.e., adenine (A)), and when using amplification, then a second primer for the reverse strand with T instead of guanine (G).

[0214] In some embodiments, a target nucleic acid obtained from a subject can be treated with an altered cytidine deaminase to produce a converted nucleic acid, and a pattern of cytosine modification can be identified in the converted nucleic acid. The pattern of cytosine modification can optionally be compared to a pattern of cytosine modification in a reference nucleic acid. In embodiments where the pattern of cytosine modification is associated with a disease or disorder, the method can be used for diagnostic or prognostic applications. For example, a subject can have a disease or disorder or be at risk of having a disease or disorder, and the reference nucleic acid can be from a normal subject, e.g., a subject who does not have a disease or disorder and is not at risk of a disease or disorder. The pattern of cytosine modification can be associated with a disease or disorder (e.g., the target nucleic acid can be a predetermined sequence), and identification of a pattern of cytosine modification in the subject that is associated with the disease or disorder can indicate that the subject has a disease or disorder or is at risk of having a disease or disorder. For example, the pattern of cytosine modification can be cis-linked to a coding region associated with a disease or disorder, and identification of the pattern or absence of the pattern in the subject can be used for diagnosis or prognosis. In one embodiment, the coding region can be a region that has transcriptional activity or does not have transcriptional activity in the reference nucleic acid. Comparison of the converted nucleic acid to the reference nucleic acid can include determining whether the pattern of cytosine modification in the converted nucleic acid indicates that the coding region in the subject has transcriptional activity or does not have transcriptional activity. When the coding region is associated with a disease or disorder, the status of transcriptional activity can be used for diagnosis or prognosis.

[0215] Comparison of the pattern of cytosine modification in a subject can also be used to identify changes in the pattern of cytosine modification in the subject over time. For example, a subject can have a disease or disorder and be undergoing treatment, or a subject can have had a disease or disorder and been cured (e.g., the subject has been treated and there are no signs of the disease or disorder) or be in remission (e.g., the subject has been treated and the signs of the disease or disorder have decreased). Target nucleic acids from the subject at different times (e.g., before treatment begins, during treatment, after treatment has stopped) can be compared, and the pattern of cytosine modification of a sequence (e.g., a predetermined sequence) can be compared, and the pattern can be used to determine the progress of treatment or the status of the subject's disease or disorder.

[0216] In some embodiments where detection of 5mC nucleotides uses amplification, the use of a polymerase that does not prefer uracil can help reduce amplification of treated target nucleic acids that contain potential false C-to-U conversions resulting from the use of an altered cytidine deaminase. B-family polymerases are known to exhibit "uracil pre-reading" which results in the polymerase stalling at uracil residues (Greagg et al., 1999, PNAS USA; 96(16):9045-50). Examples of uracil-non-preferring B-family polymerases include archaeal B-family polymerases from Pyrococcus furiosus (Pfu), Thermococcus kodakarensis (KOD), Thermococcus litoralis (Tli / Vent), Pyrococcus woesei (Pwo), and Thermococcus fumicolans (Tfu). Other examples of uracil-non-preferring polymerases include Phusion TM , and Kapa HiFi TM . In other embodiments where amplification of nucleic acids containing uracil nucleotides is desired, uracil-tolerant polymerases can be used. Examples of uracil-tolerant polymerases include Phusion UTM, Kapa UTM, Taq, and Dpo4.

[0217] Wild-type cytidine deaminases typically function at near-neutral pH (e.g., pH 7). The altered cytidine deaminases described herein can have increased activity at sub-neutral pH. In some embodiments, the pH of a reaction comprising an altered cytidine deaminase described herein can be no greater than pH 6.9, no greater than pH 6.7, no greater than pH 6.5, no greater than pH 6.3, no greater than pH 6.1, no greater than pH 6.0, no greater than pH 5.9, no greater than pH 5.7, no greater than pH 5.5, no greater than pH 5.3, no greater than pH 5.2, or no greater than pH 5.1. In some embodiments, the pH of a reaction comprising an altered cytidine deaminase described herein can be at least pH 5.1, at least pH 5.3, at least pH 5.5, at least pH 5.7, at least pH 5.9, at least pH 6.1, at least pH 6.3, at least pH 6.5, at least pH 6.7, or at least pH 6.9. In some embodiments, the pH of a reaction comprising an altered cytidine deaminase described herein can be no greater than pH 7.5, no greater than pH 7.3, or no greater than pH 7.1. Examples of pH ranges in the reaction include at least 5.1 to no greater than 6.9, at least 5.1 to no greater than 6.5, at least 5.1 to no greater than 6.3, or at least 5.1 to no greater than 6.1. The activity of the altered cytidine deaminase to deaminate 5mC oligonucleotide substrates can be increased catalytic activity, which is at least 10-fold greater, at least 50-fold greater, or at least 100-fold greater when activity is compared (see Example 3).

[0218] It is expected that cytidine deaminases, including altered cytidine deaminases, can function in substantially any buffer. Examples of useful buffers include, but are not limited to: citrate buffer, such as the citrate buffer from Thermo Fisher Scientific (catalog number #005000); sodium acetate buffer, Bis Tris-propane HCl; and Tris-HCl Tris. Examples of other buffers include, but are not limited to, Bicine, DIPSO, glycylglycine, HEPES, imidazole, malonate, MES, MOPS, PB, phosphate, PIPES, SPG, succinate, TAPS, TAPSO, tricine. In some embodiments, a reducing agent (such as dithiothreitol (DTT)) can be present. In some embodiments, divalent cations are not included.

[0219] The deamination reaction can be carried out at a temperature of 25 °C to 37 °C (such as 37 °C). Some of the cytidine deaminases described herein preferentially deaminate modified cytosine to thymine at a faster rate than cytosine deaminates to uracil. Thus, in some embodiments, the reaction time can be used to maximize the difference between the deamination of modified cytosine and the deamination of cytosine. In one embodiment, the reaction can be carried out for at least 15 minutes, at least 30 minutes, at least 45 minutes, at least 60 minutes, at least 90 minutes, at least 120 minutes, or at least 150 minutes, and not more than 15 minutes, not more than 30 minutes, not more than 45 minutes, not more than 60 minutes, not more than 90 minutes, not more than 120 minutes, not more than 150 minutes, or not more than 180 minutes.

[0220] In one embodiment, the deamination reaction can comprise a cytidine deaminase at a concentration of at least 0.5 micromolar (μM) to not more than 5 μM. For example, the concentration of the enzyme can be at least 0.5 μM, at least 1 μM, at least 2 δμM, at least 3 μM, at least 4 μM, or 5 μM, and / or not more than 5 μM, not more than 4 μM, not more than 3 δμM, not more than 2 μM, not more than 1 μM, or 0.5 μM. In one embodiment, the deamination reaction can comprise a nucleic acid at a concentration of at least 1 picomolar (pM) to not more than 2 μM. For example, the concentration of the nucleic acid can be at least 1 pM, at least 3 pM, at least 6 pM, at least 10 pM, at least 100 pM, at least 1 nanomolar (nM), at least 40 nM, at least 400 nM, at least 500 nM, at least 600 nM, at least 700 nM, at least 800 nM, at least 900 nM, or 1 δμM, and / or not more than 1 μM, not more than 900 nM, not more than 800 nM, not more than 700 nM, not more than 600 nM, not more than 500 nM, not more than 400 nM, not more than 40 nM, not more than 1 nM, not more than 100 nM, not more than 6 nM, or not more than 3 nM.

[0221] In one embodiment, the deamination reaction can include ribonuclease. Ribonuclease A is involved in an increase in the activity of cytidine deaminase (Bransteitter et al., Proceedings of the National Academy of Sciences of the United States of America 100, no. 7 (2003): 4102-7. doi.org / 10.1073 / pnas.0730835100). The opposite was observed when determining the activity of the altered cytidine deaminase of the present disclosure in the presence of ribonuclease A. When ribonuclease A is included in the reaction, the altered cytidine deaminase with cytosine-deficient deaminase activity (i.e., converting 5mC to T at a higher rate than converting C to U) has reduced activity, and for off-target cytosine deamination, the reduced activity is more significant. Thus, ribonuclease A has greater selectivity for the deamination of 5mC compared to C. Ribonuclease A can be included in the deamination reaction at a concentration of at least 1 microgram per milliliter (μg / ml) up to no more than 20 μM. For example, the concentration of ribonuclease A can be at least 1 μg / ml, at least 2 μg / ml, at least 3 μg / ml, at least 4 μg / ml, 5 μg / ml, 6 μg / ml, 7 μg / ml, 8 μg / ml, or 9 μg / ml, and / or no more than 50 μg / ml, no more than 40 μg / ml, no more than 30 μg / ml, no more than 20 μg / ml, no more than 19 μg / ml, no more than 18 μg / ml, no more than 17 μg / ml, no more than 16 μg / ml, no more than 15 μg / ml, no more than 14 μg / ml, no more than 13 μg / ml, no more than 12 μg / ml, or no more than 11 μg / ml. In one embodiment, the concentration of ribonuclease A is from 2 μg / ml to 13 μg / ml, or from 5 μg / ml to 10 μg / ml.

[0222] Target nucleic acid

[0223] The target nucleic acid that contacts the protein complex and is used in the methods, compositions, and kits provided herein can generally be any nucleic acid with a known or unknown sequence. Sequencing can determine the sequence of all or part of the target molecule. In one embodiment, the target nucleic acid can be processed into a template suitable for amplification by placing a universal amplification sequence (e.g., the sequence present in a universal adaptor) at the end of each target fragment.

[0224] Target nucleic acids typically originate from primary nucleic acids present in a sample, such as a biological sample. The primary nucleic acids can be of DNA or RNA origin. DNA primary nucleic acids can be derived from the double-stranded DNA (dsDNA) form of the sample (e.g., genomic DNA, genomic DNA fragments, cell-free DNA, etc.), or can be derived from the single-stranded form of the sample. RNA primary nucleic acids can be mRNA or non-coding RNA, e.g., microRNA or small interfering RNA. The exact sequence of the polynucleotide molecule from the primary nucleic acid sample is generally not important for the present disclosure and can be known or unknown.

[0225] The target nucleic acids of the present disclosure typically adopt a double-helical arrangement. The target nucleic acids can form a double-stranded double helix of B-DNA. In some embodiments, the single strands of the nucleic acid can form a double helix. Alternative arrangements of the nucleic acid are also possible, including A-DNA and Z-DNA. In an embodiment, the nucleic acid of interest can be circular, such as plasmid DNA. The nucleic acid of interest can have a complex tertiary structure, such as the tertiary structure found in telomeres. The tertiary structure of the nucleic acid can affect which helicases are suitable for unwinding the nucleic acid.

[0226] The primary nucleic acid molecule can represent the entire genetic complement of an organism, e.g., a genomic DNA molecule containing both intron sequences and exon sequences, as well as non-coding regulatory sequences such as promoter sequences and enhancer sequences. The primary nucleic acid molecule can represent the entire genetic complement of a particular cell of an organism, e.g., from a tumor cell, a genomic DNA molecule containing both intron sequences and exon sequences, as well as non-coding regulatory sequences such as promoter sequences and enhancer sequences. In one embodiment, a particular subset of genomic DNA can be used, e.g., a particular chromosome, DNA associated with open chromatin, DNA associated with closed chromatin, or a region of one or more particular sequences such as a particular gene (e.g., targeted sequencing). In one embodiment, the primary nucleic acid molecule can represent a particular subset of DNA, e.g., DNA having a particular sequence that anneals to a primer (such as a primer for target sequence or target enrichment). In one embodiment, a particular subset of DNA can be used, e.g., cell-free DNA, which can include DNA of a subject, including DNA from normal cells, DNA from diseased cells (such as tumor cells), and / or DNA from fetal cells.

[0227] The primary nucleic acid molecule can represent the entire transcriptome of a cell of an organism, e.g., mRNA molecules. The primary nucleic acid molecule can represent the entire transcriptome of a particular cell of an organism, e.g., from a tumor cell or a cell of a tissue, for example. In one embodiment, the primary nucleic acid molecule can represent a particular subset of mRNA, e.g., mRNA having a particular sequence that anneals to a primer (such as a primer for target sequence or target enrichment).

[0228] Samples, such as biological samples, can include nucleic acid molecules obtained from biopsy tissues, tumors, scrapings, swabs, blood, mucus, urine, feces, plasma, semen, hair, laser capture microdissection, surgically resected and other clinical or laboratory-derived samples. In some embodiments, the sample can be an epidemiological sample, an agricultural sample, a forensic sample or a pathogenic sample. In some embodiments, the sample can include cultured cells. In some embodiments, the sample can include nucleic acid molecules obtained from animals (such as human or mammalian sources). In another embodiment, the sample can include nucleic acid molecules obtained from non-mammalian sources (such as plants, bacteria, viruses or fungi). In some embodiments, the source of the nucleic acid molecules can be an archived or extinct sample or species.

[0229] Other non-limiting examples of biological sample sources can include whole organisms and samples obtained from patients. Biological samples can be obtained from any biological fluid or tissue and can be in a variety of forms (including fluids (e.g., liquids or gases), tissues, solid tissues) and preservation forms (such as dried, frozen and fixed forms). The sample can be any biological tissue, cell or fluid. Such samples include, but are not limited to, sputum, blood, serum, plasma, blood cells (e.g., white blood cells), ascites, urine, saliva, tears, sputum, vaginal fluid (secretions), flush fluids obtained during medical procedures (e.g., pelvic flush fluid or other flush fluids obtained during biopsy, endoscopy or surgery), tissues, nipple aspirates, core or fine needle biopsy tissue samples, body fluids containing cells, peritoneal fluid and pleural effusion or cells therefrom, and free-floating nucleic acids such as cell-free circulating DNA. Biological samples can also include tissue sections, such as frozen or fixed sections taken for histological purposes, or microdissected cells or their extracellular portions. In some embodiments, the sample can be a blood sample, such as a whole blood sample. In another example, the sample is an untreated dried blood spot (DBS) sample. In yet another example, the sample is a formalin-fixed paraffin-embedded (FFPE) sample. In yet another example, the sample is a saliva sample. In yet another example, the sample is a dried saliva spot (DSS) sample.

[0230] Exemplary biological samples from which target nucleic acids can be derived include, for example, those biological samples from eukaryotes, such as mammals, such as rodents, mice, rats, rabbits, guinea pigs, ungulates, horses, sheep, pigs, goats, cows, cats, dogs, primates, humans or non-human primates; plants, such as Arabidopsis thaliana, corn, sorghum, oats, wheat, rice, rapeseed or soybeans; algae, such as Chlamydomonas reinhardtii; nematodes, such as Caenorhabditis elegans; insects, such as Drosophila melanogaster, mosquitoes, fruit flies, bees or spiders; fish, such as zebrafish; reptiles; amphibians, such as frogs or Xenopus laevis; Dictyostelium discoideum; fungi, such as Pneumocystis carinii, Takifugu rubripes, yeast, Saccharamoyces cerevisiae or Schizosaccharomyces pombe; or protozoa, such as Plasmodium falciparum. Target nucleic acids can also be derived from prokaryotes, such as bacteria, Escherichia coli, Staphylococcus or Mycoplasma pneumoniae; archaea; viruses, such as hepatitis C virus or human immunodeficiency virus; or viroids. Target nucleic acids can be derived from a homogeneous culture or population of organisms described herein, or alternatively from a collection of several different organisms (e.g., in a community or ecosystem).

[0231] In some embodiments, the biological sample comprises tissue that has been processed to obtain the desired primary nucleic acid. In some embodiments, cells are used to obtain the desired primary nucleic acid. In some embodiments, cell nuclei are used to obtain the desired primary nucleic acid. The method can also include dissociating the cells and / or isolating the cell nuclei from the cells. Methods for isolating cells and cell nuclei from tissue can be used (WO 2019 / 236599).

[0232] In some embodiments, nucleic acids present in tissues, cells, or isolated cell nuclei can be processed according to a desired read. For example, the nucleic acids can be fixed during processing, and an effective fixation method can be used (WO 2019 / 236599). Fixation can be used to preserve the sample or maintain the contact of the analyte with the sample, cell, or cell nucleus. Fixation methods preserve and stabilize tissue, cell, and cell nucleus morphology and structure, inactivate proteolytic enzymes, strengthen the sample, cell, and cell nucleus, so that they can withstand further processing and staining, and prevent contamination. Examples of methods capable of effective fixation include, but are not limited to, whole genome sequencing of isolated cell nuclei and chromosome conformation capture methods (such as Hi-C). Common fixation methods include perfusion, immersion, freezing, and drying (Srinivasan et al., Am J Pathol. December 2002; Vol. 161, No. 6: pp. 1961-1971, doi: 10.1016 / S0002-9440(10)64472-0). In some embodiments such as whole genome sequencing, isolated cell nuclei can be processed to dissociate nucleosomes from DNA while keeping the cell nuclei intact, and methods for generating nucleosome-free cell nuclei can be used (WO2018 / 018008).

[0233] In some embodiments, a large amount of primary nucleic acids (e.g., from multiple cells) can be used to generate a sequencing library as described herein. In other embodiments, a single cell or cell nucleus can be used as a source of primary nucleic acids to obtain sequence information from single cells and cell nuclei. Many different single-cell library preparation methods are known in the art, including but not limited to Drop-seq, Seq-well, and single-cell combinatorial indexing (“sci-”) methods. Companies that provide single-cell products and related technologies include, but are not limited to: Illumina TM , 10X Genomics TM , Takara Biosciences TM , BD Biosciences TM , Bio-Rad Laboratories TM , 1cellbio TM , isoplexis TM , CellSee TM , nanoselect TM and DolomiteBio TM. Sci-seq is a methodological framework for uniquely tagging the nucleic acid contents of a large number of single cells or cell nuclei using split-pool barcodes. Generally, the number of cell nuclei or cells can be at least two. The upper limit depends on the practical limitations of the equipment used in other steps of the methods described herein (e.g., multi-well plates, indexing numbers). The number of cell nuclei or cells that can be used is not intended to be limited and can instead be in the billions.

[0234] The target nucleic acids used in the methods and compositions of the present disclosure can be obtained by fragmentation. Random fragmentation refers to the fragmentation of polynucleotide molecules in a disordered manner from a primary nucleic acid sample by enzymatic, chemical, or mechanical means. Such fragmentation methods are known in the art and use standard methods (Sambrook and Russell, Molecular Cloning, A Laboratory Manual, Third Edition). Additionally, random fragmentation is designed to produce fragments that are independent of sequence identity or position with respect to the sequences containing and / or surrounding the broken nucleotides. In one embodiment, random fragmentation is by mechanical means (such as nebulization or sonication) to produce fragments ranging in length from about 50 base pairs to about 1500 base pairs (more specifically, from 50 base pairs to 700 base pairs, and additionally specifically from 50 base pairs to 400 base pairs). Most particularly, the method is used to produce smaller fragments ranging in length from 50 to 150 base pairs.

[0235] Fragmentation of polynucleotide molecules by mechanical means (e.g., nebulization, sonication, and Hydroshear) produces a heterogeneous mixture of fragments with 3′ overhangs and 5′ overhangs. Thus, it is desirable to use methods or kits known in the art (e.g., Lucigen TM DNA Terminator End Repair Kit) to repair the fragment ends to generate ends that are most suitable for insertion into, for example, the blunt-ended sites of a cloning vector. In a specific embodiment, the fragment ends of the nucleic acid population are blunt-ended. More specifically, the fragment ends are blunt-ended and phosphorylated. The phosphate moiety can be introduced via enzymatic treatment, such as using polynucleotide kinase.

[0236] In a particular embodiment, the target fragment sequence is prepared with a single overhanging nucleotide by the activity of certain types of DNA polymerases, such as Taq polymerase or Klenow exo minus polymerase, which have nontemplate-dependent terminal transferase activity that adds a single deoxynucleotide (e.g., deoxyadenosine (A)) to the 3′ end of a DNA molecule (e.g., a PCR product). Such enzymes can be used to add a single nucleotide "A" to the blunt 3′ end of each strand of a double-stranded target fragment. Thus, 'A' can be added to the 3′ end of each end-repair strand of a double-stranded target fragment by reaction with Taq or Klenow exo- polymerase, and the universal adapter polynucleotide construct can be a T construct having compatible 'T' overhangs present at the 3′ end of each region of the double-stranded nucleic acid of the universal adapter. This end modification also prevents self-ligation of both the vector and the target, thus creating a bias towards formation of the target nucleic acid with the universal adapter at each end.

[0237] In one embodiment, fragmentation can be achieved using a method commonly referred to as tagmentation. Tagmentation uses a transpososome complex and combines fragmentation and ligation in a single step to add a universal adapter (WO 2016 / 130704). A transpososome complex is a transposase bound to a transposase recognition site and can insert the transposase recognition site into the target nucleic acid, a process sometimes referred to as "tagmentation of the fragment". In some such insertion events, one strand of the transposase recognition site can be transferred into the target nucleic acid. This strand is referred to as the "transferred strand". In one embodiment, the transpososome complex comprises a dimeric transposase having two subunits and two noncontiguous transposon sequences. In another embodiment, the transposase comprises a dimeric transposase having two subunits and contiguous transposon sequences.

[0238] Some embodiments may include using a hyperactive Tn5 transposase and a Tn5-type transposase recognition site (Goryshin and Reznikoff, J. Biol. Chem., 273:7367 (1998)), or a MuA transposase and a Mu transposase recognition site comprising R1 and R2 end sequences (Mizuuchi, K., Cell, 35:785, 1983; Savilahti, H et al., EMBO J., 14:4893, 1995). Those skilled in the art may also use Tn5 mosaic end (ME) sequences.

[0239] Examples of transposon sequences that can be used with the methods and compositions described herein are provided in U.S. Patent Application Publication 2012 / 0208705, U.S. Patent Application Publication 2012 / 0208724, and International Patent Application Publication WO 2012 / 061832. In some embodiments, the transposon sequence includes a first transposase recognition site and a second transposase recognition site.

[0240] Some transposome complexes useful herein include a transposase having two transposon sequences. In some such embodiments, the two transposon sequences are not linked to each other; in other words, the transposon sequences are not contiguous with each other. Examples of such transposomes are known in the art (see, e.g., U.S. Patent Application Publication 2010 / 0120098).

[0241] In one embodiment, tag fragmentation is used to generate a target nucleic acid that includes different universal sequences at each end. This can be achieved by using two types of transposome complexes, each of which includes a different nucleotide sequence as part of the transfer strand.

[0242] In some embodiments, tag fragmentation results in the formation of 5′ overhangs at one or both ends. The overhang is a region of ssDNA that may not be a suitable helicase substrate. In some embodiments, a population of target nucleic acids with tag fragmentation can be treated with a polymerase and dNTPs to fill in the overhangs, producing a completely double-stranded nucleic acid. In some of these embodiments, the resulting nucleic acid may have a single adenosine overhang at the 3′ end. Optionally, the A overhang can be removed.

[0243] The population of target nucleic acids can have an average strand length that is desired or suitable for a particular application of the methods, compositions, or kits described herein. For example, the average strand length can be less than about 100,000 nucleotides, 50,000 nucleotides, 10,000 nucleotides, 5,000 nucleotides, 1,000 nucleotides, 500 nucleotides, 100 nucleotides, or 50 nucleotides. Alternatively or in addition, the average strand length can be greater than about 10 nucleotides, 50 nucleotides, 100 nucleotides, 500 nucleotides, 1,000 nucleotides, 5,000 nucleotides, 10,000 nucleotides, 50,000 nucleotides, or 100,000 nucleotides. The average strand length of the population of target nucleic acids can be within the range between the maximum and minimum values described herein. It should be understood that the amplicons generated (or otherwise prepared or used herein) at the amplification site can have an average strand length within the range between the upper and lower limits exemplified above.

[0244] In some cases, a population of target nucleic acids can be generated under conditions or otherwise configured to have a maximum length of its components. For example, the maximum length of one or more steps of the methods described herein or of the members present in a particular composition can be less than 100,000 nucleotides, less than 50,000 nucleotides, less than 10,000 nucleotides, less than 5,000 nucleotides, less than 1,000 nucleotides, less than 500 nucleotides, less than 100 nucleotides, or less than 50 nucleotides. Alternatively or in addition, a population of target nucleic acids can be generated under conditions or otherwise configured to have a minimum length of its components. For example, the minimum length of one or more steps of the methods described herein or of the members present in a particular composition can be greater than 10 nucleotides, greater than 50 nucleotides, greater than 100 nucleotides, greater than 500 nucleotides, greater than 1,000 nucleotides, greater than 5,000 nucleotides, greater than 10,000 nucleotides, greater than 50,000 nucleotides, or greater than 100,000 nucleotides. The maximum and minimum chain lengths of the target nucleic acids in the population can be within a range between the above-mentioned maximum and minimum values. It should be understood that the amplicons generated (or otherwise prepared or used herein) at the amplification site can have a maximum and / or minimum chain length within a range between the upper and lower limits exemplified above.

[0245] In some embodiments, a sample can be enriched for a sequence of interest, e.g., a predetermined sequence. For example, a subset of genes or a region of a genome can be isolated and sequenced, or a subset of genes or a region of a genome can be interrogated by other methods such as locus-specific in vitro diagnostic methods. The predetermined sequence can be, for example, a sequence that can have a pattern of cytosine modification.

[0246] In some embodiments, target enrichment is performed by capturing a genomic region of interest by hybridization to a target-specific probe, which can be used to physically separate the target DNA that has hybridized to the bait probe from all other DNA in solution, and then washing away the other DNA. For example, some enrichment methods use biotinylated probes and then separate the biotinylated probes by magnetic precipitation using magnetic particles coated with streptavidin. In another example, some enrichment methods use an analyte array, also known as a microarray, which allows hybridization of a predetermined sequence.

[0247] Enrichment can be performed, for example, prior to treatment with cytidine deaminase. In these embodiments, enriching the target nucleic acid or a fragment thereof, such as enriching DNA in a sample, can include any suitable enrichment technique. In some embodiments, enriching DNA can include enrichment by molecular inversion probes, solution capture, pull-down probes, bait sets, standard PCR, multiplex PCR, hybridization capture, endonuclease digestion, DNase I hypersensitivity, and selective circularization. Enrichment can be achieved by negative selection of nucleic acids by eliminating unwanted material. Such enrichment includes "footprinting" techniques or "subtractive" hybridization capture. During the former, the target sample is protected from nuclease activity by a protective protein or by single-stranded and double-stranded arrangements. During the latter, nucleic acids that bind to 'bait' probes are eliminated.

[0248] In some embodiments, enrichment can include amplification using target-specific primers. In some embodiments, amplification is performed after another form of enrichment. However, typically, in embodiments where amplification is used for enrichment, the amplification step is performed after deaminase treatment to maintain the methylation status of the target DNA. In some such embodiments, amplification can include PCR amplification or whole-genome amplification.

[0249] In some embodiments, enrichment can be performed after treatment with cytidine deaminase. Typically, methods for identifying methylated cytosines result in a loss of DNA complexity due to the conversion of unmethylated DNA bases to uracil, generating a low-complexity genome and limiting the use of sequences for hybridization to a predetermined sequence specificity. Thus, after the conversion of methylated cytosines, typical methods for identifying methylated cytosines are more difficult to use in methods that include enrichment, such as hybridization enrichment sequencing and amplicon-based targeted sequencing. In contrast, due to (i) the altered cytidine deaminase converting 5mC to T, and (ii) only a small fraction of cytosines being methylated and expected to be converted by the altered cytidine deaminase. Examples of enrichment-based methods that can be used after treating the target nucleic acid with the altered cytidine deaminase include, but are not limited to, analyte arrays, the use of primers for selective amplification, CRISPR-Cas systems, and molecular cytogenetic techniques (such as FISH). Examples of arrays include, for example, methylation arrays for querying selected methylation sites across the genome (e.g., Infinium MethylationEPIC BeadChip, Illumina).

[0250] Attachment of universal adaptor

[0251] In some embodiments, the target nucleic acid used in the methods, compositions, or kits described herein may comprise universal adapters attached to each end. A target nucleic acid having universal adapters at each end may be referred to as a "modified target nucleic acid". Methods for attaching universal adapters to each end of the target nucleic acid used in the methods described herein are known to those of skill in the art. Attachment can be carried out by tag fragmentation using a transposase complex (WO 2016 / 130704) or by using standard library preparation techniques for ligation (U.S. Patent Publication 2018 / 0305753). Attachment of the universal adapter to the ends of the target nucleic acid can be carried out before or after treating the target nucleic acid with cytidine deaminase.

[0252] In some embodiments, an adapter compatible with the methods and compositions described herein may comprise additional regions that are not typically included in sequencing adapters. For some helicases of interest, certain motifs may preferably improve the recruitment of the helicase to the target nucleic acid. Examples of motifs that may be included in the adapter sequence include Y-shaped adapter structures, partial duplexes, ssDNA bubbles, forks, flat duplexes, triple junctions, etc. In some embodiments, the adapter may include more than one motif to recruit the helicase. The 3′ end or the 5′ end may comprise a motif. In some embodiments, both the 3′ end and the 5′ end comprise motifs.

[0253] In one embodiment, a double-stranded target nucleic acid from a sample (e.g., a fragmented sample that has been contacted with cytidine deaminase and converted from single-stranded to double-stranded nucleic acid) is treated by first ligating the same universal adapter molecule to the 5′ end and the 3′ end of the double-stranded target nucleic acid. In one embodiment, the universal adapter is a "matched" adapter or Y-adapter because the two strands of the adapter are formed by annealing complementary polynucleotide strands. In one embodiment, the universal adapter used in the methods of the present disclosure is referred to as a "mismatched" adapter because the adapter contains regions of sequence mismatch, i.e., they are not formed by annealing fully complementary polynucleotide strands. The general characteristics of mismatched adapters are further described in Gormley et al., U.S. Patent 7,741,463, and Bignell et al., U.S. Patent 8,053,192. The universal adapter typically comprises a universal capture binding sequence that facilitates immobilization of the target nucleic acid on an array for subsequent sequencing, and a universal primer binding site that can be used for sequencing. In another embodiment, the sample is contacted with a protein complex, such as a cytidine deaminase and a helicase or a fusion protein as described herein. The double-stranded target nucleic acid from the sample is then tag fragmented with a transpososome complex that inserts a universal adapter or a sequence that can be used to add a universal adapter into the target double-stranded nucleic acid.

[0254] The universal adapter may optionally include at least one index. The index can be used as a characteristic marker of a specific target nucleic acid source on the flow cell (U.S. Patent 8,053,192). Generally, the index is a synthetic sequence of nucleotides that is part of the universal adapter added to the target nucleic acid as part of the library preparation step. Thus, the index is a nucleic acid sequence attached to each of the target molecules in a particular sample, the presence of which indicates or is used to identify the sample or source from which these target molecules were isolated.

[0255] Preferably, the index can be up to 20 nucleotides in length, more preferably 1 to 10 nucleotides, and most preferably 4 to 6 nucleotides in length. A four-nucleotide index provides the possibility of multiplexing 256 samples on the same array, while a six-base index enables the handling of 4096 samples on the same array.

[0256] The exact nucleotide sequence of the universal adapter is generally not critical for the present disclosure and can be selected by the user such that the desired sequence elements are ultimately included in the common sequence of a plurality of different modified target nucleic acids, e.g., to provide a universal capture binding sequence for immobilizing the target nucleic acid on an array for subsequent sequencing, and binding sites for a set of universal amplification primers and / or sequencing primers. Additional sequence elements can be included, e.g., to provide a binding site for a sequencing primer that will ultimately be used for sequencing the target nucleic acid in the library, sequencing the index, or the product resulting from the amplification of the target nucleic acid from the library, e.g., on a solid support.

[0257] To prepare a deaminase-treated DNA library for analysis using a sequencing platform, it may be useful to perform additional modifications to the target DNA before or after treatment with cytidine deaminase. In some embodiments, single-stranded library preparation methods known in the art are used to prepare single-stranded deaminase-treated DNA for sequencing. Such methods include, but are not limited to, template-switching-based second-strand synthesis, adapters containing single-stranded splint overhangs, etc. Reagents for performing single-stranded library preparation methods are commercially available. Examples include xGen ssDNA and Low Input DNA Library Preparation Kits (Integrated DNA Technologies catalog number 10009859) (previously sold as Accel-NGS (Swift Biosciences)), NGS Single-Stranded DNA Library Preparation Kits (BioDynamics catalog number 30082). Another example includes the Single Reaction Single-Stranded Library (SRSLY) as described by Troll et al., BMC Genomics 20, 1023 (2019).

[0258] In some embodiments, prior to treatment with the altered cytidine deaminase, library preparation modifications are performed on the double-stranded target DNA. Methods for library preparation of double-stranded DNA templates are known in the art and include Y-adapter ligation, transposase-based tagmentation, and the like. Those skilled in the art will understand that methods for double-stranded library preparation typically include one or more amplification steps, such as by PCR. In such methods, the amplification step can be postponed until after the cytidine deaminase treatment to maintain the methylation state of the template strand. For example, in the Y-adapter ligation method, the Y-adapter can be ligated to the double-stranded template, and then the adapter-ligated template DNA can be denatured and treated with the cytidine deaminase described elsewhere herein. After treatment with the cytidine deaminase, the resulting treated single-stranded DNA molecules can be amplified using PCR, bridge amplification, and other methods known in the art.

[0259] Preparation of immobilized sample for sequencing

[0260] A library of modified target nucleic acids (e.g., target nucleic acids having universal adapters at each end) can be prepared for sequencing. Methods for attaching the modified target nucleic acids to a substrate are known in the art. In one embodiment, a plurality of capture oligonucleotides specific for the modified fragments are used to enrich the modified fragments, and these capture oligonucleotides can be immobilized on the surface of a solid substrate, such as a flow cell or beads. For example, the capture oligonucleotides can include a first number of universal binding pairs, and a second number of the binding pairs are immobilized on the surface of the solid substrate. Similarly, methods for amplifying the immobilized target nucleic acids include, but are not limited to, bridge amplification and exclusion amplification (also known as kinetic exclusion amplification (KEA)). Methods for immobilization and amplification prior to sequencing are described, for example, in Binoell et al. (US 8,053,192), Gunderson et al. (WO2016 / 130704), Shen et al. (US 8,895,249), and Pipenburg et al. (US 9,309,502).

[0261] The pooled samples can be immobilized in preparation for sequencing. Sequencing can be performed as a single molecule array or can be amplified prior to sequencing. Amplification can be performed using one or more immobilized primers. The immobilized primers can be, for example, a primer lawn on a flat surface or on a bead pool. The bead pool can be separated into an emulsion with a single bead in each "compartment" of the emulsion. When the concentration is only one template per "compartment", only a single template is amplified on each bead.

[0262] As used herein, the term "solid-phase amplification" refers to any nucleic acid amplification reaction that is performed on or associated with a solid support such that all or a portion of the amplification product is immobilized on the solid support upon formation. Specifically, the term encompasses solid-phase polymerase chain reaction (solid-phase PCR) and solid-phase isothermal amplification, which are reactions similar to standard solution-phase amplification, except that one or both of the forward and reverse amplification primers are immobilized on the solid support. Solid-phase PCR includes systems such as emulsions, where one primer is anchored to beads and the other primer is in free solution; and colony formation in a solid-phase gel matrix, where one primer is anchored to the surface and one primer is in free solution.

[0263] In some embodiments, the solid support includes a patterned surface. A "patterned surface" refers to an arrangement of different regions in or on an exposed layer of the solid support. For example, one or more of these regions can be features where one or more amplification primers are present. The features can be separated by gap regions where no amplification primers are present. In some embodiments, the pattern can be an x-y format of features in the form of rows and columns. In some embodiments, the pattern can be a repeating arrangement of features and / or gap regions. In some embodiments, the pattern can be a random arrangement of features and / or gap regions. Exemplary patterned surfaces useful in the methods and compositions described herein are in U.S. Patents 8,778,848, 8,778,849, and 9,079,148 and U.S. Patent Application Publication 2014 / 0243224.

[0264] In some specific implementations, the solid support includes an array of pores or depressions in the surface. This can be fabricated using a variety of techniques as are commonly known in the art, including but not limited to photolithography, imprinting techniques, molding techniques, and microetching techniques. Those skilled in the art will know that the technique used will depend on the composition and shape of the array substrate.

[0265] Features in the patterned surface can be pores (e.g., micropores or nanopores) in an array of pores in a solid support of glass, silicon, plastic, or other suitable solid support having a patterned and covalently linked gel, such as poly(N-(5-azidoacetamidopentyl)acrylamide-co-acrylamide) (PAZAM, see, e.g., U.S. Publication 2013 / 184796, WO 2016 / 066586, and WO 2015 / 002813). The method produces a gel pad for sequencing that can be stable during a sequencing run having a large number of cycles. Covalent linking of the polymer to the pores helps to maintain the gel as a structured feature during various uses and throughout the life of the structured substrate. However, in many embodiments, the gel need not be covalently linked to the pores. For example, under some conditions, silane-free acrylamide (SFA, see, e.g., U.S. Patent 8,563,477), which is not covalently attached to any part of the structured substrate, can be used as the gel material.

[0266] In certain embodiments, a structured substrate can be made by patterning a solid support material to have pores (e.g., micropores or nanopores), coating the patterned support with a gel material (e.g., PAZAM, SFA, or a chemically modified variant thereof, such as an azidated version of SFA (azido-SFA)), and polishing the coated gel, e.g., by chemical or mechanical polishing, to retain the gel in the pores while removing substantially all of the gel or inactivating substantially all of the gel from the interstitial regions on the surface of the structured substrate between the pores. Primer nucleic acids can be attached to the gel material. Then a solution of modified target nucleic acids can be contacted with the polished substrate such that a single modified target nucleic acid will be seeded into a single pore by interaction with the primer attached to the gel material; however, because there is no gel material or the gel material is inactivated, the target nucleic acids will not occupy the interstitial regions. Amplification of the modified target nucleic acids will be restricted to the pores because the absence of gel or gel inactivation in the interstitial regions prevents outward migration of the growing nucleic acid colony. The process can be conveniently and scalably manufactured and utilizes conventional micro- or nano-fabrication methods.

[0267] While the present disclosure encompasses "solid-phase" amplification methods in which only one amplification primer is immobilized (the other primer is typically present in free solution), in one embodiment, the solid support is provided with both a forward primer and a reverse primer that are immobilized. In practice, there will be'multiple' identical forward primers and / or'multiple' identical reverse primers immobilized on the solid support because the amplification process requires an excess of primers to sustain amplification. References herein to forward and reverse primers should be construed accordingly as covering multiple such primers unless the context indicates otherwise.

[0268] Skilled readers will appreciate that any given amplification reaction requires at least one type of forward primer and at least one type of reverse primer that are specific for the template to be amplified. However, in certain embodiments, the forward and reverse primers may include template-specific portions of the same sequence and may have exactly the same nucleotide sequence and structure (including any non-nucleotide modifications). In other words, a single type of primer can be used for solid-phase amplification, and such single-primer methods are encompassed within the scope of the present disclosure. Other embodiments may use forward and reverse primers that contain the same template-specific sequence but differ in some other structural features. For example, one type of primer may contain non-nucleotide modifications that are not present in the other type.

[0269] Primers for solid-phase amplification are preferably immobilized to a solid support at or near the 5′ end of the primer by a single-point covalent attachment such that the template-specific portion of the primer is free to anneal to its homologous template while the 3′ hydroxyl is free to undergo primer extension. Any suitable covalent attachment method known in the art can be used for this purpose. The attachment chemistry selected will depend on the nature of the solid support and any derivatization or functionalization applied to it. The primer itself may contain moieties that can be non-nucleotide chemical modifications to facilitate attachment. In one specific embodiment, the primer may contain a sulfur nucleophile at the 5′ end, such as phosphorothioate or thiophosphate. In the case of a solid-supported polyacrylamide hydrogel, this nucleophile will bind to the bromoacetamide groups present in the hydrogel. A more specific way of attaching the primer and template to the solid support is via a 5′ phosphorothioate attachment to a hydrogel composed of polymerized acrylamide and N-(5-bromoacetamidopentyl)acrylamide (BRAPA), as described in International Publication WO 05 / 065814.

[0270] Certain embodiments of the present invention may utilize a solid support that includes an inert substrate or matrix (e.g., a glass slide, polymer beads, etc.) that has been "functionalized" by, for example, applying an intermediate material layer or coating that contains reactive groups that permit covalent attachment to biomolecules such as polynucleotides. Examples of such supports include, but are not limited to, polyacrylamide hydrogels supported on an inert substrate such as glass. In such embodiments, the biomolecule (e.g., polynucleotide) may be directly covalently attached to the intermediate material (e.g., hydrogel), but the intermediate material itself may be non-covalently attached to the substrate or matrix (e.g., glass substrate). The term "covalently attached to a solid support" should be interpreted accordingly to cover this type of arrangement.

[0271] Samples can be amplified and combined on beads, where each bead contains a forward amplification primer and a reverse amplification primer. In one embodiment, a library of modified target nucleic acids is used to prepare a cluster array of nucleic acid populations, similar to those cluster arrays prepared by solid-phase amplification as described in U.S. Publication 2005 / 0100900, U.S. Patent 7,115,400, WO 00 / 18957, and WO 98 / 44151, and more specifically by isothermal solid-phase amplification. The terms 'cluster' and 'colony' are used interchangeably herein and refer to discrete sites on a solid support that include multiple identical immobilized nucleic acid strands and multiple identical immobilized complementary nucleic acid strands. The term 'cluster array' refers to an array formed by such clusters or colonies. In this context, the term 'array' should not be understood as requiring an ordered arrangement of the clusters.

[0272] The term "solid phase" or "surface" is used to denote a planar array where primers are attached to a flat surface such as a glass, silica, or plastic microscope slide or a similar flow cell device; beads where one or both primers are attached to the beads and the beads are amplified; or an array of beads on a surface after the beads have been amplified.

[0273] A cluster array can be prepared using a thermal cycling process as described in WO 98 / 44151 or a process that keeps the temperature constant, and the cycles of extension and denaturation are performed by changing the reagents. Such isothermal amplification methods are described in patent application numbers WO 02 / 46456 and U.S. Publication 2008 / 0009420.

[0274] It should be understood that any of the amplification methods described herein or commonly known in the art can be used with universal or target-specific primers to amplify immobilized DNA fragments. Suitable amplification methods include, but are not limited to, polymerase chain reaction (PCR), strand displacement amplification (SDA), transcription-mediated amplification (TMA), and nucleic acid sequence-based amplification (NASBA), as described in U.S. Patent 8,003,354. The above amplification methods can be used to amplify one or more nucleic acids of interest. For example, PCR (including multiplex PCR), SDA, TMA, NASBA, etc. can be used to amplify immobilized DNA fragments. In some embodiments, primers specific for the polynucleotide of interest are included in the amplification reaction.

[0275] Other suitable polynucleotide amplification methods can include oligonucleotide extension and ligation, rolling circle amplification (RCA) (Lizardi et al., Nat. Genet. 19:225-232 (1998)), and oligonucleotide ligation assay (OLA) (see generally U.S. Patent Nos. 7,582,420, 5,185,243, 5,679,524, and 5,573,907; EP 0 320 308 B1; EP 0 336 731 B1; EP 0 439 182 B1; WO 90 / 01069; WO 89 / 12696; and WO 89 / 09835) techniques. It should be understood that these amplification methods can be designed to amplify immobilized DNA fragments. For example, in some embodiments, the amplification method can include ligation probe amplification or an oligonucleotide ligation assay (OLA) reaction containing primers specific for the nucleic acid of interest. In some embodiments, the amplification method can include a primer extension-ligation reaction that contains primers specific for the nucleic acid of interest. As non-limiting examples of primer extension and ligation primers that can be specifically designed to amplify a nucleic acid of interest, amplification can include primers for the Golden Gate assay (Illumina, Inc., San Diego, CA), as exemplified in U.S. Patent Nos. 7,582,420 and 7,611,869.

[0276] DNA nanospheres can also be used in combination with the methods, systems, compositions, and kits described herein. Methods for forming and using DNA nanospheres for genomic sequencing can be found, for example, in U.S. Patents and Publications 7,910,354, 2009 / 0264299, 2009 / 0011943, 2009 / 0005252, 2009 / 0155781, 2009 / 0118488, and as described, for example, by Drmanac et al., (Science, Vol. 327, No. 5961, pp. 78-81, 2010). Briefly, after generating modified target nucleic acids, these modified target nucleic acids are circularized and amplified by rolling circle amplification (Lizardi et al., Nat. Genet. 19:225-232, 1998; US2007 / 0099208A1). The extended tail-to-tail structure of the amplicons promotes coiling, resulting in compact DNA nanospheres. The DNA nanospheres can be captured on a substrate, preferably to produce an ordered or patterned array such that the distance between each nanosphere is maintained, allowing for sequencing of individual DNA nanospheres. In some embodiments, such as those used by Complete Genomics, Inc. (Mountain View, Calif.), successive rounds of adapter addition, amplification, and digestion are performed prior to circularization to produce a head-to-tail construct having a number of target nucleic acids separated by adapter sequences.

[0277] Exemplary isothermal amplification methods that can be used in the methods of the present disclosure include, but are not limited to, multiple displacement amplification (MDA), such as that exemplified by Dean et al., Proc. Natl. Acad. Sci. USA, Vol. 99: pp. 5261-5266 (2002), or isothermal strand displacement nucleic acid amplification, such as that exemplified by U.S. Patent 6,214,587. Other non-PCR-based methods that can be used in the present disclosure include, for example, strand displacement amplification (SDA), which is described, for example, in Walker et al., Molecular Methods for Virus Detection, Academic Press, Inc., 1995; U.S. Patents 5,455,166 and 5,130,238, and Walker et al., Nucl. Acids Res. Vol. 20: pp. 1691-1696 (1992); or hyperbranched strand displacement amplification, which is described, for example, in Lage et al., Genome Res., 13: 294-307 (2003). The isothermal amplification methods can be used for random primer amplification of genomic DNA, for example, with a strand displacement Phi 29 polymerase or the large fragment of Bst DNA polymerase 5'->3'exo-. The use of these polymerases takes advantage of their high processive synthesis ability and strand displacement activity. The high processive synthesis ability allows the polymerase to generate fragments with lengths of 10 kb - 20 kb. As described above, a polymerase with low processive synthesis ability and strand displacement activity (such as Klenow polymerase) can be used to generate smaller fragments under isothermal conditions. Additional descriptions of the amplification reactions, conditions, and components are detailed in the disclosure of U.S. Patent 7,670,810.

[0278] In some embodiments, kinetic exclusion amplification (KEA), also known as exclusion amplification (ExAmp), can be used to perform isothermal amplification. The nucleic acid libraries of the present disclosure can be made using a method comprising the steps of reacting amplification reagents to generate a plurality of amplification sites, each of the plurality of amplification sites comprising a substantially clonal population of amplicons from a single target nucleic acid at an inoculated site. In some embodiments, the amplification reaction continues until a sufficient number of amplicons are generated to fill the capacity of the corresponding amplification site. Filling the inoculated site to capacity in this manner inhibits the landing and amplification of the target nucleic acid at that site, thereby producing a clonal population of amplicons at that site. In some embodiments, apparent clonality can be achieved even if the amplification site is not filled to capacity before the second target nucleic acid reaches the site. Under some conditions, the amplification of the first target nucleic acid can proceed to the point where a sufficient number of copies are prepared to effectively outnumber or overwhelm the production of copies of the second target nucleic acid transported to the site. For example, in embodiments of bridge amplification processes using circular features with diameters less than 500 nm, it has been determined that after 14 cycles of exponential amplification of the first target nucleic acid, contamination from the second target nucleic acid at the same site will produce an insufficient number of contaminating amplicons that will not adversely affect the sequencing-by-synthesis analysis on an Illumina sequencing platform.

[0279] In some embodiments, the amplification sites in the array can be, but need not be, completely clonal. Instead, for some applications, a single amplification site can be predominantly filled with amplicons from a first modified target nucleic acid and can also have a low level of contaminating amplicons from a second modified target nucleic acid. As long as the contamination level does not have an unacceptable effect on the subsequent use of the array, the array can have one or more amplification sites with low levels of contaminating amplicons. For example, when the array is to be used for a detection application, an acceptable contamination level will be a level that does not affect the signal-to-noise ratio or resolution of the detection technique in an unacceptable manner. Thus, apparent clonality will generally be related to the specific use or application of the array prepared by the methods described herein. Exemplary contamination levels that can be acceptable at a single amplification site for a particular application include, but are not limited to, up to 0.1%, 0.5%, 1%, 5%, 10%, or 25% contaminating amplicons. The array can include one or more amplification sites having these exemplary levels of contaminating amplicons. For example, up to 5%, 10%, 25%, 50%, 75%, or even 100% of the amplification sites in the array can have some contaminating amplicons. It should be understood that in an array or other collection of sites, up to 50%, 75%, 80%, 85%, 90%, 95%, or 99% or more of the sites can be clonal or apparently clonal.

[0280] In some embodiments, kinetic exclusion can occur when a process occurs at a rate fast enough to effectively preclude another event or process from occurring. Taking the preparation of a nucleic acid array as an example, where the sites of the array are randomly seeded with modified target nucleic acids from solution and copies of the modified target nucleic acids are generated during an amplification process to fill each of the seeded sites to capacity. According to the kinetic exclusion method of the present disclosure, the seeding and amplification processes can occur simultaneously under conditions where the amplification rate exceeds the seeding rate. Thus, the relatively rapid rate of generating copies at a site that has already been seeded with a first target nucleic acid will effectively preclude a second nucleic acid from seeding the site for amplification. The kinetic exclusion amplification method can be carried out as detailed in the disclosure of U.S. Patent Application Publication 2013 / 0338042.

[0281] Kinetic exclusion can initiate amplification by taking advantage of a relatively slow rate (e.g., the slower rate of preparing the first copy of a modified target nucleic acid) compared to the relatively rapid rate of preparing subsequent copies of the modified target nucleic acid (or the first copy of the modified target nucleic acid). In the example of the previous paragraph, kinetic exclusion occurs due to the relatively slow rate of seeding with the modified target nucleic acid (e.g., relatively slow diffusion or transport) compared to the relatively rapid rate at which amplification occurs to fill the site with copies of the modified target nucleic acid. In another exemplary embodiment, kinetic exclusion can occur due to a delay (e.g., delayed or slow activation) in the formation of the first copy of the modified target nucleic acid at a seeded site compared to the relatively rapid rate of preparing subsequent copies to fill the site. In this example, a single site may have been seeded with several different modified target nucleic acids (e.g., several modified target nucleic acids that may be present at each site prior to amplification). However, the formation of the first copy of any given modified target nucleic acid can be randomly activated such that the average rate of first copy formation is relatively slow compared to the rate of subsequent copy generation. In such a case, although a single site may have been seeded with several different modified target nucleic acids, kinetic exclusion will only allow amplification of one of these modified target nucleic acids. More specifically, once the first modified target nucleic acid has been activated for amplification, the site will be rapidly filled to capacity with its copies, thereby preventing the preparation of copies of a second modified target nucleic acid at that site.

[0282] In one embodiment, the method is performed to simultaneously (i) transport a modified target nucleic acid to an amplification site at an average transport rate and (ii) amplify the modified target nucleic acids at these amplification sites at an average amplification rate, where the average amplification rate exceeds the average transport rate (U.S. Patent 9,169,513). Thus, in such embodiments, kinetic exclusion can be achieved by using a relatively slow transport rate. For example, a sufficiently low concentration of the modified target nucleic acid can be selected to achieve a desired average transport rate, with a lower concentration resulting in a slower average transport rate. Alternatively or in addition, a high-viscosity solution and / or the presence of a molecular crowding reagent in the solution can be used to reduce the transport rate. Examples of available molecular crowding reagents include, but are not limited to, polyethylene glycol (PEG), ficoll, dextran, or polyvinyl alcohol. Exemplary molecular crowding reagents and formulations are described in U.S. Patent 7,399,590, which is incorporated herein by reference. Another factor that can be adjusted to achieve a desired transport rate is the average size of the target nucleic acid.

[0283] The amplification reagent can also include components that facilitate amplicon formation and, in some cases, increase the rate of amplicon formation. One example is a recombinase. The recombinase can facilitate amplicon formation by allowing repeated invasion / extension. More specifically, the recombinase can facilitate the invasion of the modified target nucleic acid by a polymerase and the extension of a primer by the polymerase, which uses the modified target nucleic acid as a template for amplicon formation. This process can be repeated as a chain reaction, where the amplicons generated by each round of invasion / extension are used as templates in subsequent rounds. Since denaturation cycles (e.g., via heating or chemical denaturation) are not required, the process can occur more rapidly than standard PCR. Thus, recombinase-promoted amplification can be performed isothermally. It is generally desirable to include ATP or other nucleotides (or in some cases their non-hydrolyzable analogs) in the recombinase-promoted amplification reagent to facilitate amplification. A mixture of a recombinase and a single-strand binding (SSB) protein is particularly useful because the SSB can further facilitate amplification. Exemplary formulations for recombinase-promoted amplification include those commercially available from TwistDx (Cambridge, UK) as the TwistAmp kit. The available components and reaction conditions for recombinase-promoted amplification reagents are described in US 5,223,414 and US 7,399,590.

[0284] Another example of a component that can be included in an amplification reagent to facilitate amplicon formation and in some cases increase the rate of amplicon formation is a helicase. The helicase can facilitate amplicon formation by allowing a chain reaction for amplicon formation. Since denaturation cycles (e.g., via heating or chemical denaturation) are not required, the process can occur more rapidly than standard PCR. Thus, helicase - promoted amplification can be carried out isothermally. A mixture of helicase and single - strand binding (SSB) protein is particularly useful because the SSB can further facilitate amplification. Exemplary formulations for helicase - promoted amplification include those commercially available as the IsoAmp kit from Biohelle (Beverly, MA). In addition, examples of available formulations that include helicase protein are described in US 7,399,590 and US 7,829,284. It should be noted that in the methods of the present disclosure, multiple helicases can be used for multiple functions. For example, a first helicase can be provided as a fusion protein with an altered APOBEC3A to convert modified cytosine to thymidine, and subsequently a second helicase can be provided to facilitate amplicon formation.

[0285] Another example of a component that can be included in an amplification reagent to favor amplicon formation and in some cases increase the rate of amplicon formation is an origin - binding protein. The replication origin of a gene can be bound by an origin recognition complex (ORC) that contains many origin - binding proteins. Inclusion of an origin - binding protein in an amplicon formation reaction can improve the recruitment of additional enzymes required or available for amplification.

[0286] Sequencing method

[0287] After attaching the modified target nucleic acid to a surface, the sequence of the immobilized and amplified modified target nucleic acid is determined. Sequencing can be performed using any suitable sequencing technology, and methods for determining the sequence of immobilized and amplified modified target nucleic acids (including strand resynthesis) are known in the art and are described, for example, in Bignell et al. (US 8,053,192), Gunderson et al. (WO2016 / 130704), Shen et al. (US 8,895,249), and Pipenburg et al. (US 9,309,502).

[0288] The methods described herein can be used in combination with a variety of nucleic acid sequencing techniques. Particularly suitable techniques are those in which the nucleic acids are attached to fixed positions in an array such that their relative positions do not change and in which the array is repeatedly imaged. Embodiments in which images are obtained in different color channels (e.g., corresponding to different labels used to distinguish one nucleotide base type from another) are particularly suitable. In some embodiments, the process of determining the nucleotide sequence of the modified target nucleic acid can be an automated process. Preferred embodiments include sequencing by synthesis (“SBS”) techniques.

[0289] SBS techniques generally involve the enzymatic extension of a nascent nucleic acid strand by repeated addition of nucleotides to a template strand. In conventional SBS methods, a single nucleotide monomer can be provided to the target nucleotide in the presence of polymerase in each delivery. However, in the methods described herein, more than one type of nucleotide monomer can be provided to the target nucleic acid in the presence of polymerase in a delivery.

[0290] In one embodiment, the nucleotide monomer includes locked nucleic acid (LNA) or bridged nucleic acid (BNA). The use of LNA or BNA in the nucleotide monomer increases the hybridization strength between the nucleotide monomer and the sequencing primer sequence present on the immobilized modified target nucleic acid.

[0291] SBS techniques can use nucleotide monomers having a terminator moiety or nucleotide monomers lacking any terminator moiety. Methods using nucleotide monomers lacking a terminator include, for example, pyrosequencing and sequencing using γ-phosphate labeled nucleotides, as described in further detail herein. In methods using nucleotide monomers lacking a terminator, the number of nucleotides added in each cycle is typically variable and depends on the template sequence and the manner of nucleotide delivery. For SBS techniques using nucleotide monomers having a terminator moiety, the terminator can be effectively irreversible under the sequencing conditions used, as in the case of traditional Sanger sequencing using dideoxynucleotides, or the terminator can be reversible, as in the case of the sequencing method developed by Solexa (now Illumina, Inc.).

[0292] The SBS technique can use nucleotide monomers with a labeled moiety or nucleotide monomers lacking a labeled moiety. Thus, incorporation events can be detected based on: the properties of the label, such as the fluorescence of the label; the properties of the nucleotide monomer, such as molecular weight or charge; by-products of incorporated nucleotides, such as the release of pyrophosphate; and so on. In embodiments where there are two or more different nucleotides present in the sequencing reagent, the different nucleotides can be distinguishable from one another, or alternatively two or more different labels can be indistinguishable under the detection techniques used. For example, the different nucleotides present in the sequencing reagent can have different labels and they can be distinguished using appropriate optics, as exemplified by the sequencing method developed by Solexa (now Illumina, Inc.).

[0293] Preferred embodiments include pyrosequencing technology. Pyrosequencing detects the release of inorganic pyrophosphate (PPi) when a specific nucleotide is incorporated into a nascent strand (Ronaghi, M., Karamohamed, S., Pettersson, B., Uhlen, M., and Nyren, P. (1996), “Real-time DNA sequencing using detection of pyrophosphate release.”, Analytical Biochemistry 242(1), 84-9; Ronaghi, M. (2001), “Pyrosequencing sheds light on DNA sequencing.”, Genome Res. 11(1), 3-11; Ronaghi, M., Uhlen, M., and Nyren, P. (1998) “A sequencing method based on real-time pyrophosphate.” Science 281(5375), 363; U.S. Patents 6,210,891; 6,258,568 and 6,274,320). In pyrosequencing, the released PPi can be detected by immediately converting it into adenosine triphosphate (ATP) by ATP sulfurylase, and the level of the generated ATP can be detected via photons generated by luciferase. The nucleic acid to be sequenced can be attached to features in an array, and the array can be imaged to capture chemiluminescence signals generated due to nucleotide incorporation at the features of the array. An image can be obtained after treating the array with a specific nucleotide type (e.g., A, T, C, or G). The images obtained after adding each nucleotide type will differ in terms of which features in the array are detected. These differences in the images reflect the different sequence contents of the features on the array. However, the relative positions of each feature will remain unchanged in the image. The images can be stored, processed, and analyzed using the methods described herein. For example, the images obtained after treating the array with each different nucleotide type can be processed in the same manner as illustrated herein for images obtained from different detection channels for reversible terminator-based sequencing methods.

[0294] In another exemplary type of SBS, cycle sequencing is accomplished by stepwise addition of reversible terminator nucleotides that include, for example, cleavable or photobleachable dye labels as described, for example, in WO 04 / 018497 and U.S. Patent 7,057,026. This method was commercialized by Solexa (now Illumina Inc.) and is also described in WO 91 / 06678 and WO 07 / 123,744. The availability of fluorescently labeled terminators (where termination can be reversible and the fluorescent label can be cleaved) facilitates efficient cycle reversible termination (CRT) sequencing. The polymerase can also be co-engineered to efficiently incorporate these modified nucleotides and extend from these modified nucleotides.

[0295] In some reversible terminator-based sequencing embodiments, the label does not substantially inhibit extension under the SBS reaction conditions. However, the detection label can be removable, for example, by cleavage or degradation. An image can be captured after incorporation of the label into the arrayed nucleic acid features. In a specific embodiment, each cycle involves simultaneous delivery of four different nucleotide types to the array, and each nucleotide type has a spectrally distinct label. Four images can then be obtained, each using a detection channel selective for one of the four different labels. Alternatively, the different nucleotide types can be added sequentially, and an image of the array can be obtained between each addition step. In such embodiments, each image will show the nucleic acid features into which a particular type of nucleotide has been incorporated. Due to the different sequence content of each feature, different features will be present or absent in different images. However, the relative position of the features will remain unchanged in the images. Images obtained by such reversible terminator-SBS methods can be stored, processed, and analyzed as described herein. After the image capture step, the label can be removed and the reversible terminator moiety can be removed for subsequent cycles of nucleotide addition and detection. Removing these labels after they have been detected in a particular cycle and prior to subsequent cycles can provide the advantage of reducing background signal and crosstalk between cycles. Examples of available labels and removal methods are described herein.

[0296] In certain embodiments, some or all of the nucleotide monomers may include reversible terminators. In such embodiments, the reversible terminator / cleavable fluorophore may include a fluorophore attached to the ribose moiety via a 3′ ester bond (Metzker, Genome Res., 15:1767-1776 (2005)). Other methods have separated terminator chemistry from fluorophore cleavage (Ruparel et al., Proc Natl Acad Sci USA, vol. 102: pp. 5932-5937 (2005)). Ruparel et al. describe the development of reversible terminators that use small 3′ allyl groups to block extension, but can be readily deblocked by brief treatment with a palladium catalyst. The fluorophore is attached to the base via a photocleavable linker that can be readily cleaved by exposure to long wavelength ultraviolet light for 30 seconds. Thus, disulfide reduction or photocleavage can be used as the cleavable linker. Another method of reversible termination is to use natural termination, which occurs following placement of a bulky dye on the dNTP. The presence of a charged bulky dye on the dNTP can act as an efficient terminator by steric and / or electrostatic hindrance. The presence of one incorporation event prevents further incorporation unless the dye is removed. Cleavage of the dye removes the fluorophore and effectively reverses termination. Examples of modified nucleotides are also described in U.S. Patent Nos. 7,427,673 and 7,057,026.

[0297] Additional exemplary SBS systems and methods that can be used with the methods and systems described herein are described in U.S. Patent Application Publication Nos. 2007 / 0166705, 2006 / 0188901, 2006 / 0240439, 2006 / 0281109, 2012 / 0270305 and 2013 / 0260372, U.S. Patent Nos. 7,057,026, PCT Publication No. WO 05 / 065814, U.S. Patent Application Publication No. 2005 / 0100900 and PCT Publication Nos. WO 06 / 064199 and WO 07 / 010,251.

[0298] Some embodiments may use fewer than four different labels to effect detection of the four different nucleotides. For example, SBS may be performed using the methods and systems described in the incorporated materials of U.S. Patent Application Publication 2013 / 0079232. As a first example, a pair of nucleotide types may be detected at the same wavelength, but distinguished based on the intensity difference of one member of the pair relative to the other member, or based on a change in one member of the pair that results in a distinct signal appearance or disappearance compared to the signal of the other member of the detected pair (e.g., by chemical modification, photochemical modification, or physical modification). As a second example, three of the four different nucleotide types can be detected under specific conditions, while the fourth nucleotide type lacks a label that can be detected or is minimally detected under those conditions (e.g., minimal detection due to background fluorescence, etc.). Incorporation of the first three nucleotide types into the nucleic acid can be determined based on the presence of their respective signals, and incorporation of the fourth nucleotide type into the nucleic acid can be determined based on the absence of any signal or minimal detection of any signal. As a third example, one nucleotide type may include a label detected in two different channels, while the other nucleotide types are detected in no more than one channel. The three above-described exemplary configurations are not considered mutually exclusive and can be used in various combinations. An exemplary embodiment that combines all three examples is a fluorescence-based SBS method that uses a first nucleotide type detected in a first channel (e.g., dATP having a label detected in the first channel when excited by a first excitation wavelength), a second nucleotide type detected in a second channel (e.g., dCTP having a label detected in the second channel when excited by a second excitation wavelength), a third nucleotide type detected in both the first channel and the second channel (e.g., dTTP having at least one label detected in both channels when excited by the first excitation wavelength and / or the second excitation wavelength), and a fourth nucleotide type lacking a label detected or minimally detected in either channel (e.g., dGTP having no label).

[0299] In addition, as described in U.S. Publication 2013 / 0079232, a single channel may be used to obtain sequencing data. In such so-called single-dye sequencing methods, the first nucleotide type is labeled, but the label is removed after generating the first image, and only the second nucleotide type is labeled after generating the first image. The third nucleotide type retains its label in both the first image and the second image, and the fourth nucleotide type remains unlabeled in both images.

[0300] Some embodiments may use sequencing by ligation techniques. Such techniques utilize DNA ligases to incorporate oligonucleotides and determine the incorporation of such oligonucleotides. Oligonucleotides typically have different labels related to the identity of specific nucleotides in the sequences to which the oligonucleotides hybridize. As with other SBS methods, an image may be obtained after treating an array of nucleic acid features with labeled sequencing reagents. Each image will show nucleic acid features into which a particular type of label has been incorporated. Due to the different sequence content of each feature, different features will be present in or absent from different images, but the relative positions of the features will remain unchanged in the images. Images obtained by ligation-based sequencing methods may be stored, processed, and analyzed as described herein. Exemplary SBS systems and methods that may be used with the methods and systems described herein are described in U.S. Patents 6,969,488, 6,172,218, and 6,306,597.

[0301] Some embodiments may use nanopore sequencing (Deamer, D.W. and Akeson, M., "Nanopores and nucleic acids: prospects for ultrarapid sequencing.", Trends Biotechnol. 18, 147-151 (2000); Deamer, D. and D. Branton, "Characterization of nucleic acids by nanopore analysis", Acc. Chem. Res. 35: 817-825 (2002); Li, J., M. Gershow, D. Stein, E. Brandin and J.A. Golovchenko, "DNA molecules and configurations in a solid-state nanopore microscope", Nat. Mater., 2: 611-615 (2003)). In such embodiments, the modified target nucleic acid passes through the nanopore. The nanopore can be a synthetic pore or a biomembrane protein such as alpha-hemolysin. When the modified target nucleic acid passes through the nanopore, each base pair can be identified by measuring fluctuations in the electrical conductivity of the pore. (U.S. Patent No. 7,001,792; Soni, G.V. and Meller, "A. Progress toward ultrafast DNA sequencing using solid-state nanopores.", Clin. Chem. 53, 1996-2001 (2007); Healy, K., "Nanopore-based single-molecule DNA analysis.", Nanomed., 2, 459-481 (2007); Cockroft, S.L, Chu, J., Amorin, M. and Ghadiri, M.R., "A single-molecule nanopore device detects DNA polymerase activity with single-nucleotide resolution.", J. Am. Chem. Soc. Vol. 130, pp. 818-820 (2008)). Data obtained from nanopore sequencing can be stored, processed, and analyzed as described herein. Specifically, according to the exemplary processing of the optical images and other images described herein, the data can be processed as if it were an image.

[0302] Some embodiments may use methods involving real-time monitoring of DNA polymerase activity. Nucleotide incorporation can be detected by fluorescence resonance energy transfer (FRET) interactions between a polymerase carrying a fluorophore and a γ-phosphate-labeled nucleotide (as described, for example, in U.S. Pat. Nos. 7,329,492 and 7,211,414), or nucleotide incorporation can be detected using zero-mode waveguides (as described, for example, in U.S. Pat. No. 7,315,019), and nucleotide incorporation can be detected using fluorescent nucleotide analogs and engineered polymerases (as described, for example, in U.S. Pat. Nos. 7,405,281 and U.S. Publication 2008 / 0108082). Illumination can be limited to zeptoliter-scale volumes surrounding surface-tethered polymerases, such that incorporation of fluorescently labeled nucleotides can be observed at low background (Levene, M.J. et al., “Zero-mode waveguides for single-molecule analysis at high concentrations.”, Science 299, 682-686 (2003); Lundquist, P.M. et al., “Parallel confocal detection of single molecules in real time.”, Opt. Lett. 33, 1026-1028 (2008); Korlach, J. et al., “Selective aluminum passivation for targeted immobilization of single DNA polymerase molecules in zero-mode waveguide nanostructures.”, Proc. Natl. Acad. Sci. USA, Vol. 105, pp. 1176-1181 (2008)). Images obtained by such methods can be stored, processed, and analyzed as described herein.

[0303] Some SBS embodiments include detecting protons released upon nucleotide incorporation into an extension product. For example, sequencing based on the detection of released protons can use an electrical detector and related techniques commercially available from Ion Torrent (Guilford, CT, a subsidiary of Life Technologies) or sequencing methods and systems described in U.S. Publications 2009 / 0026082, 2009 / 0127589, 2010 / 0137143, and 2010 / 0282617. The methods for amplifying a target nucleic acid using kinetic exclusion described herein can be readily applied to substrates for detecting protons. More specifically, the methods described herein can be used to generate an amplicon clone population for detecting protons.

[0304] The above-described SBS method can advantageously be performed in a variety of formats such that multiple different modified target nucleic acids are manipulated simultaneously. In certain embodiments, different modified target nucleic acids can be processed in a common reaction vessel or on the surface of a particular substrate. This allows for the convenient delivery of sequencing reagents, removal of unreacted reagents, and detection of incorporation events in a variety of ways. In embodiments using surface-bound target nucleic acids, the modified target nucleic acids can be in an array format. In an array format, the modified target nucleic acids can typically be bound to the surface in a spatially distinguishable manner. The modified target nucleic acids can be bound by direct covalent attachment, attachment to beads or other particles, or binding to a polymerase or other molecule attached to the surface. The array can include a single copy of the modified target nucleic acid at each site (also referred to as a feature), or multiple copies having the same sequence can be present at each site or feature. The multiple copies can be generated by amplification methods such as bridge amplification or emulsion PCR, as described in further detail herein.

[0305] The methods described herein can use arrays having features at any of a variety of densities, including, for example, at least about 10 features / cm 2 、100 features / cm 2 、500 features / cm 2 、1,000 features / cm 2 、5,000 features / cm 2 、10,000 features / cm 2 、50,000 features / cm 2 、100,000 features / cm 2 、1,000,000 features / cm 2 、5,000,000 features / cm 2 or higher.

[0306] An advantage of the methods described herein is that these methods provide for multiple cm 2Parallel rapid and efficient detection. Accordingly, the present disclosure provides an integrated system capable of preparing and detecting nucleic acids using techniques known in the art (such as those exemplified herein). Thus, the integrated system of the present disclosure may include fluidic components capable of delivering amplification reagents and / or sequencing reagents to one or more immobilized modified target nucleic acids, and the system includes components such as pumps, valves, reservoirs, fluid lines, etc. The flow cell may be configured in the integrated system for and / or for detecting target nucleic acids. Exemplary flow cells are described in, for example, U.S. Patent 8,241,573 and U.S. Patent 8,951,781. As exemplified for the flow cell, one or more fluidic components of the integrated system may be used for amplification methods and detection methods. Taking a nucleic acid sequencing embodiment as an example, one or more fluidic components of the integrated system may be used for the amplification methods described herein and for delivering sequencing reagents in sequencing methods (such as those exemplified above). Alternatively, the integrated system may include separate fluidic systems to perform the amplification method and to perform the detection method. Examples of integrated sequencing systems capable of generating amplified nucleic acids and also determining nucleic acid sequences include, but are not limited to, the MiSeq TM platform (Illumina, Inc., San Diego, CA) and the device described in U.S. Patent 8,951,781.

[0307] Although the embodiments provided herein are generally described using a sequencing platform (such as a sequencing-by-synthesis platform) as the readout, those of ordinary skill in the art will recognize that nucleic acids modified by the protein complexes provided herein can also be detected using any other suitable readout method. For example, a microarray can be used to evaluate the position and identity of modified cytosines. Any of a variety of analyte arrays (also referred to as "microarrays") known in the art can be used in the methods or systems described herein. A typical array contains analytes, each analyte having a separate probe or group of probes. In the latter case, the group of probes at each analyte is typically homogeneous, having a single type of probe. For example, in the case of a nucleic acid array, each analyte may have multiple nucleic acid molecules, each nucleic acid molecule having a common sequence. However, in some embodiments, the group of probes at each analyte of the array can be heterogeneous. Similarly, a protein array may have analytes containing a single protein or group of proteins, the single protein or group of proteins typically but not always having the same amino acid sequence. The probes can be attached to the surface of the array, for example, by covalent bonding of the probe to the surface or by non-covalent interactions of the probe with the surface. In some embodiments, probes such as nucleic acid molecules can be attached to the surface via a gel layer, as described, for example, in U.S. Patent Application Serial No. 13 / 784,368 and U.S. Patent Application Publication No. 2011 / 0059865A1.

[0308] Exemplary arrays include, but are not limited to, those obtained from Illumina, Inc.TM (BeadChip of (San Diego, Calif.)) TM Arrays or other arrays, such as those in which probes are attached to beads present on a surface (e.g., beads in holes on a surface), such as U.S. Patent Nos. 6,266,459, 6,355,431, 6,770,441, 6,859,570, or 7,622,294, or PCT Publication WO 00 / 63437. Other examples of commercially available microarrays that can be used include, for example, Affymatrix TM Genechip TM Microarrays or other microarrays synthesized according to what is sometimes referred to as VLSIPS TM (Very Large Scale Immobilized Polymer Synthesis) technology. Spot microarrays can also be used in the methods or systems of some specific embodiments according to the present disclosure. An example spot microarray is CodeLink TM obtained from Amersham Biosciences. Another microarray that can be used is a microarray manufactured using an inkjet printing method (such as SurePrint TM technology) obtained from Agilent Technologies.

[0309] In some embodiments, the protein complex provided herein can be used to convert 5-methylcytosine (5mC) to thymidine (T) by deamination as described herein, such as by providing a DNA sample suspected of containing double-stranded DNA, the double-stranded DNA containing at least one 5-methylcytosine (5mC), at least one 5-hydroxymethylcytosine (5hmC), at least one 5-formylcytosine (5fC), at least one 5-carboxylcytosine (5CaC), or a combination thereof; contacting the DNA with the protein complex under conditions suitable for converting 5-methylcytosine (5mC) to thymidine (T) by deamination at a rate greater than the rate of converting cytosine (C) to uracil (U) by deamination to produce a converted double-stranded DNA, wherein 5mC, 5hmC, 5fC, and / or 5CaC are converted to T.

[0310] Depending on the desired result, the sample can be treated with the protein complex before or during the preparation of the sequencing library. In some embodiments, the sample is contacted with the protein complex before tag fragmentation. Before treatment with the protein complex described herein, it may be necessary to fragment the sample tags and add adapters. Since the protein complex described herein is active against double-stranded nucleic acids, the sample can be treated before the step of separating the two strands of the sample (such as bridge amplification). It may be desirable to keep the two strands of the sample in close proximity during the preparation of the sample for sequencing.

[0311] In some embodiments, the sample is contacted with the protein complex after tag fragmentation. In some embodiments, the sample is contacted with the protein complex after tag fragmentation and after treatment with a polymerase to produce a dsDNA sample. As described above, helicases, such as some helicases consistent with those described herein, may be more active against double-stranded nucleic acids. To this end, it may be desirable to convert the sample to a fully double-stranded form prior to treatment with the protein complex described herein.

[0312] In one specific embodiment, the cytidine deaminases of the present disclosure can be used to detect 5hmC as described herein, for example by providing a DNA sample suspected of containing single-stranded DNA having at least one 5-hydroxymethylcytosine (5hmC); contacting the DNA with an altered cytidine deaminase under conditions suitable to convert unmodified cytosine to uracil and 5mC to thymidine and where conversion of 5hmC to 5hmU is undetectable.

[0313] The transformed double-stranded DNA can then be processed as needed to facilitate hybridization to the microarray. For example, the transformed DNA can be amplified. Any of a variety of amplification methods known in the art can be performed. For example, whole genome amplification can be used or amplification using universal primers that hybridize to common regions (such as adapter sequences) in the transformed DNA. Alternatively or in addition, the transformed DNA can be fragmented. Fragmentation can be performed before or after amplification, or without amplification. Any of a variety of fragmentation methods known in the art can be performed. As an example, enzymatic methods can be used for fragmentation, such as restriction endonucleases or other enzymes capable of cleaving the transformed DNA. As another example, mechanical means can be used for fragmentation, such as shearing (using, for example, an ultrasonic device (such as those provided by Covaris)). The fragmented transformed DNA can then be precipitated and / or resuspended in a buffer suitable for hybridization to the microarray. After hybridization, the methylation status of the region of interest (such as one or more specific CpG loci) can be queried at specific locations on the microarray. Methods for preparing transformed DNA for microarray analysis are known in the art. An example of such a method is described in the methylation protocol guide for the Infinium HD assay from Illumina (San Diego, CA). While such a protocol guide may describe the use of a microarray designed to query bisulfite-converted DNA, it should be understood that the array features, particularly the probe sequences, can be specifically designed for DNA that has not been bisulfite-converted. As an example, commercially available microarrays (such as the Infinium MethylationEPIC BeadChip (Illumina)) are specifically designed to hybridize to DNA fragments with reduced complexity, as found in bisulfite-converted DNA, where most (if not all) cytosines are converted to thymidine. Thus, for example, the same CpG sites can be queried in non-bisulfite-converted DNA by using a microarray that includes probes designed to hybridize to the same regions of native, non-bisulfite-converted DNA. Such microarrays can be readily obtained by those skilled in the art. In one embodiment, by using the "forward sequence" to identify probe sequences that contain native DNA sequences, a custom array can be designed using the manifest of the array (such as the Infinium MethylationEPIC BeadChip), where the native DNA sequences cover similar or identical sequence regions of allele-specific probe sequences that are designed to hybridize to DNA sequences in which most or all of the cytosines have been converted to thymidine.Using such arrays designed to hybridize to native (non-bisulfite converted) DNA sequences, methylation of CpG sites in sample DNA can be identified following the methods and analysis methods described in the methylation protocol guidelines for the Infinium HD assay from Illumina (San Diego, CA).

[0314] Composition

[0315] The present disclosure also provides compositions comprising protein complexes. The protein complexes generally comprise a cytidine deaminase and a helicase as described herein, or a helicase-cytidine deaminase complex. In addition to the cytidine deaminase and the helicase, the composition can further comprise one or more additional other components. For example, the other components can include a double-stranded DNA or RNA substrate that comprises or is suspected of comprising at least one modified cytosine. Typically, the modified cytosine is 5-methylcytosine. The modified cytosine can also be 5-hydroxymethylcytosine, 5-formylcytosine (5fC), 5-carboxylcytosine (5CaC), or a combination thereof. In another example, the double-stranded DNA or RNA substrate can be a double-stranded DNA or RNA substrate that comprises one or more known modified cytosines, e.g., a single-stranded DNA or RNA substrate that can be used as a control for measuring conversion efficiency. In another example, the other components can include a buffer having a pH as described herein. In another example, the other components can include a buffer as described herein, such as a citrate buffer, a sodium acetate buffer, or a Bis-Tris buffer. In another example, the other components can include reducing agents, including but not limited to DTT and / or TCEP, and zinc and zinc salts.

[0316] The composition can comprise a polynucleotide encoding a cytidine deaminase as described herein. The polynucleotide can be present in a vector such as a plasmid or a viral vector. The vector comprising the polynucleotide can be present in a host cell such as Escherichia coli.

[0317] Kit

[0318] The present disclosure also provides a kit for determining the methylation status of DNA or RNA. The kit includes at least one cytidine deaminase and helicase or protein complex as described herein and one or more other components in a suitable packaging material in an amount sufficient to perform at least one reaction. Examples of other components include positive control polynucleotides, such as double-stranded DNA containing one or more known modified cytosines for measuring conversion efficiency, or negative control polynucleotides, such as double-stranded DNA containing unmodified cytosines. Another component can be a glucosyltransferase, such as T4-β glucosyltransferase. Optionally, other reagents required for using the cytidine deaminase and nucleotide solution are also included, such as buffers and solutions. Instructions for using the packaging components are generally included.

[0319] As used herein, the phrase "packaging material" refers to one or more physical structures for containing the contents of the kit. The packaging material is constructed by known methods and is preferably used to provide a sterile, contaminant-free environment. The packaging material has a label that indicates that the components can be used to determine the methylation status of DNA or RNA. In addition, the packaging material contains instructions on how to use the materials in the kit to perform a reaction using the cytidine deaminase. As used herein, the term "package" refers to a solid matrix or material such as glass, plastic, paper, foil, etc., that can contain a polypeptide within a fixed range. The "instructions for use" generally include a tangible expression describing reagent concentrations or at least one assay method parameter, such as the relative amounts of reagents and samples to be mixed, the maintenance period of the reagent / sample mixture, temperature, buffer conditions, etc.

[0320] The invention is defined in the claims. However, a non-exhaustive list of non-limiting exemplary aspects is provided below. Any one or more of the features of these aspects can be combined with any one or more of the features of another example, embodiment, or aspect described herein.

[0321] Exemplary aspects

[0322] Aspect 1 is a protein complex comprising a cytidine deaminase and a helicase.

[0323] Aspect 2 is the protein complex according to aspect 1, wherein the cytidine deaminase and the helicase are attached.

[0324] Aspect 3 is the protein complex according to aspect 1 or 2, wherein the cytidine deaminase and the helicase comprise a fusion protein.

[0325] Aspect 4 is the protein complex according to any one of the preceding aspects, wherein the fusion protein comprises a linker.

[0326] Aspect 5 is the protein complex according to any one of the preceding aspects, wherein the cytidine deaminase and the helicase are biochemically conjugated.

[0327] Aspect 6 is the protein complex according to any one of the preceding aspects, wherein the cytidine deaminase is an altered cytidine deaminase, and the altered cytidine deaminase comprises an amino acid substitution mutation at a position functionally equivalent to (Tyr / Phe)130 in the wild-type APOBEC3A protein.

[0328] Aspect 7 is the protein complex according to any one of the preceding aspects, wherein the (Tyr / Phe)130 is Tyr130, and the wild-type APOBEC3A protein is SEQ ID NO: 3.

[0329] Aspect 8 is the protein complex according to any one of the preceding aspects, wherein the substitution mutation at a position functionally equivalent to Tyr130 comprises a mutation to Ala, Val or Trp.

[0330] Aspect 9 is the protein complex according to any one of the preceding aspects, wherein the altered cytidine deaminase converts 5-methylcytosine (5mC) to thymidine (T) by deamination at a rate greater than the rate of converting cytosine (C) to uracil (U) by deamination.

[0331] Aspect 10 is the protein complex according to any one of the preceding aspects, wherein the rate is at least 100-fold greater.

[0332] Aspect 11 is the protein complex according to any one of the preceding aspects, wherein the altered cytidine deaminase is a member of the AID subfamily, APOBEC1 subfamily, APOBEC2 subfamily, APOBEC3A subfamily, APOBEC3B subfamily, APOBEC3C subfamily, APOBEC3D subfamily, APOBEC3F subfamily, APOBEC3G subfamily, APOBEC3G subfamily, APOBEC3H subfamily or APOBEC4 subfamily.

[0333] Aspect 12 is the protein complex according to any one of the preceding aspects, wherein the altered cytidine deaminase comprises the ZDD motif H-[P / A / V]-E-X [23-28] -P-C-X [2-4] -C (SEQ ID NO: 12).

[0334] Aspect 13 is the protein complex according to any of the preceding aspects, wherein the altered cytidine deaminase is a member of the APOBEC3A subfamily and comprises the ZDD motif HXEX24SW(S / T)PCX[2-4]CX6FX8LX5R(L / I)YX[8-11]LX2LX

[10] M (SEQ ID NO: 13).

[0335] Aspect 14 is the protein complex according to any of the preceding aspects, wherein the helicase unwinds double-stranded DNA (dsDNA).

[0336] Aspect 15 is the protein complex according to any of the preceding aspects, wherein the helicase is a member of helicase superfamily 1, superfamily 2, superfamily 3, or family 4.

[0337] Aspect 16 is the protein complex according to any of the preceding aspects, wherein the helicase comprises a Walker A motif (SEQ ID NO: 68), a Walker B motif (SEQ ID NO: 69), or both.

[0338] Aspect 17 is the protein complex according to any of the preceding aspects, wherein the helicase is a member of the RecQ family.

[0339] Aspect 18 is the protein complex according to any of the preceding aspects, wherein the helicase comprises the RecQ helicase from Escherichia coli (SEQ ID NO: 92).

[0340] Aspect 19 is the protein complex according to any of the preceding aspects, wherein the protein complex converts 5mC to T by deamination in dsDNA at a rate greater than that of a comparable protein complex without a helicase.

[0341] Aspect 20 is the protein complex according to any of the preceding aspects, wherein the attachment protein converts 5mC to T by deamination in dsDNA at a rate greater than that of a comparable protein complex without the fusion protein.

[0342] Aspect 21 is a composition comprising: the protein complex according to any one of aspects 1 to 20; and a DNA sample comprising single-stranded DNA, the single-stranded DNA comprising at least one 5mC.

[0343] Aspect 22 is the composition according to any of the preceding aspects, wherein the sample comprises genomic DNA.

[0344] Aspect 23 is the composition according to any of the preceding aspects, wherein the genomic DNA is from a single cell or a mixture of multiple cells.

[0345] Aspect 24 is a composition comprising: a protein complex comprising a cytidine deaminase and a helicase; and a sample comprising double-stranded DNA containing at least one 5mC.

[0346] Aspect 25 is the composition according to any one of the preceding aspects, wherein the cytidine deaminase and the helicase are attached.

[0347] Aspect 26 is the composition according to any one of the preceding aspects, wherein the cytidine deaminase and the helicase comprise a fusion protein.

[0348] Aspect 27 is the composition according to any one of the preceding aspects, wherein the fusion protein comprises a linker.

[0349] Aspect 28 is the composition according to any one of the preceding aspects, wherein the cytidine deaminase and the helicase are chemically conjugated.

[0350] Aspect 29 is the composition according to any one of the preceding aspects, wherein the sample comprises a helicase recruitment motif.

[0351] Aspect 30 is the composition according to any one of the preceding aspects, wherein the helicase recruitment motif comprises a Y-shaped adaptor structure, a partial duplex, a single-stranded DNA (ssDNA) bubble, a fork, a flat duplex, a three-way junction, or a combination thereof.

[0352] Aspect 31 is the composition according to any one of the preceding aspects, wherein the sample comprises genomic DNA, cell-free DNA.

[0353] Aspect 32 is the composition according to any one of the preceding aspects, wherein the genomic DNA is from a single cell or a mixture of multiple cells.

[0354] Aspect 33 is the composition according to any one of the preceding aspects, wherein the cytidine deaminase comprises an altered cytidine deaminase comprising an amino acid substitution mutation at a position functionally equivalent to (Tyr / Phe)130 in the wild-type APOBEC3A protein.

[0355] Aspect 34 is the composition according to any one of the preceding aspects, wherein the (Tyr / Phe)130 is Tyr130, and the wild-type APOBEC3A protein is SEQ ID NO: 3.

[0356] Aspect 35 is the composition according to any one of the preceding aspects, wherein the substitution mutation at a position functionally equivalent to Tyr130 includes a mutation to Ala, Val, or Trp.

[0357] Aspect 36 is the composition according to any of the foregoing aspects, wherein the altered cytidine deaminase converts 5-methylcytosine (5mC) to thymidine (T) by deamination at a rate greater than the rate of converting cytosine (C) to uracil (U) by deamination.

[0358] Aspect 37 is the composition according to any of the foregoing aspects, wherein the rate is at least 100-fold greater.

[0359] Aspect 38 is the composition according to any of the foregoing aspects, wherein the altered cytidine deaminase is a member of the AID subfamily, APOBEC1 subfamily, APOBEC2 subfamily, APOBEC3A subfamily, APOBEC3B subfamily, APOBEC3C subfamily, APOBEC3D subfamily, APOBEC3F subfamily, APOBEC3G subfamily, APOBEC3G subfamily, APOBEC3H subfamily or APOBEC4 subfamily.

[0360] Aspect 39 is the composition according to any of the foregoing aspects, wherein the altered cytidine deaminase comprises the ZDD motif H-[P / A / V]-E-X [23-28] -P-C-X [2-4] -C (SEQ ID NO: 12).

[0361] Aspect 40 is the composition according to any of the foregoing aspects, wherein the altered cytidine deaminase is a member of the APOBEC3A subfamily and comprises the ZDD motif HXEX24SW(S / T)PCX[2-4]CX6FX8LX5R(L / I)YX[8-11]LX2LX

[10] M (SEQ ID NO: 13).

[0362] Aspect 41 is the composition according to any of the foregoing aspects, wherein the substitution mutation at the position functionally equivalent to Tyr130 includes a mutation to alanine, glycine, phenylalanine, histidine, glutamine, methionine, asparagine, lysine, valine, aspartic acid, glutamic acid, serine, cysteine, proline, arginine or threonine.

[0363] Aspect 42 is the composition according to any of the foregoing aspects, wherein the substitution mutation at the position functionally equivalent to (Tyr / Phe)130 includes a mutation to alanine.

[0364] Aspect 43 is the composition according to any of the foregoing aspects, wherein the cytidine deaminase converts cytosine (C) to uracil (U) by deamination at a rate greater than the rate of converting 5-methylcytosine (5mC) to thymidine (T) by deamination.

[0365] Aspect 44 is the composition according to any of the foregoing aspects, wherein the substitution mutation at the position functionally equivalent to Tyr130 includes a mutation to leucine or tryptophan.

[0366] Aspect 45 is the composition according to any of the foregoing aspects, wherein the helicase unwinds double-stranded DNA (dsDNA).

[0367] Aspect 46 is the composition according to any of the foregoing aspects, wherein the helicase is a member of helicase superfamily 1, superfamily 2, superfamily 3 or family 4.

[0368] Aspect 47 is the composition according to any of the foregoing aspects, wherein the helicase comprises a Walker A motif (SEQ ID NO: 68), a Walker B motif (SEQ ID NO: 69), or both.

[0369] Aspect 48 is the composition according to any of the foregoing aspects, wherein the helicase is a member of the RecQ family.

[0370] Aspect 49 is the composition according to any of the foregoing aspects, wherein the helicase comprises the RecQ helicase from Escherichia coli (SEQ ID NO: 92).

[0371] Aspect 50 is the composition according to any of the foregoing aspects, wherein the sample comprises a helicase recruitment motif.

[0372] Aspect 51 is the composition according to any of the foregoing aspects, wherein the helicase recruitment motif comprises a Y-shaped adaptor structure, a partial duplex, a single-stranded DNA (ssDNA) bubble, a fork, a flat duplex, a triple junction, or a combination thereof.

[0373] Aspect 52 is a method, the method comprising: providing a DNA sample suspected of containing dsDNA, the dsDNA comprising at least one 5-methylcytosine (5mC); contacting the dsDNA with the protein complex according to any one of aspects 1 to 20 under conditions suitable for converting 5-methylcytosine (5mC) to thymidine (T) to produce a converted double-stranded DNA, wherein 5mC is converted to T; and processing the converted double-stranded DNA to produce a sequencing library.

[0374] Aspect 53 is the method according to any of the foregoing aspects, the method further comprising: providing a surface comprising a plurality of amplification sites, wherein the amplification sites comprise at least two populations of attached single-stranded capture oligonucleotides having free 3′ ends, and contacting the surface comprising the amplification sites with the sequencing library under conditions suitable for producing a plurality of amplification sites, the plurality of amplification sites each comprising a clonal population of amplicons from a single member of the sequencing library.

[0375] Aspect 54 is the method according to any of the foregoing aspects, wherein the attached single-stranded capture oligonucleotide comprises a helicase recruitment motif.

[0376] Aspect 55 is the method according to any of the foregoing aspects, wherein the helicase recruitment motif comprises a Y-shaped adaptor structure, a partial duplex, a single-stranded DNA (ssDNA) bubble, a fork, a flat duplex, a triple junction, or a combination thereof.

[0377] Aspect 56 is the method according to any of the foregoing aspects, wherein processing the transformed double-stranded DNA to generate a sequencing library occurs after contacting the dsDNA with the protein complex.

[0378] Aspect 57 is the method according to any of the foregoing aspects, wherein processing the transformed double-stranded DNA to generate a sequencing library occurs before contacting the dsDNA with the protein complex.

[0379] Aspect 58 is the method according to any of the foregoing aspects, wherein the processing comprises fragmentation or tagmentation of the double-stranded DNA and addition of universal sequences to the double-stranded DNA fragments.

[0380] Aspect 59 is the method according to any of the foregoing aspects, wherein the universal sequence is part of an adaptor added to the double-stranded DNA fragments.

[0381] Aspect 60 is the method according to any of the foregoing aspects, wherein the sample is a biological sample.

[0382] Aspect 61 is the method according to any of the foregoing aspects, wherein the biological sample comprises cell-free DNA.

[0383] Aspect 62 is the method according to any of the foregoing aspects, wherein the biological sample comprises a fluid selected from blood or serum.

[0384] Aspect 63 is the method according to any of the foregoing aspects, wherein the sample comprises single cells or isolated cell nuclei.

[0385] Aspect 64 is the method according to any of the foregoing aspects, wherein the biological sample comprises tissue.

[0386] Aspect 65 is the method according to any of the foregoing aspects, wherein the tissue comprises tumor tissue.

[0387] Aspect 66 is a protein complex comprising: a pathway for converting 5-methylcytosine to thymine; and a pathway for directionally separating double-stranded nucleic acids.

[0388] Examples

[0389] The present disclosure is illustrated by the following examples. It should be understood that specific examples, materials, amounts, and procedures should be interpreted broadly in accordance with the scope and spirit of the present disclosure as described herein.

[0390] Example 1

[0391] Experimental determination of cytidine deaminase activity

[0392] Use a gel-based assay to monitor deamination by cytidine deaminase using the restriction enzyme SwaI ( Figure 4 ).

[0393] Deamination of C to U or 5mC to T produces a mismatched DNA substrate that can be cleaved by SwaI. These products appear as two distinct species on a denaturing polyacrylamide gel.

[0394] After incubation with APOBEC3A wild-type or its mutants and a 5′-fluorescein amidite (FAM)-labeled oligonucleotide substrate containing C or 5mC, deamination occurs or not, depending on the substrate preference of the enzyme. Introduction of a complementary oligonucleotide thereafter can result in perfectly matched or mismatched base pairing. If a mismatch is present, SwaI will cleave the double-stranded oligonucleotide. The original or cleaved 5′-FAM-labeled oligonucleotide can be visualized by the degree of migration measured using 15% urea-PAGE and a FAM filter. Although FAM is used here, essentially any label can be used, and either the 5′ or 3′ end can be labeled.

[0395] This assay was modified and adapted from Schutsky et al., Nucleic Acid Research, 45, 7655 - 7665, 2017. doi:10.1093 / nar / gkx345. Modifications to Schutsky et al. include the following. As an alternative to performing DNA precipitation and redissolving the DNA substrate in SwaI - compatible buffer, 1 μL of the altered cytidine deaminase APOBEC3A(Y130A) deamination reaction mixture was aliquoted into 9 μL of SwaI - compatible buffer for restriction enzyme digestion and thus for our gel assay. Appropriate controls were performed to determine that the SwaI restriction enzyme digestion efficiency was not affected by the APOBEC reaction buffer. As an alternative to introducing a 1.5 - fold excess of the complementary strand prior to overnight SwaI restriction enzyme digestion, a 3 - fold excess of the complementary strand was introduced. As an alternative to running the pre - warmed 20% acrylamide / Tris - borate - EDTA (TBE) / urea gel reported by Schutsky et al., the gel run was performed at room temperature with a 15% acrylamide / Tris - borate - EDTA (TBE) / urea gel, and good separation between the cut (deaminated) and uncut (unreacted) oligonucleotide substrates was observed.

[0396] The SwaI assay was first validated using FAM - labeled DNA oligonucleotides containing C, 5mC, U, or T residues (Figure 5). The oligonucleotides were purchased from Integrated DNA Technologies (IDT) and visualized by 15% urea - PAGE and FAM filter. The synthetic oligonucleotide oLB1609 contained the substrate C, and oLB1610 contained its corresponding deaminated product U. oLB1611 contained the substrate 5mC, and oLB1612 contained its corresponding deaminated product T. oLB1679 was the complementary oligonucleotide of oLB1609, oLB1610, oLB1611, and oLB1612 ( Figure 5A ). As Figure 5B shown, annealing of oLB1609 / oLB1679 and oLB1611 / oLB1679 resulted in complete base - pairing of the oligonucleotides. Addition of SwaI did not result in cleavage of this substrate. However, single - base - mismatched substrates formed by annealing of oLB1610 / oLB1679 (U / G mismatch) or oLB1612 / oLB1679 (T / G mismatch) produced cleavage products upon addition of SwaI. Since specific cleavage products were observed in the presence of U or T but not in the presence of C or 5mC, the SwaI assay can serve as a read - out for C - to - U and 5mC - to - T deamination by APOBEC3A.

[0397] As another control, DNA oligonucleotides oLB1609 (C oligonucleotide) and oLB1612 (5mC oligonucleotide) containing C and 5mC were incubated with APOBEC3A enzyme (NEBNext TM Enzymatic Methyl-seq kit (Catalog No. E7120)) purchased from New England Biolabs, which is reported to efficiently deaminate C and 5mC ( Figure 6 ). Then, the deamination reaction mixtures (5 μL, 2 μL, and 1 μL) were directly added to SwaI assay buffer containing SwaI to give a total volume of 10 μL. 1 μL of the deamination reaction was sufficient to observe the cleavage band, thus allowing quantification of the degree of deamination. Treatment with NEB APOBEC3A and subsequent digestion with SwaI formed cleavage products with similar mobilities to the Figure 5A substrates containing U and T. These data further support the view that SwaI cleavage can be used to monitor the deaminase activities of APOBEC3A and other cytidine deaminases.

[0398] Example 2

[0399] Purification of APOBEC3A(Y130X) mutant protein

[0400] The effects of all possible amino acid substitutions at position 130 of APOBEC3A on the deaminase activity of the enzyme were systematically evaluated. For this purpose, 19 different His-tagged APOBEC3A constructs were cloned, each encoding a different amino acid relative to wild-type tyrosine at position 130. The corresponding proteins were expressed in BL21(DE3) cells, purified using Ni-NTA agarose beads, and desalted / concentrated to the storage buffer (50 mM Tris pH 7.5, 200 mM NaCl, 5% (v / v) glycerol, 0.01% (v / v) Tween-20, 0.5 mM DTT) using a centrifugal column. As judged by SDS-PAGE analysis, this yielded APOBEC3A(Y130X) mutant protein preparations with 80% to 85% purity ( Figure 7 A).

[0401] Example 3

[0402] DNA deaminase activity of APOBEC3A(Y130X) mutant protein

[0403] Then, the deaminase activities of all purified APOBEC3A(Y130X) proteins were analyzed using the SwaI assay under the conditions of a 37 °C / 2-hour reaction time, and NEB APOBEC3A was used as a positive control ( Figure 8)。Incubate the Y130X recombinase at a final concentration of 10 μM to 20 μM with oLB1609 (C oligomer, upper panel) and oLB1612 (5mC oligomer, lower panel) at 37 °C for 2 hours. The NEB APOBEC3A enzyme was purchased from the NEBNext TM Enzymatic Methyl-seq Kit (Catalog No. E7120). Wild-type APOBEC3A completely deaminates 5mC and C substrates, which is consistent with previous literature. Different mutants exhibit a wide range of reactivity towards 5mC and C substrates, some of which show a preference for either substrate. Notably, APOBEC3A(Y130A) (first panel) almost completely deaminates the 5mC substrate (94.2%), while it deaminates the corresponding C substrate to a lesser extent (29.4%). Other mutants, such as APOBEC3A(Y130P) and APOBEC3A(Y130T), also show more complete 5mC deamination than the C substrate, although to a lesser extent than APOBEC3A(Y130A). In contrast, APOBEC3A(Y130L) (second panel) deaminates approximately half of the C substrate (56%), but hardly deaminates the 5mC substrate (6.8%). The deaminase activities of all APOBEC3A(Y130X) mutants were quantified and summarized in Figure 8 , Figure 9 and Figure 10 .

[0404] Since these SwaI assays were performed as single-endpoint measurements (2 hours), it is possible that the corresponding deamination reactions have saturated. Therefore, a time-course analysis of the APOBEC3A(Y130A) deaminase activity was performed. By incubating approximately 10 μM to 20 μM APOBEC(Y130A) with 500 nM C and 5mC oligonucleotide substrates, the extent of C and 5mC deamination was monitored at 0 minutes, 5 minutes, 10 minutes, 30 minutes, 60 minutes, and 120 minutes ( Figure 10 ). A large difference in the extent of 5mC and C deamination was observed at t ≤ 30 minutes.

[0405] A quantitative comparison of the kinetics of deamination by wild-type APOBEC3A and mutant APOBEC3A(Y130A) was performed. The initial deamination reaction rates were measured at a range of DNA substrate concentrations, and these rates were used to construct Michaelis curves for the 5mC and C substrates, respectively. Then, the resulting Km and Kcat values were derived from these data. The catalytic efficiency of APOBEC3A(Y130A) for 5mC is approximately 100-fold higher than that for the C substrate (Figure 11), which confirms the Figure 8 , Figure 9 and Figure 10 endpoint SwaI assays shown in

[0406] Example 4

[0407] Purification of APOBEC3A(Y130A-Y132H) double mutant protein

[0408] The APOBEC3A (Y130A - Y132H) protein was expressed in BL21(DE3) cells, purified using Ni - NTA agarose beads, and desalted / concentrated to the storage buffer (50 mM Tris pH 7.5, 200 mM NaCl, 5% (v / v) glycerol, 0.01% (v / v) Tween - 20, 0.5 mM DTT) using a centrifugal column. As judged by SDS - PAGE analysis, this yielded an APOBEC3A (Y130A - Y132H) mutant protein preparation with 90% to 95% purity ( Figure 7 B).

[0409] Example 5

[0410] DNA deaminase activity of APOBEC3A(Y130A-Y132H) double mutant protein

[0411] The deaminase activity of the purified APOBEC3A (Y130A - Y132H) double - mutant protein was then analyzed using the SwaI assay, with a reaction time of 37 °C / 2 h, and NEB APOBEC3A as a positive control. The conditions used were the same as those described in Example 3, but the SwaI assay used the following reaction conditions: 40 mM sodium acetate pH 5.2, 37 °C for 1 h to 16 h. The DNA substrate is shown in Figure 12 . After the deaminase reaction, the deaminated oligonucleotide substrate was subjected to PCR amplification, sequencing, and the number of C and 5mC deamination events for each read was counted. The DNA oligonucleotide substrates used for the APOBEC3A (Y130A - Y132H) experiment are shown in Figure 12 . APOBEC3A (Y130A - Y132H) exhibited a higher level of deamination at all methylated sites compared to unmethylated sites. This was consistent in both CpG and non - CpG contexts and was robust to changes in reaction time Figure 13 、 14 . The difference in deamination levels between methylated and unmethylated sites of APOBEC3A (Y130A - Y132H) was significantly higher than that of APOBEC3A (Y130A), indicating that APOBEC3A (Y130A - Y132H) achieved better discrimination of methylated sites than APOBEC3A (Y130A). In addition, among all xCpGx motifs, APOBEC3A (Y130A - Y132H) was able to deaminate methylated sites more efficiently compared to unmethylated sites ( Figure 15 ).

[0412] Example 6

[0413] Purification of APOBEC3A(Y130W) mutant protein

[0414] The recombinant human His-tagged APOBEC3A (Y130W) protein was expressed in Escherichia coli BL21(DE3) cells, purified using Ni-NTA affinity chromatography, and desalted / concentrated to the storage buffer (50 mM Tris pH 7.5, 200 mM NaCl, 5% (v / v) glycerol, 0.01% (v / v) Tween-20, 0.5 mM DTT) using a centrifugal column.

[0415] Example 7

[0416] DNA deaminase activity of APOBEC3A(Y130W) mutant protein

[0417] Using the SwaI assay, we measured the activity of APOBEC3A (Y130W) on C, 5mC, and 5hmC substrates at 90-minute intervals. While most C and 5mC were deaminated, no detectable activity was observed for 5hmC. ( Figure 17 )

[0418] To better understand the substrate preference of APOBEC3A (Y130W), we determined the deaminase activity against C, 5mC, and all oxidized derivatives of 5mC ( Figure 18 ). A 90-minute reaction time of the SwaI restriction enzyme assay was performed to measure the deaminase activity of the APOBEC3A wild-type enzyme and the Y130W mutant enzyme. A protein-free control was included to account for potential degradation of the oligonucleotide substrate and to account for non-specific activity of SwaI during the assay. APOBECY130A (Y130W) did not show detectable deamination of 5hmC, 5fC, and 5caC, while the wild-type APOBEC retained significant deaminase activity against 5hmC and showed residual activity against 5fC and 5caC. Both enzymes effectively deaminated C and 5mC.

[0419] Example 8

[0420] Detection of 5mC

[0421] The following examples generally follow the method described by Schutsky et al. Nature biotechnology, 10.1038 / nbt.4204. October 8, 2018, doi: 10.1038 / nbt.4204. Genomic DNA (gDNA) is provided from an organism suspected of having 5hmC. Then, a 20 ng gDNA mixture is processed as described below.

[0422] To provide single-stranded DNA and promote deamination activity, 1 μL of DMSO is added, and the sample is denatured at 95 °C for 5 minutes and rapidly cooled by transferring to a PCR tube rack pre-warmed at -80 °C. Before thawing, the reaction buffer is overlaid to a final concentration of 20 mM MES pH 6.0 + 0.1% Tween, and the modified cytidine deaminase described herein is added to the sample to a final concentration of 5 μM in a total volume of 10 μL. The deamination reaction is incubated from 4 °C to 50 °C over 2 hours under linearly increasing temperature conditions.

[0423] After deamination, the Accel Methyl-NGS kit (Swift Biosciences) is used to prepare samples for Illumina sequencing. Specifically, the reaction is purified using the Zymo Oligo Clean and Concentrator kit and eluted in 15 μL of elution buffer (10 mM Tris pH 8.0). Then, library preparation for single-stranded DNA is performed using the Accel-NGS Methyl-Seq kit (Swift Biosciences) according to the manufacturer's instructions. After purifying the library, 1 ng to 5 ng of the amplified DNA is run on an Agilent Bioanalyzer high-sensitivity DNA chip to confirm appropriate library fragment size. The resulting ACE-Seq library is sequenced in single-end mode at 1.9 pM on a NextSeq 500 sequencer (Illumina) using the NextSeq 500 / 500 High Output kit v2 (150 cycles).

[0424] The sequencing data from the samples is analyzed to detect 5mC at 500 CpG alleles suspected of having 5mC modification. Those CpG sites that sequence as thymidine are characterized as potentially containing 5mC modification in the original sample.

[0425] Example 9

[0426] Detection of 5hmC

[0427] The following examples generally follow the method described by Schutsky et al., Nature biotechnology, 10.1038 / nbt.4204. October 8, 2018, doi: 10.1038 / nbt.4204. Genomic DNA (gDNA) from an organism suspected of having 5hmC is provided. Two aliquots of 20 ng each are provided. One aliquot of the 20 ng gDNA mixture is glucosylated using UDP-glucose and T4 β-glucosyltransferase (βGT, NEB). Specifically, 0.5 μL of 10X Cutsmart buffer (NEB), 0.1 μL of 50X UDP-glucose, 0.5 μL of T4 βGT (NEB) are combined with 20 ng gDNA and water is added to a total volume of 5 μL. An aliquot of the other 20 ng gDNA is used to assemble a control reaction and the reaction components are mixed as described above, but 0.5 μL of water is added in place of T4βGT. The reactions are incubated at 37 °C for 1 hour.

[0428] To provide single-stranded DNA and facilitate deamination activity, 1 μL of DMSO is added and the samples are denatured at 95 °C for 5 minutes and quickly cooled by transferring to a PCR tube rack pre-warmed at -80 °C. Before thawing, the reaction buffer is overlaid to a final concentration of 20 mM MES pH 6.0 + 0.1% Tween and the modified cytidine deaminase is added to each sample to a final concentration of 5 μM in a total volume of 10 μL. The deamination reaction is incubated from 4 °C to 50 °C over 2 hours under linearly increasing temperature conditions.

[0429] After deamination, the Accel Methyl-NGS kit (Swift Biosciences) is used to prepare samples for Illumina sequencing. Specifically, the reactions are purified using the Zymo Oligo Clean and Concentrator kit and eluted in 15 μL of elution buffer (10 mM Tris pH 8.0). Library preparation is performed using the Accel-NGS Methyl-Seq kit (Swift Biosciences) according to the manufacturer's instructions. After purifying the library, 1 ng to 5 ng of the amplified DNA is run on an Agilent Bioanalyzer high-sensitivity DNA chip to confirm proper library fragment size. The resulting ACE-Seq library is sequenced in single-end mode at 1.9 pM on a NextSeq 500 sequencer (Illumina) using the NextSeq 500 / 500 High Output kit v2 (150 cycles).

[0430] Sequencing data from samples and controls were analyzed to detect 5hmC and 5mC at 500 CpG alleles suspected of having 5hmC modification. Specifically, any CpG site that sequenced as cytosine in the control sample was compared to the same CpG site in the sample treated with βGT. Those sites that sequenced as thymidine were characterized as potentially containing 5hmC modification in the original sample.

[0431] Example 10

[0432] Detection of 5mC and 5hmC

[0433] The following examples generally follow the method described by Schutsky et al. Nature biotechnology, 10.1038 / nbt.4204. October 8, 2018, doi: 10.1038 / nbt.4204. Genomic DNA (gDNA) was provided from an organism suspected of having 5mC and 5hmC. Two aliquots of 20 ng each were provided. Then, aliquots of each 20 ng gDNA mixture were processed as described below.

[0434] To provide single-stranded DNA and promote deamination activity, 1 μL of DMSO was added and the sample was denatured at 95 °C for 5 minutes and rapidly cooled by transferring to a PCR tube rack pre-incubated at -80 °C. Before thawing, the reaction buffer was overlaid to a final concentration of 20 mM MES pH 6.0 + 0.1% Tween. For one aliquot, the modified cytidine deaminase (Y130W) described herein was added to the sample to a final concentration of 5 μM in a total volume of 10 μL. For the other aliquot, wild-type cytidine deaminase was added to the sample to a final concentration of 5 μM in a total volume of 10 μL. The deamination reaction was incubated from 4 °C to 50 °C over 2 hours under linearly increasing temperature conditions.

[0435] After deamination, the Accel Methyl-NGS kit (Swift Biosciences) was used to prepare samples for Illumina sequencing. Specifically, the Zymo Oligo Clean and Concentrator kit was used to purify the reaction and elute in 15 μL of elution buffer (10 mM Tris pH 8.0). Then, the Accel-NGS Methyl-Seq kit (Swift Biosciences) was used to perform library preparation for single-stranded DNA according to the manufacturer's instructions. After purifying the library, 1 ng to 5 ng of amplified DNA was run on an Agilent Bioanalyzer High Sensitivity DNA chip to confirm appropriate library fragment size. The resulting APOBEC-coupled epigenetic sequencing (ACE-Seq) library was sequenced at 1.9 pM in single-end mode on a NextSeq 500 sequencer (Illumina) using the NextSeq 500 / 500 High Output kit v2 (150 cycles).

[0436] Sequencing data from the samples was analyzed to detect 5mC at 500 CpG alleles suspected of having 5mC and / or 5hmC modifications. Those CpG sites that were predominantly sequenced as thymidine in wild-type samples but as cytosine in Y130W samples were characterized as potentially containing 5hmC modifications in the original samples.

[0437] Example 11

[0438] Detection of 5mC and 5hmC

[0439] Human genomic DNA was combined with fully unmethylated λ control DNA (New England Biolabs) and enzymatically CpG-methylated pUC19 control DNA and mechanically sheared to generate fragments of approximately 300 bp. Then, according to the standard Illumina library preparation procedure, this sheared DNA (10 ng to 100 ng) was subjected to end repair, A-tailing, and adapter ligation. The sample was then split into 2 aliquots for different treatments.

[0440] Glucosylate an aliquot of adapter-ligated DNA using UDP-glucose and T4 β-glucosyltransferase (βGT, NEB). Specifically, treat the sample with 0.5 U / μL of T4 βGT (NEB) in 1X CutSmart buffer (NEB) containing 40 μM UDP-glucose at 37 °C for 3 hours. Use another aliquot of the DNA sample to assemble a control reaction, omitting the T4 βGT enzyme. Subsequently, optionally subject the sample to SPRI purification and then denature it by incubating at 50 °C in 0.02 N sodium hydroxide for 10 minutes. Subsequently, incubate the ssDNA sample with cytidine deaminase (200 nM) in 50 mM Bis-Tris (pH 6.5), 10 μg / mL RNase A at 37 °C for 25 minutes for enzymatic deamination. Then perform PCR amplification of the library using unique dual-index primers and Q5U (New England Biolabs) for 9 cycles. Sequence the sample on a NovaSeq 6000 and analyze it using the DRAGEN methylation pipeline. Compare the methylation calls between the two treatment conditions (+ / - T4 βGT) to determine which sites are hydroxymethylated. Bases detected as methylated in both sample treatments are designated as mC, while bases detected as methylated under the -βGT condition but unmethylated under the +βGT condition are designated as hmC( Figure 19 ).

[0441] Example 12

[0442] General method for generating methylation sequencing library

[0443] According to standard library preparation procedures, first subject the sheared DNA or cfDNA (10 ng to 200 ng) to end repair, A-tailing, and adapter ligation. Suitable adapters for this method can have unmodified C or pyrrolo-C modifications to discourage deamination of the adapter sequence. Then denature the adapter-ligated DNA into ssDNA. Subsequently, incubate the ssDNA sample in a buffer solution with engineered deaminase (APOBEC3A-Y130A-Y132H, 50 nM to 1000 nM) for enzymatic deamination for 5 minutes to 3 hours at an incubation temperature in the range of 20 °C to 55 °C. Use unique dual-index primers before PCR amplification and use uracil-tolerant polymerase (e.g., KapaU TM ) or uracil-intolerant polymerase (e.g., Q5 KAPA HiFi TM)Optionally, SPRI purify the deaminated library using 9 to 12 cycles of PCR. Then sequence the library on a NextSeq 550 or NovaSeq 6000 and analyze it using the DRAGEN methylation pipeline.

[0444] Method for denaturation

[0445] A variety of methods for denaturation are known to those skilled in the art. These methods include, but are not limited to: (i) heating at a moderate temperature in the presence of NaOH or a high pH buffer (e.g., 0.02 N sodium hydroxide at 50 °C for 10 minutes); (ii) heating to a high temperature (e.g., 95 °C for 10 minutes); (iii) heating in the presence of DMF (e.g., 50% DMF at 95 °C for 10 minutes); (iv) heating in the presence of formamide (e.g., 50% formamide at 95 °C for 10 minutes); and (v) heating in the presence of DMSO (e.g., 50% DMSO at 95 °C for 10 minutes).

[0446] Suitable buffer for deamination

[0447] Typically, a deamination buffer is added to the denatured DNA sample, so any additives that facilitated denaturation will be present at a lower (diluted) concentration in the deamination reaction.

[0448] Several buffer systems in which deamination can occur have been identified. Buffer 50 mM Bis Tris, pH 6.5 can be used, and other pH levels between 5 and 7.5 are feasible. Additionally, other buffer strengths (concentrations) are feasible. Those skilled in the art will recognize that if NaOH is used for denaturation, an appropriate buffer strength / pH can be used to ensure that the final pH is within the ideal range for the deaminase. Alternative buffer systems containing MES (e.g., tcichemicals.com / US / en / c / 10367) can also give good results.

[0449] In some embodiments, when used for deamination of genomic DNA (gDNA) samples, other buffer systems have been shown to have lower performance. Examples include citrate buffer and sodium acetate buffer ( Figure 20)。These buffer systems may have lower performance due to their sodium content, which is believed to be detrimental to the activity of the altered deaminase. APOBEC3A is a zinc-dependent enzyme (Marx et al., Scientific Reports 5, no. December 2015: 1-9. https: / / doi.org / 10.1038 / srep18191), and thus the presence of metal cations (e.g., sodium, magnesium) can generally be harmful. Another explanation is that metal cations can affect the formation of secondary structures in nucleic acids (Einert et al., Biophysical Journal 100, no. 11 (2011): 2745-53, doi.org / 10.1016 / j.bpj.2011.04.038), which can also affect the activity of the deaminase towards its substrate, as APOBEC3A cannot deaminate C in double-stranded DNA such as hairpins. Therefore, any buffer system that maintains the pH within the desired range is expected to be usable, and in some embodiments, a buffer with minimal sodium can be used.

[0450] Selectivity of modified cytidine deaminase for genomic DNA

[0451] The altered cytidine deaminase maintains some activity towards deaminating unmethylated cytidine, and thus it can be used to modulate the activity to maintain the desired selectivity for methylated cytosine. For example, high enzyme concentrations, long incubation times at high temperatures, or buffers that promote high enzyme activity may result in undesired levels of deamination of unmethylated cytidine. Typically, the enzyme concentration can be from 50 nM to 1000 nM, and the reaction can be carried out at an incubation temperature in the range of 20 °C to 55 °C for 5 minutes to 3 hours. Specifically, it can be helpful to adjust the enzyme concentration based on the purification method of the deaminase and how much activity the given enzyme preparation contains.

[0452] As an example, a library is prepared, deamination is carried out according to the general method described above, and sequencing is performed. Varying the concentration of APOBEC-Y130A-Y132H results in observable differences in the activity level and the off-target cytidine deamination level ( Figure 20 A). In addition, while it can be expected that commercial buffers for APOBEC activity would be well-suited for this application, testing of the APOBEC buffer included in the EM-seq TM kit (New England Biolabs) results in an undesired high level of activity towards the cytidine substrate, which causes an elevated conversion of unmethylated C nucleobases in λDNA ( Figure 20 B).

[0453] Method for determining suitable conditions for deamination

[0454] To determine whether the conditions are suitable for selective deamination on genomic DNA, a DNA mixture containing NA12878 (human) DNA, fully CpG-methylated pUC19, and fully unmethylated λ was prepared. After adapter ligation, using the general method described above, the samples were placed under various deamination conditions and prepared as sequencing libraries. The methylation levels of pUC19 and λ were used to select conditions with the desired 5mC activity and selectivity.

[0455] Example 13

[0456] Specific example of conditions for library preparation and deamination by modified cytidine deaminase

[0457] In all the methods described in Examples 13 to 20 below, the altered cytidine deaminase specifically refers to APOBEC3A-Y130A-Y132H. A variety of human genomic DNA samples were tested for different application evaluations.

[0458] Method A: Human genomic DNA was combined with fully unmethylated λ control DNA and enzymatically CpG-methylated pUC19 control DNA, and mechanically sheared to generate fragments of approximately 300 bp. Then, according to the standard Illumina library preparation procedure, the sheared DNA (10 ng to 100 ng) was subjected to end repair, A-tailing, and adapter ligation. The adapter-ligated DNA was denatured by incubation in 0.02 N sodium hydroxide at 50 °C for 10 minutes. Subsequently, the ssDNA samples were subjected to enzymatic deamination with cytidine deaminase (200 nM) at 37 °C for 25 minutes in 50 mM Bis-Tris (pH 6.5), 10 μg / mL RNase A. Then, the library was PCR amplified using unique dual-index primers and Q5U (New England Biolabs) with 9 cycles of PCR. The samples were sequenced on a NovaSeq6000 and analyzed using the DRAGEN methylation pipeline.

[0459] Method B: Combine human genomic DNA with completely unmethylated λ control DNA and enzymatically CpG-methylated pUC19 control DNA, and perform mechanical shearing to generate fragments of approximately 300 bp. Then, according to the standard Illumina library preparation procedure, subject this sheared DNA (10 ng to 100 ng) to end repair, A-tailing, and adapter ligation. Denature the adapter-ligated DNA by incubating at 50 °C in 0.02 N sodium hydroxide for 10 minutes. Subsequently, perform enzymatic deamination of the ssDNA sample for 15 minutes at 37 °C in 50 mM Bis-Tris (pH 6.5), 10 μg / mL RNase A together with the modified cytidine deaminase (200 nM). Then amplify the library by PCR using unique dual-index primers and Q5U (New England Biolabs) for 9 cycles. Sequence the sample on the NovaSeq 6000 and analyze it using the DRAGEN methylation pipeline.

[0460] Method C: Combine human genomic DNA with completely unmethylated λ control DNA and enzymatically CpG-methylated pUC19 control DNA, and perform mechanical shearing to generate fragments of approximately 300 bp. Then, according to the standard Illumina library preparation procedure, subject this sheared DNA (10 ng to 100 ng) to end repair, A-tailing, and adapter ligation. Denature the adapter-ligated DNA by incubating at 50 °C in 0.02 N sodium hydroxide for 10 minutes. Subsequently, perform enzymatic deamination of the ssDNA sample for 15 minutes at 37 °C in 50 mM Bis-Tris (pH 6.5), 10 μg / mL RNase A together with the modified cytidine deaminase (200 nM). Then amplify the library by PCR using unique dual-index primers and Q5 HiFi (New England Biolabs) for 9 cycles. Sequence the sample on the NovaSeq 6000 and analyze it using the DRAGEN methylation pipeline.

[0461] Method D: Combine human genomic DNA with completely unmethylated λ control DNA and enzymatically CpG-methylated pUC19 control DNA, and perform mechanical shearing to generate fragments of approximately 300 bp. Then, according to the standard Illumina library preparation protocol, subject this sheared DNA (10 ng to 100 ng) to end repair, A-tailing, and adapter ligation. Denature the adapter-ligated DNA by incubating it in 0.02 N sodium hydroxide at 50 °C for 10 minutes. Subsequently, perform enzymatic deamination of the ssDNA sample in 50 mM Bis-Tris (pH 6.5) with the modified cytidine deaminase (200 nM) at 37 °C for 25 minutes. Then, perform PCR amplification of the library using unique dual-index primers and Q5U (New England Biolabs) with 9 cycles of PCR. Sequence the sample on a NovaSeq 6000 and analyze it using the DRAGEN methylation pipeline.

[0462] Use of RNase A

[0463] The literature indicates that ribonuclease can increase the activity of cytidine deaminase by removing contaminating RNA (Bransteitter et al., Proceedings of the National Academy of Sciences of the United States of America 100, no. 7 (2003): 4102-7. https: / / doi.org / 10.1073 / pnas.0730835100). However, testing of ribonuclease A showed that, compared to this hypothesis, ribonuclease A decreased the activity of the modified cytidine deaminase, with a more significant effect on reducing off-target cytosine deamination, resulting in a more 5mC-selective result ( Figure 21 ).

[0464] Example 14

[0465] Deamination of 5hmC

[0466] To evaluate 5hmC deamination, oligonucleotides modified with C, 5mC, or 5hmC were designed in a defined context. Assemble these oligonucleotides together using ligation, and construct the oligonucleotides such that the resulting ligated fragments contain the handles required for subsequent amplification and sequencing ( Figure 22A ). Incorporate the assembled control oligonucleotides into the adapter-ligated DNA library and process according to Method A (Example 12). Analysis of the reported methylation from this control oligonucleotide showed that 5mC is the most preferred substrate for APOBEC Y130A-Y132H, yet still has significant activity towards 5hmC (Figure 22B )。

[0467] Example 15

[0468] Determination of methylation on CpG islands

[0469] Using NA12878 gDNA, libraries were prepared and deaminated according to Method A (Example 13). EM-Seq was also used TM transformation (New England Biolabs) to generate a comparative dataset. To evaluate methylation performance against the human genome, methylpy was used to calculate the regional methylation values of CpG islands across the entire human genome. The methylation values for each region were plotted against the methylation values for each region from the EM-Seq TM dataset ( Figure 23 ). Regional analysis showed a high correlation between the use of the Y130A-Y132H deaminase and the use of EM-Seq TM (a commercial methylation detection method). This indicates that the 5mC-selective deamination assay described herein can effectively detect and report methylation in CpG islands.

[0470] Example 16

[0471] Interpretation of methylation library variants

[0472] Libraries were prepared from human genome samples NA12878, NA24385, and NA24631 and deaminated according to Method A (Example 13). A comparative dataset without transformation was also generated by performing PCR directly after adapter ligation. After running variant calling analysis, comparison with the truth set for each genome using hap.py showed that libraries subjected to methylation transformation (altered cytidine deaminase) provided good SNV / indel calling performance, approaching that of the no-transformation control ( Figure 24 ). Although SNV / indel calling is discussed herein, other types of variant calling (including copy number variation (CNV), short tandem repeat (STR), and structural variant (SV)) are also feasible.

[0473] Example 17

[0474] Differentially methylated region (DMR) interpretation

[0475] Differential methylation analysis is commonly used to compare methylomes of different diseases, tissues, and cell types (Chen et al., Briefings in Functional Genomics 15, no. 6 (2016): 485-90. https: / / doi.org / 10.1093 / bfgp / elw018). Using genomic DNA isolated from HCC2218 tumor and normal cell types (CRL-2363D, CRL-2343D, ATCC) as well as HCC1187 tumor and normal cell types (CRL-2323D, CRL-2322D, ATCC), libraries were prepared and deaminated according to Method A (Example 13). EM-Seq TM conversion (New England Biolabs) and bisulfite conversion (EZ DNA Methylation-Gold Kit-Zymo Research) were also used to generate comparative datasets. For differential methylation analysis, the program HOME for identifying DMRs was used (Srivastava et al., BMC Bioinformatics 20, no. 1 (2019): 1-15, doi.org / 10.1186 / s12859-019-2845-y). For each methylation assay, methylation between tumor and normal samples was compared. As Figure 25 shown, this assay was able to detect relevant differentially methylated regions in the ZNF-154 gene (Almeida et al., BMC Cancer 19, no. 1 (2019): 1-12. https: / / doi.org / 10.1186 / s12885-019-5403-0). In addition, quantitative evaluation of the DMR interpretation performance using Method A, Method B, and Method C (Example 13) showed that the modified cytidine deaminase-based methods detected the expected DMRs with high precision and re-interpretation( Figure 26 ).

[0476] Example 18

[0477] Methylation at promoter

[0478] The methylation status of the promoter region can be correlated with both histone status and gene expression activity. To evaluate whether the modified cytidine deaminase assay can be used to probe the methylation status of the promoter, libraries were prepared from the human genomic sample NA12878 and deaminated according to Method A (Example 13). EM-Seq TMConversion was performed to generate a comparison library. Methylation levels at known histone marker sites (including H3K36me3 and H3K27ac) were quantified from the resulting data. H3K36me3 sites are expected to be inactive promoters with high histone and DNA methylation, while H3K27ac sites are expected to be active promoters with high histone acetylation and low DNA methylation. As expected, the modified cytidine deaminase assay was able to report these methylation trends( Figure 27 ).

[0479] Example 19

[0480] Detection of tumor samples in the background of normal samples

[0481] Libraries were prepared from human genomic samples with HCC2218 normal DNA or 10% spike-ins of HCC2218 tumor DNA added to the background of HCC2218 normal DNA using appropriate adapters. Methylation levels were evaluated for the 10% spike-in samples compared to methylation methods established for CpGs within regions also targeted by the PanSeer cancer panel (Chen et al. Nature Communications 11, no. 1 (2020): 1 - 10, https: / / doi.org / 10.1038 / s41467-020-17316-z). Libraries subjected to the modified cytidine deaminase reflected methylation changes between HCC2218 normal DNA and 10% spike-ins of HCC2218 tumor DNA added to HCC2218 normal DNA. Compared to samples without tumor DNA spike-ins, methods A, B, and C (Example 13) produced higher tumor signals in the 10% tumor spike-in samples( Figure 28 ). This reflects the ability of the modified cytidine deaminase to detect low levels of tumor DNA, and the modified cytidine deaminase can be applied to cell-free DNA (cfDNA) samples for early cancer detection (Jamshidi et al., Cancer Cell 40, no. 12 (2022): 1537 - 1549.e12.doi.org / 10.1016 / j.ccell.2022.10.022.) or minimal residual disease (MRD) testing (Jin et al., Proceedings of the National Academy of Sciences of the United States of America 118, no. 5 (2021): 1 - 8, doi.org / 10.1073 / pnas.2017421118.).

[0482] Example 20

[0483] Enrichment

[0484] Most commercially available methylation assays perform C>T conversion, resulting in extensive C>T conversion and low-complexity sequences. This makes enrichment very challenging because probes must be designed for all possible methylation states / strands (www.twistbioscience.com / sites / default / files / resources / 2020-02 / AppNote_Methylation-singles.pdf). This makes hybridization enrichment sequencing of methylated-converted libraries expensive and more difficult to optimize. An alternative is amplicon-based targeted sequencing; however, advantages of hybridization-based targeted sequencing include the ability to easily sequence both unenriched samples (whole-genome sequences (WGS)) and enriched samples, and to target sequence multiple contiguous regions of interest (Singh et al., Diagnostics (Basel, Switzerland) 12, no. 7 (June 24, 2022): 1539. doi.org / 10.3390 / diagnostics12071539.). Since the altered cytidine deaminase can perform 5mC>T conversion, enrichment of methylated-converted...

Claims

1. A protein complex, said protein complex comprising a cytidine deaminase and a helicase.

2. The protein complex according to claim 1, wherein the cytidine deaminase and the helicase are attached.

3. The protein complex according to claim 2, wherein the cytidine deaminase and the helicase comprise a fusion protein.

4. The protein complex according to claim 3, wherein the fusion protein comprises a linker.

5. The protein complex according to claim 2, wherein the cytidine deaminase and the helicase are biochemically conjugated.

6. The protein complex according to claim 1, wherein the cytidine deaminase is an altered cytidine deaminase, said altered cytidine deaminase comprising an amino acid substitution mutation at a position functionally equivalent to (Tyr / Phe)130 in the wild-type APOBEC3A protein.

7. The protein complex according to claim 6, wherein the (Tyr / Phe)130 is Tyr130, and the wild-type APOBEC3A protein is SEQ ID NO:

3.

8. The protein complex according to claim 1, wherein the cytidine deaminase is an altered cytidine deaminase, said altered cytidine deaminase comprising amino acid substitution mutations at positions functionally equivalent to (Tyr / Phe)130 and Tyr132 in the wild-type APOBEC3A protein.

9. The protein complex according to any one of claims 6 to 8, wherein the substitution mutation at the position functionally equivalent to Tyr130 comprises a mutation to Ala, Val or Trp, and the substitution mutation at the position functionally equivalent to Tyr132 comprises a mutation to Arg, His, Leu or Gln, or a combination thereof.

10. The protein complex according to any one of claims 6 to 9, wherein the altered cytidine deaminase converts 5-methylcytosine (5mC) to thymidine (T) by deamination at a rate greater than the rate of converting cytosine (C) to uracil (U) by deamination.

11. The protein complex according to claim 10, wherein the rate is at least 100 times greater.

12. The protein complex according to any one of claims 6 to 9, wherein the altered cytidine deaminase is a member of the AID subfamily, APOBEC1 subfamily, APOBEC2 subfamily, APOBEC3A subfamily, APOBEC3B subfamily, APOBEC3C subfamily, APOBEC3D subfamily, APOBEC3F subfamily, APOBEC3G subfamily, APOBEC3G subfamily, APOBEC3H subfamily or APOBEC4 subfamily.

13. The protein complex according to claim 12, wherein the altered cytidine deaminase comprises the ZDD motif H-[P / A / V]-E-X [23-28] -P-C-X [2-4] -C (SEQ ID NO: 12).

14. The protein complex according to claim 13, wherein the altered cytidine deaminase is a member of the APOBEC3A subfamily and comprises the ZDD motif HXEX2 4 SW(S / T)PCX [2 - 4] CX 6 FX 8 LX 5 R(L / I)YX [8-11] LX 2 LX [10] M (SEQ ID NO: 13).

15. The protein complex according to any one of claims 1 to 14, wherein the helicase unwinds double-stranded DNA (dsDNA).

16. The protein complex according to any one of claims 1 to 15, wherein the helicase is a member of helicase superfamily 1, superfamily 2, superfamily 3 or family 4.

17. The protein complex according to claim 16, wherein the helicase comprises a Walker A motif (SEQ ID NO: 68), a Walker B motif (SEQ ID NO: 69), or both.

18. The protein complex according to claim 17, wherein the helicase is a member of the RecQ family.

19. The protein complex according to claim 18, wherein the helicase comprises RecQ helicase from Escherichia coli (SEQ ID NO: 92).

20. The protein complex according to any one of claims 1 to 14, wherein the protein complex converts 5mC to T by deamination in dsDNA at a rate greater than that of a comparable protein complex without a helicase.

21. The protein complex according to any one of claims 2 to 14, wherein the attachment protein converts 5mC to T by deamination in dsDNA at a rate greater than that of a comparable protein complex without the fusion protein.

22. A composition comprising: The protein complex according to any one of claims 1 to 21; and A DNA sample comprising double-stranded DNA containing at least one 5mC.

23. The composition according to claim 22, wherein the sample comprises genomic DNA.

24. The composition according to claim 23, wherein the genomic DNA is from a single cell or a mixture of multiple cells.

25. A composition comprising: A protein complex comprising: A cytidine deaminase, and A helicase; and A sample comprising double-stranded DNA containing at least one 5mC.

26. The composition according to claim 25, wherein the cytidine deaminase and the helicase are attached.

27. The composition according to claim 26, wherein the cytidine deaminase and the helicase comprise a fusion protein.

28. The composition according to claim 27, wherein the fusion protein comprises a linker.

29. The composition according to claim 26, wherein the cytidine deaminase and the helicase are chemically conjugated.

30. The composition according to any one of claims 25 to 29, wherein the sample comprises a helicase recruitment motif.

31. The composition according to claim 30, wherein the helicase recruitment motif comprises a Y-shaped adaptor structure, a partial duplex, a single-stranded DNA (ssDNA) bubble, a fork, a flat duplex, a triple junction, or a combination thereof.

32. The composition according to any one of claims 25 to 29, wherein the sample comprises genomic DNA, cell-free DNA.

33. The composition according to claim 32, wherein the genomic DNA is from a single cell or a mixture of multiple cells.

34. The composition according to any one of claims 25 to 29, wherein the cytidine deaminase comprises an altered cytidine deaminase, and the altered cytidine deaminase comprises an amino acid substitution mutation at a position functionally equivalent to (Tyr / Phe)130 in the wild-type APOBEC3A protein.

35. The composition according to claim 34, wherein the (Tyr / Phe)130 is Tyr130, and the wild-type APOBEC3A protein is SEQ ID NO:

3.

36. The composition according to any one of claims 34 or 35, wherein the substitution mutation at a position functionally equivalent to Tyr130 comprises a mutation to Ala, Val or Trp.

37. The composition according to any one of claims 34 to 36, wherein the altered cytidine deaminase converts 5-methylcytosine (5mC) to thymidine (T) by deamination at a rate greater than the rate of converting cytosine (C) to uracil (U) by deamination.

38. The composition according to claim 37, wherein the rate is at least 100 times greater.

39. The composition according to any one of claims 34 to 38, wherein the altered cytidine deaminase is a member of the AID subfamily, APOBEC1 subfamily, APOBEC2 subfamily, APOBEC3A subfamily, APOBEC3B subfamily, APOBEC3C subfamily, APOBEC3D subfamily, APOBEC3F subfamily, APOBEC3G subfamily, APOBEC3G subfamily, APOBEC3H subfamily or APOBEC4 subfamily.

40. A composition according to any one of claims 34 to 39, wherein the altered cytidine deaminase comprises the ZDD motif H-[P / A / V]-E-X [23-28] -P-C-X [2-4] -C (SEQ ID NO: 12).

41. A composition according to any one of claims 34 to 40, wherein the altered cytidine deaminase is a member of the APOBEC3A subfamily and comprises the ZDD motif HXEX 24 SW(S / T)PCX [2-4] CX 6 FX 8 LX 5 R(L / I)YX [8-11] LX 2 LX [10] M (SEQ ID NO: 13).

42. The composition according to any one of claims 34 to 41, wherein the substitution mutation at a position functionally equivalent to Tyr130 comprises a mutation to alanine, glycine, phenylalanine, histidine, glutamine, methionine, asparagine, lysine, valine, aspartic acid, glutamic acid, serine, cysteine, proline, arginine or threonine.

43. The composition according to claim 42, wherein the substitution mutation at a position functionally equivalent to (Tyr / Phe)130 comprises a mutation to alanine.

44. The composition according to any one of claims 34 to 43, wherein the cytidine deaminase converts cytosine (C) to uracil (U) by deamination at a rate greater than the rate of converting 5-methylcytosine (5mC) to thymidine (T) by deamination.

45. The composition according to any one of claims 34 to 44, wherein the substitution mutation at a position functionally equivalent to Tyr130 comprises a mutation to leucine or tryptophan.

46. The composition according to any one of claims 25 to 45, wherein the helicase unwinds double-stranded DNA (dsDNA).

47. The composition according to any one of claims 25 to 46, wherein the helicase is a member of helicase superfamily 1, superfamily 2, superfamily 3, or family 4.

48. The composition according to claim 47, wherein the helicase comprises a Walker A motif (SEQ ID NO: 68), a Walker B motif (SEQ ID NO: 69), or both.

49. The composition according to claim 48, wherein the helicase is a member of the RecQ family.

50. The composition according to claim 49, wherein the helicase comprises the RecQ helicase from Escherichia coli (SEQ ID NO: 92).

51. The composition according to any one of claims 25 to 50, wherein the sample comprises a helicase recruitment motif.

52. The composition according to claim 51, wherein the helicase recruitment motif comprises a Y-shaped adaptor structure, a partial duplex, a single-stranded DNA (ssDNA) bubble, a fork, a flat duplex, a triple junction, or a combination thereof.

53. A method, the method comprising: providing a DNA sample suspected of containing dsDNA, the dsDNA comprising at least one 5-methylcytosine (5mC); contacting the sample with the protein complex according to any one of claims 1 to 7 under conditions suitable for converting 5-methylcytosine (5mC) to thymidine (T) to produce a converted dsDNA, wherein 5mC is converted to T; and processing the converted dsDNA to produce a sequencing library.

54. The method according to claim 53, the method further comprising: providing a surface comprising a plurality of amplification sites, wherein the amplification sites comprise at least two populations of attached single-stranded capture oligonucleotides having free 3′ termini, and contacting the surface comprising the amplification sites with the sequencing library under conditions suitable for producing a plurality of amplification sites, each of the plurality of amplification sites comprising a clonal population of amplicons from a single member of the sequencing library.

55. The method according to claim 54, wherein the attached single-stranded capture oligonucleotide comprises a helicase recruitment motif.

56. The method according to claim 55, wherein the helicase recruitment motif comprises a Y-shaped adaptor structure, a partial duplex, an ssDNA bubble, a fork, a flat duplex, a triple junction, or a combination thereof.

57. The method according to claim 53, wherein processing the converted double-stranded DNA to produce a sequencing library occurs after contacting the dsDNA with the protein complex.

58. The method according to claim 53, wherein processing the converted double-stranded DNA to produce a sequencing library occurs before contacting the dsDNA with the protein complex.

59. The method according to claim 53, wherein the processing comprises fragmentation or tagmentation of the double-stranded DNA and addition of a universal sequence to the double-stranded DNA fragments.

60. The method according to claim 59, wherein the universal sequence is part of an adaptor added to the double-stranded DNA fragments.

61. The method according to claim 53, wherein the sample is a biological sample.

62. The method according to claim 59, wherein the biological sample comprises cell-free DNA.

63. The method according to claim 59, wherein the biological sample comprises a fluid selected from blood or serum.

64. The method according to claim 53, wherein the sample comprises single cells or isolated cell nuclei.

65. The method according to claim 61, wherein the biological sample comprises tissue.

66. The method according to claim 65, wherein the tissue comprises tumor tissue.

67. A protein complex, the protein complex comprising: a pathway for converting 5-methylcytosine to thymine; and a pathway for directionally separating double-stranded nucleic acids.

Citation Information

Patent Citations

  • Ammonia refrigeration rapid defueling apparatus

    CN2630757Y

  • Method for detecting a target nucleic acid sequence

    EP0320308B1

  • Method of amplifying and detecting nucleic acid sequences

    EP0336731B1

  • Improved method of amplifying target nucleic acids applicable to both polymerase and ligase chain reactions

    EP0439182B1

  • Hyperactive AID / APOBEC and hmC dominant TET enzymes

    US10961525B2