New method
The method simplifies nucleic acid interaction identification by using biotin labeling and recombinase enzymes for one-step processing, overcoming the limitations of existing technologies in cell sample size and library complexity, achieving efficient and resolved interaction analysis.
Patent Information
- Application Number
- JP2022520576
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-10-04
- Filing Date
- 2020-10-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2040-10-05
AI Technical Summary
Current methods for identifying nucleic acid interactions, such as Hi-C and Capture Hi-C, require large numbers of cells and are inefficient for studying rare cell types or patient samples, and struggle with complex libraries that hinder resolution of specific interactions.
A method involving crosslinking, endonuclease fragmentation, biotin labeling, one-step fragmentation and oligonucleotide insertion using recombinase enzymes, followed by targeted enrichment and sequencing to identify nucleic acid segments interacting with target segments.
Enables efficient identification of nucleic acid interactions with reduced sample requirements, simplified processing, and enhanced resolution, capturing over 22,000 promoters and genomic loci in a single experiment with a significantly more quantitative readout.
Smart Images

Figure 0007702937000011 
Figure 0007702937000012 
Figure 0007702937000001
Abstract
Description
Technical Field
[0001] (Field of the Invention) The present invention relates to a method for identifying a nucleic acid segment that interacts with one or more target nucleic acid segments and a kit for carrying out this method. The present invention also relates to a method for identifying one or more interacting nucleic acid segments indicative of a particular disease.
Background Art
[0002] (Background of the Invention) Regulatory elements play a central role in the genetic control of organisms and have been shown to contribute to health and disease (e.g., cancer and autoimmune disorders). It has been demonstrated that such regulatory elements (e.g., enhancers) can be located (on a linear scale) a significant genomic distance away from their target genes. Approaches have been developed to capture these regulatory elements and their target genes and are widely applied to investigate the influence of regulatory landscape dynamics on gene expression and phenotype establishment and to investigate the role of genetic alterations in disease development. However, when studying with a small number of cells, determining which target genes these regulatory elements regulate is a major challenge.
[0003] One of the first methods developed to identify interactions between genomic loci was chromosome conformation capture (3C) technology (Dekker et al., Science (2002) 295:1306-1311). This method required the creation of a 3C library by crosslinking the nuclear composition so that spatially very close genomic loci became ligated, removing the DNA loops intervening between the crosslinks by digestion, and ligating and reversing the crosslinks of the interaction regions to create the 3C library. Subsequently, this library can be used to detect / identify the frequency of interactions between known sequences. However, this method requires that the interacting regions of interest be known in advance in order to detect them. Subsequently, further technological developments have been made to overcome the limitations of the 3C method.
[0004] Hi-C is a genome-wide method that does not require any prior knowledge about the interactome of interest. This method uses junction markers to isolate all ligated interaction sequences in a cell (see WO 2010 / 036323 and Lieberman-Aiden et al., 2009). This method provides information about all interactions occurring in the nuclear composition at a particular time point, but the resulting library is extremely complex, which hinders its analysis at the resolution required to identify significant interactions between specific elements, such as promoters and enhancers. To overcome this limitation, Capture Hi-C technology has been developed that involves a capture step to enrich the Hi-C library for chromosomal interactions at at least one end containing the region of interest (WO 2015 / 033134, Dryden et al., 2014 and Schoenfelder et al., 2015). WO 2015 / 033134 discloses methods and kits for identifying nucleic acid segments that interact with a target nucleic acid segment by using nucleic acid molecules for isolation. However, this method requires starting with a large number of cells (30 to 40 million cells), which is impossible when studying rare cell types, early stages of biogenesis, or cells derived from patient / biopsy samples.
[0005] Accordingly, there is a need to provide an improved method for identifying nucleic acid interactions that overcomes the limitations of currently available methodologies. SUMMARY OF THE INVENTION
[0006] (Summary of the Invention) According to a first aspect of the present invention, a method for identifying nucleic acid segments that interact with one or more target nucleic acid segments, comprising: (a) obtaining a nucleic acid composition comprising one or more target nucleic acid segments; (b) crosslinking the nucleic acid composition; (c) fragmenting the crosslinked nucleic acid composition using an endonuclease enzyme; (d) Filling the ends of the fragmented cross-linked nucleic acid segments with one or more nucleotides containing a biotin moiety covalently linked thereto; (e) Ligating the fragmented nucleic acid segments obtained in step (d) to produce a ligation fragment; (f) Using a recombinase enzyme to perform fragmentation and insertion of oligonucleotides on the ligation fragment in one step; (g) Enriching the fragment containing the biotin moiety of step (d); (h) Enriching the fragment containing the one or more target nucleic acid segments; (i) Sequencing the enriched fragment obtained in step (h) to identify the nucleic acid segment that interacts with the one or more target nucleic acid segments, thereby providing the said method.
[0007] According to a further aspect of the present invention, a method for identifying one or more interacting nucleic acid segments indicative of a specific disease state, comprising: (a) Performing the method defined herein on a nucleic acid composition obtained from an individual having a specific disease; (b) Quantifying the frequency of interaction between the nucleic acid segment and the one or more target nucleic acid segments; and (c) Comparing the frequency of interaction in the nucleic acid composition derived from the individual having the disease state with the frequency of interaction in a normal control nucleic acid composition derived from a healthy subject, wherein a difference in the frequency of interaction in the nucleic acid composition indicates a specific disease, thereby providing the said comparing.
[0008] According to yet a further aspect of the present invention, a kit for identifying a nucleic acid segment that interacts with one or more target nucleic acid segments, comprising a buffer and reagents capable of performing the method defined herein, thereby providing the said kit.
Brief Description of the Drawings
[0009] (Brief Description of the Drawings)
Figure 1
Figure 2
Mode for Carrying Out the Invention
[0010] (Detailed Description of the Invention) According to a first aspect of the present invention, a method for identifying a nucleic acid segment that interacts with one or more target nucleic acid segments, comprising: (a) obtaining a nucleic acid composition comprising one or more target nucleic acid segments; (b) crosslinking the nucleic acid composition; (c) fragmenting the crosslinked nucleic acid composition using an endonuclease enzyme; (d) filling the ends of the fragmented crosslinked nucleic acid segments with one or more nucleotides comprising a biotin moiety covalently linked; (e) ligating the fragmented nucleic acid segments obtained in step (d) to produce a ligation fragment; (f) using a transposase enzyme to perform fragmentation and insertion of oligonucleotides on the ligation fragment in one step; (g) enriching the fragments containing the biotin moiety of step (d); (h) enriching the fragments containing the one or more target nucleic acid segments; (i) sequencing the enriched fragments obtained in step (h) to identify the nucleic acid segments that interact with the one or more target nucleic acid segments, the method as described above is provided.
[0011] The method of the present invention provides a means for identifying interaction sequences and interaction nucleic acid segments by using targeted amplification or an isolation nucleic acid molecule for isolating one or more target nucleic acid segments. Such a method has the advantage of focusing on data of specific interactions in a very complex library. Furthermore, by using a method that includes the addition of an isolation nucleic acid molecule that binds to targeted amplification or one or more target nucleic acid segments, information can be organized into various subpopulations according to the type of reagent used for selection or the type of isolation nucleic acid molecule used (for example, a promoter for identifying promoter interactions). Subsequently, detailed information regarding chromosomal interactions within a specific target group of interest can be obtained.
[0012] The method of the present invention further provides for one-step fragmentation and insertion of oligonucleotides using recombinase enzymes. Such one-step fragmentation and oligonucleotide insertion has the advantage of providing a significant reduction in the total number of steps in the method and a reduction in the manipulation of nucleic acid compositions. For example, in certain embodiments of the method, one tube can be used from the acquisition of the nucleic acid composition (step (a) as defined herein) to the ligation of the fragmented nucleic acid segments (step (e) as defined herein), and also from the performance of one-step fragmentation and oligonucleotide insertion (step (f) as defined herein) to the enrichment of fragments containing biotin (step (g) as defined herein). In methods that do not include one-step fragmentation and oligonucleotide insertion, fragmentation by physical or enzymatic means (e.g., sonication or restriction enzyme digestion), end repair of library fragments, addition of dATP to the 3'-ends of library fragments, size selection, ligation of oligonucleotide sequences, and purification of fragments from unligated oligonucleotides need to be performed separately (such as the method disclosed in WO 2015 / 033134). Thus, it will be appreciated that the present invention provides a method for identifying nucleic acid segments that interact with one or more target nucleic acid segments, which includes simpler and fewer steps and may include a shortened time frame to completion. Therefore, the method of the present invention takes less time than conventional protocols and further reduces the overall cost of library preparation. It will also be appreciated that such advantages can reduce the loss of nucleic acid compositions. Such a reduction in the loss of nucleic acid compositions can reduce the amount of starting material, e.g., the number of cells from which the nucleic acid composition is obtained, or increase the nucleic acid composition available for subsequent analysis as a result.
[0013] Furthermore, while existing techniques, such as 4C, were able to capture genome-wide interactions of one or a few promoters, the method described herein can capture over 22,000 promoters and the genomic loci that interact with them in a single experiment. Additionally, the method yields a significantly more quantitative readout.
[0014] Genome-Wide Association Studies (GWAS) have identified thousands of single nucleotide polymorphisms (SNPs) associated with diseases. However, many of these SNPs are located very far from genes, making it extremely difficult to predict which genes these SNPs act on. Therefore, the present method provides a method for identifying interacting nucleic acid segments even when they are located far apart from each other in the genome.
[0015] References to "nucleic acid segments" as used herein are equivalent to references to "nucleic acid sequences" and refer to polymers of any nucleotide (i.e., for example, adenine (A), thymidine (T), cytosine (C), guanosine (G) and / or uracil (U)). This polymer may or may not result in a functional genomic fragment or gene. Combinations of nucleic acid sequences can ultimately include chromosomes. Nucleic acid sequences containing deoxyribonucleosides are called deoxyribonucleic acid (DNA). Nucleic acid sequences containing ribonucleosides are called ribonucleic acid (RNA). RNA is further characterized and can be divided into several types, such as protein-coding RNA, messenger RNA (mRNA), transfer RNA (tRNA), long non-coding RNA (lnRNA), long intergenic non-coding RNA (lincRNA), antisense RNA (asRNA), microRNA (miRNA), short interfering RNA (siRNA), small nuclear (snRNA), and small nucleolar RNA (snoRNA).
[0016] A "single nucleotide polymorphism" or "SNP" is a change in a single nucleotide (i.e., A, C, G, or T) in the genome that differs between members of a biological species or between paired chromosomes.
[0017] The term "target nucleic acid segment(s) of 1 or more" will be understood to refer to 1 or more sequences that are the subject of knowledge of the user. Isolating only ligated fragments that contain 1 or more target nucleic acid segments helps to focus on data that identify specific interactions with a particular gene or gene segment of interest. Alternatively, performing targeted amplification to enrich for fragments that contain 1 or more target nucleic acid segments helps to focus on data that identify specific interactions with a particular gene or gene segment of interest by increasing the ratio of fragments in a composition that contain 1 or more target nucleic acid segments.
[0018] References herein to the term "interaction" or "interact" refer to an association between two elements, such as a genomic interaction between a nucleic acid segment and a target nucleic acid segment in the present method. An interaction can result in one interacting element acting on the other, such as silencing or activating the element to which it binds. An interaction can occur between two nucleic acid segments that are either in close proximity to each other on a linear genomic sequence or are located far apart. Thus, in one embodiment, 1 or more nucleic acid segments that interact with 1 or more target nucleic acid segments are in very close proximity to the 1 or more target nucleic acid segments on a linear genomic sequence, such as being relatively close to each other on the same chromosome. In a further embodiment, 1 or more nucleic acid segments that interact with 1 or more target nucleic acid segments are located far from the 1 or more target nucleic acid segments on a linear genomic sequence, such as being present on different chromosomes or being even further apart on the same chromosome.
[0019] References herein to the term "nucleic acid composition" refer to any composition that contains nucleic acids and proteins. The nucleic acids in a nucleic acid composition are organized into chromosomes, where regulatory proteins (i.e., such as histones) can be in an associated state with the chromosomes. In one embodiment, the nucleic acid composition includes a nuclear composition. Such nuclear compositions can typically include nuclear genomic organization or chromatin.
[0020] As used herein, references to "cross-linking" or "cross-linked" refer to any stable chemical association between two compounds such that they can be further processed as one unit. Such stability can be based on covalent and / or non-covalent bonds (e.g., ionic). For example, nucleic acids and / or proteins can be cross-linked by chemicals (i.e., e.g., fixation reagents), heat, pressure, pH changes, or radiation, such that they maintain their spatial relationships during routine laboratory procedures (i.e., e.g., extraction, washing, centrifugation, etc.). Cross-linking as used herein is equivalent to the terms "fixing" or "fixed" applied to any method or process that immobilizes any and all cellular processes. Thus, cross-linked / fixed cells accurately maintain the spatial relationships between components in the nucleic acid composition at the time of fixation. Many chemicals, including but not limited to formaldehyde, formalin, or glutaraldehyde, can provide fixation.
[0021] As used herein, references to the term "fragment" refer to any nucleic acid sequence that is shorter than the sequence from which it is derived. Fragments can be of any size, ranging from several megabases and / or kilobases in nucleotide length to several nucleotides in length. Fragments are preferably longer than 5 nucleotide bases, e.g., 10, 15, 20, 25, 30, 40, 50, 100, 250, 500, 750, 1000, 2000, 5000, or 10000 nucleotide bases in length. Fragments can be even longer, e.g., 1, 5, 10, 20, 25, 50, 75, 100, 200, 300, 400, or 500 kilobases in nucleotide length. For example, methods such as restriction enzyme digestion, sonication, acid incubation, base incubation, microfluidization, etc. can all be used for fragmenting nucleic acid compositions.
[0022] In some embodiments, fragmentation (i.e., step (c)) is performed using an endonuclease enzyme. Examples of suitable endonuclease enzymes include, but are not limited to, sequence-specific endonucleases, such as restriction enzymes, and sequence-nonspecific endonucleases, such as MNase or DNase.
[0023] Accordingly, in one embodiment, the endonuclease enzyme is a sequence-specific endonuclease, such as a restriction enzyme. As used herein, the term "restriction enzyme" refers to any protein that cleaves nucleic acid at a specific base pair sequence. Cleavage can result in blunt or sticky ends, depending on the type of restriction enzyme selected. Examples of restriction enzymes include, but are not limited to, Eco RI, Eco RII, Bam HI, Hind III, Dpn II, Bgl II, Nco I, Taq I, Not I, Hinf I, Sau 3A, Pvu II, Sma I, Hae III, Hga I, Alu I, Eco RV, Kpn I, Pst I, Sac I, Sal I, Sca I, Spe I, Sph I, Stu I, Xba I. In a further embodiment, fragmentation (i.e., step (c)) is performed using a restriction enzyme. In one embodiment, the restriction enzyme is Hind III. In a further embodiment, the restriction enzyme is Dpn II.
[0024] In another embodiment, the endonuclease enzyme is a sequence-nonspecific endonuclease. As used herein, the term "sequence-nonspecific endonuclease" refers to any protein that cleaves nucleic acids and is not restricted to the sequence of the nucleic acid. For example, a sequence-nonspecific endonuclease can cleave nucleic acids in any region where no protein (e.g., nucleosome and / or transcription regulator) is bound. Examples of sequence-nonspecific endonucleases are known in the art and include, but are not limited to, DNase, RNase, and MNase. MNase is a nonspecific endo-exonuclease derived from the bacterium Staphylococcus aureus and binds to and cleaves the protein-unbound regions of DNA on chromatin - DNA bound to histone or other chromatin-binding proteins remains undigested. In yet a further embodiment, fragmentation (i.e., step (c)) is performed using a sequence-nonspecific endonuclease.
[0025] In another embodiment, fragmentation (i.e., step (c)) is performed using sonication.
[0026] References herein to "filling in the ends" of a term, fragment or nucleic acid segment refer to adding nucleotides to the 3' end of a crosslinked nucleic acid composition or segment after fragmentation. Such filling includes adding dATP, dCTP, dGTP, and / or dTTP nucleotides to the 3' end of the nucleic acid composition or segment. To enable enrichment of nucleic acid fragments or segments that are ligated and thus contain ligation junctions, one or more nucleotides used for filling as described herein may include a biotin moiety covalently linked. Thus, in one embodiment, filling in the ends of fragmented crosslinked nucleic acid segments includes adding a biotin moiety to the ends of the crosslinked nucleic acid fragments. In a further embodiment, filling in the ends of fragmented crosslinked nucleic acid segments includes "marking" the ends of the crosslinked nucleic acid fragments using a "junction marker". Such "marking" of the ends of the crosslinked nucleic acid fragments or addition of a biotin moiety to the ends enables subsequent selection or enrichment of nucleic acid fragments and / or segments ligated according to step (e) and the methods defined herein.
[0027] The junction marker enables purification of the fragments ligated prior to the enrichment step (h), and thus only the ligated sequences, rather than unligated (i.e., non-interacting) fragments, are reliably enriched.
[0028] In certain embodiments, the junction marker comprises a labeled nucleotide linker (i.e., a nucleotide that includes a biotin moiety covalently linked). In further embodiments, the junction marker comprises biotin. In one embodiment, the junction marker may comprise a modified nucleotide. In one embodiment, the junction marker may comprise an oligonucleotide linker sequence.
[0029] References herein to the terms "ligated" or "ligating" refer to the optional joining of two nucleic acid segments, typically including phosphodiester bonds. Ligation is typically facilitated by the presence of a cofactor reagent and an energy source (i.e., in the presence of, for example, adenosine triphosphate (ATP)) and a catalytic enzyme (i.e., a ligase such as, for example, T4 DNA ligase). In the methods described herein, to generate a single ligated fragment, fragments of two cross-linked nucleic acid segments are ligated and integrated.
[0030] In one embodiment, ligating fragmented nucleic acid segments to generate a ligated fragment (i.e., step (e)) utilizes intranuclear ligation. Thus, in certain embodiments, ligation of fragmented nucleic acid segments is performed by intranuclear ligation. Such intranuclear ligation has the advantage of requiring small amounts of reagents used, thereby reducing loss of the nucleic acid composition and thus allowing intranuclear ligation to also reduce the amount of starting material. For example, the number of cells from which the nucleic acid composition is obtained can be reduced, or the nucleic acid composition available for subsequent analysis resulting therefrom can be increased.
[0031] References herein to "one-step fragmentation and oligonucleotide insertion" refer to performing fragmentation of a ligated fragment and insertion of an oligonucleotide sequence in one step. Such methods utilize recombinase enzymes that bind to the oligonucleotide sequences and insert them into the fragment. This process is also known as "tagmentation". Thus, in one embodiment, one-step fragmentation and oligonucleotide insertion includes tagmentation.
[0032] Advantages of fragmentation of a project and ligation of oligonucleotides include, in addition to those described above, that there is no need to remove any binding pair elements (e.g., biotin) incorporated into the nucleic acid composition from the non-ligated fragments, no need to perform size selection of the ligated fragments, enzymatic fragmentation by recombinase eliminates the need for end repair and the addition of A tails since sonication is not performed. Further, the insertion of oligonucleotides that may contain barcode sequences and / or unique molecular identifiers and / or the insertion of adapter sequences is performed simultaneously with fragmentation. Such barcode sequences or unique molecular identifiers enable the identification of specific nucleic acid compositions in subsequent analysis and processing and the combination of multiple nucleic acid compositions in subsequent steps. On the other hand, the ability to identify and analyze individual nucleic acid compositions is retained. Thus, in one embodiment, the oligonucleotide sequence is an "adapter" sequence that enables or facilitates subsequent library preparation and sequencing of nucleic acid fragments containing the adapter. In a further embodiment, the adapter contains a barcode sequence and / or a unique molecular identifier.
[0033] In yet a further embodiment, fragmentation of a project and insertion of oligonucleotides includes inserting a barcode sequence into the ligated fragments. In one embodiment, the paired end adapter sequences contain a barcode sequence and / or a unique molecular identifier.
[0034] In yet further advantages of the method of the present invention, the one-step fragmentation and oligonucleotide ligation (e.g., tagmentation) presented herein results in a significantly enriched library of fragments containing one or more target nucleic acid segments as compared to previously published protocols. For example, enrichment values of at least 5-fold to 20-fold, or at least 5-fold to 80-fold can be generated as compared to libraries generated according to existing known or conventional Hi-C protocols. In one embodiment, a library enriched at least 5-fold, at least 10-fold, at least 15-fold, or at least 20-fold can be generated according to the method defined herein as compared to a library generated according to a conventional Hi-C protocol. In a further embodiment, a library enriched at least 10-fold, at least 11-fold, at least 12-fold, at least 13-fold, at least 14-fold, or at least 15-fold can be generated according to the method defined herein. In yet a further embodiment, a library enriched at least 50-fold, at least 55-fold, or at least 60-fold can be generated according to the method defined herein. It will be understood that any enrichment value of the library obtained when practicing the method defined herein as compared to a library generated according to a conventional Hi-C protocol can be determined depending on the nature of the endonuclease enzyme used for fragmentation of the crosslinked nucleic acid composition. For example, when using the restriction enzyme Hind III, an enrichment value of up to 20-fold can be obtained. Alternatively, when using the restriction enzyme Dpn II, an enrichment value of up to 80-fold can be obtained.
[0035] As used herein, the term "paired-end adapter" refers to any set of primer pairs that enables automated high-throughput sequencing reading from both ends. For example, such high-throughput sequencing devices compatible with these adapters include, but are not limited to, Solexa (Illumina), 454 System, and / or ABI SOLiD. For example, the method can include using universal primers together with polyA tails.
[0036] It will be understood that the recombinase enzymes suitable for use in the present method include any enzyme that can remove (or cleave) a sequence and insert the sequence into an oligonucleotide or nucleic acid fragment. Examples of such recombinase enzymes include retroviral integrases, and transposase enzymes such as MuA, Tn5, Tn7, and Tc1 / mariner type transposases. Thus, in one embodiment of the method, the recombinase enzyme is a retroviral integrase. In a further embodiment, the recombinase enzyme is a transposase enzyme, such as Tn5 transposase. To activate the recombinase, integrase, or transposase enzyme in the methods presented herein, the enzyme may be mutated to overcome the low level of activity of such enzymes in their native state. Thus, in yet a further embodiment, the recombinase enzyme is a mutant transposase, such as a hyperactive transposase. Such a hyperactive transposase can be a mutant Tn5 transposase. In one embodiment, the recombinase is Tn5 transposase, such as hyperactive Tn5 transposase.
[0037] Tn5 transposase is a member of the recombinase protein of the RNase superfamily, including retroviral integrase, and catalyzes the movement of a part of nucleic acid to another part of the genome or another genome (known as a transposon) by a so-called "cut and paste" mechanism. Recombinases, such as transposase enzymes and transposon elements, are found in certain bacteria and are involved in the acquisition of antibiotic resistance.
[0038] Transposase enzymes are generally inactive, and when mutations occur in the active site or other sites of the protein, hyperactive enzymes can be generated. Methods for producing Tn5 transposase enzymes are known in the art (Picelli et al., (2014) Genome Research 24:2033-2040). However, when purifying Tn5 transposase enzymes, these methods can be further adapted by using oligonucleotide sequences such as adapter sequences.
[0039] Oligonucleotide sequences such as adapter sequences used when purifying recombinase enzymes (e.g., Tn5 transposase) integrate with the enzyme and are subsequently inserted into nucleic acid fragments or segments by the recombinase. Such sequences can be diverse and can include additional elements that allow for further processing of the nucleic acid fragments or segments into which they are inserted. For example, an oligonucleotide integrated with a purified recombinase enzyme can include an adapter sequence and / or barcode sequence for sequencing. However, it will be understood that all such oligonucleotides include a transposon sequence or element that allows for integration with the enzyme. Examples of transposon sequences or elements include Tn5 transposase-compatible mosaic end (ME) sequences as well as sequences that are sterically compatible with the binding pockets of recombinase and / or transposase enzymes.
[0040] Accordingly, in one embodiment, the recombinase enzyme of the method includes a mosaic end double-stranded (MEDS) oligonucleotide, which includes half of a pair of terminal adapter sequences. In a further embodiment, the recombinase enzyme includes a pair of terminal adapter sequences for sequencing. In yet a further embodiment, the transposase enzyme may include an oligonucleotide including a pair of terminal adapter sequences for sequencing that further includes a barcode sequence. In a further embodiment, the oligonucleotide sequence is selected from SEQ ID NO:1, SEQ ID NO:2, and / or SEQ ID NO:3 as defined herein. In an alternative embodiment, the oligonucleotide sequence includes any sequence that enables subsequent library preparation and sequencing. It will be understood that such sequences can be ligated to the nucleic acid segment for amplification and isolation of the nucleic acid segment and for sequence analysis by high-throughput or next-generation sequencing. Examples of next-generation sequencing platforms include: Roche 454 (i.e., Roche 454 GS FLX), the SOLiD system from Applied Biosystems (i.e., SOLiDv4), the GAIIx from Illumina, the HiSeq 2000, and the MiSeq sequencing instrument, the Ion Torrent semiconductor-based sequencing instrument from Life Technologies, the PacBio RS from Pacific Biosciences, and the MinION from Oxford Nanopore.
[0041] As used herein, references to "enriching" or "enrichment" refer to the isolation of any nucleic acid segment, or an increase in the ratio of a nucleic acid segment of interest or a target nucleic acid segment to other nucleic acid segments in a nucleic acid composition. It will be understood that such references include terms such as "isolating", "isolation", "separating", "removing", and "purifying". For example, the concentration or isolation of a nucleic acid segment of interest or a target nucleic acid segment may include positive methods such as "pulling out" the nucleic acid segment of interest or the target nucleic acid segment, or negative methods such as the exclusion of nucleic acid segments that are not of interest or do not contain the target nucleic acid segment. Alternatively, enrichment or isolation may include selective or targeted amplification of the nucleic acid segment of interest or the target nucleic acid segment. Such selective or targeted amplification of the nucleic acid segment of interest or the target nucleic acid segment increases the ratio of such segments in the nucleic acid composition (i.e., the segment is enriched).
[0042] In one embodiment, the enrichment step (h) includes performing targeted amplification to enrich a fragment containing the one or more target nucleic acid segments.
[0043] In another embodiment, the concentration step (h) is for enriching a fragment containing the one or more target nucleic acid segments by: (i) adding an isolation nucleic acid molecule that binds to the one or more target nucleic acid segments, wherein the isolation nucleic acid molecule is labeled with the first half of a binding pair; and (ii) isolating the fragment containing the one or more target nucleic acid segments that is bound to the isolation nucleic acid molecule by using the second half of the binding pair. In certain embodiments, steps (i) and (ii) above can be performed sequentially, i.e., step (i) can be performed followed by step (ii). In further embodiments, steps (i) and (ii) above can be performed simultaneously.
[0044] Accordingly, the enrichment step (h) of the present method involves the enrichment of a nucleic acid fragment or segment or a target nucleic acid segment that contains a specific target segment or sequence.
[0045] References herein to "targeted amplification" refer to amplification using a method that preferentially amplifies a specific nucleic acid segment or target nucleic acid segment of interest. Such targeted amplification can utilize specific primer sequences that are complementary to sequences present in the target nucleic acid segment or target nucleic acid segment (e.g., promoter or silencer sequences). Accordingly, in one embodiment, the primer sequence is complementary to a promoter sequence. In another embodiment, the primer sequence is complementary to a sequence containing an SNP. Primer sequences utilized in the methods presented herein can include additional elements involved in the processing or analysis of the subsequent amplified nucleic acid segment. For example, the primer sequence can include an adapter sequence for sequencing as described herein or a unique molecular identifier useful for the identification of a nucleic acid segment or group of segments (e.g., those derived from a particular sample). Alternatively or additionally, targeted amplification can utilize specific conditions that are advantageous for the amplification of the target nucleic acid segment or a fragment containing the target nucleic acid segment. It will be understood that the amplification can be carried out by any method known in the art, such as polymerase chain reaction (PCR). It will be further understood that the targeted amplification described herein can include the amplification of nucleic acid segments in solution or on a support portion used for enrichment, such as the surface of beads. Also, the extension of the primer sequence can be carried out prior to amplification such that amplification on the support portion can further include a step of extending the primer sequence prior to the amplification of the extended sequence and nucleic acid segment.
[0046] References herein to "nucleic acid molecules for isolation" refer to molecules formed from nucleic acids configured to bind to one or more target nucleic acid segments. For example, a nucleic acid molecule for isolation may contain a sequence complementary to one or more target nucleic acid segments, which sequence then forms an interaction (i.e., forms base pairs (bp)) with the nucleotide bases of the one or more target nucleic acid segments. It will be understood that nucleic acid molecules for isolation, such as biotinylated RNA, do not need to contain the entire complementary sequence of one or more target nucleic acid segments to form complementary interactions and isolate the one or more target nucleic acid segments from a nucleic acid composition. Nucleic acid molecules for isolation can be at least 10 nucleotide bases in length, such as at least 20, 30, 40, 50, 60, 70, 80, 90, 100, 130, 150, 170, 200, 300, 400, 500, 750, 1000, 2000, 3000, 4000, or 5000 nucleotide bases in length.
[0047] In one embodiment, the addition of the nucleic acid molecule for isolation that binds to one or more target nucleic acid segments is carried out at 65°C to 72°C. In certain embodiments, the addition of the isolated nucleic acid molecule is carried out at 65°C. Thus, in a further embodiment, step (i) of the enrichment step (h) is carried out at 65°C to 72°C, such as 65°C. In another embodiment, the isolation of the fragment containing one or more target nucleic acid segments bound to the nucleic acid molecule for isolation using the second half of the binding pair is carried out at 68°C to 72°C. In certain embodiments, the isolation of the fragment using the second half of the binding pair is carried out at 68°C. Thus, in yet a further embodiment, step (ii) of the enrichment step (h) is carried out at 68°C to 72°C, such as 68°C.
[0048] In one embodiment, the isolation nucleic acid molecule is added in the presence of a blocking or blocker sequence. Such a blocker sequence prevents the ligation fragment containing the adapter sequence from binding to other ligation fragments containing the adapter sequence through any sequence complementarity of the adapter sequence. Thus, such a blocker sequence prevents a fragment that does not contain one or more target nucleic acid segments from binding to a fragment that contains one or more target nucleic acid segments. In certain embodiments, the blocker sequence is added to the ligation fragment before adding the isolation nucleic acid molecule. In another embodiment, the blocker sequence is added to the ligation fragment simultaneously with or combined with the isolation nucleic acid molecule. Thus, in one embodiment, the blocker sequence comprises an adapter sequence ligated to the fragment, for example, any sequence that is compatible with a sequence complementary to a particular adapter sequence. In further embodiments, the blocker sequence comprises any sequence that is compatible with, for example, a sequence complementary to a MEDS oligonucleotide that comprises half of a pair of end adapter sequences. In some embodiments, the blocker sequence is selected from SEQ ID NO: 10, SEQ ID NO: 11, SEQ ID NO: 12, SEQ ID NO: 13, SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, and / or SEQ ID NO: 17 as defined herein.
[0049] Furthermore, the enrichment step (h) of the present method can be carried out according to methods known in the art and using known reagents. For example, when the enrichment step (h) involves an isolation nucleic acid molecule that binds to one or more target nucleic acid segments described herein, the present method or the steps of the present method can be carried out, for example, in the presence of a buffer containing a high concentration divalent cation salt of 100 mM to 600 mM. The salt can be present in a molar ratio of 2.5:1 to 60:1. Also, a volume exclusion / thickening agent can also be present, for example, at a concentration of 0.002% to 0.1%. Alternatively, the method or the steps of the method can include incubating the nucleic acid composition in the presence of a buffer. The incubation can optionally be carried out over a period of 8 hours or less at two different temperatures, where the two different temperatures are cycled 2 to 100 times. Examples of such buffers and methods are described in US 9,587,268. Thus, when the enrichment step (h) is carried out according to the specific embodiments disclosed herein, it will be understood that the enrichment containing the isolated nucleic acid molecule can be carried out more rapidly than when using conventional reagents and reaction methods.
[0050] References herein to "binding pairs" refer to at least two moieties (i.e., a first half and a second half) that specifically recognize each other to form a bond. Suitable binding pairs include, for example, biotin and avidin, or derivatives of biotin and avidin, such as streptavidin and neutravidin.
[0051] References herein to "labeling" or "labeled" refer to the process of identifying a target by binding a marker, where the marker includes a specific moiety (i.e., an affinity tag) having an inherent affinity for a ligand. For example, the label can serve to selectively purify an isolation nucleic acid sequence (i.e., by, for example, affinity chromatography). Such labels can include, but are not limited to, biotin labels, histidine labels (e.g., 6His), or FLAG labels.
[0052] In one embodiment, the nucleic acid molecule for isolation contains biotin. In a further embodiment, the nucleic acid molecule for isolation is labeled with biotin. In still a further embodiment, the nucleic acid molecule for isolation is labeled with a histidine label or a FLAG label. Thus, according to a particular embodiment, the binding pair may include a label (e.g., a histidine or FLAG label) and an antibody.
[0053] In one embodiment, one or more target nucleic acid segments are selected from a promoter, an enhancer, a silencer, or an insulator. In a further embodiment, one or more target nucleic acid segments are a promoter. In yet another embodiment, one or more target nucleic acid segments are an insulator.
[0054] References to the terms "promoter" and "multiple promoters" herein refer to nucleic acid sequences that facilitate the initiation of transcription of an operably linked coding region. A promoter is sometimes referred to as a "transcription initiation region". Regulatory elements often interact with a promoter in order to activate or inhibit transcription.
[0055] The inventors of the present invention used the method of the present invention to identify thousands of promoter interactions, and 10 to 20 interactions per promoter were found. By the method described herein, some cell-specific interactions or interactions associated with various disease states were identified. A wide range of isolation distances between interacting nucleic acid segments were also identified. Most interactions are within 100 kilobases, but some interactions extend up to 2 megabases or more. Interestingly, use of the method also shows that both active and inactive genes form interactions.
[0056] The nucleic acid segments that have been identified as interacting with a promoter are candidates for regulatory elements required for proper gene regulation. Their disruption can change transcriptional output and contribute to disease, and thus the association of these elements with their target genes can provide potential new drug targets for novel therapies.
[0057] Identifying which regulatory elements interact with a promoter is extremely important for understanding genetic interactions. Also, since this method provides a snapshot of the interactions in a nucleic acid composition at a particular point in time, it is envisioned that this method can be carried out over a series of time points, or developmental states, or experimental conditions to construct a picture of the changes in interactions in the nucleic acid composition of a cell.
[0058] In one embodiment, it will be understood that the target nucleic acid segment interacts with a nucleic acid segment containing a regulatory element. In a further embodiment, the regulatory element includes an enhancer, a silencer, or an insulator.
[0059] As used herein, the term "regulatory gene" refers to any nucleic acid sequence encoding a protein, where the protein binds to the same or a different nucleic acid sequence to regulate the rate of transcription or otherwise affect the expression level of the same or a different nucleic acid sequence. As used herein, the term "regulatory element" refers to any nucleic acid sequence that affects the activity state of another genomic element. For example, various regulatory elements can include, but are not limited to, enhancers, activators, repressors, insulators, promoters, or silencers.
[0060] In one embodiment, the target nucleic acid molecule is a genomic site identified via chromatin immunoprecipitation (ChIP) sequencing. In a ChIP sequencing experiment, protein-DNA interactions are analyzed by cross-linking protein-DNA complexes in a nucleic acid composition. After the protein-DNA complexes are subsequently isolated (by immunoprecipitation), the genomic regions to which the protein is bound are sequenced.
[0061] In some embodiments, it is envisioned that the nucleic acid segment is located on the same chromosome as the target nucleic acid segment. Alternatively, the nucleic acid segment is located on a different chromosome from the target nucleic acid segment.
[0062] Using this method, long-range interactions, short-range interactions, or very short-range interactions can be identified. As used herein, the term "long-range interaction" refers to the detection of nucleic acid segments that interact at distant locations in a linear genomic sequence. This type of interaction can identify, for example, two genomic regions located on different arms of the same chromosome or on different chromosomes.
[0063] As used herein, the term "short-range interaction" refers to the detection of nucleic acid segments that interact at relatively proximal positions in the genome. As used herein, the term "very short-range interaction" refers to the detection of nucleic acid segments that are very proximal to each other in a linear genome, for example, segments that interact within a part of the same gene.
[0064] The inventors have shown that SNPs were located in nucleic acid segments that interacted at a higher frequency than probabilistically expected. Thus, by using the method of the present invention, SNPs that interact with a specific gene and are likely to regulate that gene can be identified.
[0065] Thus, it will be understood from the disclosure presented herein that this method can be used to identify any nucleic acid interaction, particularly DNA-DNA interactions in a nucleic acid composition.
[0066] In one embodiment, the nucleic acid molecule for isolation is obtained from a bacterial artificial chromosome (BAC), fosmid, or cosmid. In a further embodiment, the nucleic acid molecule for isolation is obtained from a bacterial artificial chromosome (BAC).
[0067] In one embodiment, the nucleic acid molecule for isolation is DNA, cDNA, or RNA. In a further embodiment, the nucleic acid molecule for isolation is RNA.
[0068] The nucleic acid molecule for isolation can be utilized by suitable methods such as solution hybridization selection (see WO 2009 / 099602). In this method, a set of "bait" sequences is generated and a hybridization mixture is formed. This can be used to isolate a subgroup of target nucleic acids from a sample (i.e., the "pond").
[0069] In one embodiment, the first half of the binding pair contains biotin and the second half of the binding pair contains streptavidin.
[0070] In one embodiment, the method further includes reversing the crosslinking prior to step (f). There are several methods known in the art for reversing crosslinking, which will be understood to vary depending on the method by which the crosslinking was initially formed. For example, the crosslinked nucleic acid composition can be subjected to high heat, such as 50°C, 55°C, 60°C, 65°C, 70°C, 75°C, 80°C, above 85°C or higher, to reverse the crosslinking. Further, the crosslinked nucleic acid composition needs to be subjected to high heat for more than 1 hour, such as at least 5 hours, 6 hours, 7 hours, 8 hours, 9 hours, 10 hours, or 12 hours or more. In one embodiment, reversing the crosslinking prior to step (f) includes incubating the crosslinked nucleic acid composition at 65°C for at least 8 hours (i.e., overnight) in the presence of proteinase K.
[0071] In one embodiment, the method further includes purifying the nucleic acid composition to remove any fragments that do not contain the junction marker prior to step (f).
[0072] References herein to "purification" refer to nucleic acid compositions that have been subjected to a process (i.e., fractionation, for example) to remove various other components, and the compositions retain substantially their expressed biological activity. When the term "substantially purified" is used, this designation refers to a composition in which the nucleic acid forms a major component of the composition, for example, about 50%, about 60%, about 70%, about 80%, about 90%, or about 95% or more of the composition (i.e., for example, weight / weight (w / w), volume / volume (v / v), and / or weight / volume (w / v)).
[0073] In one embodiment, the method further comprises amplifying the ligated fragment of the target isolated prior to step (i). In a further embodiment, the amplification is carried out by polymerase chain reaction (PCR).
[0074] In one embodiment, the nucleic acid composition is derived from a mammalian cell nucleus. In a further embodiment, the mammalian cell nucleus can be a human cell nucleus. Many human cells available in the art can be used in the methods described herein, such as GM12878 (human lymphoblastoid cell line) or CD34+ (human ex vivo hematopoietic progenitor cells).
[0075] It will be understood that the methods described herein are found to be useful not only in humans but also in a variety of organisms. Also, for example, the methods can be used to identify genomic interactions in plants and animals.
[0076] Accordingly, in another embodiment, the nucleic acid composition is derived from a non-human cell nucleus. In one embodiment, the non-human cell is selected from the group including, but not limited to, plants, yeast, mouse, bovine, porcine, equine, canine, feline, caprine, or ovine. In one embodiment, the non-human cell nucleus is a mouse cell nucleus or a plant cell nucleus.
[0077] From the advantages of the invention referred to herein, it will be appreciated that the method provides a reduction in the loss of nucleic acid compositions during the processes referred to herein. Such a reduction in the loss of nucleic acid compositions makes it possible to reduce the amount of starting material, for example the number of cells from which the nucleic acid composition is obtained. Thus, in one embodiment, the nucleic acid composition can be derived from fewer cells than existing promoter capture or conformation capture techniques. In a further embodiment, the nucleic acid composition is derived from 1 million or fewer cells, 500,000 or fewer cells, 200,000 or fewer cells, 50,000 or fewer cells, or 10,000 or fewer cells. In yet a further embodiment, the nucleic acid composition is derived from 1 million cells, 500,000 cells, 200,000 cells, 50,000 cells, or 10,000 cells. In one embodiment, the nucleic acid composition is derived from 1 million cells, 50,000 cells, or 10,000 cells.
[0078] In one embodiment, the method as defined herein: (i) crosslinking a nucleic acid composition comprising one or more target nucleic acid segments; (ii) fragmenting the crosslinked nucleic acid composition; (iii) marking the ends of the fragments with biotin; (iv) ligating the fragmented nucleic acid segments to produce a ligation fragment; (v) reversing the crosslinking; (vi) using a transposase enzyme to perform fragmentation and adapter insertion in one step on the ligation fragment; (vii) pulling down the ligation fragment using streptavidin; (viii) performing targeted amplification of the fragments comprising one or more target nucleic acid segments; and (ix) sequencing to identify nucleic acid segments that interact with one or more target nucleic acid segments.
[0079] In another embodiment, the method as defined herein: (i) Crosslinking a nucleic acid composition comprising a target nucleic acid segment of 1 or more; (ii) Fragmenting the crosslinked nucleic acid composition; (iii) Marking the ends of the fragments with biotin; (iv) Ligating the fragmented nucleic acid segments to produce a ligation fragment; (v) Reversing the crosslinking; (vi) Using a transposase enzyme to perform fragmentation and adapter insertion on the ligation fragment in one step; (vii) Performing pull - down of the fragments using streptavidin and amplification using PCR; (viii) Performing promoter capture by adding an isolation nucleic acid molecule that binds to 1 or more target nucleic acid segments, labeling the isolation nucleic acid molecule with the first half of a binding pair, said step; and (ix) Isolating a ligation fragment comprising 1 or more target nucleic acid segments bound to the isolation nucleic acid molecule by using the second half of the binding pair; (x) Performing amplification using PCR; and (xi) Sequencing to identify nucleic acid segments that interact with 1 or more target nucleic acid segments.
[0080] According to a further aspect of the present invention, a method for identifying 1 or more interacting nucleic acid segments indicative of a particular disease state, comprising: a) Performing the method as defined herein on a nucleic acid composition obtained from an individual having a particular disease state; b) Quantifying the frequency of interaction between nucleic acid segments and 1 or more target nucleic acid segments; c) Comparing the frequency of interaction in a nucleic acid composition from an individual having the disease state with the frequency of interaction in a normal control nucleic acid composition from a healthy subject, wherein the difference in the frequency of interaction in the nucleic acid composition indicates a particular disease, said comparing.
[0081] References to "frequency of interaction" or "interaction frequency" as used herein refer to the number of times a specific interaction is found in a nucleic acid composition (i.e., a sample). In some cases, a decrease in the interaction frequency in a nucleic acid composition as compared to a normal control nucleic acid composition from a healthy subject indicates a specific disease state (i.e., because the nucleic acid segments interact less frequently). Alternatively, an increase in the interaction frequency in a nucleic acid composition as compared to a normal control nucleic acid composition from a healthy subject indicates a specific disease state (i.e., because the nucleic acid segments interact more frequently). In some cases, the difference is at least a 0.5-fold difference, such as a 1-fold, 1.5-fold, 2-fold, 2.5-fold, 3-fold, 4-fold, 5-fold, 7-fold, or 10-fold difference.
[0082] In one aspect of the invention, the frequency of interaction can be used to determine the spatial proximity of two different nucleic acid segments. As the interaction frequency increases, the probability that two genomic regions are physically close to each other within the 3D nuclear space increases. Conversely, as the interaction frequency decreases, the probability that two genomic regions are physically close to each other within the 3D nuclear space decreases.
[0083] Quantification can be performed by calculating the frequency of interaction of a nucleic acid composition from a patient or by any method suitable for the purification or extraction of a nucleic acid composition sample or a dilution thereof. Also, for example, based on the results of high-throughput sequencing, it becomes possible to examine the frequency of a specific interaction. In the method of the present invention, quantification can be performed by measuring the concentration of a target nucleic acid segment or a ligation product in one or more samples. The nucleic acid composition can be obtained from cells in a biological sample that can include cerebrospinal fluid (CSF), whole blood, serum, plasma, or an extract or purification thereof, or a dilution thereof. In one embodiment, the biological sample can be cerebrospinal fluid (CSF), whole blood, serum, or plasma. Also, biological samples include tissue homogenates, tissue sections, and biopsy samples taken from a living subject or postmortem. The sample can be prepared, for example, diluted, concentrated, and stored in a normal manner if appropriate.
[0084] In one embodiment, the disease state is selected from cancer, autoimmune disease, developmental disorder, genetic disorder, diabetes, cardiovascular disease, kidney disease, lung disease, liver disease, neurological disease, viral infection or bacterial infection. In a further embodiment, the disease state is cancer or an autoimmune disease. In still a further embodiment, the disease state is cancer, such as breast cancer, intestinal cancer, bladder cancer, bone cancer, brain cancer, cervical cancer, colon cancer, endometrial cancer, esophageal cancer, kidney cancer, liver cancer, lung cancer, ovarian cancer, pancreatic cancer, prostate cancer, skin cancer, gastric cancer, testicular cancer, thyroid cancer, or uterine cancer, leukemia, lymphoma, myeloma or melanoma.
[0085] As used herein, the reference to "autoimmune disease" includes diseases resulting from an immune reaction that targets the human body itself, such as acute disseminated encephalomyelitis (ADEM), ankylosing spondylitis, Behcet's disease, celiac disease, Crohn's disease, type 1 diabetes, Graves' disease, Guillain-Barré syndrome (GBS), psoriasis, rheumatoid arthritis, rheumatic fever, Sjögren's syndrome, ulcerative colitis, and vasculitis.
[0086] As used herein, the reference to "developmental disorder" includes diseases typically rooted in childhood, such as learning disability, communication disorder, autism, attention deficit hyperactivity disorder (ADHD), and developmental coordination disorder.
[0087] As used herein, the reference to "genetic disorder" includes diseases caused by one or more abnormalities in the genome, such as Angelman syndrome, Canavan disease, Charcot-Marie-Tooth disease, color vision disorder, cat cry syndrome, cystic fibrosis, Down syndrome, Duchenne muscular dystrophy, hemochromatosis, hemophilia, Klinefelter syndrome, neurofibromatosis, phenylketonuria, polycystic kidney, Prader-Willi syndrome, sickle cell disease, Tay-Sachs disease, and Turner syndrome.
[0088] According to a further aspect of the invention, there is provided a kit for identifying a nucleic acid segment that interacts with one or more target nucleic acid segments, the kit comprising a buffer and reagents capable of performing the methods defined herein.
[0089] The kit may include one or more articles and / or reagents for carrying out the method. For example, oligonucleotide probes, pairs of amplification primers, and / or recombinase enzyme-related oligonucleotides for use in the methods described herein may be provided in isolated form and may be in a suitable container, such as a vial, that protects the contents from the external environment and is part of the kit. The kit may include instructions for use according to the protocol of the methods described herein. Kits intended for use of nucleic acids in PCR may include one or more other reagents required for the reaction, such as polymerase, nucleotides, buffer, and the like.
[0090] In one embodiment, the kit includes a recombinase enzyme. In a further embodiment, the recombinase enzyme included in the kit as defined herein is a transposase enzyme, such as a hyperactive mutant transposase enzyme, such as a hyperactive mutant Tn5 transposase.
[0091] In accordance with yet a further aspect of the invention, there is provided a recombinase enzyme as defined herein that enables one-step fragmentation and adapter insertion. Accordingly, there is also provided herein a recombinase enzyme capable of tagmentation.
[0092] In one embodiment, the recombinase enzyme provided herein is a hyperactive mutant transposase enzyme. In a further embodiment, the transposase enzyme is a hyperactive mutant Tn5 transposase. In yet a further embodiment, the transposase enzyme includes paired end adapter sequences.
[0093] Examples of types of buffers and reagents other than the buffers and reagents already described that may be included in the kit can be read from the examples described herein.
[0094] The following studies and protocols illustrate embodiments of the methods described herein.
Example
[0095] (Example) (Abbreviations)
Table 1
[0096] (Fixation of cells) 1. For each experiment, at least 50,000 cells must be fixed with 2% formaldehyde at room temperature for 10 minutes. · Quench with glycine at a final concentration of 0.125 M. · Centrifuge at 4°C, 1500 rpm (400×g) for 5 minutes. · Discard the supernatant and carefully resuspend the pellet in 100 μl of cold 1×PBS. Centrifuge at 4°C, 1500 rpm (400×g) for 5 minutes. · Discard the supernatant and either flash freeze in liquid nitrogen or proceed directly to the next step.
[0097] (Permeabilization of cells and restriction digestion) 2. Resuspend the fixed cell pellet obtained in step 1 in 100 μL of ice-cold lysis buffer. Incubate the tube on ice for 30 minutes. 3. Centrifuge the tube at 4°C at approximately 600 g for 5 minutes. 4. Remove the supernatant and leave approximately 20 μl of the solution containing the nuclear pellet. 5. Wash the pellet twice with 100 μL of 1.2×NEBuffer3 (when using Dpn II) or NEBuffer2 (when using Hind III). 6. Remove the supernatant leaving approximately 20 μL. Add 334 μl of 1.2×NEBuffer3 (or NEBuffer2 if working with Hind III). 7. Add 12 μl of 10% SDS (final concentration 0.3%, w / v); agitate on a thermomixer at 37°C, 950 rpm for 1 hour. Add 8.80 μl of 10% Triton (final concentration 1.8%, v / v). Incubate with shaking at 37 °C and 950 rpm for 1 hour on a thermomixer. Add 9.30 μl of Dpn II (50 U / μl) (or 15 μl of Hind III - 100 U / μl and 15 μl of H2O), and incubate with shaking at 37 °C and 950 rpm for 12 - 16 hours on a thermomixer.
[0098] (Biotin labeling and Hi-C ligation) Spin the digestion mix briefly. Add 4.5 μl of dCTP, dTTP, and dGTP (10 mM mix), 37.5 μl of biotin - dATP, and 10 μl of Klenow (5 U / μl). Incubate at 37 °C for 45 minutes with shaking at 700 rpm for 10 seconds every 30 seconds. Centrifuge at 4 °C and 600 g for 6 minutes. Leave approximately 50 μl including the pellet and remove the supernatant. Add 835 μl of H2O / 100 μl of T4 DNA ligase buffer / 5 μl of BSA (20 mg / ml) / 10 μl of T4 DNA ligase (Invitrogen). Incubate at 16 °C for at least 4 hours (or up to 12 hours).
[0099] (Purification of Hi-C DNA) Centrifuge the tube at 4 °C and 600 g for 6 minutes. Leave 200 μl in the tube and remove 800 μl of the supernatant. Add 15 μl of proteinase K (10 mg / ml). Incubate at 65 °C for 4 hours (optional). Add 15 μl of proteinase K (10 mg / ml). Incubate at 65 °C overnight. 20. Purify using 1× volume of SPRI beads (Beckman Coulter Ampure XP beads A63881) according to the manufacturer's instructions. Since the recovery of long DNA fragments may decrease, ensure that the beads do not dry out too much. Incubate in nuclease-free water for 10 minutes.
[0100] (Tagmentation) 21. Prepare several tagmentation reaction mixtures as follows according to (the total amount of collected DNA): · X μl of DNA · 4 μl of tagmentation buffer (5×) · Y μl of Tn5 · 16 - X - Y of nuclease-free water. Incubate at 55°C for 7 minutes without mixing. Make the DNA fragments have a distribution of approximately 400 bp. As a guideline: When working with approximately 50 ng of DNA, use 0.5 - 1 μl of 12.3 uM Tn5. When working with 100 - 300 ng of DNA, use 1 μl of 24.6 uM Tn5. To obtain better results, adjust the amount of Tn5 to obtain an appropriate fragment distribution. 22. Check the distribution of DNA fragments using TapeStation or Bioanalyzer: · Use 1 μl from the tagmentation mix, add 3 μl of H20 and 1 μl of 0.2% SDS. Incubate at 55°C for 7 minutes. · Add 5 μl of 0.2% SDS and incubate at 55°C for 7 minutes to detach the transposase. · Use 2 μl from this mixture to load onto TapeStation or Bioanalyzer. If the distribution is accurate, add 1 μl of nuclease-free water to the first tagmentation mix obtained in step 21, add 5 μl of 0.2% SDS, and incubate at 55°C for 7 minutes to detach Tn5. Combine 23.25 μl of this tagmentation mix with the remaining 3 μl of the mixture obtained in step 22.
[0101] (Pull-down of Hi-C ligation products) 24. Use 25 μl of streptavidin MyOne C1 Dynabeads per sample for pull-down of ligation events. To prepare the beads, wash the beads twice with 400 μl of TB buffer (rotate for 3 minutes per wash). Resuspend 25 μl of the beads in 50 μl of 2× NTB buffer. 25. Mix the (previously obtained) beads with 22 μl of TLE and 28 μl of the tagmentation mix (obtained in step 24). Incubate at room temperature for 45 minutes with gentle rotation. 26. Wash the beads 4 times with 100 μl of 1× NTB, followed by 2 washes with 50 μl of TLE. Resuspend the beads in 25 μl of nuclease-free water.
[0102] (Library preparation) 27. Prepare five reactions as follows: · 5 μl from the mixture obtained in the previous step · 29.5 μl of H2O · 10 μl of KAPA HiFi buffer (5×) · 1.5 μl dNTP (10 mM) · 1 μl KAPA HiFi DNA polymerase · 3 μl of i7 / i5 primer (10 μM) mix. The PCR conditions are: · 3' at 72 °C · 4 - 7 cycles of {10'' at 95 °C, 30'' at 55 °C, 30'' at 72 °C} · 5' at 72 °C 28. Combine the reactions and purify using Ampure SPRI beads (1× ratio). Confirm the quality and quantity of the captured Hi-C library using TapeStation / Bioanalyzer and Qubit.
[0103] (Capture hybridization of Hi-C library using biotin-RNA - Method 1) Prepare three PCR strips: "DNA", "Hybridization", and "RNA". 29a. Preparation of Hi-C library: Transfer 300 ng to 1 μg, especially a volume corresponding to 500 ng, of the Hi-C library to a 1.5 ml Eppendorf tube and dry it using a SpeedVac (45 °C, approximately 15 minutes). Resuspend the Hi-C DNA pellet in 4 μl of nuclease-free water. 30a. Prepare the blocker mix. For each sample: · Blocker #1 - 2.5 μl (Agilent Technologies) · Blocker #2 - 2.5 μl (Agilent Technologies) · Custom blocker - 1 μl. 31a. Mix the blocker mix obtained in the previous step with the DNA library obtained in step 29. Transfer 10 μl of the DNA library to the wells of the corresponding PCR strip. Keep on ice. 32a. Prepare the hybrid mix. Keep this at room temperature. · HBI - 25 μl · HBII - 1 μl · HBIII - 10 μl · HBIV - 13 μl. Mix well; if a precipitate forms, heat at 65 °C for 5 minutes. Dispense 30 μl per capture into each well of the "Hybridization" PCR strip (Agilent 410022), close the lid of the PCR strip tube (Agilent optical cap 8x strip 401425), and keep at room temperature. 33a. Prepare a 1:4 RNase block solution (e.g., 3 μl RNase block + 9 μl water). 34a. Prepare biotin-RNA. For each capture: Mix 5 μl of the custom bait (or 2 μl of the custom bait + 3 μl of nuclease-free water if the capture system size is <3 Mb) with 2 μl of RNase block dilution. Transfer these 7 μl to an “RNA” PCR strip. Keep on ice. 35a. Hybridization reaction: Set the PCR thermocycler to the following program: 5' at 95°C, ∞ at 65°C.
[0104] The lid of the PCR device must be heated. Throughout the procedure, work quickly and strive to keep the lid of the PCR device open for the shortest possible time. If the sample evaporates, the hybridization conditions will not be optimal. 36a. Transfer the “DNA” PCR strip containing the Hi-C library to the following black-marked positions in the PCR device and start the PCR program. Incubate the DNA at 95°C for 5 minutes.
Chemical formula
Chemical formula
Chemical formula
Chemical formula
Chemical formula
Chemical formula
[0105] (Streptavidin-biotin pulldown and washing for use in Method 1 above) 41a. Buffer preparation: Binding buffer (BB, Agilent Technologies) at room temperature Washing buffer I (WB I, Agilent Technologies) at room temperature Washing buffer II (WB II Agilent Technologies) at 65 °C to 72 °C, especially 65 °C NEB2 1× (NEB B7002S) at room temperature 42a. Wash the magnetic beads: After thoroughly mixing Dynabeads MyOne Streptavidin T1 (Life Technologies 65601), add 60 μl per captured Hi-C sample to a 1.5 ml low-binding (lobind) Eppendorf tube. Wash the beads as follows (the same procedure is used for all subsequent washing steps): · Add 200 μl of BB. · Mix by vortexing for 5 seconds (low to medium speed setting). · Place the tube in a Dynal magnetic separator (Life Technologies). · Regenerate the beads and discard the supernatant. Repeat steps a) - d) for a total of three washes. 43a. Biotin - streptavidin pulldown: Using Dynabeads MyOne streptavidin T1 beads in 200 μl BB in a new low - binding Eppendorf tube, open the lid of the PCR machine (while operating the PCR machine) and pipette all of the hybridization reaction solution into the tube containing the streptavidin beads. Incubate at room temperature for 30 minutes on a rotator. 44a. Wash: After 30 minutes, place the sample in the magnetic separator and discard the supernatant. Resuspend the beads in 500 μl WB I and transfer to a new tube. Incubate at room temperature for 15 minutes. Vortex for 5 seconds every 2 - 3 minutes. Separate the beads and buffer on the magnetic separator and remove the supernatant. Resuspend in 500 μl WB II (pre - warmed to 65 - 72 °C, especially 65 °C) and transfer to a new tube. Incubate at 65 - 72 °C, especially 65 °C for 10 minutes and vortex for 5 seconds every 2 - 3 minutes (low - to - medium speed setting). Repeat the wash with WB II at 65 °C - 72 °C, especially 65 °C, a total of three times. Resuspend in 200 μl of Neb2 1×. Attach directly to the magnet. Remove the supernatant and resuspend in 30 μl of Neb2 1×. At this point, the preparation for PCR amplification of the RNA / DNA mixture hybrid "catch" on the bead surface is complete (step 45).
[0106] (Capture hybridization of Hi - C library using biotin - RNA - Method 2 (using a buffer containing a high concentration of divalent cation salts, referred to herein as "rapid hybridization") As described according to the embodiments utilizing this method herein, the preparation time can be significantly reduced (e.g., approximately 2 hours and 45 minutes). 29b. Thaw the rapid hybridization buffer and pre-warm it at room temperature until ready for use, and keep it at room temperature until ready. 30b. Prepare the blocker mix: · 2.5 μl of 1 mg / ml Cot-1 DNA obtained from the same species as the species from which the 2.5 μl of nucleic acid composition is derived, e.g., human Cot-1 DNA · 2.5 μl of 10 mg / ml salmon sperm DNA · 1 μl of custom blocker. 31b. Prepare the blocking reaction at room temperature as follows: · Add 6 μl of the prepared blocker mix to 11 μl of the prepared DNA sample (approximately 100 ng - 1 μg, e.g., 500 ng). · Pipette up and down to mix. Spin down briefly. 32b. Program the thermal cycler as shown below. Start the program and immediately press the pause button. This will heat the lid while adding the blocker mix to the pre-prepared genomic DNA fragment library. · Denaturation - 95°C for 5 minutes · Blocking - 65°C for 10 minutes · Hybridization - 65°C for 1 minute; 37°C for 3 seconds - 50 cycles · Storage - Fix at 65°C. 33b. Place the sample in the thermal cycler, resume the program, and perform denaturation and blocking. 34b. While the sample is incubating on the thermal cycler, prepare the capture bait mix on ice. 35b. Dilute the SureSelect RNase block for capture (1 part RNase block: 3 parts water): · Mix 1 μl of RNase block (Agilent Technologies). · 3 μl of water. 36b. Prepare the hybridization mix: · 2 μl of diluted SureSelect RNase block · 5 μl of SureSelect custom bait (or if the capture system is <3 Mb, 2 μl of SureSelect custom bait + 3 μl of nuclease-free water) · 6 μl of room temperature 5× Rapid Hybridization Buffer. 37b. When the thermal cycle reaches the first hybridization cycle at 65°C, press the pause button. At this point, the thermal cycler is maintained at 65°C. Open the lid of the thermal cycler and pipette 13 μl of the hybridization mix into each corresponding blocking reaction solution. Mix well by pipetting up and down slowly 8 - 10 times. The hybridization reaction solution is 30 μl at this point. 38b. Seal the wells with caps, close the lid, and press the play button to resume the program. A hybridization profile that cycles with the heated lid activated will be executed. 39b. Prepare magnetic beads (Dynabeads MyOne Streptavidin T1, Invitrogen). · Vigorously resuspend Dynal (Invitrogen) magnetic beads with a vortex mixer. · Use 60 μl of Dynabeads T1 magnetic beads for each hybridization sample. · Wash the beads: (a) Add 200 μl of SureSelect Binding Buffer (Agilent Technologies). (b) Mix the beads by pipetting up and down 10 times. (c) Place the tube on a magnetic stand. (d) Wait for 2 - 5 minutes and discard the supernatant. (e) Repeat steps (a) - (d) for a total of 3 washes. (f) Resuspend the beads in 20 μl of SureSelect Binding Buffer. 40b. Capture the DNA hybridized using streptavidin beads. · After incubation, remove the sample from the thermal cycler and briefly spin at room temperature to collect the liquid. · Add all hybridization mixtures of each sample to the corresponding prepared Dynal MyOne T1 streptavidin bead solution and invert the strip tubes placed on the plate 3 - 5 times to mix. · Incubate the hybrid capture / bead solution at room temperature for 30 minutes on a rotator or shaker. · By dispensing 1500 μl per sample, pre - warm Wash Buffer #2 at 68 °C - 72 °C, especially at 68 °C. · After 30 minutes, briefly spin - down the hybrid capture / bead solution. 41b. Wash the beads: (a) Separate the beads and buffer on a magnetic separation device and remove the supernatant. (b) Resuspend the beads in 500 μl of Wash Buffer #1 by pipetting up and down 8 - 10 times, then let stand at 23 °C for 10 minutes. Separate the beads and buffer on the magnetic separation device and remove the supernatant. (c) Repeat steps (a) - (b). (d) Separate the beads and buffer on a magnetic stand for 1 minute and remove the supernatant. (e) Add 500 μl of pre - warmed Wash Buffer #2. Gently pipette up and down 10 times to resuspend the beads. When pipetting the wash buffer up and down, dispense the buffer directly onto the pelleted beads to quickly resuspend the beads. (f) Incubate the sample at 68 °C - 72 °C, especially at 68 °C for 10 minutes. (g) Repeat steps (d) - (f) for a total of 3 washes. (h) Separate the beads and buffer on a magnetic stand. Ensure that all of Wash Buffer #2 is removed. (i) Resuspend in 50 μl of nuclease - free water and separate the beads on a magnetic stand to remove the supernatant. (j) Resuspend the beads in 23 μl of nuclease - free water and proceed to PCR. Proceed to PCR amplification of the capture Hi-C library (Step 45).
[0107] (PCR Amplification of the Capture Hi-C Library) 45. Prepare PCR with 5 amplification cycles as follows: · 5 μl from the mix obtained in the previous step · 29.5 μl of mQ (water) · 10 μl of KAPA HiFi buffer (5×) · 1.5 μl dNTP (10 mM) · 1 μl of KAPA HiFi DNA polymerase · 3 μl of primer (10 uM) mix (P5-FCA-R and FCA-P7F) PCR conditions: · 3' at 95°C · 5 cycles of {20'' at 95°C, 30'' at 55°C, 30'' at 72°C} · 3' at 72°C 46. Pool all the individual PCR reactions obtained in the above step. Place in a magnetic separation device and transfer the supernatant to a new 1.5 ml low-binding Eppendorf tube. Purify with 1× volume of SPRI beads (Beckman Coulter Ampure XP beads A63881) according to the manufacturer's instructions. Resuspend in TLE or nuclease-free water with a final volume of 20 μl. Confirm the quality and quantity of the capture Hi-C library by TapeStation / Bioanalyzer and KAPA qPCR.
[0108] (Tn5 Transposase Adapter Sequence) Sequences used for assembly on Tn5 transposase:
Table 2
[0109] (Primers for Pre-Capture PCR)
Table 3
[0110] (Blocker array)
Table 4
[0111] (Buffer solution) (5× Rapid hybridization buffer) 1540 mM MgCl2*6H2O, 0.0417% w / w HPMC, 100 mM Tris (pH 8.0) and H2O.
[0112] (Washing buffer #1) "Low stringency buffer" - high salt concentration, low temperature (2× SSC, 0.1% SDS and H2O) for removing non-specific binding probes.
[0113] (Washing buffer #2) ( "High stringency buffer" - low salt concentration, high temperature (0.1× SSC, 0.1% SDS, and H2O) for removing low-affinity hybridization probes. The present application provides an invention in the following aspects. (Aspect 1) A method for identifying a nucleic acid segment that interacts with one or more target nucleic acid segments, comprising: (a) obtaining a nucleic acid composition comprising one or more target nucleic acid segments; (b) crosslinking the nucleic acid composition; (c) fragmenting the crosslinked nucleic acid composition using an endonuclease enzyme; (d) filling in the ends of the fragmented crosslinked nucleic acid segments with one or more nucleotides comprising a biotin moiety covalently linked; (e) ligating the fragmented nucleic acid segments obtained in step (d) to produce a ligation fragment; (f) using a recombinase enzyme to perform fragmentation and insertion of oligonucleotides in one step on the ligation fragment; (g) enriching the fragment containing the biotin moiety of step (d); (h) enriching the fragment containing the one or more target nucleic acid segments; (i) sequencing the enriched fragment obtained in step (h) to identify the nucleic acid segment that interacts with the one or more target nucleic acid segments, the method as described above. (Aspect 2) The method according to Aspect 1, wherein step (h) comprises performing targeted amplification to enrich the fragment containing the one or more target nucleic acid segments. (Aspect 3) Step (h) comprises: (i) adding an isolation nucleic acid molecule that binds to the one or more target nucleic acid segments, and labeling the isolation nucleic acid molecule with the first half of a binding pair; and (ii) isolating the fragment containing the one or more target nucleic acid segments that binds to the isolation nucleic acid molecule by using the second half of the binding pair, the method according to Aspect 1. (Aspect 4) The method according to any one of Aspects 1 to 3, wherein step (f) is performed by tagmentation. (Aspect 5) The method according to any one of Aspects 1 to 4, wherein step (e) utilizes intranuclear ligation. (Aspect 6) The method according to any one of Aspects 1 to 5, wherein the recombinase enzyme is retroviral integrase, such as a mutant transposase, such as a hyperactive Tn5 transposase. (Aspect 7) The method according to any one of Aspects 1 to 6, wherein the recombinase enzyme comprises a paired end adapter sequence for sequencing or a fragment thereof. (Aspect 8) The method according to any one of Aspects 1 to 7, wherein the oligonucleotide and / or adapter sequence comprises a barcode sequence. (Aspect 9) The method according to aspect 7 or 8, wherein the oligonucleotide and / or adapter sequence is selected from SEQ ID NO: 1, SEQ ID NO: 2, and / or SEQ ID NO: 3, or an oligonucleotide sequence that enables subsequent library preparation and sequencing. (Aspect 10) The method according to any one of aspects 1 to 9, wherein the addition of the nucleic acid molecule for isolation in step (h) is carried out in the presence of a sequence that prevents the ligation fragment from binding to other ligation fragments via the complementarity of the adapter sequence, such as a blocker sequence. (Aspect 11) The method according to any one of aspects 1 to 10, wherein the one or more target nucleic acid segments are selected from a promoter, a silencer, an enhancer, or an insulator. (Aspect 12) The method according to any one of aspects 1 to 11, wherein the nucleic acid molecule for isolation is obtained from a bacterial artificial chromosome (BAC), a fosmid, or a cosmid. (Aspect 13) The method according to any one of aspects 1 or 3 to 12, wherein the nucleic acid molecule for isolation is RNA. (Aspect 14) The method according to any one of aspects 3 to 13, wherein the first half of the binding pair contains biotin and the second half of the binding pair contains streptavidin. (Aspect 15) The method according to any one of aspects 1 to 14, wherein the restriction enzyme used in step (c) is Hind III or Dpn II. (Aspect 16) The method according to any one of aspects 3 to 15, further comprising amplifying the enriched fragment containing the biotin moiety in step (g). (Aspect 17) The method according to any one of aspects 1 to 16, further comprising amplifying the ligation fragment isolated before step (i). (Aspect 18) The method according to any one of aspects 1, 2, or 4 to 17, wherein the targeted amplification or amplification is carried out by PCR. (Aspect 19) The method according to any one of aspects 1 to 18, wherein the nucleic acid composition is derived from a mammalian cell nucleus, such as a human cell nucleus. (Aspect 20) The method according to any one of aspects 1 to 19, wherein the nucleic acid composition is derived from a non-human cell nucleus, such as a mouse cell nucleus or a plant cell nucleus. (Aspect 21) The method according to any one of aspects 1 to 20, wherein the nucleic acid composition is derived from 10,000, 50,000, 200,000, 500,000, or 1,000,000 cells. (Aspect 22) A method for identifying one or more interacting nucleic acid segments indicative of a specific disease state, comprising: (a) Performing the method according to any one of aspects 1 to 21 on a nucleic acid composition obtained from an individual having a specific disease; (b) Quantifying the frequency of interaction between a nucleic acid segment and one or more target nucleic acid segments; and (c) Comparing the frequency of interaction in the nucleic acid composition derived from an individual having the disease state with the frequency of interaction in a normal control nucleic acid composition derived from a healthy subject, wherein the difference in the frequency of interaction in the nucleic acid composition indicates a specific disease, said comparing. The method comprising said comparing. (Aspect 23) The method according to aspect 22, wherein the disease state is selected from: cancer, autoimmune disease, developmental disease, or genetic disorder. (Aspect 24) A kit for identifying a nucleic acid segment that interacts with one or more target nucleic acid segments, comprising a buffer and reagents capable of performing the method according to any one of aspects 1 to 21. (Aspect 25) The kit according to aspect 24, wherein the recombinase enzyme is a retroviral integrase or a transposase enzyme, such as a mutant transposase enzyme, such as a hyperactive Tn5 transposase.
Claims
1. A method for identifying a nucleic acid segment that interacts with one or more target nucleic acid segments, comprising: (a) obtaining a nucleic acid composition comprising one or more target nucleic acid segments; (b) crosslinking the nucleic acid composition; (c) fragmenting the crosslinked nucleic acid composition using an endonuclease enzyme; (d) filling in the ends of the fragmented crosslinked nucleic acid segments with one or more nucleotides comprising a biotin moiety covalently linked; (e) ligating the fragmented nucleic acid segments obtained in step (d) to produce a ligation fragment; (f) using a recombinase enzyme to perform fragmentation and insertion of oligonucleotides on the ligation fragment in one step; (g) enriching the fragments comprising the biotin moiety of step (d); (h) enriching the fragments comprising the one or more target nucleic acid segments; (i) sequencing the enriched fragments obtained in step (h) to identify the nucleic acid segments that interact with the one or more target nucleic acid segments.
2. The method according to claim 1, wherein step (h) comprises performing targeted amplification to enrich the fragments comprising the one or more target nucleic acid segments.
3. Step (h) comprises: (i) adding an isolation nucleic acid molecule that binds to the one or more target nucleic acid segments, said adding comprising labeling the isolation nucleic acid molecule with the first half of a binding pair; and (ii) isolating the fragments comprising the one or more target nucleic acid segments that are bound to the isolation nucleic acid molecule by using the second half of the binding pair.
4. The method according to any one of claims 1 to 3, wherein step (f) is performed by tagmentation.
5. The method according to any one of claims 1 to 4, wherein step (e) utilizes intranuclear ligation.
6. The method according to any one of claims 1 to 5, wherein the recombinase enzyme is retroviral integrase.
7. The method according to claim 6, wherein the recombinase enzyme is a mutant transposase.
8. The method according to claim 7, wherein the mutant transposase is a hyperactive Tn5 transposase.
9. The method according to any one of claims 1 to 8, wherein the recombinase enzyme comprises a paired-end adapter sequence for sequencing or a fragment thereof.
10. The method according to any one of claims 1 to 9, wherein the oligonucleotide and / or adapter sequence comprises a barcode sequence. **Claim 11** The method according to claim 9 or 10, wherein the oligonucleotide and / or adapter sequence is selected from SEQ ID NO: 1, SEQ ID NO: 2, and / or SEQ ID NO: 3, or an oligonucleotide sequence that enables subsequent library preparation and sequencing. **Claim 12** The method according to any one of claims 1 to 11, wherein the addition of the nucleic acid molecule for isolation in step (h) is carried out in the presence of a sequence that prevents the ligation fragment from binding to other ligation fragments via the complementarity of the adapter sequence. **Claim 13** The method according to claim 12, wherein the sequence that prevents the ligation fragment from binding to other ligation fragments via the complementarity of the adapter sequence is a blocker sequence. **Claim 14** The method according to any one of claims 1 to 13, wherein the one or more target nucleic acid segments are selected from a promoter, a silencer, an enhancer, or an insulator. **Claim 15** The method according to any one of claims 1 to 14, wherein the nucleic acid molecule for isolation is obtained from a bacterial artificial chromosome (BAC), a fosmid, or a cosmid. **Claim 16** The method according to any one of claims 1 or 3 to 15, wherein the nucleic acid molecule for isolation is RNA. **Claim 17** The method according to any one of claims 3 to 16, wherein the first half of the binding pair contains biotin and the second half of the binding pair contains streptavidin. **Claim 18** The method according to any one of claims 1 to 17, wherein the restriction enzyme used in step (c) is Hind III or Dpn II. **Claim 19** The method according to any one of claims 3 to 18, further comprising amplifying the enriched fragment containing the biotin moiety in step (g). **Claim 20** The method according to any one of claims 1 to 19, further comprising amplifying the ligation fragment isolated before step (i). **Claim 21** The method according to any one of claims 1, 2, or 4 to 20, wherein the targeted amplification or amplification is carried out by PCR. **Claim 22** The method according to any one of claims 1 to 21, wherein the nucleic acid composition is derived from a mammalian cell nucleus. **Claim 23** The method according to claim 22, wherein the mammalian cell nucleus is a human cell nucleus. **Claim 24** The method according to any one of claims 1 to 23, wherein the nucleic acid composition is derived from a non-human cell nucleus.
25. The method according to claim 24, wherein the non-human cell nucleus is a mouse cell nucleus or a plant cell nucleus.
26. The method according to any one of claims 1 to 25, wherein the nucleic acid composition is derived from 10,000, 50,000, 200,000, 500,000, or 1,000,000 cells.
27. A method for identifying one or more interacting nucleic acid segments indicative of a particular disease state, comprising: (a) performing the method according to any one of claims 1 to 26 on a nucleic acid composition obtained from an individual having a particular disease; (b) quantifying the frequency of interaction between a nucleic acid segment and one or more target nucleic acid segments; and (c) comparing the frequency of interaction in the nucleic acid composition from the individual having the disease state with the frequency of interaction in a normal control nucleic acid composition from a healthy subject, wherein a difference in the frequency of interaction in the nucleic acid composition indicates a particular disease, said comparing.
28. The method according to claim 27, wherein the disease state is selected from cancer, autoimmune disease, developmental disease, or genetic disorder.
29. A kit for identifying a nucleic acid segment that interacts with one or more target nucleic acid segments, comprising a buffer and reagents capable of performing the method according to any one of claims 1 to 26.
30. The kit according to claim 29, wherein the recombinase enzyme is a retroviral integrase or a transposase enzyme.
31. The kit according to claim 30, wherein the transposase enzyme is a mutant transposase enzyme.
32. The kit according to claim 31, wherein the mutant transposase enzyme is a hyperactive Tn5 transposase.
Citation Information
Patent Citations
Chromosome conformation capture method including selection and enrichment steps
JP2016533758A