New method

By employing cross-linking, endonuclease fragmentation, biotin filling, and recombinase single-step fragmentation, the problem of requiring a large number of cell samples in existing technologies has been solved, achieving efficient identification of nucleic acid segment interactions in a small number of cell samples.

CN114787378BActive Publication Date: 2026-04-24BABRAHAM INST
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BABRAHAM INST
Filing Date
2020-10-05
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing technologies require large numbers of cell samples to identify interactions between genomic loci, making them difficult to apply to rare cell types or early developmental stages. Furthermore, conventional methods often involve complex libraries that make it difficult to distinguish interactions between specific regulatory elements.

Method used

An improved method is employed, comprising cross-linking of nucleic acid compositions, endonuclease fragmentation, biotin filling, recombinase single-step fragmentation and oligonucleotide insertion, enrichment of biotin-containing fragments, and sequencing, to identify interactions between target nucleic acid segments.

Benefits of technology

This method can efficiently identify nucleic acid segment interactions in fewer cell samples, providing higher resolution and fewer steps, reducing nucleic acid composition loss, and capturing long-distance genomic interactions, making it suitable for rare cell types and early developmental stage samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114787378B_ABST
    Figure CN114787378B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for identifying a nucleic acid segment that interacts with a target nucleic acid segment or a plurality of target nucleic acid segments and a kit for performing the method. The present invention also relates to a method for identifying one or more interacting nucleic acid segments indicative of a particular disease.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for identifying nucleic acid segments that interact with one or more target nucleic acid segments, and a kit for performing the method. The invention also relates to a method for identifying one or more interacting nucleic acid segments that indicate a specific disease. Background Technology

[0002] Regulatory elements play a central role in the genetic control of organisms and have been shown to contribute to health and disease (e.g., in cancer and autoimmune diseases). These regulatory elements (e.g., enhancers) have been shown to exist at considerable genomic distances (on a linear scale) from their target genes. Methods for capturing these regulatory elements and their target genes have been developed and are widely used to study regulatory landscape dynamics influencing gene expression and phenotypic development, as well as the role of gene modifications in disease formation. However, identifying which target genes these regulatory elements regulate presents a significant challenge when using low-cell-count operations.

[0003] One of the first methods developed for identifying interactions between genomic loci was chromosome conformation capture (3C) technology (Dekker et al., Science (2002) 295:1306-1311). This method involves creating a 3C library by cross-linking nuclear compositions, thereby connecting spatially close genomic loci; removing intercalated DNA loops between the cross-links through a digestion process; and ligating and reversing the cross-links of interacting regions to generate a 3C library. This library can then be used to detect / identify the frequency of interactions between known sequences. However, this method requires prior knowledge of the interactions in order to detect the target interacting regions. Since then, the technique has been further developed to overcome the limitations of the accompanying 3C method.

[0004] Hi-C is a genome-wide approach that requires no prior knowledge of the target interactome of interest. This approach uses junction markers to isolate all connected interacting sequences in the cell (see WO 2010 / 036323 and Lieberman-Aiden et al., 2009). While this provides information on all interactions present at a specific time point in the nuclear composition, the resulting library is extremely complex, hindering the resolution required to analyze them to identify obvious interactions between specific elements such as promoters and enhancers. To overcome this limitation, capture Hi-C techniques have been developed, involving a capture step of enriching Hi-C libraries for chromosome interactions, the library containing the target region at at least one end (see WO 2015 / 033134, Dryden et al., 2014, and Schoenfelder et al., 2015). WO 2015 / 033134 discloses a method and kit for identifying nucleic acid segments that interact with one or more target nucleic acid segments by using isolated nucleic acid molecules. However, this method requires starting with a huge number of cells (30-40 million cells), which is not possible when working with rare cell types, cells from early stages of biological development, or patient / biopsy samples.

[0005] Therefore, there is a need to provide an improved method for identifying nucleic acid interactions that overcomes the limitations of existing available methodologies. Invention Overview

[0007] According to a first aspect of the present invention, a method is provided for identifying a nucleic acid segment that interacts with one or more target nucleic acid segments, the method comprising the following steps:

[0008] (a) Obtaining a nucleic acid composition containing one or more target nucleic acid segments;

[0009] (b) Crosslinking the nucleic acid composition;

[0010] (c) Fragmenting the cross-linked nucleic acid composition using an endonuclease;

[0011] (d) Fill the ends of the fragmented cross-linked nucleic acid segment with one or more nucleotides containing covalently linked biotin moieties;

[0012] (e) Connect the fragmented nucleic acid segment obtained from step (d) to produce a ligated fragment;

[0013] (f) The ligated fragment was fragmented in a single step and oligonucleotides were inserted using a recombinase;

[0014] (g) Enrichment of fragments containing the biotin moiety from step (d);

[0015] (h) Enrichment of fragments containing one or more target nucleic acid segments;

[0016] (i) Sequencing the enriched fragments obtained in step (h) to identify nucleic acid segments that interact with the one or more target nucleic acid segments.

[0017] According to another aspect of the present invention, a method is provided for identifying one or more interacting nucleic acid segments indicating a specific disease state, the method comprising:

[0018] (a) Applying the method described herein to a nucleic acid composition obtained from an individual with a specific disease;

[0019] (b) Quantifying the frequency of interactions between nucleic acid segments and one or more target nucleic acid segments; and

[0020] (c) The interaction frequencies in the nucleic acid composition from individuals with the said disease state are compared with the interaction frequencies in a normal control nucleic acid composition from healthy subjects, such that the difference in interaction frequencies in the nucleic acid composition indicates a specific disease.

[0021] According to another aspect of the invention, a kit is provided for identifying nucleic acid segments that interact with one or more target nucleic acid segments, comprising buffers and reagents capable of performing the methods described herein. Brief description of the attached diagram

[0023] Figure 1 A comparative overview of conventional protocols compared to the miniaturized promoter capture Hi-C protocol presented in this article. Abbreviations: Tn5 – recombinase / transposase, B – biotin, NGS – next-generation sequencing, UMI – unique molecular identifier.

[0024] Figure 2 Results obtained using the methods described herein. Significant interactions were elicited when CHiCAGO was used in samples of 50,000 and 1 million human CD4+ T cells (1 / 600 and 1 / 30 of the starting materials used in the conventional Hi-C protocol, respectively). Invention Details

[0026] According to a first aspect of the present invention, a method is provided for identifying a nucleic acid segment that interacts with one or more target nucleic acid segments, the method comprising the following steps:

[0027] (a) Obtaining a nucleic acid composition containing one or more target nucleic acid segments;

[0028] (b) Crosslinking the nucleic acid composition;

[0029] (c) Fragmenting the cross-linked nucleic acid composition using an endonuclease;

[0030] (d) Fill the ends of the fragmented cross-linked nucleic acid segment with one or more nucleotides containing covalently linked biotin moieties;

[0031] (e) Connect the fragmented nucleic acid segment obtained from step (d) to produce a ligated fragment;

[0032] (f) Use transposases to perform single-step fragmentation and oligonucleotide insertion on the ligated fragment;

[0033] (g) Enrichment of fragments containing the biotin moiety from step (d);

[0034] (h) Enrichment of fragments containing one or more target nucleic acid segments;

[0035] (i) Sequencing the enriched fragments obtained in step (h) to identify nucleic acid segments that interact with the one or more target nucleic acid segments.

[0036] The method of this invention provides a means of identifying interacting sequences and nucleic acid segments by using targeted amplification or isolation of one or more target nucleic acid segments. Such methods have the advantage of focusing data on specific interactions within highly complex libraries. Furthermore, methods including targeted amplification or addition of isolation nucleic acid molecules that bind to one or more target nucleic acid segments can also be used to organize information into multiple subsets depending on the type of reagent used or the type of isolation nucleic acid molecule used (e.g., a promoter for identifying promoter interactions). Detailed information about chromosome interactions within a specific group of the target can then be obtained.

[0037] The method of the present invention further provides single-step fragmentation and oligonucleotide insertion using recombinases. This single-step fragmentation and oligonucleotide insertion has the advantage of providing a method with significantly fewer overall steps and reduced manipulation of the nucleic acid composition. For example, in a specific embodiment of the method of the present invention, a single tube can be used from obtaining the nucleic acid composition (as described herein, step (a)) to ligating the fragmented nucleic acid segment (as described herein, step (e)) and also from performing single-step fragmentation and oligonucleotide insertion (as described herein, step (f)) to enriching the biotin-containing fragment (as described herein, step (g)). Methods that do not include single-step fragmentation and oligonucleotide insertion require separate fragmentation, library fragment end repair, addition of dATP to the 3' end of the library fragment, size selection, ligation of oligonucleotide sequences, and purification of unligated oligonucleotide fragments (as those methods disclosed in WO 2015 / 033134) using physical or enzymatic means (e.g., sonication or restriction enzyme digestion). Therefore, it will be understood that the present invention provides a method for identifying nucleic acid segments that interact with one or more target nucleic acid segments, a method that is simpler, includes fewer steps, and can have a shorter completion timeframe. Thus, the method of the present invention is faster and reduces the overall cost of library generation compared to conventional methods. It will also be understood that such advantages can lead to reduced loss of nucleic acid composition. This reduced loss of nucleic acid composition allows for a reduction in the amount of starting material (e.g., the number of cells from which the nucleic acid composition is obtained) or an increase in the amount of resulting nucleic acid composition available for subsequent analysis.

[0038] Furthermore, while prior art techniques such as 4C allow for the capture of genome-wide interactions of one or a few promoters, the method described herein can capture over 22,000 promoters and their interacting genomic loci in a single experiment. Additionally, the method of this invention produces significantly more quantitative reads.

[0039] Genome-wide association studies (GWAS) have identified thousands of disease-linked single nucleotide polymorphisms (SNPs). However, many of these SNPs are located at great distances from genes, making it very challenging to predict which genes they act on. Therefore, the method of this invention provides a way to identify interacting nucleic acid segments, even if they are geographically dispersed within the genome.

[0040] As used herein, the reference to “nucleic acid segment” is equivalent to the reference to “nucleic acid sequence”, and refers to any polymer of nucleotides (i.e., adenine (A), thymidine (T), cytosine (C), guanosine (G), and / or uracil (U)). Such polymers may or may not produce functional genomic segments or genes. Combinations of nucleic acid sequences can ultimately comprise chromosomes. Nucleic acid sequences containing deoxyribonucleosides are called deoxyribonucleic acid (DNA). Nucleic acid sequences containing ribonucleosides are called ribonucleic acid (RNA). RNA can be further characterized into several types, such as protein-coding RNA, messenger RNA (mRNA), transfer RNA (tRNA), non-coding long RNA (lnRNA), intergenic non-coding long RNA (lincRNA), antisense RNA (asRNA), microRNA (miRNA), short interfering RNA (siRNA), small nucleoRNA (snRNA), and nucleolar small RNA (snoRNA).

[0041] A single nucleotide polymorphism (SNP) is a single nucleotide variation (i.e., A, C, G, or T) that differs between members of a biological species or between paired chromosomes in the genome.

[0042] It should be understood that the term "target nucleic acid segment" refers to one or more target sequences known to the user. Isolating only the ligated fragments containing the target nucleic acid segment helps to aggregate data to identify specific interactions with a particular gene or target gene segment. Alternatively, targeted amplification to enrich fragments containing the one or more target nucleic acid segments helps to aggregate data to identify specific interactions with a particular gene or target gene segment by increasing the proportion of fragments containing the one or more target nucleic acid segments in the composition.

[0043] References to the terms "interaction" or "interactivity" herein refer to an association between two elements, such as, in the method of this invention, a genomic interaction between a nucleic acid segment and a target nucleic acid segment. An interaction can cause one interacting element to influence another element, for example, silencing or activating the element it binds to. Interactions can occur between two nucleic acid segments that are close or distant on a linear genomic sequence. Thus, in one embodiment, the nucleic acid segment or segments interacting with the target nucleic acid segment or segments are adjacent to the target nucleic acid segment or segments on a linear genomic sequence, for example, relatively close to each other on the same chromosome. In yet another embodiment, the nucleic acid segment or segments interacting with the target nucleic acid segment or segments are distant from the target nucleic acid segment on a linear genomic sequence, for example, existing on a different chromosome; or, if on the same chromosome, further distant.

[0044] As used herein, the term "nucleic acid composition" refers to any composition comprising nucleic acids and proteins. The nucleic acids in a nucleic acid composition can be organized into chromosomes, and the proteins (i.e., histones, for example) can bind to chromosomes with regulatory functions. In one embodiment, the nucleic acid composition comprises a nuclear composition. Such a nuclear composition may generally comprise a nuclear genome organization structure or chromatin.

[0045] As used herein, references to “crosslinking” or “cross-linking” refer to any stable chemical bond between two compounds, allowing them to be further processed as a unit. This stability can be based on covalent and / or non-covalent bonds (e.g., ionic bonds). For example, nucleic acids and / or proteins can be crosslinked by chemicals (i.e., immobilizers), heat, pressure, pH changes, or radiation, thereby maintaining their spatial relationships during routine laboratory procedures (i.e., extraction, washing, centrifugation, etc.). As used herein, crosslinking is equivalent to the term “fixation” or “fixation method,” which applies to any method or process that fixes any and all cellular processes. Crosslinked / fixed cells thus precisely maintain the spatial relationships between the components in a nucleic acid composition upon fixation. Many chemicals can provide fixation, including but not limited to formaldehyde, formalin, or glutaraldehyde.

[0046] As used herein, the term "fragment" refers to any nucleic acid sequence shorter than the sequence from which it is derived. Fragments can be of any size, ranging from trillions and / or thousands of bases to just a few nucleotides in length. Fragments are suitably longer than 5 nucleotides, for example, 10, 15, 20, 25, 30, 40, 50, 100, 250, 500, 750, 1000, 2000, 5000, or 10000 nucleotides. Fragments can be even longer, for example, 1, 5, 10, 20, 25, 50, 75, 100, 200, 300, 400, or 500 thousand nucleotides. Methods such as restriction enzyme digestion, sonication, acid incubation, alkaline incubation, and microfluidization can be used to fragment nucleic acid compositions.

[0047] In some implementations, fragmentation is performed using a nuclease (i.e., step (c)). Suitable examples of nucleases include, but are not limited to, sequence-specific nucleases such as restriction enzymes, and non-sequence-specific nucleases such as MNase or DNase.

[0048] Therefore, in one embodiment, the endonuclease is a sequence-specific endonuclease, such as a restriction enzyme. As used herein, the term "restriction enzyme" refers to any protein that cleaves nucleic acids at a specific base pair sequence. Depending on the type of restriction enzyme chosen, cleavage can produce blunt ends or sticky ends. Examples of restriction enzymes include, but are not limited to, Eco RI, EcoRII, Bam HI, Hind III, Dpn II, Bgl II, Nco I, Taq I, Not I, Hinf I, Sau 3A, Pvu II, SmaI, Hae III, Hga I, Alu I, Eco RV, Kpn I, Pst I, Sac I, Sal I, Sca I, Spe I, Sph I, Stu I, and Xba I. In yet another embodiment, fragmentation is performed using a restriction enzyme (i.e., step (c)). In one embodiment, the restriction enzyme is Hind III. In yet another embodiment, the restriction enzyme is Dpn II.

[0049] In one alternative embodiment, the endonuclease is a non-sequence-specific endonuclease. As used herein, the term "non-sequence-specific endonuclease" refers to any protein that cleaves nucleic acids and is not limited to the sequence of said nucleic acids; for example, it can cleave nucleic acids in any region where proteins (e.g., nucleosomes and / or transcription factors) are not bound. Examples of non-sequence-specific endonucleases are known in the art and include, but are not limited to, DNases, RNases, and MNases. MNase is a non-specific endonuclease derived from the bacterium Staphylococcus aureus that binds to and cleaves regions of DNA on chromatin that are not bound to proteins—DNA bound to histones or other chromatin-binding proteins remains undigested. In yet another embodiment, fragmentation is performed using a non-sequence-specific endonuclease (i.e., step (c)).

[0050] In another embodiment, fragmentation is performed using ultrasonic processing (i.e., step (c)).

[0051] The term "filling the end of a fragment or the end of a nucleic acid segment" as used herein refers to the addition of a nucleotide to the 3' end of a cross-linked nucleic acid composition or segment after fragmentation. This filling includes the addition of dATP, dCTP, dGTP, and / or dTTP nucleotides to the 3' end of the nucleic acid composition or segment. To allow enrichment of already linked and thus containing linker-bound nucleic acid fragments or segments, one or more nucleotides used for filling, as described herein, may contain a covalently linked biotin moiety. Thus, in one embodiment, filling the end of the fragmented cross-linked nucleic acid segment includes adding a biotin moiety to the end of the cross-linked nucleic acid fragment. In yet another embodiment, filling the end of the fragmented cross-linked nucleic acid segment includes "labeling" the end of the cross-linked nucleic acid fragment with a "boundary marker." This "labeling" end or the addition of a biotin moiety to the end of the cross-linked nucleic acid fragment allows for subsequent selection or enrichment of nucleic acid fragments and / or segments already linked according to step (e) and as described herein.

[0052] Boundary markers allow for the purification of ligated fragments prior to the enrichment step (h), thus ensuring that only ligated sequences are enriched, and not unligated (i.e., non-interacting) fragments are enriched.

[0053] In some embodiments, the boundary marker comprises a labeled nucleotide linker (i.e., a nucleotide containing a covalently linked biotin moiety). In yet another embodiment, the boundary marker comprises biotin. In one embodiment, the boundary marker may comprise a modified nucleotide. In one embodiment, the boundary marker may comprise an oligonucleotide linker sequence.

[0054] In this document, the terms “linked” or “joined” refer to any bonding of two nucleic acid segments, which typically involves a phosphodiester bond. This bonding is usually facilitated in the presence of a catalytic enzyme (e.g., a ligase such as T4 DNA ligase) in the presence of a cofactor reagent and an energy source (i.e., adenosine triphosphate (ATP)). In the methods described herein, fragments of two already cross-linked nucleic acid segments are joined together to produce a single ligated fragment.

[0055] In one embodiment, the ligation of fragmented nucleic acid segments to produce ligated fragments (i.e., step (e)) utilizes in-nuclear ligation. Therefore, in some embodiments, the ligation of fragmented nucleic acid segments is performed via in-nuclear ligation. This in-nuclear ligation has the advantage of allowing the use of small volumes of reagents, resulting in reduced loss of the nucleic acid composition and thus allowing for a reduction in the amount of starting material. For example, the number of cells from which the nucleic acid composition is obtained can be reduced, or the amount of resulting nucleic acid composition available for subsequent analysis can be increased.

[0056] The reference to "single-step fragmentation and oligonucleotide insertion" in this document refers to the fragmentation of a ligated fragment and the insertion of an oligonucleotide sequence in a single step. These methods utilize recombinases that bind to oligonucleotide sequences and insert them onto the fragment. This process is also known as "tagmentation." Therefore, in one embodiment, single-step fragmentation and oligonucleotide insertion includes tagging.

[0057] In addition to the advantages mentioned above, the advantages of single-step fragmentation and oligonucleotide ligation further include the elimination of the need to remove any binding pair elements (such as biotin) already incorporated into the nucleic acid composition from the unligated fragment, the elimination of the need for size selection of the ligated fragment, the elimination of the need for end repair due to the absence of sonication, and the elimination of the need for A-tail addition. Furthermore, the insertion of oligonucleotide and / or adapter sequences (which may include barcode sequences and / or unique molecular identifiers) occurs simultaneously with fragmentation. Such barcode sequences or unique molecular identifiers can allow for the identification of specific nucleic acid compositions in subsequent analysis and processing and allow for the merging of multiple nucleic acid compositions in subsequent steps while retaining the ability to identify and analyze individual nucleic acid compositions. Thus, in one embodiment, the oligonucleotide sequence is an "adaptor sequence" that allows or enables subsequent library preparation and sequencing of nucleic acid fragments containing the adaptor. In yet another embodiment, the adaptor contains a barcode sequence and / or a unique molecular identifier.

[0058] In yet another embodiment, single-step fragmentation and oligonucleotide insertion includes inserting a barcode sequence into the already linked fragment. In one embodiment, the paired end-adaptor sequence contains a barcode sequence and / or a unique molecular identifier.

[0059] Another advantage of the method of the present invention, which utilizes single-step fragmentation and oligonucleotide ligation (e.g., tagging) as illustrated herein, is the generation of significantly enriched fragment libraries containing one or more target nucleic acid segments compared to previously disclosed methods. For example, enrichment values ​​between at least 5-fold and 20-fold, or between at least 5-fold and 80-fold, can be generated compared to libraries generated according to previously known or conventional Hi-C protocols. In one embodiment, libraries enriched at least 5-fold, at least 10-fold, at least 15-fold, or at least 20-fold can be generated according to the method described herein compared to libraries generated according to conventional Hi-C protocols. In yet another embodiment, libraries enriched at least 10-fold, at least 11-fold, at least 12-fold, at least 13-fold, at least 14-fold, or at least 15-fold can be generated according to the method described herein. In yet another embodiment, libraries enriched at least 50-fold, at least 55-fold, or at least 60-fold can be generated according to the method described herein. It will be understood that any enrichment values ​​obtained when implementing the methods described herein, compared to libraries generated according to conventional Hi-C protocols, can depend on the identity of the endonuclease used to fragment the cross-linked nucleic acid composition. For example, when using the restriction enzyme Hind III, enrichment values ​​up to 20-fold can be obtained. Alternatively, when using the restriction enzyme Dpn II, enrichment values ​​up to 80-fold can be obtained.

[0060] As used herein, the term "paired-end adaptor" refers to any set of primer pairs that allow automated high-throughput sequencing to read from both ends. Examples of such high-throughput sequencing devices compatible with these adaptors include, but are not limited to, Solexa (Illumina), 454System, and / or ABI SOLiD. For example, the method may include the use of universal primers linked to polyadenylated tails.

[0061] It will be understood that the recombinases suitable for use in the methods of this invention include any enzyme capable of removing (or cleaving) a sequence and inserting it into an oligonucleotide or nucleic acid fragment. Examples of such recombinases include retroviral integrases and transposases such as MuA, Tn5, Tn7, and Tc1 / sailor transposases. Thus, in one embodiment of the methods of this invention, the recombinase is a retroviral integrase. In yet another embodiment, the recombinase is a transposase, such as a Tn5 transposase. For recombinases, integrases or transposases that are active in the methods shown herein can be mutated to overcome the naturally occurring low activity levels of such enzymes. Thus, in yet another embodiment, the recombinase is a mutant transposase, such as a hyperactive transposase. This hyperactive transposase may be a mutant Tn5 transposase. In one embodiment, the recombinase is a Tn5 transposase, such as a hyperactive Tn5 transposase.

[0062] Tn5 transposases are members of the RNase superfamily of recombinase proteins, which includes retroviral integrase and catalyzes the movement of a portion of nucleic acid (called a transposon) to another part of the genome or another genome via a so-called "cut-and-stick" mechanism. Recombinases such as transposases and transposon elements are found in some bacteria and are involved in the acquisition of antibiotic resistance. Transposases are usually inactive, and mutations at or elsewhere in the active site of the protein can lead to the production of hyperactive enzymes. Methods for producing Tn5 transposases are known in the art (Picelli et al., (2014) Genome Research 24:2033-2040). However, these methods can be further modified when purifying Tn5 transposases by using oligonucleotide sequences, such as adaptor sequences.

[0063] The oligonucleotide sequences (such as adaptor sequences) used in purifying recombinases (e.g., Tn5 transposases) are incorporated with the enzyme and subsequently inserted by the recombinase into nucleic acid fragments or regions. These sequences can be sequence-diverse and contain additional elements that enable further processing of the nucleic acid fragments or regions into which these sequences are inserted. For example, oligonucleotides incorporated with purified recombinases may contain sequencing adaptor sequences and / or barcode sequences. However, it will be understood that all such oligonucleotides contain transposon sequences or elements that allow incorporation with the enzyme. Examples of transposon sequences or elements include Tn5 transposase-compatible mosaic end (ME) sequences and sequences spatially compatible with the binding pocket of the recombinase and / or transposase.

[0064] Therefore, according to one embodiment, the recombinase of this method comprises a mosaic terminal double-stranded (MEDS) oligonucleotide containing half of a paired-end adaptor sequence. In yet another embodiment, the recombinase comprises a paired-end adaptor sequence for sequencing. In further embodiments, the transposase may comprise an oligonucleotide containing a paired-end adaptor sequence for sequencing, additionally comprising a barcode sequence. In other embodiments, the oligonucleotide sequence is selected from SEQ ID NO:1, SEQ ID NO:2, and / or SEQ ID NO:3 as described herein. In an alternative embodiment, the oligonucleotide sequence comprises any sequence that enables subsequent library preparation and sequencing. It will be understood that such sequences enable the amplification and isolation of nucleic acid segments and the binding of said nucleic acid segments for sequence analysis by high-throughput or next-generation sequencing methods. Examples of next-generation sequencing platforms include: Roche 454 (i.e., Roche 454GS FLX), Applied Biosystems' SOLiD system (i.e., SOLiDv4), Illumina's GAIIx, HiSeq 2000 and MiSeq sequencers, Life Technologies' Ion Torrent semiconductor-based sequencers, Pacific Biosciences' PacBio RS and Oxford Nanopore's MinION.

[0065] The terms "enrichment" or "concentration" in this document refer to any separation of nucleic acid segments or an increase in the proportion of a target nucleic acid segment relative to other nucleic acid segments in the nucleic acid composition. It will be understood that such terms include "separation," "isolation," "separation," "removal," "purification," etc. For example, enriching or separating a target nucleic acid segment may include positive methods, such as "separating" the target nucleic acid segment, or negative methods, such as excluding nucleic acid segments not intended for use or excluding nucleic acid segments that do not contain the target nucleic acid segment. Alternatively, enrichment or separation may include selective or directed amplification of the target nucleic acid segment. Such selective or directed amplification of the target nucleic acid segment will increase the proportion of such segments in the nucleic acid composition (i.e., enriching the segment).

[0066] In one implementation, the enrichment step (h) includes the step of: performing targeted amplification to enrich fragments containing one or more target nucleic acid segments.

[0067] In an alternative implementation, the enrichment step (h) includes the following steps:

[0068] (i) Adding a separation nucleic acid molecule that binds to the target nucleic acid segment, wherein the separation nucleic acid molecule is labeled with the first half of a binding pair; and

[0069] (ii) By using the second half of the binding pair, a fragment containing a target nucleic acid segment that binds to the nucleic acid molecule used for separation is isolated.

[0070] This enriches fragments containing one or more target nucleic acid segments. In some embodiments, steps (i) and (ii) above can be performed sequentially, i.e., step (i) is performed, followed by step (ii). In other embodiments, steps (i) and (ii) above can be performed simultaneously.

[0071] Therefore, the enrichment step (h) of the method of the present invention includes enriching nucleic acid fragments or target segments or target nucleic acid segments containing specific target segments or sequences.

[0072] The term "directed amplification" as used herein refers to amplification using a method that preferentially amplifies a specific target nucleic acid segment or segment. This directed amplification may utilize specific primer sequences complementary to the target nucleic acid segment or sequence (e.g., promoter or silencer sequence) present in the target nucleic acid segment. Thus, in one embodiment, the primer sequence is complementary to a promoter sequence. In another embodiment, the primer sequence is complementary to a sequence containing an SNP. The primer sequences used in the methods shown herein may contain additional elements involved in subsequent processing or analysis of the amplified nucleic acid segment. For example, the primer sequence may contain an adaptor sequence, as described herein, for sequencing or as a unique molecular identifier for identifying the single nucleic acid segment or group of nucleic acid segments (e.g., those derived from a specific sample). Alternatively or additionally, directed amplification may utilize specific conditions that may favor the amplification of the target nucleic acid segment or fragment containing that target nucleic acid segment. It will be understood that amplification can be performed by any method known in the art, such as polymerase chain reaction (PCR). It will be further understood that directed amplification as described herein may include amplifying nucleic acid segments in solution or on a support portion (e.g., beads) used for enrichment. Primer sequence extension can also be performed before amplification, so that amplification of the support portion can additionally include the step of extending the primer sequence before amplifying the extended sequence and nucleic acid segment.

[0073] The reference to "nucleic acid molecule for isolation" in this document refers to a molecule formed from nucleic acids configured to bind to a target nucleic acid segment. For example, a nucleic acid molecule for isolation may contain a complementary sequence to the target nucleic acid segment, which subsequently interacts with the nucleotide bases of the target nucleic acid segment (i.e., forms a base pair (bp)). It should be understood that nucleic acid molecules for isolation, such as biotinylated RNA, do not need to contain a complete complementary sequence to the target nucleic acid segment to form a complementary interaction and to isolate it from the nucleic acid composition. Nucleic acid molecules for isolation may be at least 10 nucleotide bases long, for example, at least 20, 30, 40, 50, 60, 70, 80, 90, 100, 130, 150, 170, 200, 300, 400, 500, 750, 1000, 2000, 3000, 4000, or 5000 nucleotide bases long.

[0074] In one embodiment, a separation nucleic acid molecule binding to one or more target nucleic acid segments is added between 65°C and 72°C. In a specific embodiment, the separation nucleic acid molecule is added at 65°C. Therefore, in yet another embodiment, step (i) of the above enrichment step (h) is performed between 65°C and 72°C (e.g., at 65°C). In another embodiment, a fragment containing the target nucleic acid segment bound to the separation nucleic acid molecule is separated using the second half of the binding pair between 68°C and 72°C. In a specific embodiment, the fragment is separated using the second half of the binding pair at 68°C. Therefore, in yet another embodiment, step (ii) of the enrichment step (h) is performed between 68°C and 72°C (e.g., at 68°C).

[0075] In one embodiment, a nucleic acid molecule for isolation is added in the presence of a blocking or blocking sequence. Such blocking sequences prevent the linked fragment containing an adaptor sequence from binding to other linked fragments containing an adaptor sequence via any complementarity in the adaptor sequence. Therefore, such blocking sequences prevent fragments that do not contain the one or more target nucleic acid segments from binding to fragments that do contain the one or more target nucleic acid segments. In some embodiments, the blocking sequence is added to the linked fragment prior to the addition of the nucleic acid molecule for isolation. In an alternative embodiment, the blocking sequence is added to the linked fragment synchronously or together with the nucleic acid molecule for isolation. It will thus be understood that in one embodiment, the blocking sequence contains any sequence compatible with the adaptor sequence linked to the fragment, such as a sequence complementary to a particular adaptor sequence. In yet another embodiment, the blocking sequence contains any sequence compatible with (or complementary to) a MEDS oligonucleotide containing a half of a paired-terminal adaptor sequence. In some implementations, the blocking sequence is selected from SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:13, SEQ ID NO:14, SEQ ID NO:15, SEQ ID NO:16 and / or SEQ ID NO:17 as described herein.

[0076] Additionally, the enrichment step (h) of the method of the present invention can be performed according to methods known in the art and using reagents known in the art. For example, where the enrichment step (h) includes a nucleic acid molecule for isolation that binds to the target nucleic acid segment as described herein, the method or the step thereof can be performed in the presence of a buffer with a high concentration of divalent cation salt (e.g., between 100 mM and 600 mM). The salt can be present at a molar ratio between 2.5:1 and 60:1. A volume exclusion agent / thickening agent may also be present, for example at a concentration between 0.002% and 0.1%. Additionally or alternatively, the method or the step thereof may include incubating the nucleic acid composition in the presence of a buffer. Incubation can last for 8 hours or less, optionally at two different temperatures, wherein the two different temperatures are cycled between 2 and 100 times. Examples of such buffers and methods are described in US 9,587,268. Therefore, it will be understood that when performing the enrichment step (h) according to certain embodiments disclosed herein, the enrichment of nucleic acid molecules for isolation is more rapid than when using conventional reagents and reaction methods.

[0077] The term "binding pair" as used herein refers to at least two parts (i.e., the first half and the second half) that specifically recognize each other to form a bond. Suitable binding pairs include, for example, biotin and avidin or biotin and avidin derivatives such as streptavidin and neutral avidin.

[0078] The references to “tagged” or “marked” in this document refer to the process of distinguishing targets by linking markers, where the markers contain a specific portion (i.e., an affinity tag) that has a unique affinity for the ligand. For example, a marker can serve to selectively purify and isolate nucleic acid sequences (i.e., for example, using affinity chromatography). Such markers may include, but are not limited to, biotinylate markers, histidine markers (i.e., 6His), or FLAG markers.

[0079] In one embodiment, the nucleic acid molecule for isolation contains biotin. In yet another embodiment, the nucleic acid molecule for isolation is labeled with biotin. In still another embodiment, the nucleic acid molecule for isolation is labeled with a histidine label or a FLAG label. Thus, according to certain embodiments, the binding pair may contain a label (such as a histidine or FLAG label) and an antibody.

[0080] In one embodiment, the target nucleic acid segment is selected from promoters, enhancers, silencers, or insulators. In yet another embodiment, the target nucleic acid segment is a promoter. In a further alternative embodiment, the target nucleic acid segment is an insulator.

[0081] In this article, references to the terms "one promoter" and "multiple promoters" refer to the nucleic acid sequences that facilitate the initiation of transcription in coding regions that promote efficient linking. Promoters are sometimes called "transcription initiation regions." Regulatory elements frequently interact with promoters to activate or repress transcription.

[0082] The inventors have used the method of this invention to identify thousands of promoter interactions, with each promoter exhibiting ten to twenty interactions. The method described herein has identified some interactions as cell-specific or associated with different disease states. Wide ranges of spacing between interacting nucleic acid segments have also been identified – most interactions are in the 100 kilobase range, but some may extend to 2 megabases and beyond. Interestingly, this method has also been used to show that both active and inactive genes form interactions.

[0083] Nucleic acid segments identified as interacting with promoters are candidates for regulatory elements required for proper genetic control. Disruption of these elements can alter transcriptional output and contribute to disease; therefore, linking these elements to their target genes may provide potential new drug targets for novel therapies.

[0084] Identifying which regulatory elements interact with promoters is crucial for understanding gene interactions. The method of this invention also provides a snapshot look at interactions within the nucleic acid composition at specific time points, thus allowing the method to be implemented across a series of time points, developmental stages, or experimental conditions to construct a picture of changes in interactions within the cellular nucleic acid composition.

[0085] It should be understood that in one implementation, the target nucleic acid segment interacts with a nucleic acid segment containing a regulatory element. In yet another implementation, the regulatory element includes an enhancer, a silencer, or an insulator.

[0086] As used herein, the term "regulatory gene" refers to any nucleic acid sequence that encodes a protein, which binds to the same or different nucleic acid sequences and thus regulates the transcription rate of the same or different nucleic acid sequences or otherwise affects their expression levels. As used herein, the term "regulatory element" refers to any nucleic acid sequence that affects the active state of another genomic element. For example, various regulatory elements may include, but are not limited to, enhancers, activators, repressors, insulators, promoters, or silencers.

[0087] In one implementation, the target nucleic acid molecule is a genomic region identified by chromatin immunoprecipitation (ChIP) sequencing. ChIP sequencing experiments analyze protein-DNA interactions by cross-linking protein-DNA complexes in the nucleic acid composition. The protein-DNA complexes are then separated (by immunoprecipitation), followed by sequencing of the protein-bound genomic regions.

[0088] In some implementations, the conceived nucleic acid segment is located on the same chromosome as the target nucleic acid segment. Alternatively, the nucleic acid segment is located on a different chromosome than the target nucleic acid segment.

[0089] This method can be used to identify long-range interactions, short-range interactions, or close-adjacent interactions. As used herein, "long-range interactions" refers to the detection of distantly interacting nucleotide regions within a linear genomic sequence. This type of interaction can identify, for example, two genomic regions located on different arms of the same chromosome or on different chromosomes. As used herein, "short-range interactions" refers to the detection of interacting nucleotide regions located relatively close to each other within the genome. As used herein, "close-adjacent interactions" refers to the detection of interacting nucleotide regions that are very close to each other within a linear genome or, for example, parts of the same gene.

[0090] The inventors have shown that SNPs are more frequently located in interacting nucleic acid regions and are more often found by chance than expected. Therefore, the method of the present invention can be used to identify which SNPs interact with specific genes and thus may regulate those genes.

[0091] Therefore, it will be understood from the disclosure presented herein that the method of the present invention can be used to identify any nucleic acid interactions in a nucleic acid composition, particularly DNA-DNA interactions.

[0092] In one embodiment, the nucleic acid molecule for isolation is obtained from a bacterial artificial chromosome (BAC), an F sclerotium, or a sclerotium. In yet another embodiment, the nucleic acid molecule for isolation is obtained from a bacterial artificial chromosome (BAC).

[0093] In one embodiment, the nucleic acid molecule used for isolation is DNA, cDNA, or RNA. In yet another embodiment, the nucleic acid molecule used for isolation is RNA.

[0094] Nucleic acid molecules for isolation can be used in suitable methods, such as solution hybridization selection (see WO 2009 / 099602). In this method, a set of 'bait' sequences is generated to form a hybridization mixture that can be used to isolate target nucleic acid subpopulations from a sample (i.e., a 'pond').

[0095] In one embodiment, the first half of the binding pair contains biotin and the second half of the binding pair contains streptavidin.

[0096] In one embodiment, the method additionally includes reversing crosslinking prior to step (f). It should be understood that several methods for reversing crosslinking are known in the art, and it will depend on how the crosslinking was initially formed. For example, crosslinking can be reversed by subjecting the crosslinked nucleic acid composition to high temperatures, such as above 50°C, 55°C, 60°C, 65°C, 70°C, 75°C, 80°C, 85°C, or higher. Alternatively, it may be necessary to subject the crosslinked nucleic acid composition to high temperatures for more than 1 hour, for example, at least 5 hours, 6 hours, 7 hours, 8 hours, 9 hours, 10 hours, or 12 hours or more. In one embodiment, reversing crosslinking prior to step (f) includes incubating the crosslinked nucleic acid composition at 65°C for at least 8 hours (i.e., overnight) in the presence of proteinase K.

[0097] In one embodiment, the method additionally includes purifying the nucleic acid composition prior to step (f) to remove any fragments that do not contain boundary markers.

[0098] The reference to “purified” in this document may refer to a nucleic acid composition that has been treated (i.e., fractionated) to remove multiple other components and the composition substantially retains its exhibited biological activity. When the term “substantially purified” is used, this designation will refer to a composition in which the nucleic acid constitutes the major component of the composition, such as about 50%, about 60%, about 70%, about 80%, about 90%, about 95% or more (i.e., weight / weight (w / w), volume / volume (v / v), and / or weight / volume (w / v)) of the composition.

[0099] In one embodiment, the method additionally includes amplifying the isolated target ligation fragment prior to step (i). In yet another embodiment, amplification is performed by polymerase chain reaction (PCR).

[0100] In one embodiment, the nucleic acid composition is derived from a mammalian cell nucleus. In yet another embodiment, the mammalian cell nucleus may be a human cell nucleus. Many human cells are available in the art for use in the methods described herein, such as GM12878 (a lymphoblastoid human cell line) or CD34+ (human ex vivo hematopoietic progenitor cells).

[0101] It will be understood that the methods described herein are applicable to a range of organisms, not just humans. For example, the methods of this invention can also be used to identify genome interactions in plants and animals.

[0102] Therefore, in an alternative embodiment, the nucleic acid composition is derived from a non-human cell nucleus. In one embodiment, the non-human cell is selected from, but is not limited to, plants, yeast, mice, cows, pigs, horses, dogs, cats, goats, or sheep. In one embodiment, the non-human cell nucleus is a mouse cell nucleus or a plant cell nucleus.

[0103] As will be understood from the advantages of the invention as described herein, the method of the invention reduces nucleic acid composition loss during the steps mentioned herein. This reduced loss of nucleic acid composition can allow for a reduction in the amount of starting material (e.g., the number of cells from which the nucleic acid composition is obtained). Thus, in one embodiment, the nucleic acid composition can be derived from a smaller number of cells than with previous promoter capture or conformation capture techniques. In yet another embodiment, the nucleic acid composition is derived from 1 million or fewer cells, 0.5 million or fewer cells, 0.2 million or fewer cells, 50,000 or fewer cells, or 10,000 or fewer cells. In yet another embodiment, the nucleic acid composition is derived from 1 million, 0.5 million, 0.2 million, 50,000 or 10,000 cells. In some embodiments, the nucleic acid composition is derived from 1 million, 50,000 or 10,000 cells.

[0104] In one implementation, the method described herein includes the following steps:

[0105] (i) Crosslinking a nucleic acid composition containing one or more target nucleic acid segments;

[0106] (ii) Fragmenting the cross-linked nucleic acid composition;

[0107] (iii) Label the ends of the fragment with biotin;

[0108] (iv) Connect fragmented nucleic acid segments to generate connected fragments;

[0109] (v) Reverse crosslinking;

[0110] (vi) Use transposases to perform single-step fragmentation and adaptor insertion on the ligated fragment;

[0111] (vii) Reduce the number of ligated fragments using streptavidin;

[0112] (viii) Perform targeted amplification of fragments containing one or more target nucleic acid segments; and

[0113] (ix) Sequencing to identify nucleic acid segments that interact with one or more target nucleic acid segments.

[0114] In another implementation, the method described herein includes the following steps:

[0115] (i) Crosslinking a nucleic acid composition containing one or more target nucleic acid segments;

[0116] (ii) Fragmenting the cross-linked nucleic acid composition;

[0117] (iii) Label the ends of the fragment with biotin;

[0118] (iv) Connect fragmented nucleic acid segments to generate connected fragments;

[0119] (v) Reverse crosslinking;

[0120] (vi) Use transposases to perform single-step fragmentation and adaptor insertion on the ligated fragment;

[0121] (vii) Reduce the fragment with streptavidin and amplify using PCR;

[0122] (viii) A promoter is captured by adding a separation nucleic acid molecule that binds to the one or more target nucleic acid segments, wherein the separation nucleic acid molecule is labeled with the first half of the binding pair; (ix) A ligated fragment containing the one or more target nucleic acid segments that bind to the separation nucleic acid molecule is isolated by using the second half of the binding pair.

[0123] (x) PCR amplification was used; and

[0124] (xi) Sequencing to identify nucleic acid segments that interact with one or more target nucleic acid segments.

[0125] According to another aspect of the present invention, a method is provided for identifying one or more interacting nucleic acid segments indicating a specific disease state, the method comprising:

[0126] a) The method described herein is applied to nucleic acid compositions obtained from individuals with a specific disease state;

[0127] b) Quantify the frequency of interactions between nucleic acid regions and one or more target nucleic acid regions;

[0128] c) The interaction frequencies in the nucleic acid composition from individuals with the said disease state are compared with the interaction frequencies in a normal control nucleic acid composition from healthy subjects, so that the difference in the interaction frequencies in the nucleic acid composition indicates a specific disease.

[0129] As used herein, the term "interaction frequency" refers to the number of times a specific interaction occurs in the nucleic acid composition (i.e., the sample). In some cases, a lower interaction frequency in a nucleic acid composition compared to a normal control nucleic acid composition from healthy subjects indicates a specific disease state (i.e., due to less frequent interaction between nucleic acid segments). Alternatively, a higher interaction frequency in a nucleic acid composition compared to a normal control nucleic acid composition from healthy subjects indicates a specific disease state (i.e., due to more frequent interaction between nucleic acid segments). In some cases, the difference will be represented by at least a 0.5-fold difference, such as a 1-fold, 1.5-fold, 2-fold, 2.5-fold, 3-fold, 4-fold, 5-fold, 7-fold, or 10-fold difference.

[0130] In one aspect of the invention, the frequency of interactions can be used to determine the spatial proximity of two different nucleic acid regions. As the interaction frequency increases, the probability that two genomic regions are physically close to each other in 3D nuclear space increases. Conversely, as the interaction frequency decreases, the probability that two genomic regions are physically close to each other in 3D nuclear space decreases.

[0131] Quantification can be performed using any method suitable for calculating the frequency of interactions in a nucleic acid composition derived from a patient or purified or extracted sample of the nucleic acid composition or a dilution thereof. For example, high-throughput sequencing results can also make it possible to examine the frequency of specific interactions. In the method of the present invention, quantification can be performed by measuring the concentration of a target nucleic acid segment or ligation product in one or more samples. The nucleic acid composition can be obtained from cells in a biological sample, which may include cerebrospinal fluid (CSF), whole blood, serum, plasma, or extracts or purifications derived therefrom, or dilutions thereof. In one embodiment, the biological sample may be cerebrospinal fluid (CSF), whole blood, serum, or plasma. The biological sample also includes tissue homogenates, tissue sections, and biopsy specimens obtained from a living subject or post-mortem. Samples can be prepared, for example, diluted or concentrated as needed, and preserved in a conventional manner.

[0132] In one implementation, the disease state is selected from: cancer, autoimmune disease, developmental disease, genetic disease, diabetes, cardiovascular disease, kidney disease, lung disease, liver disease, nervous system disease, viral infection, or bacterial infection. In another implementation, the disease state is cancer or an autoimmune disease. In yet another implementation, the disease state is cancer, such as breast cancer, colorectal cancer, bladder cancer, bone cancer, brain cancer, cervical cancer, colon cancer, endometrial cancer, esophageal cancer, kidney cancer, liver cancer, lung cancer, ovarian cancer, pancreatic cancer, prostate cancer, skin cancer, stomach cancer, testicular cancer, thyroid cancer, or uterine cancer, leukemia, lymphoma, myeloma, or melanoma.

[0133] The mention of "autoimmune diseases" in this article includes diseases that originate from an immune response targeting the body itself, such as acute disseminated encephalomyelitis (ADEM), ankylosing spondylitis, etc. Diseases such as cerebrospinal fluid malabsorption, Crohn's disease, type 1 diabetes, Graves' disease, Guillain-Barré syndrome (GBS), psoriasis, rheumatoid arthritis, and rheumatic fever. Syndrome, ulcerative colitis, and vasculitis.

[0134] The reference to “developmental disorders” in this article includes a variety of conditions that typically originate in childhood, such as learning disabilities, communication disorders, autism, attention deficit hyperactivity disorder (ADHD), and developmental ataxia.

[0135] The reference to “genetic diseases” in this article includes diseases originating from one or more abnormalities in the genome, such as Angelman syndrome, Canavan disease, Charcot-Marie-Tooth disease, color blindness, Cri du Chat syndrome, cystic fibrosis, Down syndrome, Duchenne muscular dystrophy, hemochromatosis, hemophilia, congenital testicular hypoplasia syndrome, neurofibromatosis, phenylketonuria, polycystic kidney disease, Prader-Willi syndrome, sickle cell disease, Tay-Sachs disease, and Turner syndrome.

[0136] According to another aspect of the invention, a kit is provided for identifying nucleic acid segments that interact with one or more target nucleic acid segments, the kit comprising buffers and reagents capable of performing the methods described herein.

[0137] The kit may include one or more items and / or reagents for performing the method. For example, oligonucleotide probes, amplification primer pairs, and / or recombinase-binding oligonucleotides used in the methods described herein may be provided in isolated form and may be part of the kit, for example, in a suitable container such as a vial in which the contents are protected from the external environment. The kit may include instructions for use according to the protocol of the methods described herein. Kits in which nucleic acids are intended for use in PCR may include one or more other reagents required for the reaction, such as polymerases, nucleotides, buffer solutions, etc.

[0138] In one embodiment, the kit contains a recombinase. In yet another embodiment, the recombinase contained in the kit as described herein is a transposase, such as a highly active mutant transposase, for example, a highly active mutant Tn5 transposase.

[0139] According to another aspect of the invention, a recombinase as described herein is provided that can perform single-step fragmentation and adaptor insertion. Therefore, a recombinase capable of being tagged is also provided herein.

[0140] In one embodiment, the recombinase provided herein is a highly active mutant transposase. In yet another embodiment, the transposase is a highly active mutant Tn5 transposase. In still another embodiment, the transposase comprises a paired end adaptor sequence.

[0141] It should be understood that, in addition to those mentioned above, examples of the types of buffer solutions and reagents to be included in the kit can also be seen in the embodiments described herein.

[0142] The following research and schemes illustrate the implementation of the methods described in this paper: Example

[0143] Abbreviations:

[0144] BB binding buffer

[0145] BSA (Bovine Serum Albumin)

[0146] dd dideoxy

[0147] EDTA (ethylenediaminetetraacetic acid)

[0148] HB Agilent hybridization buffers (HBI, HBII, HBIII, and HBIV)

[0149] NaCl (sodium chloride)

[0150] NTB (Non-Tween Buffer)

[0151] PBS (Phosphate Buffered Saline)

[0152] PCR Polymerase Chain Reaction

[0153] PE mating ends

[0154] rpm (revolutions per minute)

[0155] SDS Sodium lauryl sulfate

[0156] SPRI beads reversibly immobilized in solid phase

[0157] TB Tween Buffer

[0158] Tn5 transposase

[0159] Tris-HCl Tris(hydroxymethyl)aminomethane hydrochloride

[0160] WB washing buffer

[0161] Cell fixation

[0162] 1. For a single experiment, at least 50,000 cells must be fixed at room temperature with a final formaldehyde concentration of 2% for 10 minutes.

[0163] • Quenched with a final concentration of glycine at 0.125 M.

[0164] Centrifuge at 1500 rpm (400 x g) at 4°C for 5 minutes.

[0165] • Discard the supernatant and carefully resuspend the precipitate in 100 μl of cold 1xPBS. Centrifuge at 1500 rpm (400 x g) at 4 °C for 5 minutes.

[0166] • Discard the supernatant and flash-freeze it in liquid nitrogen or proceed directly to the next step.

[0167] Cell permeation and restriction enzyme digestion

[0168] 2. Resuspend the fixed cell pellet from step 1 in 100 μl of ice-cold lysis buffer. Incubate the tube on ice for 30 minutes.

[0169] 3. Centrifuge the tube at approximately 600g at 4°C for 5 minutes.

[0170] 4. Remove the supernatant, leaving approximately 20 μl of solution containing the nuclear precipitate.

[0171] 5. Wash the precipitate twice with 100 μl of 1.2x NE buffer 3 (if using Dpn II) or NE buffer 2 (if using Hind III).

[0172] 6. Remove the supernatant, leaving approximately 20 μl. Add 334 μl of 1.2x NE buffer 3 (or NE buffer 2 if using Hind III).

[0173] 7. Add 12 μl of 10% SDS (final concentration 0.3%, w / v); shake at 950 rpm on a hot mixer at 37°C for 1 hour.

[0174] 8. Add 80 μl of 10% Triton (final concentration 1.8%, v / v); shake at 950 rpm on a hot mixer at 37°C for 1 hour.

[0175] 9. Add 30 μl Dpn II (50 U / μl) (or 15 μl Hind III – 100 U / μl and 15 μl H2O) and shake at 950 rpm on a hot mixer at 37°C for 12–16 hours.

[0176] Biotin labeling and Hi-C linker

[0177] 10. Briefly centrifuge the digestion mixture.

[0178] 11. Add 4.5 μl of dCTP, dTTP and dGTP (10 mM mixture), 37.5 μl of biotin-dATP and 10 μl of Klenow (5 U / μl). Incubate at 37 °C for 45 minutes, shaking at 700 rpm for 10 seconds every 30 seconds.

[0179] 12. Centrifuge at 600g at 4℃ for 6 minutes.

[0180] 13. Remove the supernatant, leaving approximately 50 μl containing the precipitate.

[0181] 14. Add 835 μl H2O / 100 μl T4 DNA ligase buffer / 5 μl BSA (20 mg / ml) / 10 μl T4 DNA ligase (Invitrogen).

[0182] 15. Incubate at 16°C for a minimum of 4 hours (or up to 12 hours).

[0183] Purification of Hi-C DNA

[0184] 16. Centrifuge the tube at 600g at 4℃ for 6 minutes.

[0185] 17. Remove 800 μl of supernatant, leaving 200 μl in the tube.

[0186] 18. Add 15 μl of proteinase K (10 mg / ml). Incubate at 65°C for 4 hours (optional).

[0187] 19. Add 15 μl of proteinase K (10 mg / ml). Incubate overnight at 65°C.

[0188] 20. Follow the manufacturer's instructions to purify using 1x volume of SPRI beads (Beckman Coulter Ampure XP beads A63881). Do not allow the beads to dry out completely, as this may reduce the recovery of long DNA fragments. Incubate in nuclease-free water for 10 minutes.

[0189] Tagging

[0190] 21. Several labeling reactions are set up as follows (based on the total amount of DNA collected):

[0191] ·Xμl DNA

[0192] • 4 μl of tagged buffer (5X)

[0193] ·Yμl Tn5

[0194] ·16-XY nuclease-free water

[0195] Incubate at 55°C without mixing for 7 minutes.

[0196] The target is the distribution of DNA fragments of approximately 400 bp.

[0197] As a guideline: If working with approximately 50 ng of DNA, use 0.5–1 μl of 12.3 μM Tn5. If working with 100–300 ng of DNA, use 1 μl of 24.6 μM Tn5.

[0198] For better results, titrate the amount of Tn5 to obtain the correct fragment distribution.

[0199] 22. Examine the distribution of DNA fragments on TapeStation or Bioanalyzer:

[0200] • Using 1 μl of the labeled mixture, add 3 μl of H2O and 1 μl of 0.2% SDS. Incubate at 55°C for 7 minutes.

[0201] • The transposase was removed by adding 5 μl of 0.2% SDS and incubating at 55°C for 7 minutes.

[0202] • Load this mixture into TapeStation or Bioanalyzer using 2 μl.

[0203] If the distribution is correct—add 1 μl of nuclease-free water to the initial tagging mixture from step 21 and remove Tn5 by adding 5 μl of 0.2% SDS and incubating at 55°C for 7 minutes.

[0204] 23. Combine 25 μl of this labeled mixture with the remaining 3 μl of the mixture from step 22.

[0205] Reduce Hi-C linker products

[0206] 24. To minimize ligation events, use 25 μl of streptavidin MyOne C1 Dyna beads per sample. To prepare the beads, wash them twice with 400 μl of TB buffer (rotating for 3 minutes each time). Resuspend 25 μl of beads in 50 μl of 2xNTB buffer.

[0207] 25. Combine the beads (from the previous step) with 22 μl of TLE and 28 μl of the tagging mixture (from step 24). Incubate at room temperature for 45 minutes, rotating slowly.

[0208] 26. Wash the beads four times with 100 μl of 1xNTB, followed by two washes with 50 μl of TLE. Resuspend the beads in 25 μl of nuclease-free water.

[0209] Library preparation

[0210] 27. Perform the following 5 reactions:

[0211] ·5μl from the previous mixture

[0212] 29.5 μl H2O

[0213] • 10μl KAPA HiFi buffer (5x)

[0214] ·1.5 μl dNTPs (10 mM)

[0215] 1 μl KAPA HiFi DNA Polymerase

[0216] • 3 μl i7 / i5 primer (10 μM) mixture

[0217] PCR conditions:

[0218] • At 72℃ for 3 minutes

[0219] • 4-7 cycles or less {10 seconds 95°C; 30 seconds 55°C; 30 seconds 72°C}

[0220] • At 72℃ for 5 minutes

[0221] 28. Combine the reactions and purify them using Ampure SPRI beads (1x ratio). Check the quality and quantity of the captured Hi-C library using TapeStation / Bioanalyzer and Qubit.

[0222] Biotin-RNA capture hybridization Hi-C library – Method 1

[0223] Prepare three PCR strips: "DNA", "hybridization", and "RNA".

[0224] 29a. Preparation of Hi-C libraries: Transfer volumes of Hi-C libraries between 300 ng and 1 μg, especially 500 ng, into 1.5 ml microcentrifuge tubes and dry using SpeedVac (45°C, approximately 15 minutes). Resuspend the Hi-C DNA precipitate in 4 μl of nuclease-free water.

[0225] 30a. Prepare the blocking mixture. Per sample:

[0226] • Blocker #1 – 2.5 μl (Agilent Technologies)

[0227] • Blocker #2 – 2.5 μl (Agilent Technologies)

[0228] Customized blocking agent – ​​1μl

[0229] 31a. Mix the blocking mixture from the previous step with the DNA library from step 29. Transfer 10 μl of the DNA library into the well of the appropriate PCR strip. Keep on ice.

[0230] 32a. Prepare the hybridization mixture. Keep it at room temperature.

[0231] HBI – 25μl

[0232] ·HBII–1μl

[0233] ·HBIII–10μl

[0234] ·HBIV–13μl

[0235] Mix thoroughly; if precipitate has formed, heat at 65°C for 5 minutes. Aliquot each 30 μl of the capture into each well of the “Hybridization” PCR strip (Agilent 410022), cap with the PCR tube cap (Agilent Optical Cap 8x Strip 401425) and keep at room temperature.

[0236] 33a. Prepare an RNase blocking solution at a ratio of 1:4 (e.g., 3 μl RNase blocker + 9 μl water).

[0237] 34a. Preparation of Biotin-RNA. For each capture: Mix 5 μl of custom bait (or 2 μl of custom bait + 3 μl of nuclease-free water if the capture system size is <3 Mb) with 2 μl of RNase blocking dilution. Transfer 7 μl of this mixture to an “RNA” PCR strip. Keep on ice.

[0238] 35a. Hybridization reaction: The PCR thermal cycling program was set as follows: 95°C for 5 minutes, 65°C - ∞

[0239] The PCR instrument lid must be heated. Throughout the process, operate quickly and try to keep the PCR instrument lid open for as little time as possible. Sample evaporation will result in suboptimal hybridization conditions.

[0240] 36a. Transfer the “DNA” PCR strip containing the Hi-C library to the PCR instrument, positioned as shown in the black mark below, and start the PCR program. Incubate the DNA at 95°C for 5 minutes.

[0241]

[0242] 37a. Once the temperature has reached 65°C, transfer the "hybridization" PCR strip containing the hybridization buffer to the PCR instrument, positioned as shown by the gray marker below. Incubate at 65°C for 5 minutes.

[0243]

[0244] 38a. Transfer the “DNA” PCR strip containing biotinylated RNA bait to the PCR instrument, positioned at the location marked by the cross-shaded line below. Incubate for 2 minutes.

[0245]

[0246] 39a. Open the "Hybridization" and "RNA" strips. Pipette 13 μl of hybridization buffer into 7 μl of RNA bait (from gray to the cross-shaded area). Discard the PCR strip containing the hybridization buffer. Proceed immediately to the next step.

[0247]

[0248] 40a. Remove the cap from the “DNA” PCR strip containing the Hi-C library. Transfer 10 μl of the Hi-C library into 20 μl of RNA bait (from the black line to the cross-shaded line) containing hybridization buffer. Verify that there is no residue in the DNA PCR strip and discard it.

[0249]

[0250] Immediately seal the remaining “RNA” PCR strips (now containing Hi-C library / hybridization buffer / RNA bait) with unused PCR tube caps and incubate at 65°C for 24 hours.

[0251]

[0252] Streptavidin-Biotin Reduction and Washing – to be used in conjunction with Method 1 above.

[0253] 41a. Preparation of buffer solution:

[0254] Prepare the binding buffer (BB, Agilent Technologies) at room temperature.

[0255] Prepare Wash Buffer I (WB I, Agilent Technologies) at room temperature.

[0256] Washing buffer II (WB II Agilent Technologies) is prepared between 65°C and 72°C, especially at 65°C.

[0257] Prepare NEB2 1x (NEB B7002S) at room temperature.

[0258] 42a. Washing magnetic beads:

[0259] Thoroughly mix Dyna beads with MyOne streptavidin T1 (Life Technologies 65601), then add 60 μl of the captured Hi-C sample to a 1.5 ml low-binding microcentrifuge tube. Wash the beads as follows (all subsequent washing steps follow the same procedure):

[0260] Add 200μl BB

[0261] Mix for 5 seconds on a vortex mixer (set to low to medium).

[0262] • Place the tube on the Dynal magnetic separator (Life Technologies)

[0263] • Recycle the beads and discard the supernatant.

[0264] Repeat steps a) through d) for a total of 3 washes.

[0265] 43a. Biotin-streptavidin reduction:

[0266] With 200 μl of BB in an unused, low-binding microcentrifuge tube containing Dyna MyOne streptavidin T1 beads, open the PCR instrument (while the PCR instrument is running) and transfer the entire hybridization reaction into the tube containing the streptavidin beads. Incubate at room temperature on a rotor for 30 minutes.

[0267] 44a. Washing:

[0268] After 30 minutes, place the sample on the magnetic separator and discard the supernatant.

[0269] Resuspend the beads in 500 μl of WBI and transfer to a new tube. Incubate at room temperature for 15 minutes. Vortex mix for 5 seconds every 2 to 3 minutes.

[0270] Separate the beads and buffer solution on a magnetic separator and remove the supernatant. Resuspend in 500 μl of WB II (pre-warmed to between 65°C and 72°C, especially 65°C) and transfer to a new tube. Incubate at 65°C to 72°C, especially 65°C, for 10 minutes, vortexing for 5 seconds every 2 to 3 minutes (set to low to medium). Repeat the washing in WB II a total of 3 times, all at 65°C to 72°C, especially 65°C.

[0271] Resuspend in 200 μl Neb2 1X. Place directly on a magnet. Remove the supernatant and resuspend in 30 μl Neb2 1X.

[0272] The RNA / DNA hybrid 'catch' on the bead is now ready for PCR amplification (step 45).

[0273] Biotin-RNA capture hybridization of Hi-C libraries – Method 2 (using a buffer containing a high concentration of divalent cation salts – referred to in this paper as “rapid hybridization”)

[0274] As described in this article, depending on the implementation of this method, preparation time can be significantly reduced (e.g., to approximately 2 hours and 45 minutes).

[0275] 29b. Pre-warm the rapid hybridization buffer to room temperature until thawed and keep at room temperature until ready for use.

[0276] 30b. Preparation of the blocking mixture:

[0277] • 2.5 μl 1 mg / ml Cot-1 DNA from the same species from which the nucleic acid composition was derived, such as human Cot-1 DNA

[0278] • 2.5 μl 10 mg / ml salmon sperm DNA

[0279] ·1μl custom-made blocking agent

[0280] 31b. The blocking reaction is set up at room temperature as follows:

[0281] • Add 6 μl of the blocking mixture prepared above to 11 μl of the prepared DNA sample (approximately 100 ng - 1 μg, e.g., 500 ng).

[0282] • Mix by suction from top to bottom. Briefly centrifuge.

[0283] 32b. Program the thermal cycler as shown below. Start the program and immediately press the pause button. This will heat the lid while the blocking mixture is added to the pre-prepared genomic DNA fragment library.

[0284] • Denaturation – 95℃ for 5 minutes

[0285] • Block – 65℃ for 10 minutes

[0286] • Hybridization – 50 cycles – 65°C for 1 minute; 37°C for 3 seconds

[0287] • Storage – Keep at 65℃

[0288] 33b. Place the sample into the thermal cycler and resume the program to denature and block.

[0289] 34b. While incubating the sample on a thermal cycler, prepare the trap bait mixture on ice.

[0290] 35b. Dilute SureSelect RNase Block to capture (1 part RNase Block: 3 parts water):

[0291] • Mix 1 μl of RNase Block (Agilent Technologies Inc.)

[0292] 3μl water

[0293] 36b. Preparation of hybridization mixtures:

[0294] • 2 μl diluted SureSelect RNase Block

[0295] • 5 μl SureSelect custom bait (or 2 μl SureSelect custom bait + 3 μl nuclease-free water, if the capture system is <3 Mb)

[0296] · 6 μl room temperature 5X rapid hybridization buffer

[0297] 37b. When the thermal cycle reaches the first hybridization cycle at 65°C, press the pause button. The thermal cycler is now maintained at 65°C. Open the thermal cycler lid and aspirate 13 μl of the hybridization mixture into each corresponding blocking reaction. Mix thoroughly by slowly pumping up and down 8 to 10 times. The hybridization reaction volume is now 30 μl.

[0298] 38b. Seal all holes with the caps, close the thermal cycler cover, and press the run button to restore the program running cycle hybridization conditions, while simultaneously activating the heating cover.

[0299] 39b. Prepare magnetic beads (Dyna Beads MyOne Streptavidin T1, Invitrogen)

[0300] • Vigorously suspend Dynal (Invitrogen) magnetic beads on a vortex mixer

[0301] • For each hybridization sample, use 60 μl of Dyna Beads T1 magnetic beads.

[0302] • Washing the beads:

[0303] (a) Add 200 μl of SureSelect binding buffer (Agilent Technologies Inc.)

[0304] (b) Mix the beads by pumping them up and down 10 times.

[0305] (c) Place the tube on the magnetic base

[0306] (d) Wait 2-5 minutes and discard the supernatant.

[0307] (e) Repeat steps (a) through (d) for a total of 3 washes.

[0308] (f) Resuspend the beads in 20 μl SureSelect binding buffer.

[0309] 40b. Capturing hybridized DNA using streptavidin beads

[0310] • After incubation, remove the sample from the thermal cycler and briefly centrifuge at room temperature to collect the liquid.

[0311] • Add the entire hybridization mixture of each sample to the washed and prepared corresponding Dynal MyOne T1 streptavidin bead solution and invert the strip-tube / plate to mix 3 to 5 times.

[0312] • Incubate the hybrid-capture / bead solution at room temperature for 30 minutes on a rotating or shaking table.

[0313] • Dispense 1500 μl / sample into equal portions and pre-warm the wash buffer #2 between 68°C and 72°C, especially at 68°C.

[0314] • Briefly centrifuge the hybridization-capture / bead solution after 30 minutes.

[0315] 41b. Washing the beads:

[0316] (a) Separate the beads and buffer solution on a magnetic separator and remove the supernatant.

[0317] (b) Resuspend the beads in 500 μl of wash buffer #1 by aspiration 8-10 times, then incubate at 23°C for 10 minutes. Separate the beads and buffer using a magnetic separator and remove the supernatant.

[0318] (c) Repeat steps (a) to (b).

[0319] (d) Separate the beads and buffer solution on the magnetic base for 1 minute and remove the supernatant.

[0320] (e) Add 500 μl of pre-warmed wash buffer #2. Slowly aspirate up and down 10 times to resuspend the beads. When aspirating the wash buffer, allow the buffer to directly disperse the precipitated beads for faster resuspending.

[0321] (f) Incubate the sample between 68°C and 72°C, especially at 68°C for 10 minutes.

[0322] (g) Repeat steps (d) through (f) for a total of 3 washes.

[0323] (h) Separate the beads and buffer solution on the magnetic stand. Ensure that the entire wash buffer #2 has been removed.

[0324] (i) Resuspend the beads in 50 μl of nuclease-free water, separate the beads on a magnetic stand, and remove the supernatant.

[0325] (j) The beads were resuspended in 23 μl of nuclease-free water and then introduced into PCR.

[0326] Proceed to PCR amplification of the captured Hi-C library (step 45).

[0327] PCR amplification of Hi-C libraries

[0328] 45. Set up PCR with the following 5 amplification cycles:

[0329] ·5μl from the previous mixture

[0330] ·29.5μl mQ (water)

[0331] • 10μl KAPA HiFi buffer (5x)

[0332] ·1.5 μl dNTPs (10 mM)

[0333] 1 μl KAPA HiFi DNA Polymerase

[0334] • 3 μl primer (10 μM) mixture (P5-FCA-R and FCA-P7F)

[0335] PCR conditions:

[0336] • 3 minutes at 95℃

[0337] • 5 or fewer cycles {20 seconds 95°C; 30 seconds 55°C; 30 seconds 72°C}

[0338] • At 72℃ for 3 minutes

[0339] 46. ​​Combine all PCR reactions from the above steps. Place on a magnetic separator and transfer the supernatant to an unused 1.5 ml low-binding microcentrifuge tube. Purify using 1X volume of SPRI beads (Beckman Coulter Ampure XP beads A63881) following the manufacturer's instructions. Resuspend in a final volume of 20 μl TLE or nuclease-free water.

[0340] The quality and quantity of the captured Hi-C library were examined using TapeStation / Bioanalyzer and KAPA qPCR.

[0341] Tn5 transposase adaptor sequence

[0342] Sequences used for assembly on Tn5 transposase:

[0343] Tn5MErev 5'-[phos]CTGTCTCTTATACACATCT-3' SEQ ID NO:1 FC-A 5'-GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG-3' SEQ ID NO:2 FC-B 5'-TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG-3' SEQ ID NO:3

[0344] Primers for pre-capture PCR

[0345]

[0346] “rc” indicates that the barcode sequences are inversely complementary.

[0347] Blocker sequence

[0348] i5Rdd TCGTCGGCAGCGTCAGATGTGTATAAGAGA / 3ddC / SEQ ID NO:10 ddi7F GTCTCGTGGGCTCGGAGATGTGTATAAGAGA / 3ddC / SEQ ID NO:11 i5F CTGTCTCTTATACACATCTGACGCTGCCGACGA SEQ ID NO:12 P5-FCA-F GTGTAGATCTCGGTGGTCGCCGTATCATT SEQ ID NO:13 P5-FCA-R AATGATACGGCGACCACCGAGATCTACAC SEQ ID NO:14 i7R CTGTCTCTTATACACATCTCCGAGCCCACGAGAC SEQ ID NO:15 FCA-P7F CAAGCAGAAGACGGCATACGAGAT SEQ ID NO:16 FCA-P7R ATCTCGTATGCCGTCTTCTGCTTG SEQ ID NO:17

[0349] Buffer solution

[0350] 5X Rapid Hybridization Buffer

[0351] 1540mM MgCl2*6H2O, 0.0417% w / w HPMC, 100mM Tris (pH 8.0) and H2O.

[0352] Washing buffer #1

[0353] (“Low-tightness buffer” – high salt concentration and low temperature to remove non-specifically bound probes) 2X SSC, 0.1% SDS and H2O.

[0354] Washing buffer #2

[0355] (“High-tightness buffer” – low salt concentration and high temperature to remove low-affinity hybridization probes) 0.1X SSC, 0.1% SDS and H2O.

Claims

1. A method for identifying a nucleic acid segment that interacts with one or more target nucleic acid segments, the method comprising the following steps: (a) Obtaining a nucleic acid composition containing one or more target nucleic acid segments; (b) Crosslinking the nucleic acid composition; (c) Fragmenting the cross-linked nucleic acid composition using an endonuclease; (d) Fill the ends of the fragmented cross-linked nucleic acid segment with one or more nucleotides containing covalently linked biotin moieties; (e) Connect the fragmented nucleic acid segment obtained from step (d) to produce a ligated fragment; (f) The ligated fragment was fragmented in a single step and oligonucleotides were inserted using a recombinase; (g) Enrichment of fragments containing the biotin moiety from step (d); (h) Enrichment of fragments containing one or more target nucleic acid segments; (i) Sequencing the enriched fragments obtained in step (h) to identify nucleic acid segments that interact with the one or more target nucleic acid segments.

2. The method of claim 1, wherein step (h) comprises performing targeted amplification to enrich fragments containing the one or more target nucleic acid segments.

3. The method according to claim 1, wherein step (h) comprises: (i) Adding a separation nucleic acid molecule that binds to the one or more target nucleic acid segments, wherein the separation nucleic acid molecule is labeled with the first half of a binding pair; and (ii) By using the second half of the binding pair, a fragment containing one or more target nucleic acid segments that are bound to the nucleic acid molecule for separation is isolated.

4. The method according to any one of claims 1 to 3, wherein step (f) is performed by labeling.

5. The method according to any one of claims 1 to 4, wherein step (e) utilizes intranuclear linkage.

6. The method according to any one of claims 1 to 5, wherein the recombinase is a retroviral integrase.

7. The method according to claim 6, wherein the retroviral integrase is a mutant transposase.

8. The method according to claim 7, wherein the mutant transposase is a highly active Tn5 transposase.

9. The method according to any one of claims 1 to 8, wherein the recombinase comprises a sequencing-compatible end adapter sequence or a fragment thereof.

10. The method of claim 9, wherein the oligonucleotide and / or adaptor sequence comprises a barcode sequence.

11. The method according to claim 9 or claim 10, wherein the oligonucleotide and / or adaptor sequence is selected from: SEQ ID NO:1, SEQ ID NO:2 and / or SEQ ID NO:

3.

12. The method according to any one of claims 3 to 11, wherein the addition of the nucleic acid molecule for separation in step (h) is carried out in the presence of a sequence that prevents the ligated fragment from binding to other ligated fragments by means of the complementarity of the adapter sequence.

13. The method of claim 12, wherein the linker sequence is a blocking sequence.

14. The method according to any one of claims 1 to 13, wherein the one or more target nucleic acid segments are selected from: promoters, silencers, enhancers or insulators.

15. The method according to any one of claims 3 to 14, wherein the nucleic acid molecule for isolation is obtained from a bacterial artificial chromosome (BAC), F sclerotium, or sclerotium.

16. The method according to any one of claims 3 to 15, wherein the nucleic acid molecule used for isolation is RNA.

17. The method according to any one of claims 3 to 16, wherein the first half of the binding pair comprises biotin and the second half of the binding pair comprises streptavidin.

18. The method according to any one of claims 1 to 17, wherein the endonuclease used in step (c) is Hind III or Dpn II.

19. The method according to any one of claims 3 to 18, further comprising step (g) amplifying the enriched fragment containing the biotin moiety.

20. The method according to any one of claims 1 to 19, further comprising amplifying the enriched fragment from step (h) prior to step (i).

21. The method according to any one of claims 2 or 4 to 20, wherein directional amplification or amplification is performed by PCR.

22. The method according to any one of claims 1 to 21, wherein the nucleic acid composition is derived from the nucleus of a mammalian cell.

23. The method of claim 22, wherein the mammalian cell nucleus is a human cell nucleus.

24. The method according to any one of claims 1 to 23, wherein the nucleic acid composition is derived from a non-human cell nucleus.

25. The method of claim 24, wherein the non-human cell nucleus is a mouse cell nucleus or a plant cell nucleus.

26. The method according to any one of claims 1 to 24, wherein the nucleic acid composition is derived from 10,000, 50,000, 200,000, 500,000, or 1 million cells.

27. A method for identifying one or more interacting nucleic acid segments, comprising: (a) Performing the method of any one of claims 1 to 26 on a nucleic acid composition obtained from an individual; and (b) Quantify the frequency of interactions between nucleic acid segments and one or more target nucleic acid segments.

28. A kit for identifying nucleic acid segments that interact with one or more target nucleic acid segments, comprising buffers and reagents capable of performing the method according to any one of claims 1 to 26.

29. The kit according to claim 28, wherein the recombinase is a retroviral integrase or transposase.

30. The kit according to claim 29, wherein the transposase is a mutant transposase.

31. The kit according to claim 30, wherein the mutant transposase is a highly active Tn5 transposase.

Citation Information

Patent Citations

  • Fast hybridization for next generation sequencing target enrichment

    US9587268B2

  • Selection of nucleic acids by solution hybridization to oligonucleotide baits

    WO2009099602A1

  • Method of identifing interactions between genomic loci

    WO2010036323A1

  • Chromosome conformation capture method including selection and enrichment steps

    WO2015033134A1

  • Chromosome conformation capture method including selection and enrichment steps

    CN105658813A