A Hi-C library construction method

By using special sequencing adapters to insert reconnected fragments in Hi-C technology, the problems of long library construction cycle, large sample volume and noisy data are solved, realizing efficient and low-cost Hi-C library construction and data analysis.

CN117887809BActive Publication Date: 2025-11-28THE PEOPLES HOSPITAL OF GUANGXI ZHUANG AUTONOMOUS REGION
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410031543.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-09
Publication Date
2025-11-28
Estimated Expiration
2044-01-09

AI Technical Summary

Technical Problem

Hi-C technology has problems such as long library preparation time, large sample volume, high proportion of invalid fragments, a lot of noisy data, high sequencing cost and complex data analysis.

Method used

By using special sequencing adapters inserted into the reconnected fragments, only molecules containing the reconnected fragments are detected by the sequencer, eliminating the need for biotin labeling and retrieving the reconnected fragments. By setting the sequencing adapters in the sequencing primer sequences, the proportion of invalid fragments is reduced, simplifying experimental operations and analysis.

Benefits of technology

This approach enables the construction of Hi-C libraries with short preparation time, low cost, and a high proportion of effective data, reducing the requirements for sample and sequencing volume and improving data analysis efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117887809B_ABST
    Figure CN117887809B_ABST
Patent Text Reader

Abstract

The application provides a Hi-C library construction method, which comprises the following steps: cross-linking DNA and proteins in sample cells, enzyme cutting the DNA in the cross-linking body, adding a sequencing adaptor, re-connecting the DNA fragments, de-cross-linking, adding a library adaptor, and library amplification. By using the Hi-C library construction method of the application, a Hi-C library with short library construction time, low cost, convenient experimental operation and high effective data proportion can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to a Hi-C library construction method and belongs to the technical field of gene sequencing. BACKGROUND

[0002] Chromatin conformation capture is a molecular biology technique for studying the spatial organization of the three-dimensional structure of chromosomes. It has important significance in revealing the three-dimensional structure of the genome, understanding the mechanism of gene regulation, studying the mechanism of disease occurrence, and promoting drug research. Since 2002, Deker proposed 3C, and with the improvement of research level, 4C, 5C, CHIA-PET, Hi-C, Capture Hi-C and other technologies have been gradually derived. These methods have revealed the three-dimensional structure and interaction of chromosomes to different extents, providing an important tool for studying gene regulation, genome stability and disease occurrence.

[0003] Among these methods, the Hi-C technology is the most widely used and has derived in situ Hi-C, micro-C, pore-C and other different branches. Hi-C is developed based on the principle of 3C. It cross-links chromatin, cuts DNA and connects the interacting DNA fragments, uses biotin labeling and retrieves the reconnected fragments, and then performs high-throughput sequencing on the reconnected fragments to obtain sequencing data for studying the three-dimensional structure of chromosomes. The data generated by Hi-C sequencing can provide a chromosome interaction map in the whole genome range and be used to analyze the spatial organization and structure of chromatin.

[0004] However, since the Hi-C technology usually needs to retrieve the reconnected fragments by biotin, there are usually problems of long library construction period and large sample amount. Moreover, the fragments that are added with biotin but not reconnected will be retrieved together with the reconnected fragments by biotin and detected in the subsequent sequencing process. The unreconnected fragments are invalid fragments irrelevant to the analysis of the spatial organization and structure of chromatin and will become noise data such as dangling end after being detected. Hi-C studies the chromatin conformation in the whole genome range, and the technical difficulty lies in how to analyze and process massive data, which usually requires a large amount of samples and sequencing amount to achieve satisfactory detection results, and the sequencing cost is extremely expensive. Therefore, the proportion of invalid fragments and noise data has been an important factor affecting the detection cost and quality of the Hi-C technology. In addition, the Hi-C technology detects the DNA fragments after reconnection of chromosomes, and the reconnection position is not determined. When analyzing the sequencing results, the reconnection site needs to be found first, otherwise the sequence cannot be aligned with the reference genome, making the analysis and processing complicated and time-consuming. SUMMARY

[0005] In view of the deficiencies in the prior art described above, the purpose of the present application is to provide a Hi-C library construction method with less time, low proportion of invalid fragments, and convenient experimental operation and analysis.

[0006] The present application is the result of in-depth research to solve the above problems. It has been found that by improving the Hi-C technology, using a special sequencing adapter, especially a sequencing adapter provided with a sequencing primer sequence, and inserting the sequencing adapter into the reconnected fragment, only the molecules containing the reconnected fragment will be detected by the sequencer, obtaining a library with a lower proportion of invalid fragments. The library construction method of the present application can reduce the requirements of the library for the starting amount of the sample and the sequencing amount, and also eliminates the steps of biotin labeling and retrieving the reconnected fragments, thereby reducing the Hi-C library construction cost and detection cost as a whole, thereby completing the present application. By using the Hi-C library construction method of the present application, a Hi-C library with short library construction time, low cost, and high proportion of valid data can be achieved.

[0007] That is, the present application comprises,

[0008] 1. A Hi-C library construction method, comprising the following steps:

[0009] Cross-linking: cross-linking the DNA in the sample cells with proteins to obtain a cross-linking body;

[0010] Enzymatic digestion: enzymatically digesting the DNA in the cross-linking body to obtain cross-linked DNA fragments;

[0011] Sequencing adapter addition: adding a sequencing adapter to the cross-linked DNA fragments to obtain DNA fragments with added sequencing adapters, wherein the sequencing adapter comprises at least one sequencing primer sequence;

[0012] Reconnection of DNA fragments: reconnecting the DNA fragments with added sequencing adapters to obtain cross-linked reconnected DNA fragments.

[0013] Decross-linking: decross-linking the cross-linked reconnected DNA fragments from the proteins to obtain reconnected DNA fragments;

[0014] Library adapter addition, wherein the library adapter comprises a first library structure sequence and a second library structure sequence;

[0015] Library amplification: amplifying the DNA fragments with added library adapters using primers with the first library structure sequence and primers with the second library structure sequence to obtain a Hi-C library.

[0016] 2. The Hi-C library construction method according to item 1, wherein the DNA fragments with added sequencing adapters are reconnected by the added sequencing adapters, and the reconnected DNA fragments internally comprise two sequencing adapters.

[0017] 3. The Hi-C library construction method of item 1, wherein the length of the sequencing adaptor is 10-40 bp, preferably 10-30 bp, more preferably 15-25 bp.

[0018] 4. The Hi-C library construction method of item 1, wherein the sequencing adaptor is composed of two DNA single strands, the sequencing primer sequence is set on the single strand connected with the cross-linked DNA fragment at the 3' end, and the sequencing primer sequence is set at the 3' end of the single strand where it is located.

[0019] 5. The Hi-C library construction method of item 4, wherein the 5' end of the single strand where the sequencing primer sequence is set is a protruding nucleotide sequence of SEQ ID NO: 1, or the 3' end of the single strand where the sequencing primer sequence is not set is a protruding nucleotide sequence of SEQ ID NO: 1.

[0020] SEQ ID NO: 1: 5'-N 1… N m N' 1… N' m -3', wherein N is any one of A, T, C, G deoxyribonucleotide, N 1… N m and N' m… N'1is the reverse complementary sequence, and m is 1-4, preferably 1-3, more preferably 1-2, and most preferably 1.

[0021] 6. The Hi-C library construction method of item 1, wherein end repair and A-tailing are performed on the cross-linked DNA fragments after the enzyme digestion step.

[0022] 7. The Hi-C library construction method of item 6, wherein the 3' end of the single strand where the sequencing primer sequence is set is a protruding T.

[0023] 8. The Hi-C library construction method of item 7, wherein the sequencing adaptor is composed of the nucleotide sequences of SEQ ID NO: 2 and SEQ ID NO: 3,

[0024] SEQ ID NO: 2: 5'-CGTGTGCTGTGACTGGAGT-3',

[0025] SEQ ID NO: 3: 5'-CTCCAGTCACAGCACA-3'.

[0026] 9. The Hi-C library construction method of item 1, wherein fragmentation and A-tailing are performed on the reconnected DNA fragments after the decross-linking step.

[0027] 10. The Hi-C library construction method of item 1, wherein the library adapter consists of the nucleotide sequences of SEQ ID NO: 4 and SEQ ID NO: 5.

[0028] SEQ ID NO: 4: 5'-GTGAAGATCTCGTATGCCGTCTTCTGCTTG-3',

[0029] SEQ ID NO: 5: 5'-AATGATACGGCGACCACCGAGATCTACAC-NN

[0030] NNNN-CTTCACT-3'.

[0031] 11. The Hi-C library construction method of item 1, wherein the reagent system used in the step of adding sequencing adapters is: 1x NEB T4 DNA ligase buffer, 400,000 U / mL NEB T4 DNA ligase, and 0.1 mM sequencing adapters.

[0032] 12. A Hi-C library obtained by the Hi-C library construction method of any one of items 1-11, wherein the interior of the re-ligated fragments comprises two sequencing adapters, and the sequencing adapters comprise at least one sequencing primer sequence.

[0033] 13. A method for detecting a Hi-C library by sequencing the Hi-C library constructed by the Hi-C library construction method of item 12, wherein the sequencing primer sequence in the interior of the re-ligated fragments is used as the sequencing starting point.

[0034] According to an aspect of the present application, there is provided a Hi-C library construction method, which comprises the steps of: cross-linking DNA and proteins in sample cells, enzyme cutting DNA in the cross-linking product, adding sequencing adapters, re-ligating DNA fragments, de-cross-linking, adding library adapters, and library amplification.

[0035] In the above-mentioned Hi-C library construction method,

[0036] The cross-linking step is to cross-link DNA and proteins in sample cells to obtain a cross-linking product. The cross-linking product can maintain the physical interaction between DNA fragments and associated proteins. Cross-linking reagents that can be used include formaldehyde.

[0037] The enzyme cutting step is to enzyme cut DNA in the cross-linking product to obtain cross-linked DNA fragments. Enzyme cutting reagents that can be used include restriction endonucleases or non-restriction endonucleases.

[0038] The step of adding sequencing adapters is to add sequencing adapters to the cross-linked DNA fragments to obtain DNA fragments with added sequencing adapters. The sequencing adapters comprise at least one sequencing primer sequence.

[0039] The adapter is usually a short nucleotide sequence known for connecting unknown sequencing fragments, which can make the sequencing product associated with the computer system. In the present application, the sequencing adapter is an adapter comprising nucleotides or nucleotide sequences related to the sequencing process. These nucleotides or nucleotide sequences related to the sequencing process can be sequencing primers. The sequencing primer refers to a short nucleotide sequence chain at the start of DNA synthesis in the sequencing process. That is, a nucleotide sequence that can serve as the starting point for sequence detection. For example, the sequencing primer of read1 or read2 provided by the Illumina sequencing platform, etc., which is usually used as the starting point of DNA synthesis, i.e., the starting point of sequence detection, in the Illumina sequencing platform sequencing. In the specific embodiments of the present application, the sequencing primer sequence can be a commercial sequence, for example, the sequencing primer sequence of read1 or read2 of the Illumina sequencing platform, etc., or a customized sequencing primer sequence according to the sequencing needs.

[0040] The step of re-ligating DNA fragments is to re-ligate the cross-linked DNA fragments to obtain cross-linked re-ligated DNA fragments. The DNA fragments cleaved by enzymes are re-bound. DNA fragments in close proximity in three-dimensional space are highly likely to bind together, i.e., DNA fragments that may have interaction relationships are re-ligated, thereby making it possible to detect which fragments have interaction relationships. In the present application, the re-ligated fragments are re-ligated via the DNA fragments with added sequencing adapters to obtain. The obtained re-ligated DNA fragments usually comprise two sequencing adapters inside.

[0041] The step of de-cross-linking is to de-cross-link the cross-linked re-ligated DNA fragments from the protein to obtain re-ligated DNA fragments, i.e., DNA fragments obtained by re-ligation of DNA fragments that may have interaction relationships, i.e., re-ligated DNA fragments. De-cross-linking can use a reagent containing proteinase K.

[0042] The step of adding library adapters is to add library adapters to both ends of the re-ligated DNA fragments. The library adapters comprise a first library structure sequence and a second library structure sequence.

[0043] In the present application, the library linker refers to a nucleotide sequence added to both ends of the DNA fragment to be detected in the library construction process, which can be matched with the sequencer for sequencing. It is an important bridge for the connection between the DNA fragment to be detected and the sequencing carrier, such as a sequencing chip (Flow cell) and the like. Because the DNA obtained from the biological sample itself has a structure that cannot be directly detected on the sequencing platform. In order to meet the requirements of the sequencing platform or the sequencing chip for the structure of the sequencing fragment, it is usually necessary to process the obtained DNA to form a library that meets the requirements of the sequencing platform or the sequencing chip, and then sequence the prepared library on the machine. The linker used in the preparation of the library is called a library linker. Unlike the library linker used in the present application, which only contains a bridge for connecting the DNA fragment to be detected and the sequencing carrier, the library linker commonly used in the prior art contains not only the bridge for connecting the DNA fragment to be detected and the sequencing carrier, but also the sequencing primer sequence for detecting the library insert.

[0044] In the present application, the library structure-related sequence refers to other sequences in the library fragment in addition to the unknown sample fragment and the sequencing primer sequence. These sequences can have different functions. Sequences with the same function can be one or more than two. Different functions can be used, for example, to associate with the sequencing platform or the sequencing chip, or to label the sample tag, etc., but do not include the function of marking the sequencing start site. For example, the P5, P7, index1, index2 sequences provided by the Illumina sequencing platform or the sequences complementary to these sequences, etc. are associated with the Illumina sequencing platform or chip. For example, the P5, P7 sequences or their reverse complementary sequences are usually matched with the fixed sequence provided on the Illumina sequencing chip, which is the commonly used associated sequence when the platform is sequenced. The index sequence can be used to distinguish different samples and achieve simultaneous detection of multiple samples. For example, the nucleotide sequence of multiple N in SEQ ID NO: 5. The index can be a nucleotide sequence of 6-10 bp. In the specific embodiments of the present application, the library structure sequence can be a commercial sequence or a customized sequence according to the sequencing needs.

[0045] In the present application, the first and / or second library structure sequence can be all or part of the library structure sequence. Among them, the part of the first and / or second library structure sequence refers to the length of which is a certain proportion of the length of the whole library structure sequence. This proportion can be more than 1 / 3, preferably more than 1 / 2, more preferably more than 3 / 4, more preferably more than 9 / 10, and most preferably 1.

[0046] Library amplification is to amplify the DNA fragments with library adapters using primers with the first library structure sequence and primers with the second library structure sequence to obtain the Hi-C library.

[0047] The reagents used in the library amplification can be selected from one or more of the following: DNA polymerase, RNA polymerase, DNA ligase, RNA ligase, primers, and buffer systems. Alternatively, a commercial kit for library amplification can be used.

[0048] In the present application, the description of "first" or "second" is for the purpose of differentiation, such as different sequencing adapters, library adapters, library structure sequences, or sequencing associated sequences, etc. which are functionally the same or similar, but have some differences in structure or other properties. The characteristics of sequencing adapters, library adapters, library structure sequences, and sequencing associated sequences, etc. have been described above and will not be repeated here.

[0049] Further, in the above-mentioned Hi-C library construction method, the DNA fragments with sequencing adapters are reconnected through the sequencing adapters they carry, and the reconnected DNA fragments internally include two sequencing adapters, i.e. two sequencing adapters are contained between the DNA fragments that can have interaction relationship. The sequencing adapters include a segment of sequencing primer sequence. That is, the inside of the reconnected fragments includes two segments of sequencing primer sequence.

[0050] Further, in the above-mentioned Hi-C library construction method, the sequencing adapter is composed of two DNA single strands, and the sequencing primer sequence is arranged on the single strand connected to the cross-linked DNA fragment at the 3' end, and the sequencing primer sequence is arranged at the 3' end of the single strand where it is located. The length of the sequencing primer sequence is 10-40 bp, preferably 10-30 bp, more preferably 15-25 bp. The ratio of the length of the sequencing primer sequence to the length of the single strand where it is located is 1 / 2 or more, preferably 3 / 4 or more, more preferably 4 / 5 or more. The sequencing primer sequence can be all or part of the sequence of the sequencing primer. Among them, the part of the sequencing primer sequence refers to a certain ratio of its length to the length of the whole sequencing primer. This ratio can be 1 / 3 or more, preferably 1 / 2 or more, more preferably 3 / 4 or more, more preferably 9 / 10 or more, and most preferably 1.

[0051] Further, in the above-mentioned Hi-C library construction method, the preferred length of the sequencing adapter used is 15-25 bp, and within the length range of ±10 bp, preferably ±5 bp, good experimental results can also be obtained.

[0052] Further, in the above-mentioned library construction method, the sequencing primer sequence is arranged at the 3' end of one of the strands of the sequencing adapter. The ratio of the length of the sequencing primer sequence to the length of the sequencing adapter is 1 / 3 or more, preferably 1 / 2 or more, more preferably 3 / 4 or more, and more preferably 9 / 10 or more.

[0053] Further, in the above-mentioned Hi-C library construction method, the DNA fragments can be added with preferred sequencing adapters. Specifically, the 5' end of the single strand of the sequencing primer sequence is set to be a protruding nucleotide sequence represented by SEQ ID NO: 1, or the 3' end of the single strand of the sequencing primer sequence is set to be a protruding nucleotide sequence represented by SEQ ID NO: 1. SEQ ID NO: 1: 5'-N 1… N m N' 1… N' m -3', wherein N is any one of A, T, C, G base deoxyribonucleotide, N 1… N m and N' m… N'1is a reverse complementary sequence, and m is 1-4, preferably m is 3, more preferably m is 2, and most preferably m is 1. The DNA fragments added with the sequencing adapters are connected by the protruding SEQ ID NO: 1 of the sequencing adapters to become reconnected fragments.

[0054] Further, in the above-mentioned Hi-C library construction method, the cross-linked DNA fragments can be subjected to end repair and A-tailing after the enzyme digestion step. End repair refers to repairing the ends of DNA fragments that are damaged or incomplete, such as after enzyme digestion or other breaking treatment, so that the ends can all exist in double-stranded form, i.e. blunt ends, without single-stranded free and nucleotide sequences. A-tailing generally refers to adding an adenine (A) deoxyribonucleotide to the 3' end of each strand of double-stranded DNA fragments that do not have single-stranded free nucleotide sequences at the ends.

[0055] Further, in the above-mentioned Hi-C library construction method, the A-tailed DNA fragments can be added with preferred sequencing adapters. The preferred sequencing adapters have a thymine (T) deoxyribonucleotide protruding at the 3' end of the single strand of the sequencing primer sequence. Such treatment can obtain sequencing adapters with a protruding T at the 3' end. Using 3' end protruding T sequencing adapters can facilitate the connection of the sequencing adapters to DNA fragments with A tails.

[0056] Further, in the above-mentioned Hi-C library construction method, the A-tailed DNA fragments are added with more preferred sequencing adapters. The sequencing adapters can use the product of annealing of the sequences represented by SEQ ID NO: 2 and SEQ ID NO: 3,

[0057] Further, in the above-mentioned Hi-C library construction method, the A-tailed DNA fragments are added with more preferred sequencing adapters. The sequencing adapters can use the product of annealing of the sequences represented by SEQ ID NO: 2 and SEQ ID NO: 3,

[0058] SEQ ID NO: 2: 5'-CGTGTGCTGTGACTGGAGT-3'

[0059] SEQ ID NO: 3: 5'-CTCCAGTCACAGCACA-3'.

[0060] In the present application, the sequencing primer sequence can adopt one of the sequencing primers provided by the Illumina sequencing platform, such as the sequencing primer sequence of read or read2. Other customized sequences can also be selected according to the particularity of the sample or platform.

[0061] Further, in the above-mentioned Hi-C library construction method, the re-ligated DNA fragments are fragmented and A-tailed after the de-crosslinking step. In library construction, fragmentation generally refers to fragmenting longer DNA fragments into library insert sizes suitable for a sequencer by methods such as enzyme digestion, preferably fragmenting into an average fragment length of 200-1000 bp, more preferably fragmenting into an average fragment length of 300-800 bp, and most preferably fragmenting into an average fragment length of 400-700 bp. The methods that can be used for fragmentation include, but are not limited to, non-limiting endonuclease digestion, or ultrasonic fragmentation. The A-tailing is described above and will not be repeated here.

[0062] Further, in the above-mentioned Hi-C library construction method, the library adapter is composed of the sequences shown in SEQ ID NO: 4 and SEQ ID NO: 5,

[0063] SEQ ID NO: 4: 5'-GTGAAGATCTCGTATGCCGTCTTCTGCTTG-3',

[0064] SEQ ID NO: 5: 5'-AATGATACGGCGACCACCGAGATCTACAC-NNNNNNNNNN-CTTCACT-3'.

[0065] Further, in the above-mentioned Hi-C library construction method, the ligation of the sequencing adapter can use reagents such as T4 DNA ligase, T4 DNA ligase buffer, etc. The ligation of the sequencing adapter can use the reagent system: 1x NEB T4 DNA ligase buffer, 400,000 U / mL NEB T4 DNA ligase, and 0.1 mM sequencing adapter.

[0066] Further, in the above-mentioned Hi-C library construction method, the ligation library adapter can use reagents such as T4 DNA ligase, T4 DNA ligase buffer, etc. The ligation library adapter can use reagent system: 10x NEB T4 DNA ligase buffer, 400,000 U / mL NEB T4 DNA ligase, and 0.1 mM library adapter.

[0067] Further, in the above-mentioned Hi-C library construction method, the library amplification can use PCR reagents such as NEB Q5 high-fidelity polymerase, KAPA HiFi Hot Start Ready Mix.

[0068] Further, in the above-mentioned Hi-C library construction method, a purification step is performed after the decrosslinking step. The purification step can use reagents such as SPRIselect beads, Ampure XP beads, Dynabeads.

[0069] Alternatively, in the above-mentioned Hi-C library construction method, a purification step is performed after the fragmentation step, end repair step, or A addition step. The purification step can use reagents as described above, which will not be repeated here.

[0070] According to another aspect of the present application, a Hi-C library is also provided, which is obtained by the above-mentioned library construction method. The re-ligated fragment contains two sequencing adapters in its interior. The two sequencing adapters are connected by the reverse complementary pairing of SEQ ID NO: 1 and inserted between two DNA fragments with spatial interactions. That is, the re-ligated fragment contains two sequencing adapters in its interior.

[0071] According to another aspect of the present application, a detection method of the Hi-C library is also provided, which is by sequencing the above-mentioned Hi-C library, taking the sequencing primer sequence in the interior of the re-ligated fragment as the sequencing starting point. That is, the direction of sequencing is from the interior of the re-ligated fragment to the exterior of the re-ligated fragment. That is, from the two sequencing primer sequences in the interior of the re-ligated fragment, respectively, to the two ends of the re-ligated fragment, rather than reading the sequence from the two ends of the re-ligated fragment to the interior of the re-ligated fragment. For example, the sequencing primer takes the two sequencing primer sequences in the interior of the re-ligated fragment as the starting point, respectively, reads the base sequence of the re-ligated fragment to the two ends of the re-ligated fragment.

[0072] In the present application, the operation of the sequencing adaptor can make the cross-linked DNA fragments connected with the sequencing adaptor with special sequences, and the subsequent re-ligation step can make the DNA fragments cross-linked by associated proteins connected through the sequencing adaptor. The sequencing adaptor becomes the bridge for the ligation of the DNA fragments. Such connection can make the re-ligated DNA fragments inserted with two sequencing adaptors at the re-ligation position. Only the re-ligated fragments containing two sequencing adaptors can be detected by the sequencer in double-end sequencing, thereby reducing the proportion of invalid fragments entering the library and sequencing. At the same time, because the sequencing starting point (such as sequencing primer) is located between the two re-ligated DNA fragments, when double-end sequencing is performed, only one of the two re-ligated DNA fragments will be read at each end, which can be directly aligned when compared with the reference genome sequence, saving analysis time. The reading direction of the traditional Hi-C library sequencing is from the two ends of the re-ligated fragment to the inside of the re-ligated fragment, and one end of the sequencing result may contain the connection point of the re-ligated DNA fragment. Such reads need to find the connection point and remove the sequence at the 3' end of the connection point before aligning to the reference genome, which is time-consuming and resource-consuming.

[0073] The method of the present application can make the method more convenient by connecting special sequencing adaptors to the DNA fragments before the fragment re-ligation process and other multi-step optimization, and the proportion of valid fragments in the library is greatly improved. At the same time, compared with the existing Hi-C method, the method of the present application can reduce the cost of library construction and shorten the library construction time. That is, compared with the prior art, the present application realizes a Hi-C library construction method with less time, low cost, high proportion of effective data, and convenient experimental operation and analysis. BRIEF DESCRIPTION OF DRAWINGS

[0074] Figure 1 is a schematic diagram of the library fragment structure of Example 1.

[0075] Among them, 1 is the interaction DNA fragment 1, 2 is the interaction DNA fragment 2, 3 is the sequencing adaptor, 4 is the first library structure sequence and its complementary sequence, 5 is the second library structure sequence and its complementary sequence, 6 is the library Index and its complementary sequence, and 7 is the arrow expression of the sequencing direction.

[0076] DETAILED DESCRIPTION OF THE INVENTION

[0077] EXAMPLE

[0078] The present application will be further described in detail in conjunction with the drawings and examples. It should be understood that the specific examples described herein are for the purpose of explaining the present application, and are not a limitation of the present application.

[0079] Example 1: Hi-C library construction method of the method of the present application

[0080] (I) Constructing Hi-C library

[0081] 1. Chromatin cross-linking

[0082] 100 million E14 mouse macrophage cell strain RAW264.7 cells (sample 1) were collected, centrifuged at 400G for 5 minutes, and the supernatant was removed. The cells were resuspended in 2 mL of PBS, 54 μL of 37% formaldehyde was added, and the final concentration of formaldehyde was 1%. 10 minutes later, pre-cooled glutamic acid was added to a final concentration of 130 mM to terminate cross-linking. Centrifugation at 3000G for 5 minutes, remove the supernatant, and resuspend the cell nuclei with 1 mL of PBS.

[0083] 2. Enzymatic digestion, end repair and adenine nucleotide (A) addition

[0084] 3000G centrifugation for 5 minutes, remove the supernatant. Resuspend the cell nuclei in 48 μL of 1x NEBbuffer DpnII, add 0.5 μL of 50,000 unites / mL of NEB DpnII Enzymne, 0.5 μL of 5,000 U / mL of NEB Tag DNA polymerase and 1 μL of 10 nM of dNTP. Incubate at 37°C for 15 minutes for enzyme digestion, and at 65°C for 30 minutes for end repair and A addition. 3000G centrifugation for 5 minutes, remove the supernatant, and resuspend the cell nuclei with 1 mL of PBS.

[0085] 3. Ligation of sequencing adapters

[0086] 3000G centrifugation for 5 minutes, remove the supernatant. Resuspend the cell nuclei in 42 μL of 1x NEB T4 DNA ligase buffer. Add 5 μL of 400,000 U / mL of NEB T4 DNA ligase, 3 μL of 0.1 mM of sequencing adapter (prepared by gradient annealing of oligonucleotide sequences shown in SEQ ID NO: 2 and SEQ ID NO: 3). Incubate at 22°C for 1 hour. 3000G centrifugation for 5 minutes, remove the supernatant, and resuspend the cell nuclei with 1 mL of PBS.

[0087] In this example, the sequencing primer sequence is SEQ ID NO: 8, which is set in SEQ ID NO: 2. That is, in this example, SEQ ID NO: 1 is designed as CG, located at the 5' end of SEQ ID NO: 2.

[0088] 4. DNA fragment re-ligation

[0089] Prepare the following reaction solution: 10 μΐ of 10x NEB T4 DNA ligase buffer, 2.5 μΐ of 100 mM EGTA, 2 μΐ of 10,000 Units / mL NEB T4 Polynucleotide Kinase, 10 μΐ of 400,000 U / mL NEB T4 DNA ligase, and water to make up to 100 μΐ. Centrifuge the product of Step 3 at 3000G for 5 minutes, and remove the supernatant. Resuspend the nuclei in 100 μΐ of the reaction solution, incubate at 37°C for 20 minutes, and then incubate at 22°C for 30 hours.

[0090] 5. De-crosslinking

[0091] Add 2 μΐ of 120 Units / mL NEB Thermolabile Proteinase K, incubate at 37°C for 15 minutes for de-crosslinking, and incubate at 55°C for 10 minutes to inactivate the de-crosslinking enzyme.

[0092] Purify the DNA using 180 μΐ of SPRI select beads, and recover in 44 μΐ of Elution buffer.

[0093] 6. Fragment the DNA, end repair, and add A.

[0094] To the product of the previous step, add 5 μΐ of 10x NEB Micrococcal Nuclease Reaction Buffer, 0.5 μΐ of 2,000,000 gel units / mL NEB Micrococcal Nuclease, 0.5 μΐ of 5,000 U / mL NEB Tag DNA Polymerase, and 1 μΐ of 10 nM dNTP, and mix well. Incubate at 32°C for 5 minutes, and then at 65°C for 30 minutes.

[0095] Use 0.5x-0.8x SPRI select beads to fragment and purify the DNA.

[0096] 7. Add library adapters

[0097] Prepare the following adapter ligation reaction solution: add 10 μΐ of 10x NEB T4 DNA ligase buffer, 2 μΐ of 400,000 U / mL NEB T4 DNA ligase, and 1.5 μΐ of 0.1 mM library adapters (prepared by gradient annealing of the oligonucleotide sequences shown in SEQ ID NO: 4 and SEQ ID NO: 5) to the product of the previous step, add water to make up to 100 μΐ, and incubate at 20°C for 15 minutes. The multiple N's are index sequences.

[0098] The DNA was purified using 0.8x SPRI select beads. 23 μL Elution buffer was recovered.

[0099] 8. Hi-C library amplification:

[0100] To the product of the previous step, 1 μL of 0.01 mM primer with the first library structure sequence (oligonucleotide sequence shown in SEQ ID NO: 6), 1 μL of 0.01 mM primer with the second library structure sequence (oligonucleotide sequence shown in SEQ ID NO: 7), and 25 μL KAPA HiFi Hot Start Ready Mix were added.

[0101] The Hi-C library was amplified by PCR using a PCR machine.

[0102] PCR program: 1) 98 degrees Celsius for 3 minutes, 2) 98 degrees Celsius for 30 seconds, 3) 54 degrees Celsius for 30 seconds, 4) 72 degrees Celsius for 30 seconds, 5) 72 degrees Celsius for 1 minute, 6) 10 degrees Celsius stop. Steps 2-4 were repeated 12 times.

[0103] The DNA was size selected and purified using 0.5x-0.8x SPRI select beads. The Hi-C library was obtained.

[0104] The Hi-C library was sequenced and analyzed. Figure 1 As can be seen, the re-ligated fragments are composed of two interacting DNA fragments and two sequencing adapters between them. The two interacting DNA fragments are connected by the CG sequence of the sequencing adapters they carry through reverse complementation.

[0105] This example relates to sequences as shown in Table 1:

[0106] Table 1:

[0107] Name Sequence Remarks SEQ ID NO: 2 5'-CGTGTGCTGTGACTGGAGT-3' SEQ ID NO: 3 5'- [phosphorylated] CTCCAGTCACAGCACA -3' SEQ ID NO: 4 5'- [phosphorylated] GTGAAGATCTCGTATGCCGTCTTCTGCTTG -3' SEQ ID NO: 5 5'-AATGATACGGCGACCACCGAGATCTACAC-NNNNNN-CTTCACT-3'. SEQ ID NO: 6 5'-AATGATACGGCGACCACCGAGATCTACAC-3' SEQ ID NO: 7 5'-CAAGCAGAAGACGGCATACGAGAT-3' SEQ ID NO: 8 5'-CGTGTGCTGTGACTGGAGT-3'

[0108] (B) Library sequencing strategy and analysis strategy

[0109] The library obtained by the above method was sequenced and analyzed. The library was sequenced using HiSeq 2500 in PE150 double-end mode. The sequencing amount was 200M read pairs. In the step of detecting read1 and read2, custom sequencing primers were used, such as the oligonucleotide sequence shown in SEQ ID NO: 8. The library was sequenced as follows: Figure 1 As can be seen, the sequencing primer sequence between the re-ligated fragments was used as the starting point for sequencing to both ends to perform two-step sequencing of read1 and read2, respectively.

[0110] Note that if multiple libraries are detected on one lane, the sequencer can only use Index 2 to detect when using the Forward Strand Workflow for Illumina sequencers (e.g. NovaSeq 6000 with v1.0 reagent kits, MiniSeq with rapid reagent kits, MiSeq, HiSeq 2500, or HiSeq 2000).

[0111] After the data is split, the remaining Illumina library adapter sequence and sequencing primer sequence are trimmed off, and the read 1 sequencing data and read 2 sequencing data are respectively reverse complemented. The data is analyzed using the analysis method described in reference [1].

[0112] Comparative Example 1

[0113] According to the processing method described in reference [1], 25 million E14 mouse macrophage strain RAW264.7 cells (sample 2) were subjected to Hi-C library construction and sequencing analysis.

[0114] Statistical analysis of sequencing results of examples and comparative examples

[0115] The statistical data of the sequencing results of the examples and comparative examples are shown in Table 2.

[0116] Table 2:

[0117] Sample 1 Sample 2 Total Read Pairs 208965257 250090034 Unmapped Read Pairs 7473526 36751533 Mapped Read Pairs 183869028 181651094 PCR Dup Read Pairs 10769332 37629807 Dangling End 20035 42640196 valid Read Pairs 173079661 101381091 Percentage of valid data (valid Read Pairs) 82.83 40.54

[0118] According to the above statistical data, it can be seen that the proportion of valid data in the sequencing data obtained by the method of Example 1 is 82.83%. The proportion of valid data in the sequencing data obtained by the method of Comparative Example 1 is only 40.54%. Therefore, the Hi-C library construction method of the present application can greatly improve the proportion of valid data (the ratio of valid read pairs to total reads).

[0119] In addition, the sequencing amount of Example 1 is less than 210M, while the sequencing amount of the comparative example is 250M. Therefore, using the method of Example 1, only a sequencing amount comparable to or less than the comparative example is required to obtain a higher number of valid data. Therefore, under the condition of requiring a comparable amount of valid data, the library construction method of the present application requires less sequencing amount.

[0120] At the same time, by comparing the method of reference [1] with the method of Example 1, it can be seen that the library construction step of the method of the present application is less, the time is shorter, and the operation is simpler.

[0121] Reference [1]: Jon-Matthew Belton, Rachel Patton McCord, Johan Gibcus, Natalia Naumova, Ye Zhan, and Job Dekker. Hi-C: A comprehensive technique to capture the conformation of genomes. Methods. 2012 November; 58(3)

[0122] According to the present application, the cell properties or functions, as well as the functions of chromatin conformation, DNA, RNA and proteins can be studied in combination with other methods for studying DNA, RNA and proteins. These other methods for studying DNA, RNA and proteins include but are not limited to: single cell sequencing, gene chip, QPCR, first generation sequencing, second generation sequencing, third generation sequencing, fourth generation sequencing, gene sequencing, genome sequencing, metagenome sequencing, exon sequencing, intron sequencing, target gene capture sequencing, RNA sequencing, expression profile sequencing, transcriptome sequencing, small RNA transcriptome, microRNA sequencing, metatranscriptome sequencing, LncRNA sequencing, tumor gene sequencing, tumor genome sequencing, Bisulfite methylation sequencing, ChIP-DNA sequencing, MeDIP sequencing, RRBS sequencing, Target-BS sequencing, hmC sequencing.

[0123] It should also be noted that any of the technical features described in the specification as part of a certain technical solution can also be applied to other technical solutions, provided that they can be implemented and do not obviously deviate from the spirit of the present application; and, provided that they can be implemented and do not obviously deviate from the spirit of the present application, the technical features described as part of different technical solutions can also be combined in any way to constitute other technical solutions. The present application also includes technical solutions obtained by combining in the above-mentioned cases, and these technical solutions are equivalent to those described in the specification.

[0124] The above description shows and describes the preferred embodiments of the present application, as previously described, it is understood that the present application is not limited to the forms disclosed herein, but is to be accorded the full scope of the application as defined in the appended claims, and equivalents thereof.

Claims

1. A Hi-C library construction method, comprising the following steps Cross-linking: cross-linking DNA in sample cells with proteins to obtain cross-linking products; Enzymatic digestion: enzymatically digesting DNA in the cross-linking products to obtain cross-linked DNA fragments; Adding sequencing adapters: adding sequencing adapters to the cross-linked DNA fragments to obtain DNA fragments with added sequencing adapters, wherein the sequencing adapters comprise at least one sequencing primer sequence; Re-ligating DNA fragments: re-ligating the DNA fragments with added sequencing adapters, wherein the DNA fragments with added sequencing adapters are re-ligated through the sequencing adapters carried thereby to obtain cross-linked re-ligated DNA fragments, and the re-ligated DNA fragments comprise two sequencing adapters internally; De-cross-linking: de-cross-linking the cross-linked re-ligated DNA fragments from proteins to obtain re-ligated DNA fragments; Adding library adapters, wherein the library adapters comprise a first library structure sequence and a second library structure sequence; Library amplification: amplifying the DNA fragments with added library adapters using primers with the first library structure sequence and primers with the second library structure sequence to obtain a Hi-C library, wherein the sequencing adapters are composed of two DNA single strands, the sequencing primer sequence is arranged on the single strand connected to the cross-linked DNA fragments at the 3' end, and the sequencing primer sequence is arranged at the 3' end of the single strand on which it is arranged, the 5' end of the single strand on which the sequencing primer sequence is arranged is a protruding nucleotide sequence of SEQ ID NO: 1, or the 3' end of the single strand on which no sequencing primer sequence is arranged is a protruding nucleotide sequence of SEQ ID NO: 1, SEQ ID NO: 1: 5'-N 1… N m N' 1… N' m -3', wherein N is any one of A, T, C, G deoxyribonucleotide, N 1… N m and N' m… N'1is the reverse complement, m is 1-4, the sequencing adapters are composed of the oligonucleotide sequences of SEQ ID NO: 2 and SEQ ID NO:

3.

2. The Hi-C library construction method of claim 1, wherein, End repair and A-tailing are performed on the cross-linked DNA fragments after the enzymatic digestion step.

3. The Hi-C library construction method of claim 2, wherein, The 3' end of the single strand on which the sequencing primer sequence is arranged is a protruding T.

4. The Hi-C library construction method of claim 1, wherein, Fragmentation and A-tailing are performed on the re-ligated DNA fragments after the de-cross-linking step. 5.A Hi-C library obtained by the Hi-C library construction method of any one of claims 1-4, wherein the re-ligated fragments comprise two sequencing adapters internally. 6.A detection method of a Hi-C library, which comprises sequencing the Hi-C library of claim 5, and taking the sequencing primer sequence in the re-ligated fragments as a sequencing starting point.

Citation Information

Patent Citations

  • DLO Hi-C (Digestion-Ligation-Only Hi-C) chromosome conformation capture method

    CN106480178A