A PCR-free Hi-C library construction method
By using PCR free method and sequencing connector technology in Hi-C library construction, the problems of long library construction period, large sample usage and PCR preference in Hi-C technology are solved, and efficient and low-cost Hi-C library construction and data analysis are achieved.
Patent Information
- Application Number
- CN202410031539.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-09
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-01-09
AI Technical Summary
The existing Hi-C technology has problems such as long library construction period, large sample usage, high proportion of invalid fragments, and high detection costs and data noise caused by PCR preferences.
Using PCR free Hi-C library construction method, by inserting the sequencing linker into the reconnected DNA fragment, only molecules containing reconnected fragments are sequenced, reducing invalid fragments, and eliminating PCR amplification and biotin labeling steps.
It realizes the construction of Hi-C library with short library construction time, low cost, high effective data ratio and no PCR preference, reducing detection costs and improving data quality.
Smart Images

Figure CN117845338B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a PCR free Hi-C library construction method, and belongs to the technical field of gene sequencing. Background Art
[0002] Chromatin conformation capture is a molecular biology technique used to study the spatial organization of the three-dimensional structure of chromosomes. It is of great significance in revealing the three-dimensional structure of the genome, understanding gene regulation mechanisms, studying disease mechanisms, and promoting drug development. Since Deker proposed 3C in 2002, 4C, 5C, CHIA-PET, Hi-C, Capture Hi-C and other technologies have gradually emerged with the improvement of research level. These methods have revealed the three-dimensional structure and interaction of chromosomes to varying degrees, providing important tools for studying gene regulation, genome stability and disease occurrence.
[0003] Among these methods, Hi-C technology is the most widely used, and has derived different branches such as in situ Hi-C, micro-C, pore-C, etc. Hi-C is developed based on the 3C principle. It cross-links chromatin, cuts DNA and connects the interacting DNA fragments, uses biotin to label and retrieve the reconnected fragments, and then performs high-throughput sequencing on the reconnected fragments to obtain sequencing data for studying the internal three-dimensional structure of chromosomes. The data generated by Hi-C sequencing can provide a genome-wide map of chromosome interactions and is used to analyze the spatial organization and structure of chromatin.
[0004] However, since Hi-C technology usually requires biotin to retrieve reconnected fragments, there are usually problems with long library construction cycles and large sample usage. In addition, fragments that have been added with biotin but not reconnected will be retrieved by biotin together with the reconnected fragments and detected in the subsequent sequencing process. Unreconnected fragments are invalid fragments that are irrelevant to the analysis of the spatial organization and structure of chromatin. After being detected, they will become noise data during data analysis, such as dangling ends. Hi-C studies the chromosome conformation within the entire genome. The technical difficulty lies in how to analyze and process massive data. Usually a large amount of samples and sequencing volume are required to achieve satisfactory test results, and the sequencing cost is extremely expensive.
[0005] Therefore, the ratio of invalid fragments and noise data has always been an important factor affecting the detection cost and quality of Hi-C technology. At the same time, Hi-C technology detects DNA fragments after chromosome reconnection, and the reconnection position is uncertain. When analyzing the sequencing results, it is necessary to find the reconnection site first, otherwise it is impossible to compare the sequence with the reference genome, making the analysis and processing complicated and time-consuming.
[0006] In addition, since the traditional Hi-C library construction method requires PCR amplification of the reconnected DNA fragments retrieved by biotin, repeated detected fragments are introduced into the library, wasting sequencing data and reducing the proportion of valid data; at the same time, PCR will also introduce PCR preference into Hi-C detection, affecting the accuracy of the results. Summary of the invention
[0007] In view of the above-mentioned deficiencies in the prior art, the object of the present invention is to provide a Hi-C library construction method which takes less time, has a high ratio of effective data, is convenient for experimental operation and analysis, and has no PCR preference.
[0008] The present invention has conducted in-depth research to solve the above problems, and found that: by improving the Hi-C technology, using special sequencing adapters, especially sequencing adapters with sequencing primer sequences, the sequencing adapters are inserted into the reconnected fragments, so that only molecules containing reconnected fragments will be detected by the sequencer, and a library with a lower proportion of invalid fragments is obtained. At the same time, the present invention adopts a PCR free library construction method, which can avoid the bias and duplication generated by PCR. In addition, the library construction method of the present invention can also omit the steps of biotin labeling and retrieving reconnected fragments, shorten the library construction time, and reduce the Hi-C library construction cost and detection cost as a whole, thereby completing the present invention. By adopting the PCR free Hi-C library construction method of the present invention, a Hi-C library with short library construction time, low cost, high proportion of valid data and no PCR bias can be achieved.
[0009] That is, the present invention includes,
[0010] 1. A method for constructing a PCR-free Hi-C library, the method comprising the following steps:
[0011] Cross-linking: cross-linking the DNA and proteins in the sample cells to obtain a cross-linked product;
[0012] Enzyme digestion: Enzyme digestion of DNA in the cross-linked body to obtain cross-linked DNA fragments;
[0013] Adding sequencing adapters: adding sequencing adapters to the cross-linked DNA fragments to obtain DNA fragments with sequencing adapters, wherein the sequencing adapters include at least one sequencing primer sequence;
[0014] Reconnect DNA fragments: Reconnect the DNA fragments with sequencing adapters to obtain cross-linked reconnected DNA fragments.
[0015] Decrosslinking: Decrosslinking the crosslinked reconnected DNA fragments with proteins to obtain reconnected DNA fragments;
[0016] Adding library adapters, wherein the library adapters include an adapter with a first library structure sequence and an adapter with a second library structure sequence, to obtain a PCR free Hi-C library.
[0017] 2. According to the PCR free Hi-C library construction method described in item 1, the DNA fragment with added sequencing adapters is reconnected through the added sequencing adapters, and the reconnected DNA fragment includes two sequencing adapters.
[0018] 3. According to the PCR free Hi-C library construction method of item 1, the length of the sequencing adapter is 10-40 bp, preferably 10-30 bp, and more preferably 15-25 bp.
[0019] 4. According to the PCR free Hi-C library construction method described in item 1, the sequencing adapter consists of two DNA single strands, the sequencing primer sequence is set on the single strand connected to the cross-linked DNA fragment at the 3' end, and the sequencing primer sequence is set at the 3' end of the single strand where it is located.
[0020] 5. The method for constructing a PCR free Hi-C library according to item 4, wherein the 5' end of the single strand with the sequencing primer sequence is a protruding nucleotide sequence shown in SEQ ID NO: 1, or the 3' end of the single strand without the sequencing primer sequence is a protruding nucleotide sequence shown in SEQ ID NO: 1,
[0021] SEQ ID NO: 1: 5'-N 1… N m N' 1… N' m -3', wherein N is any one of A, T, C, G deoxyribonucleotides, N 1… N m With N' m… N' 1 is a reverse complementary sequence, m is 1-4, preferably m is 1-3, more preferably m is 1-2, and most preferably m is 1.
[0022] 6. The method for constructing a PCR free Hi-C library according to item 1, wherein the cross-linked DNA fragments are end-repaired and A-added after the enzyme digestion step.
[0023] 7. The method for constructing a PCR free Hi-C library according to item 6, wherein the 3' end of the single strand of the sequencing primer sequence is set to a protruding T.
[0024] 8. The method for constructing a PCR free Hi-C library according to item 7, wherein the sequencing adapter consists of the nucleotide sequences shown in SEQ ID NO: 2 and SEQ ID NO: 3,
[0025] SEQ ID NO: 2: 5'-CGTGTGCTGTGACTGGAGT-3',
[0026] SEQ ID NO: 3: 5'-CTCCAGTCACAGCACA-3'
[0027] 9. According to the PCR free Hi-C library construction method described in item 1, the reconnected DNA fragments were fragmented and A-treated after the cross-linking removal step.
[0028] 10. According to the PCR free Hi-C library construction method described in item 1, the linker with the first library structure sequence consists of the nucleotide sequences shown in SEQ ID NO: 4 and SEQID NO: 5, and the linker with the second library structure sequence consists of the nucleotide sequences shown in SEQ ID NO: 6 and SEQID NO: 7.
[0029] SEQ ID NO: 4: 5'GGTGTAGATCTCGGTGGTCGCCGTATCATT-3'
[0030] SEQ ID NO: 5: 5'-AATGATACGGCGACCACCGAGATCTACACT-3'
[0031] SEQ ID NO: 6: 5'-TCTCGTATGCCGTCTTCTGCTTG-3'
[0032] SEQ ID NO: 7: 5'-CAAGCAGAAGACGGCATACGAGAT-3'
[0033] 11. The PCR free Hi-C library construction method according to item 1, wherein the reagent system used in the step of adding sequencing adapters is: 1xNEB T4 DNA ligase buffer, 400,000 U / mL NEBT4 DNA ligase, and 0.1 mM sequencing adapter.
[0034] 12. A PCR free Hi-C library, obtained by the PCR free Hi-C library construction method according to any one of items 1 to 11, wherein the reconnected fragments include two sequencing adapters, and the sequencing adapters include at least one sequencing primer sequence.
[0035] 13. A method for detecting a PCR free Hi-C library, comprising sequencing the PCR free Hi-C library constructed by the PCR free Hi-C library construction method described in item 12, wherein the sequencing primer sequence inside the reconnected fragment is used as the sequencing start point.
[0036] According to one aspect of the present invention, a PCR free Hi-C library construction method is provided, the method comprising the steps of cross-linking DNA and protein in sample cells, enzymatically digesting DNA in the cross-linked body, adding sequencing adapters, reconnecting DNA fragments, resolving cross-links, and adding library adapters.
[0037] In the above PCR free Hi-C library construction method,
[0038] The cross-linking step is to cross-link the DNA and proteins in the sample cells to obtain a cross-linked body. The cross-linked body can maintain the physical interaction between the DNA fragments and the related proteins. The cross-linking reagent that can be used is, for example, formaldehyde.
[0039] The step of enzyme digestion is to digest the DNA in the cross-linked body to obtain cross-linked DNA fragments. The enzyme digestion reagent can be a restriction endonuclease or a non-restriction endonuclease.
[0040] Adding a sequencing adapter is to add a sequencing adapter to the cross-linked DNA fragment to obtain a DNA fragment with a sequencing adapter. The sequencing adapter includes at least one sequencing primer sequence.
[0041] The connector is usually a known short nucleotide sequence, which is used to connect unknown sequencing fragments, and can establish a connection between the sequencing product and the computer system. In the present invention, the sequencing connector is a connector including nucleotides or nucleotide sequences related to the sequencing process. These nucleotides or nucleotide sequences related to the sequencing process can be sequencing primers. Sequencing primers refer to the short nucleotide sequence chain at the starting point of DNA synthesis during the sequencing process. That is, a nucleotide sequence that can be used as the starting point of sequence detection. For example, the sequencing primers of read1 or read2 provided by the Illumina sequencing platform, etc., are usually used as the starting point of DNA synthesis when the Illumina sequencing platform is sequenced, that is, the starting point of sequencing detection. In a specific embodiment of the present invention, the sequencing primer sequence can be a commercial sequence, for example, the sequencing primer sequence of read1 or read2 of the Illumina sequencing platform, etc., or it can be a sequencing primer sequence customized according to sequencing needs.
[0042] The step of reconnecting DNA fragments is to reconnect the cross-linked DNA fragments to obtain cross-linked reconnected DNA fragments. The DNA fragments that have been cut by enzymes are rebound. DNA fragments that are adjacent in three-dimensional space are combined with a high probability, that is, DNA fragments that may have an interactive relationship are reconnected, thereby making it possible to detect which fragments have an interactive relationship. In the present invention, the reconnected fragments are obtained by reconnecting the DNA fragments with the added sequencing adapters through the added sequencing adapters. The obtained reconnected DNA fragments usually include two sequencing adapters.
[0043] The cross-linking removal step is to remove the cross-links between the cross-linked reconnected DNA fragments and proteins to obtain reconnected DNA fragments, that is, DNA fragments obtained by reconnecting DNA fragments that may have an interactive relationship, that is, reconnected DNA fragments. The cross-linking removal can be performed using a reagent containing proteinase K.
[0044] The step of adding library adapters is to add library adapters to both ends of the reconnected DNA fragments. The library adapters include adapters with a first library structure sequence and adapters with a second library structure sequence.
[0045] In the present invention, the library connector refers to a nucleotide sequence that can be sequenced with a sequencer at both ends of the DNA fragment to be tested during the library construction process. It is an important bridge connecting the DNA fragment to be tested and the sequencing vector, such as a sequencing chip (Flow cell). Because the DNA obtained from the biological sample has a structure that itself may not be directly detected on the sequencing platform. In order to meet the requirements of the sequencing platform or sequencing chip for the structure of the sequencing fragment, it is usually necessary to process the obtained DNA to form a library that meets the requirements of the sequencing platform or sequencing chip, and then the prepared library is sequenced on the machine. The connector used in the preparation of the library is called a library connector. Unlike the library connector used in the present invention, which only contains a bridge connecting the DNA fragment to be tested and the sequencing vector, the library connector commonly used in the prior art includes both a bridge connecting the DNA fragment to be tested and the sequencing vector, and also includes a sequencing primer sequence for detecting the inserted fragment of the library.
[0046] In the present invention, the library structure-related sequence refers to other sequences in the library fragment excluding the unknown sample fragment and the sequencing primer sequence. These sequences may have different functions. And the sequences with the same function may be one or more than two. Different functions may be used, for example, to associate with a sequencing platform or a sequencing chip, or to mark a label of a sample, etc., but do not include the function of marking the start site of sequencing. For example, the P5, P7, index1, index2 sequences provided by the Illumina sequencing platform, or sequences that are reverse complementary to these sequences, etc., are associated with the Illumina sequencing platform or chip. For example, the P5, P7 sequences, or their reverse complementary sequences are usually matched with the fixed sequences set on the Illumina sequencing chip, and are commonly used associated sequences for sequencing on this platform. The index sequence can be used to distinguish different samples and realize the simultaneous detection of multiple samples. In a specific embodiment of the present invention, the library structure sequence may be a commercial sequence, or a sequence customized according to sequencing needs.
[0047] In the present invention, the first and / or second library structure sequence may be the entire or partial sequence of the library structure sequence. The partial first and / or second library structure sequence refers to a length that is in a certain ratio to the length of the entire library structure sequence. This ratio may be 1 / 3 or more, preferably 1 / 2 or more, more preferably 3 / 4 or more, more preferably 9 / 10 or more, and most preferably 1.
[0048] In the present invention, the description of "first" or "second" is for distinction, such as different sequencing adapters, library adapters, library structure sequences or sequencing-associated sequences, which are the same or similar in function, but have different features in structure or other properties. The features of sequencing adapters, library adapters, library structure sequences and sequencing-associated sequences are as described above and will not be repeated here.
[0049] Furthermore, in the above-mentioned PCR free Hi-C library construction method, the DNA fragments with sequencing adapters are reconnected through the sequencing adapters they carry, and the reconnected DNA fragments include two sequencing adapters, that is, the DNA fragments that may have an interactive relationship contain two sequencing adapters.
[0050] Further, in the above-mentioned PCR free Hi-C library construction method, the sequencing adapter is composed of two single-stranded DNA chains, the sequencing primer sequence is arranged on the single-stranded chain connected to the cross-linked DNA fragment at the 3' end, and the sequencing primer sequence is arranged at the 3' end of the single-stranded chain in which it is located. The length of the sequencing primer sequence is 10-40bp, preferably 10-30bp, and more preferably 15-25bp. The ratio of the length of the sequencing primer sequence to the length of the single-stranded chain in which it is located is more than 1 / 2, preferably more than 3 / 4, and more preferably more than 4 / 5. The sequencing primer sequence can be the entire or partial sequence of the sequencing primer. Among them, the partial sequencing primer sequence refers to a certain ratio between its length and the length of all sequencing primers. This ratio can be more than 1 / 3, preferably more than 1 / 2, more preferably more than 3 / 4, more preferably more than 9 / 10, and most preferably 1.
[0051] Furthermore, in the above-mentioned PCR free Hi-C library construction method, the preferred length of the sequencing adapter used is 15-25 bp, and within the length range of ±10 bp, preferably ±5 bp, good experimental results can also be obtained.
[0052] Furthermore, in the above-mentioned library construction method, the sequencing primer sequence is set at the 3' end of one of the chains of the sequencing adapter. The ratio of the length of the sequencing primer sequence to the length of the sequencing adapter is 1 / 3 or more, preferably 1 / 2 or more, more preferably 3 / 4 or more, and more preferably 9 / 10 or more.
[0053] Furthermore, in the above-mentioned PCR free Hi-C library construction method, a sequencing adapter can be preferably added to the DNA fragment. Specifically, the 5' end of the single strand with the sequencing primer sequence is a protruding nucleotide sequence shown in SEQ ID NO: 1, or the 3' end of the single strand without the sequencing primer sequence is a protruding nucleotide sequence shown in SEQ ID NO: 1. SEQ ID NO: 1: 5'-N 1… N m N' 1… N' m -3', wherein N is any one of A, T, C, G deoxyribonucleotides, N 1… N m With N' m… N' 1 is a reverse complementary sequence, m is 1-4, preferably m is 3, more preferably m is 2, and most preferably m is 1. The DNA fragments with sequencing adapters are connected through the protruding SEQ ID NO: 1 of the sequencing adapter to form reconnected fragments.
[0054] Furthermore, in the above-mentioned PCR free Hi-C library construction method, the cross-linked DNA fragments can be end-repaired and A-added after the enzyme digestion step. End repair refers to the repair of the ends of damaged or incomplete DNA fragments, such as those that have been treated with enzyme digestion or other interruptions, so that their ends can exist in a double-stranded form, that is, a flat end, without the existence of single-stranded free and nucleotide sequences. A addition generally refers to the addition of an adenine (A) deoxyribonucleotide to the 3' end of each strand of a double-stranded DNA fragment without a single-stranded free nucleotide sequence at the end.
[0055] Furthermore, in the above-mentioned PCR free Hi-C library construction method, a sequencing adapter can be preferably added to the DNA fragment with A. Preferably, the sequencing adapter has a protruding thymine (T) deoxyribonucleotide at the 3' end of the single strand of the sequencing primer sequence. Such treatment can obtain a sequencing adapter with a protruding T at the 3' end. The use of a sequencing adapter with a protruding T at the 3' end can smoothly connect the sequencing adapter to the DNA fragment with an A tail.
[0056] Furthermore, in the above-mentioned PCR free Hi-C library construction method, a sequencing adapter is preferably added to the DNA fragment with A. Preferably, the 3' end of the sequencing adapter at the single strand without the sequencing primer sequence is not A. This can effectively avoid the connection between adapters.
[0057] Furthermore, in the above-mentioned PCR free Hi-C library construction method, a more preferred sequencing adapter is added to the DNA fragment with A added. The sequencing adapter can be a product of annealing treatment of the sequences shown in SEQ ID NO: 2 and SEQ ID NO: 3.
[0058] SEQ ID NO: 2: 5'-CGTGTGCTGTGACTGGAGT-3'
[0059] SEQ ID NO: 3: 5'-CTCCAGTCACAGCACA-3'.
[0060] In the present invention, the sequencing primer sequence can be one of the sequencing primers provided by the Illumina sequencing platform, such as the sequencing primer sequence of read or read2. Other customized sequences can also be selected according to the particularity of the sample or platform.
[0061] Furthermore, in the above-mentioned PCR free Hi-C library construction method, the reconnected DNA fragments are fragmented and A-treated after the cross-linking removal step. In library construction, fragmentation generally refers to fragmenting longer DNA fragments into library insert fragments of a length suitable for a sequencer by methods such as enzyme digestion, preferably fragmenting to an average fragment length of 200-1000bp, more preferably fragmenting to an average fragment length of 300-800bp, and most preferably fragmenting to an average fragment length of 400-700bp. Fragmentation can be performed by methods such as non-restriction endonuclease digestion, or ultrasonic shearing. The description of A-treatment is as described above and will not be repeated here.
[0062] Furthermore, in the above-mentioned PCR free Hi-C library construction method, the linker with the first library structure sequence consists of the nucleotide sequences shown in SEQ ID NO: 4 and SEQ ID NO: 5, and the linker with the second library structure sequence consists of the nucleotide sequences shown in SEQ ID NO: 6 and SEQ ID NO: 7.
[0063] SEQ ID NO: 4: 5'-GTGTAGATCTCGGTGGTCGCCGTATCATT-3'
[0064] SEQ ID NO: 5: 5'-AATGATACGGCGACCACCGAGATCTACACT-3'
[0065] SEQ ID NO: 6: 5'-TCTCGTATGCCGTCTTCTGCTTG-3'
[0066] SEQ ID NO: 7: 5'-CAAGCAGAAGACGGCATACGAGAT-3'
[0067] Furthermore, in the above-mentioned PCR free Hi-C library construction method, reagents such as T4 DNA ligase, T4 DNA ligase buffer, etc. can be used to connect sequencing adapters. The reagent system for connecting sequencing adapters can be: 1x NEB T4DNAligase buffer, 400,000 U / mL NEB T4 DNA ligase, and 0.1mM sequencing adapter.
[0068] Furthermore, in the above PCR free Hi-C library construction method, the library adapters can be connected using reagents such as T4 DNA ligase, T4 DNA ligase buffer, etc. The library adapters can be connected using a reagent system: 10x NEB T4DNAligase buffer, 400,000 U / mL NEB T4 DNA ligase, and 0.1 mM library adapter.
[0069] Furthermore, in the above-mentioned PCR free Hi-C library construction method, a purification step is performed after the cross-linking removal step. The purification step can use reagents such as SPRI select beads, Ampure XP beads, and Dynabeads.
[0070] Optionally, in the above-mentioned PCR free Hi-C library construction method, a purification step is performed after the fragmentation step, the final trimming step or the A-addition step. The reagents that can be used in the purification step are as described above and will not be described in detail here.
[0071] According to another aspect of the present invention, a PCR free Hi-C library is also provided, which is obtained by the above-mentioned library construction method. The interior of the reconnected fragment includes two sequencing adapters. The sequencing adapter includes a sequencing primer sequence. The two sequencing adapters are inserted between two DNA fragments with spatial interaction through reverse complementary pairing of SEQ ID NO: 1. That is, the interior of the reconnected fragment includes two sequencing primer sequences.
[0072] According to another aspect of the present invention, a method for detecting a PCR free Hi-C library is also provided, which is performed by sequencing the above-mentioned PCR free Hi-C library, and the sequencing primer sequence inside the reconnected fragment is used as the sequencing starting point. In other words, the direction of sequencing is to detect from the inside of the reconnected fragment to the outside of the reconnected fragment. That is, the two sequencing primer sequences inside the reconnected fragment are respectively detected to the two ends of the reconnected fragment, rather than reading the sequence from the two ends of the reconnected fragment to the inside of the reconnected fragment. For example, the sequencing primer uses the two sequencing primer sequences inside the reconnected fragment as the starting point to read the base sequence of the reconnected fragment to the two ends of the reconnected fragment.
[0073] In the present invention, the cross-linked DNA fragments can be connected to the sequencing adapters with special sequences through the operation of adding sequencing adapters, and the subsequent reconnection step allows the DNA fragments cross-linked by the associated protein to be connected through the sequencing adapters. The sequencing adapter becomes a bridge for connecting enzyme-cut DNA fragments. Such a connection method can insert two sequencing adapters at the reconnection position of the reconnected DNA fragment. Only the reconnected fragments containing two sequencing adapters can be detected by the sequencer in double-end sequencing, thereby reducing the proportion of invalid fragments entering the library and sequencing. At the same time, because the sequencing start point (such as sequencing primer) is located between the two reconnected DNA fragments, each end will only read one of the two reconnected DNA fragments when performing double-end sequencing, and can be directly compared when aligning with the reference genome sequence, saving analysis time. The reading direction during traditional Hi-C library sequencing is to read from the two segments of the reconnected fragment to the inside of the reconnected fragment, and the sequencing result at one end may contain the connection point of the reconnected DNA fragment. Such reads need to find the connection point and remove the sequence at the 3' end of the connection point before aligning to the reference genome, which takes more analysis time and resources.
[0074] The present invention does not require PCR amplification during the Hi-C library construction process, further improving the proportion of valid library data, and the generated library has no PCR preference.
[0075] The method of the present invention connects special sequencing adapters and other multi-step optimizations to the DNA fragments before the fragment reconnection process, making the method of the present invention more convenient, without the need for biotin retrieval, and greatly improving the proportion of effective fragments in the library. At the same time, compared with the existing Hi-C method, the method of the present invention can reduce the cost of library construction and shorten the library construction time. That is, compared with the prior art, the present invention realizes a PCR free Hi-C library construction method that takes less time, has low cost, has a high proportion of effective data, has no PCR preference, and is easy to operate and analyze experimentally. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] Figure 1 It is a schematic diagram of the library fragment structure of Example 1.
[0077] Among them, 1-interacting DNA fragment 1, 2-interacting DNA fragment 2, 3-sequencing adapter, 4-first library structure sequence and its complementary sequence, 5-second library structure sequence and its complementary sequence, 6-arrow represents the sequencing direction. DETAILED DESCRIPTION Example
[0078] The present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are for explaining the present invention, not for limiting the present invention.
[0079] Example 1 A method for constructing a PCR-free Hi-C library
[0080] 1. Library construction
[0081] 1. Chromatin cross-linking
[0082] Collect 50 million E14 mouse macrophage cell line RAW264.7 cells, centrifuge at 400G for 5 minutes, and remove the supernatant. Resuspend the cells in 2 ml PBS, add 54μL 37% formalin to make the final concentration of formaldehyde 1%. Add pre-cooled glutamate to a final concentration of 130mM for 10 minutes to terminate cross-linking. Centrifuge at 3000G for 5 minutes, remove the supernatant, and resuspend the cell nuclei in 1mL PBS.
[0083] 2. Enzyme digestion, end repair and addition of adenine nucleotides (A)
[0084] Centrifuge at 3000G for 5 minutes and remove the supernatant. Resuspend the nuclei in 48μL 1x NEBbuffer DpnII, add 0.5μL 50,000 units / mL NEB DpnII Enzymne, 0.5μL 5,000 U / mL NEB Tag DNA polymerase and 1μL 10nM dNTP and mix well. Incubate at 37 degrees Celsius for 15 minutes for enzyme digestion and incubate at 65 degrees Celsius for 30 minutes for end repair and A addition. Centrifuge at 3000G for 5 minutes, remove the supernatant and resuspend the nuclei in 1mL PBS.
[0085] 3. Ligation of sequencing adapters
[0086] Centrifuge at 3000G for 5 minutes and remove the supernatant. Resuspend the nuclei in 42μL 1x NEB T4 DNA ligase buffer. Add 5μL 400,000 U / mL NEB T4 DNA ligase and 3μL 0.1mM sequencing adapter (prepared by gradient annealing of the oligonucleotide sequences shown in SEQ ID NO: 2 and SEQID NO: 3). In this example, the sequencing primer sequence is SEQ ID NO: 8, which is set in SEQID NO: 2. That is, in this example, SEQ ID NO: 1 is designed as CG, which is located at the 5' end of SEQ ID NO: 2. Incubate at 22 degrees Celsius for 1 hour. Centrifuge at 3000G for 5 minutes, remove the supernatant, and resuspend the nuclei in 1ml PBS.
[0087] 4. DNA fragment religation
[0088] Prepare the following reaction solution: 10μL 10x NEB T4 DNA ligase buffer, 2.5μL 100mM EGTA, 2μL 10,000Units / mL NEB T4 Polynucleotide Kinase, 10μL 400,000U / mL NEB T4DNA ligase, and add water to make up to 100μL. Centrifuge the product of step 3 at 3000G for 5 minutes and remove the supernatant. Use 100μL of the reaction solution to resuspend the nuclei, incubate at 37 degrees Celsius for 20 minutes, and then incubate at 22 degrees Celsius for 30 hours.
[0089] 5. Decrosslinking
[0090] Add 2 μL 120 Units / mL NEB Thermolabile Proteinase K, incubate at 37°C for 15 minutes to reverse crosslinking, and incubate at 55°C for 10 minutes to inactivate the enzyme.
[0091] The DNA was purified using 180 μL of SPRIselect beads and recovered in 44 μL of Elution buffer.
[0092] 6. Fragment DNA, end repair and A addition.
[0093] Add 5 μL 10x NEB Micrococcal Nuclease Reaction Buffer, 0.5 μL 2,000,000 gel units / ml NEB Micrococcal Nuclease, 0.5 μL 5,000 U / mL NEBTag DNA polymerase and 1 μL 10 nM dNTP to the product from the previous step, mix well, incubate at 32 degrees Celsius for 5 minutes, and then incubate at 65 degrees Celsius for 30 minutes.
[0094] Use 0.5x-0.8x SPRI select beads to screen and purify DNA fragments.
[0095] 7. Ligation of library adapters
[0096] Prepare the following adapter ligation reaction solution: add 10 μL 10x NEB T4 DNA ligase buffer, 2 μL 400,000 U / mL NEB T4 DNA ligase, 0.75 μL 0.1 mM first library adapter (prepared by gradient annealing of oligonucleotide sequences shown in SEQ ID NO: 4 and SEQ ID NO: 5), 0.75 μL 0.1 mM second library adapter (prepared by gradient annealing of oligonucleotide sequences shown in SEQ ID NO: 6 and SEQ ID NO: 7) to the product of the previous step, add water to make up to 100 μL, and incubate at 20 degrees Celsius for 15 minutes.
[0097] Use 0.5x-0.8x SPRI select beads to select DNA fragments and recover them in 20μL Elution buffer.
[0098] Get PCR free Hi-C library.
[0099] Depend on Figure 1 It can be seen that the reconnected fragment consists of two interacting DNA fragments and two sequencing adapters between them. The two interacting DNA fragments are connected by reverse complementation through the CG sequences of the sequencing adapters they carry.
[0100] The sequences involved in this embodiment are shown in Table 1.
[0101] Table 1:
[0102]
[0103] 2. Library sequencing strategy and analysis strategy
[0104] The library obtained by the above method was sequenced and data analyzed. The library was sequenced with PE150 double-end sequencing using Nextseq. The sequencing amount was 200M reads. In the steps of detecting read1 and read2, customized sequencing primers such as the oligonucleotide sequence shown in SEQ ID NO: 8 were used. Figure 1 It can be seen that in this embodiment, the sequencing primer sequence between the reconnected fragments is used as the sequencing starting point to perform two-step sequencing of read1 and read2 towards both ends respectively.
[0105] Note that since there is no library index in this example, only one library can be detected in one lane. If you want to detect multiple libraries in one lane, you can introduce the library index in the library adapter and detect it in the Index 2 step of sequencing. You can only use the Illumina sequencer (such as NovaSeq6000 with v1.0 reagent kits, MiniSeq with rapid reagent kits, MiSeq, HiSeq2500, or HiSeq2000) of the Forward Strand Workflow when using Index 2 detection. You can also introduce the library index in the sequencing adapter and detect it in the Index1 step of sequencing. All Illuminate platforms are available.
[0106] After the data was split, the residual Illumina library adapter sequence and sequencing primer sequence were trimmed. Then the sequencing data of read1 and read2 were reverse complemented. The data were analyzed using the analysis method described in reference [1].
[0107] Comparative Example 1
[0108] Referring to the processing method described in reference [1], Hi-C library construction and sequencing analysis were performed on 25 million E14 mouse macrophage cell line RAW264.7 cells (sample 2).
[0109] Statistics and analysis of sequencing results of examples and comparative examples
[0110] The statistical data of the sequencing results of the embodiments and comparative examples are shown in Table 2.
[0111] Table 2:
[0112]
[0113] According to the above statistical data, the effective data accounted for 89.87% of the sequencing data obtained by the method of Example 1. However, the effective data accounted for only 40.54% of the sequencing data obtained by the method of Comparative Example 1. It can be seen that the PCR free Hi-C library construction method of the present invention can greatly improve the effective data ratio (the ratio of valid Read Pairs to Total reads).
[0114] In addition, the sequencing amount of Example 1 is less than 200M, while the sequencing amount of the comparative example is 250M. Therefore, the method of Example 1 only requires a sequencing amount equivalent to or less than the comparative example to obtain a higher amount of valid data. It can be seen that when the amount of valid data is equivalent, the library construction method of the present invention has a lower requirement for sequencing amount.
[0115] At the same time, by comparing the method of reference [1] with the method of Example 1, it can be seen that the method of the present invention has fewer steps in library construction, takes less time, and is simpler to operate.
[0116] References[1]: Jon-MatthewBelton, Rachel Patton McCord, Johan Gibcus,Natalia Naumova, Ye Zhan, and Job Dekker. Hi-C: A comprehensive technique tocapture the conformation of genomes.Methods. 2012 November; 58(3)
[0117] According to the present invention, other methods for studying DNA, RNA, and proteins can be used in combination to study cell characteristics or functions, as well as chromatin conformation, DNA, RNA, and protein functions. These other methods for studying DNA, RNA, and proteins include, but are not limited to: single cell sequencing, gene chip, QPCR, first generation sequencing, second generation sequencing, third generation sequencing, fourth generation sequencing, gene sequencing, genome sequencing, metagenomic sequencing, exon sequencing, intron sequencing, target gene capture sequencing, RNA sequencing, expression profile sequencing, transcriptome sequencing, small RNA transcriptome, microRNA sequencing, macrotranscriptome sequencing, LncRNA sequencing, tumor gene sequencing, tumor genome sequencing, Bisulfite methylation sequencing, ChIP-DNA sequencing, MeDIP sequencing, RRBS sequencing, Target-BS sequencing, and hmC sequencing.
[0118] It should also be noted that, under the premise that it is feasible and does not obviously violate the main purpose of the present invention, any technical feature or combination of technical features described as a component of a certain technical solution in this specification can also be applied to other technical solutions; and, under the premise that it is feasible and does not obviously violate the main purpose of the present invention, the technical features described as components of different technical solutions can also be combined in any way to form other technical solutions. The present invention also includes technical solutions obtained by combination under the above circumstances, and these technical solutions are equivalent to those recorded in this specification.
[0119] The above description shows and describes the preferred embodiments of the present invention. As mentioned above, it should be understood that the present invention is not limited to the form disclosed herein, and should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications and environments, and can be modified within the scope of the invention concept described herein through the above teachings or the technology or knowledge of the relevant field. Changes and variations made by those skilled in the art do not depart from the spirit and scope of the present invention, and should be within the scope of protection of the claims attached to the present invention.
Claims
1. A method for constructing a PCR-free Hi-C library, the method comprising the following steps: Cross-linking: cross-linking the DNA and proteins in the sample cells to obtain a cross-linked product; Enzyme digestion: Enzyme digestion of DNA in the cross-linked body to obtain cross-linked DNA fragments; Adding sequencing adapters: adding sequencing adapters to the cross-linked DNA fragments to obtain DNA fragments with sequencing adapters, wherein the sequencing adapters include at least one sequencing primer sequence; Reconnecting DNA fragments: reconnecting the DNA fragments with sequencing adapters to obtain cross-linked reconnected DNA fragments, wherein the DNA fragments with sequencing adapters are reconnected through the sequencing adapters they carry, and the reconnected DNA fragments include two sequencing adapters; Decrosslinking: Decrosslinking the crosslinked reconnected DNA fragments with proteins to obtain reconnected DNA fragments; Adding library adapters, wherein the library adapters include an adapter with a first library structure sequence and an adapter with a second library structure sequence, to obtain a PCR free Hi-C library, in, The sequencing adapter is composed of two single-stranded DNAs, and the sequencing primer sequence is arranged on the single-stranded DNA fragment connected to the cross-linked DNA fragment at the 3' end. The sequencing primer sequence is arranged at the 3' end of the single-stranded DNA fragment. The 5' end of the single strand of the sequencing primer sequence is a protruding nucleotide sequence shown in SEQ ID NO: 1, or the 3' end of the single strand of the sequencing primer sequence is not a protruding nucleotide sequence shown in SEQ ID NO: 1, SEQ ID NO:1:5'-N 1… N m N' 1… N' m -3', Wherein, N is any one of A, T, C, and G deoxyribonucleotides. 1… N m With N' m… N'1 is a reverse complementary sequence, and m is 1-4.
2. According to the PCR free Hi-C library construction method of claim 1, the length of the sequencing adapter is 10-40 bp.
3. The method for constructing a PCR free Hi-C library according to claim 1, wherein: After the enzyme digestion step, the cross-linked DNA fragments are end-repaired and A-added.
4. The method for constructing a PCR free Hi-C library according to claim 3, wherein: Set the sequencing primer sequence so that the 3' end of the single strand has an overhanging T.
5. The method for constructing a PCR free Hi-C library according to claim 1, wherein: After the cross-link reversal step, the religated DNA fragments were fragmented and A-treated.
6. A PCR free Hi-C library, obtained by the PCR free Hi-C library construction method according to any one of claims 1 to 5, wherein the reconnected DNA fragments include two sequencing adapters.
7. A method for detecting a PCR free Hi-C library, comprising sequencing the PCR free Hi-C library of claim 6, wherein the sequencing primer sequence inside the reconnected fragment is used as the sequencing start point.
Citation Information
Patent Citations
DLO Hi-C (Digestion-Ligation-Only Hi-C) chromosome conformation capture method
CN106480178A
Method and kit for constructing capture library with high detection performance
CN113493932A