A method for constructing a single-cell Hi-C library
By using special sequencing adapters to insert reconnected fragments in single-cell Hi-C technology, the problems of long library construction cycle, low success rate and high proportion of invalid fragments are solved, and efficient and low-cost single-cell Hi-C library construction and data analysis are achieved.
Patent Information
- Application Number
- CN202410031542.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-09
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2044-01-09
AI Technical Summary
Existing single-cell Hi-C technology has problems such as long library construction cycle, low success rate, high proportion of invalid fragments, and complex and time-consuming data analysis. In particular, unreconnected fragments are detected during the sequencing process, resulting in noisy data, which affects the detection cost and quality.
Special sequencing adapters are inserted into the reconnected fragments so that only molecules containing reconnected fragments are detected by the sequencer. Through cross-linking, enzyme digestion, addition of sequencing adapters, reconnection, isolation of single cells, decrosslinking, addition of library adapters and library amplification steps, the proportion of invalid fragments is reduced and the experimental operation and analysis process are simplified.
It achieves the construction of single-cell Hi-C libraries with short construction time, low cost, and high effective data ratio, reduces the requirements for samples and sequencing amount, and improves data analysis efficiency.
Smart Images

Figure CN117802205B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a single-cell Hi-C library construction method, belonging to the technical field of gene sequencing. Background Art
[0002] Chromatin conformation capture is a molecular biology technique used to investigate the spatial organization of the three-dimensional structure within chromosomes. It holds significant significance for revealing the three-dimensional structure of the genome, understanding gene regulation, studying disease mechanisms, and advancing drug development. Since Deker proposed 3C in 2002, advancements in research have led to the development of 4C, 5C, ChIA-PET, Hi-C, and Capture Hi-C. These methods, to varying degrees, reveal the three-dimensional structure and interactions within chromosomes, providing crucial tools for studying gene regulation, genome stability, and disease.
[0003] Traditional Hi-C techniques typically require a large number of cells to obtain sufficient DNA fragments for analysis. Due to the variability in genomic structure between different cells, single-cell Hi-C technology has emerged. Single-cell Hi-C can reveal structural differences between different cell types, providing groundbreaking opportunities for studying genomic structural variations between cell types and cellular heterogeneity in development and disease. However, because single-cell Hi-C, derived from Hi-C, often requires biotin retrieval to retrieve reconnected fragments, it often suffers from long library construction times and low success rates. Furthermore, fragments that have been biotin-treated but not reconnected are retrieved along with the reconnected fragments and subsequently detected during sequencing. Unconnected fragments are invalid fragments that are irrelevant to the analysis of chromatin spatial organization and structure. Once detected, they become noise in data analysis, such as dangling ends. Therefore, the ratio of invalid fragments to noise data has always been a significant factor affecting the cost and quality of single-cell Hi-C analysis. In addition, single-cell Hi-C technology detects DNA fragments after chromosome reconnection, and the reconnection position is uncertain. When performing data analysis on the sequencing results, it is necessary to first find the reconnection site, otherwise the sequence cannot be compared with the reference genome, making the analysis and processing complicated and time-consuming. Summary of the Invention
[0004] In view of the above-mentioned deficiencies in the prior art, the object of the present invention is to provide a method for constructing a single-cell Hi-C library that is time-saving, has a low proportion of invalid fragments, and is convenient for experimental operation and analysis.
[0005] The present invention has conducted in-depth research to solve the above problems, and found that: by improving the single-cell Hi-C technology, using special sequencing adapters, especially sequencing adapters provided with sequencing primer sequences, the sequencing adapters are inserted inside the reconnected fragments, so that only molecules containing reconnected fragments will be detected by the sequencer, and a library with a lower proportion of invalid fragments is obtained. The library construction method of the present invention can reduce the requirements of the library on the sample starting amount and sequencing amount, and at the same time, it can also save the steps of biotin labeling and retrieving reconnected fragments, shortening the library construction time and reducing the single-cell Hi-C library construction cost and detection cost as a whole, thereby completing the present invention. By adopting the single-cell Hi-C library construction method of the present invention, a single-cell Hi-C library with short library construction time, low cost, and high proportion of valid data can be achieved.
[0006] That is, the present invention includes,
[0007] 1. A method for constructing a single-cell Hi-C library, comprising the following steps:
[0008] Cross-linking: cross-linking the DNA and proteins in the sample cells to obtain a cross-linked product;
[0009] Enzyme digestion: Enzyme digestion of DNA in the cross-linked body to obtain cross-linked DNA fragments;
[0010] Adding sequencing adapters: adding sequencing adapters to the cross-linked DNA fragments to obtain DNA fragments with sequencing adapters, wherein the sequencing adapters include at least one sequencing primer sequence;
[0011] Reconnect DNA fragments: Reconnect the DNA fragments with sequencing adapters to obtain cross-linked reconnected DNA fragments.
[0012] Isolation of single cells;
[0013] Decrosslinking: Decrosslinking the cross-linked reconnected DNA fragments from proteins to obtain reconnected DNA fragments;
[0014] adding a library linker, wherein the library linker comprises a first library structure sequence and a second library structure sequence;
[0015] Library amplification: Primers with the first library structure sequence and primers with the second library structure sequence are used to amplify the DNA fragments with library adapters to obtain a single-cell Hi-C library.
[0016] 2. According to the single-cell Hi-C library construction method described in item 1, the DNA fragments with sequencing adapters are reconnected through the sequencing adapters they carry, and the reconnected DNA fragments include two sequencing adapters.
[0017] 3. According to the single-cell Hi-C library construction method of claim 1, the length of the sequencing adapter is 10-40 bp, preferably 10-30 bp, and more preferably 15-25 bp.
[0018] 4. The single-cell Hi-C library construction method according to item 1, wherein the sequencing adapter is composed of two single-stranded DNAs, and the sequencing primer sequence is arranged on the single strand connected to the cross-linked DNA fragment at its 3' end.
[0019] 5. The single-cell Hi-C library construction method according to item 4, wherein the 5' end of the single strand with a sequencing primer sequence is an overhanging nucleotide sequence, or the 3' end of the single strand without a sequencing primer sequence is an overhanging nucleotide sequence, and the overhanging nucleotide sequence is as follows:
[0020] 5'-N 1…… N m N' 1…… N' m -3',
[0021] Wherein, N is any one of A, T, C, G deoxyribonucleotides, N 1…… N m With N' m…… N'1 is a reverse complementary sequence, m is 1-4, preferably m is 1-3, more preferably m is 1-2, and most preferably m is 1.
[0022] 6. The method for constructing a single-cell Hi-C library according to item 1, wherein the cross-linked DNA fragments are end-repaired and A-added after the enzyme digestion step.
[0023] 7. The method for constructing a single-cell Hi-C library according to item 6, wherein the 3' end of the single-stranded sequencing primer sequence is set to a protruding T.
[0024] 8. The method for constructing a single-cell Hi-C library according to item 7, wherein the sequencing adapter consists of the nucleotide sequences shown in SEQ ID NO: 1 and SEQ ID NO: 2.
[0025] SEQ ID NO: 1: 5'-CGTGTGCTGTGACTGGAGT-3',
[0026] SEQ ID NO: 2: 5'-CTCCAGTCACAGCACA-3'.
[0027] 9. The method for constructing a single-cell Hi-C library according to item 1, wherein the library linker further comprises a cell barcode.
[0028] 10. The method for constructing a single-cell Hi-C library according to claim 9, wherein the cell barcode is directly linked to the first library structure sequence. Preferably, the cell barcode is positioned near one end of the reconnected DNA fragment.
[0029] 11. According to the single-cell Hi-C library construction method of item 9, the length of the cell barcode is 4-25 bp, preferably 6-10 bp, and more preferably 8-16 bp.
[0030] 12. Following the single-cell Hi-C library construction method described in item 1, fragment and A-add the reconnected DNA fragments after the cross-linking step.
[0031] 13. The method for constructing a single-cell Hi-C library according to item 1, wherein the library adapter consists of the nucleotide sequences shown in SEQ ID NO: 3 and SEQ ID NO: 4.
[0032] SEQ ID NO: 3'-GTGAAGATCTCGTATGCCGTCTTCTGCTTG-3',
[0033] SEQ ID NO: 4: 5'-AATGATACGGCGACCACCGAGATCTACAC-NN
[0034] NNNN-CTTCACT-3'.
[0035] 14. The single-cell Hi-C library construction method according to item 1, wherein the reagent system used in the step of adding sequencing adapters is: 1x NEB T4 DNA ligase buffer, 400,000 U / mL NEB T4 DNA ligase, and 0.1 mM sequencing adapters.
[0036] 15. A single-cell Hi-C library obtained by the single-cell Hi-C library construction method according to any one of items 1 to 14, wherein the reconnected fragments comprise two sequencing primer sequences, and the sequencing adapter comprises at least one sequencing primer sequence.
[0037] 16. A method for detecting a single-cell Hi-C library, comprising sequencing a single-cell Hi-C library constructed by the method for constructing a single-cell Hi-C library according to item 15, wherein the sequencing primer sequence within the reconnected fragment is used as the sequencing starting point.
[0038] According to one aspect of the present invention, a method for constructing a single-cell Hi-C library is provided, comprising the steps of cross-linking DNA and proteins within sample cells, enzymatically digesting the DNA in the cross-links, adding sequencing adapters, reconnecting the DNA fragments, isolating single cells, resolving cross-links, adding library adapters, and amplifying the library.
[0039] In the above single-cell Hi-C library construction method,
[0040] The crosslinking step involves crosslinking DNA and proteins within the sample cells to form a crosslinked complex. This crosslinked complex maintains the physical interactions between the DNA fragments and the associated proteins. A crosslinking agent such as formaldehyde can be used.
[0041] The enzymatic digestion step is to digest the DNA in the cross-linked body to obtain cross-linked DNA fragments. The enzymatic digestion reagent can be a restriction endonuclease or a non-restriction endonuclease.
[0042] Adding a sequencing adapter is to add a sequencing adapter to the cross-linked DNA fragment to obtain a DNA fragment with a sequencing adapter. The sequencing adapter includes at least one sequencing primer sequence.
[0043] A connector is typically a known short nucleotide sequence used to connect unknown sequencing fragments, allowing the sequencing product to establish a connection with the computer system. In the present invention, a sequencing connector is a connector comprising nucleotides or nucleotide sequences related to the sequencing process. These nucleotides or nucleotide sequences related to the sequencing process can be sequencing primers. A sequencing primer refers to a short nucleotide sequence chain at the starting point of DNA synthesis during the sequencing process. That is, a nucleotide sequence that can be used as the starting point of sequence detection. For example, the sequencing primers for read1 or read2 provided by the Illumina sequencing platform are typically used as the starting point of DNA synthesis when sequencing on the Illumina sequencing platform, i.e., the starting point of sequencing detection. In a specific embodiment of the present invention, the sequencing primer sequence can be a commercial sequence, for example, the sequencing primer sequence for read1 or read2 of the Illumina sequencing platform, or it can be a sequencing primer sequence customized according to sequencing needs.
[0044] The step of reconnecting DNA fragments involves reconnecting cross-linked DNA fragments to obtain cross-linked reconnected DNA fragments. The enzymatically cleaved DNA fragments are rebound. DNA fragments that are adjacent in three dimensions are likely to be bound together, meaning that DNA fragments that may interact are reconnected, thereby enabling detection of which fragments interact with each other. In the present invention, reconnected fragments are obtained by reconnecting DNA fragments to which sequencing adapters have been added using the added sequencing adapters. The resulting reconnected DNA fragments typically include two sequencing adapters.
[0045] The isolation of single nuclei involves separating individual cells or nuclei from a larger number of sample cells or nuclei for subsequent manipulation. Common methods include flow cytometric sorting, capillary aspiration under a microscope, gradient dilution, or direct dilution. For example, sample nuclei are placed in a multi-well plate at a concentration of one cell per well. Methods include flow cytometric sorting, capillary aspiration under a microscope, gradient dilution, or direct dilution to a concentration of one cell per well.
[0046] The cross-linking removal step involves removing the cross-links between the cross-linked reconnected DNA fragments and the protein to obtain reconnected DNA fragments. This is the DNA fragment obtained by reconnecting the potentially interacting DNA fragments, also known as the reconnected DNA fragments. Cross-linking removal can be performed using a reagent containing proteinase K.
[0047] The step of adding a library adapter is to add a library adapter to both ends of the reconnected DNA fragment. The library adapter includes a first library structure sequence and a second library structure sequence.
[0048] In the present invention, the library adapter refers to a nucleotide sequence that is added to both ends of the DNA fragment to be tested during the library construction process and can be sequenced with a sequencer. It is an important bridge connecting the DNA fragment to be tested and a sequencing vector, such as a sequencing chip (Flow cell). Because the structure of the DNA obtained from the biological sample itself may not be directly detected on the sequencing platform. In order to meet the requirements of the sequencing platform or sequencing chip for the structure of the sequencing fragment, it is usually necessary to process the obtained DNA to form a library that meets the requirements of the sequencing platform or sequencing chip, and then the prepared library is sequenced on the machine. The adapter used in the library preparation process is called a library adapter. Unlike the library adapter used in the present invention, which only contains a bridge connecting the DNA fragment to be tested and the sequencing vector, the library adapter commonly used in the prior art includes both a bridge connecting the DNA fragment to be tested and the sequencing vector and a sequencing primer sequence for detecting the inserted fragment of the library.
[0049] In the present invention, library structure-related sequences refer to other sequences in the library fragments excluding unknown sample fragments and sequencing primer sequences. These sequences may have different functions. There may be one or more than two sequences with the same function. Different functions may be used, for example, to associate with a sequencing platform or sequencing chip, or to mark a label for a sample, but do not include a function for marking the start site of sequencing. For example, the P5, P7, index1, index2 sequences provided by the Illumina sequencing platform, or sequences that are reverse complementary to these sequences, are associated with the Illumina sequencing platform or chip. For example, the P5, P7 sequences, or their reverse complementary sequences are usually matched with the fixed sequences set on the Illumina sequencing chip, and are commonly used associated sequences for sequencing on this platform. The index sequence can be used to distinguish different samples and to achieve simultaneous detection of multiple samples. In a specific embodiment of the present invention, the library structure sequence may be a commercial sequence or a sequence customized according to sequencing needs.
[0050] In the present invention, the first and / or second library structure sequences may be all or part of the library structure sequence. The partial first and / or second library structure sequence refers to a ratio of its length to the length of the entire library structure sequence. This ratio may be greater than 1 / 3, preferably greater than 1 / 2, more preferably greater than 3 / 4, more preferably greater than 9 / 10, and most preferably 1.
[0051] Library amplification uses primers with the first library structure sequence and primers with the second library structure sequence to amplify the DNA fragments with library adapters to obtain a single-cell Hi-C library.
[0052] The reagents used for library amplification can be selected from one or more of DNA polymerase, RNA polymerase, DNA ligase, RNA ligase, primers and buffer systems, or commercial kits for library amplification.
[0053] In the present invention, the terms "first" or "second" are used to distinguish between different sequencing adapters, library adapters, library structure sequences, or sequencing-associated sequences, which may be functionally identical or similar but differ in structure or other properties. The characteristics of sequencing adapters, library adapters, library structure sequences, and sequencing-associated sequences are described above and will not be repeated here.
[0054] Furthermore, in the above-mentioned single-cell Hi-C library construction method, the DNA fragments with sequencing adapters are reconnected through the sequencing adapters they carry. The reconnected DNA fragments include two sequencing adapters, that is, two sequencing adapters are contained between the DNA fragments that may have an interacting relationship.
[0055] Furthermore, in the above-mentioned single-cell Hi-C library construction method, the sequencing adapter is composed of two single-stranded DNA chains, and the sequencing primer sequence is arranged on the single-stranded chain connected to the cross-linked DNA fragment at the 3' end, and the sequencing primer sequence is arranged at the 3' end of the single-stranded chain in which it is located. The length of the sequencing primer sequence is 10-40bp, preferably 10-30bp, and more preferably 15-25bp. The ratio of the length of the sequencing primer sequence to the length of the single-stranded chain in which it is located is more than 1 / 2, preferably more than 3 / 4, and more preferably more than 4 / 5. The sequencing primer sequence can be the entire or partial sequence of the sequencing primer. Wherein, the partial sequencing primer sequence refers to a ratio whose length is in a certain ratio to the length of the entire sequencing primer. This ratio can be more than 1 / 3, preferably more than 1 / 2, more preferably more than 3 / 4, more preferably more than 9 / 10, and most preferably 1.
[0056] Furthermore, in the above-mentioned single-cell Hi-C library construction method, the preferred length of the sequencing adapter used is 15-25 bp, and good experimental results can also be obtained within the length range of ±10 bp, preferably ±5 bp.
[0057] Furthermore, in the above-mentioned library construction method, the sequencing primer sequence is disposed at the 3' end of one strand of the sequencing adapter. The ratio of the length of the sequencing primer sequence to the length of the sequencing adapter is greater than 1 / 3, preferably greater than 1 / 2, more preferably greater than 3 / 4, and even more preferably greater than 9 / 10.
[0058] Furthermore, in the above-mentioned single-cell Hi-C library construction method, a sequencing adapter can be added to the DNA fragment. Specifically, the 5' end of the single strand with the sequencing primer sequence is a protruding nucleotide sequence, or the 3' end of the single strand without the sequencing primer sequence is a protruding nucleotide sequence 5'-N 1… N m N' 1… N' m -3', wherein N is any one of the deoxyribonucleotides A, T, C, or G. 1… N m With N' m… N'1 is a reverse complementary sequence, m is 1-4, preferably m is 3, more preferably m is 2, and most preferably m is 1. The DNA fragments with sequencing adapters are connected through the protruding nucleotide sequences of the sequencing adapters to form reconnected fragments.
[0059] Furthermore, in the single-cell Hi-C library construction method described above, the cross-linked DNA fragments can be end-repaired and A-addition treated after the enzyme digestion step. End-repair refers to the repair of damaged or incomplete DNA fragments, such as those that have been digested or otherwise disrupted, so that the ends are double-stranded, i.e., blunt-ended, without any free single-stranded nucleotide sequences. A-addition typically involves adding an adenine (A) deoxyribonucleotide to the 3' end of each strand of a double-stranded DNA fragment that lacks free single-stranded nucleotide sequences.
[0060] Furthermore, in the above-described single-cell Hi-C library construction method, sequencing adapters can be preferably added to the A-tailed DNA fragments. Preferably, the sequencing adapters have an overhanging thymine (T) deoxyribonucleotide at the 3' end of the single strand containing the sequencing primer sequence. This treatment can produce sequencing adapters with an overhanging T at the 3' end. Using sequencing adapters with an overhanging T at the 3' end allows for smooth ligation of the sequencing adapters to DNA fragments with A tails.
[0061] Furthermore, in the above-mentioned single-cell Hi-C library construction method, sequencing adapters are preferably added to the DNA fragments with A added. Preferably, the sequencing adapters do not have an A at the 3' end of the single strand where the sequencing primer sequence is not set. This can effectively prevent ligation between adapters.
[0062] Furthermore, in the above-mentioned single-cell Hi-C library construction method, a sequencing adapter is preferably added to the DNA fragment with A added thereto. The sequencing adapter can be a product of annealing treatment of the sequences shown in SEQ ID NO: 1 and SEQ ID NO: 2.
[0063] SEQ ID NO: 1: 5'-CGTGTGCTGTGACTGGAGT-3'
[0064] SEQ ID NO: 2: 5'-CTCCAGTCACAGCACA-3'.
[0065] In the present invention, the sequencing primer sequence can be one of the sequencing primers provided by the Illumina sequencing platform, such as the sequencing primer sequence for read 1 or read 2. Other customized sequences can also be selected based on the specificity of the sample or platform.
[0066] In the present invention, a cell barcode refers to a marker used to distinguish individual cells. By introducing a unique DNA sequence tag, individual cells can be identified and tracked. In the above-mentioned single-cell Hi-C library construction method, a cell barcode is added to the library linker, such as a nucleotide sequence consisting of multiple Ns in SEQ ID NO: 4, to label individual cells, allowing for differentiation during data analysis. The cell barcode can be a nucleotide sequence of approximately 4-25 bp.
[0067] Furthermore, in the above-mentioned single-cell Hi-C library construction method, the cell barcode is directly linked to the first library structure sequence. Preferably, the cell barcode is positioned near one end of the reconnected DNA fragment. For example, in SEQ ID NO: 4, the cell barcode is positioned 3' to the first library structure sequence. Preferably, the cell barcode is 6-20 bp in length, more preferably 8-16 bp in length.
[0068] Furthermore, in the single-cell Hi-C library construction method described above, the reconnected DNA fragments are fragmented and A-addition treated after the cross-linking removal step. In library construction, fragmentation generally refers to the reduction of longer DNA fragments by enzyme digestion or other methods into fragments of a length suitable for library inserts in a sequencer. Preferably, the fragments are fragmented to an average length of 200-1000 bp, more preferably to an average length of 300-800 bp, and most preferably to an average length of 400-700 bp. Fragmentation can be achieved by methods such as non-restriction enzyme digestion or ultrasonic fragmentation. A-addition treatment has been described above and will not be repeated here.
[0069] Furthermore, in the above-mentioned single-cell Hi-C library construction method, the library adapter is composed of the sequences shown in SEQ ID NO: 3 and SEQ ID NO: 4.
[0070] SEQ ID NO: 3'-GTGAAGATCTCGTATGCCGTCTTCTGCTTG-3',
[0071] SEQ ID NO: 4: 5'-AATGATACGGCGACCACCGAGATCTACAC-NNNNNN-CTTCACT-3'.
[0072] Furthermore, in the single-cell Hi-C library construction method described above, reagents such as T4 DNA ligase and T4 DNA ligase buffer can be used to ligate sequencing adapters. A reagent system for ligating sequencing adapters can be: 1x NEB T4 DNA ligase buffer, 400,000 U / mL NEB T4 DNA ligase, and 0.1 mM sequencing adapters.
[0073] Furthermore, in the single-cell Hi-C library construction method described above, library adapters can be ligated using reagents such as T4 DNA ligase, T4 DNA ligase buffer, etc. The library adapters can be ligated using a reagent system consisting of: 10x NEB T4 DNA ligase buffer, 400,000 U / mL NEB T4 DNA ligase, and 0.1 mM library adapters.
[0074] Furthermore, in the above-mentioned single-cell Hi-C library construction method, library amplification can use PCR reagents such as NEB Q5 high-fidelity polymerase and KAPA HiFi Hot Start Ready Mix.
[0075] According to another aspect of the present invention, a single-cell Hi-C library is provided, obtained by the above-described library construction method. The reconnected fragment includes two sequencing adapters. The sequencing adapters include a sequencing primer sequence. The two sequencing adapters are inserted between the two DNA fragments that interact with each other spatially by reverse-complementary pairing of the overhanging nucleotide sequences. That is, the reconnected fragment includes two sequencing adapters.
[0076] According to another aspect of the present invention, a method for detecting a single-cell Hi-C library is provided. The method involves sequencing the single-cell Hi-C library and detecting the sequence from two sequencing primer sequences within the reconnected fragment, respectively, toward both ends of the reconnected fragment. Specifically, the sequencing primer sequences read the sequence from the interior of the DNA fragment with the sequencing adapter as the sequencing starting point, rather than from both ends of the reconnected fragment as the sequencing starting point.
[0077] In the present invention, the addition of sequencing adapters allows the cross-linked DNA fragments to be connected to sequencing adapters with special sequences. The subsequent reconnection step allows the DNA fragments cross-linked by the associated proteins to be connected through the sequencing adapters. The sequencing adapters serve as a bridge for connecting enzyme-divided DNA fragments. This connection method allows the reconnected DNA fragments to have two sequencing adapters inserted at the reconnection position. Only reconnected fragments containing two sequencing adapters can be detected by the sequencer in double-end sequencing, thereby reducing the proportion of invalid fragments entering the library and sequencing. At the same time, because the sequencing starting point (such as a sequencing primer) is located between the two reconnected DNA fragments, only one of the two reconnected DNA fragments will be read at each end during double-end sequencing. When comparing with the reference genome sequence, they can be directly compared, saving analysis time.
[0078] The method of the present invention is more convenient by connecting special sequencing adapters and other multi-step optimizations to DNA fragments before fragment reconnection, eliminating the need for biotin retrieval and significantly increasing the proportion of valid fragments in the library. At the same time, compared with existing single-cell Hi-C methods, the method of the present invention can reduce library construction costs and shorten library construction time. That is, compared with the existing technology, the present invention realizes a single-cell Hi-C library construction method that is time-saving, low-cost, has a low proportion of invalid fragments, and is easy to operate and analyze. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] Figure 1 Schematic diagram of the library fragment structure of Example 1.
[0080] Among them, 1-interacting DNA fragment 1, 2-interacting DNA fragment 2, 3-sequencing adapter, 4-first library structure sequence and its complementary sequence, 5-second library structure sequence and its complementary sequence, 6-cell barcode and its complementary sequence, 7-arrow indicates the sequencing direction.
[0081] Specific embodiments of the invention Example
[0082] The present invention will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are for explaining the present invention and are not intended to limit the present invention.
[0083] Example 1: A method for constructing a single-cell Hi-C library
[0084] (I) Library construction
[0085] 1. Chromatin Cross-linking
[0086] One million E14 mouse macrophage RAW264.7 cells (Sample 1) were harvested and centrifuged at 400g for 5 minutes. The supernatant was removed. The cells were resuspended in 2 mL of PBS and 54 μL of 37% formalin was added to a final formaldehyde concentration of 1%. Cross-linking was terminated by adding pre-chilled glutamate to a final concentration of 130 mM over 10 minutes. The cells were centrifuged at 3000g for 5 minutes, the supernatant removed, and the nuclei resuspended in 1 mL of PBS.
[0087] 2. Enzyme digestion, end repair, and addition of adenine nucleotides (A)
[0088] Centrifuge at 3000g for 5 minutes and remove the supernatant. Resuspend the nuclei in 48μL of 1x NEB buffer DpnII, add 0.5μL of 50,000 units / mL NEB DpnII Enzyme, 0.5μL of 5,000 U / mL NEB Tag DNA Polymerase, and 1μL of 10nM dNTPs. Incubate at 37°C for 15 minutes for enzyme digestion and incubate at 65°C for 30 minutes for end repair and A addition. Centrifuge at 3000g for 5 minutes, remove the supernatant, and resuspend the nuclei in 1mL of PBS.
[0089] 3. Ligation of Sequencing Adapters
[0090] Centrifuge at 3000g for 5 minutes and remove the supernatant. Resuspend the nuclei in 42 μL of 1x NEB T4 DNA ligase buffer. Add 5 μL of 400,000 U / mL NEB T4 DNA ligase and 3 μL of 0.1 mM sequencing adapters (prepared by gradient annealing of the oligonucleotide sequences represented by SEQ ID NO: 1 and SEQ ID NO: 2). Incubate at 22°C for 1 hour. Centrifuge at 3000g for 5 minutes, remove the supernatant, and resuspend the nuclei in 1 ml of PBS.
[0091] The sequencing primer sequence is SEQ ID NO: 7, which is set in SEQ ID NO: 1. That is, in this embodiment, the protruding nucleotide sequence is designed as CG, which is located in SEQ ID NO: 1.
[0092] 4. DNA Fragment Religation
[0093] Prepare the following reaction mixture: 10 μL of 10x NEB T4 DNA ligase buffer, 2.5 μL of 100 mM EGTA, 2 μL of 10,000 units / mL NEB T4 Polynucleotide Kinase, and 10 μL of 400,000 units / mL NEB T4 DNA ligase. Make up to 100 μL with water. Centrifuge the product from the previous step at 3000g for 5 minutes and remove the supernatant. Resuspend the nuclei in 100 μL of the reaction mixture and incubate at 37°C for 20 minutes, followed by incubation at 22°C for 30 hours.
[0094] 5. Isolation of Single Cells
[0095] Dilute the cell nuclei suspension to 200 nuclei / mL and transfer 5 μL / well to a 96-well PCR plate.
[0096] 6. Decrosslinking
[0097] Add 0.5 μL of 120 Units / mL NEB Thermolabile Proteinase K to each well and incubate at 37°C for 15 minutes to reverse cross-linking. Inactivate the enzyme by incubating at 55°C for 10 minutes.
[0098] 7. Fragment DNA, repair ends and add A.
[0099] Add 0.6 μL of 10x NEB Micrococcal Nuclease Reaction Buffer, 0.2 μL of 2,000,000 gel units / mL NEB Micrococcal Nuclease, 0.2 μL of 5,000 U / mL NEB Tag DNA Polymerase, and 0.1 μL of 10 nM dNTPs to each well and mix well. Incubate at 32°C for 5 minutes, then at 65°C for 30 minutes.
[0100] 8. Ligation of library adapters
[0101] Add 0.75 μL of 10x NEB T4 DNA ligase buffer, 0.5 μL of 400,000 U / mL NEB T4 DNA ligase, and 0.25 μL of 0.1 mM library adapters (prepared by gradient annealing of the oligonucleotide sequences represented by SEQ ID NO: 3 and SEQ ID NO: 4) to each well. The nucleotide sequence of multiple Ns serves as the single-cell barcode. Label the single cells in each well (including cells 1-6 in sample 1) with a library adapter containing a different single-cell barcode. Incubate at 20°C for 15 minutes.
[0102] Collect the liquid from all wells into a PCR tube and purify the DNA using 0.8x SPRI select beads. Recover the DNA in 23 μL of elution buffer.
[0103] 9. Single-cell Hi-C library amplification:
[0104] Add 1 μL 0.01 mM primer of the first library structure sequence (oligonucleotide sequence shown in SEQ ID NO: 5), 1 μL 0.01 mM primer of the second library structure sequence (oligonucleotide sequence shown in SEQ ID NO: 6), and 25 μL KAPA HiFi Hot Start Ready Mix to the product of the previous step.
[0105] The recombinant fragments containing the target region were enriched by PCR using a PCR instrument.
[0106] PCR program: 1) 98°C for 3 minutes, 2) 98°C for 30 seconds, 3) 54°C for 30 seconds, 4) 72°C for 30 seconds, 5) 72°C for 1 minute, 6) stop at 10°C. Steps 2) to 4) were cycled 18 times.
[0107] DNA fragments were screened and purified using 0.5x-0.8x SPRI select beads. Hi-C libraries (library 1) were generated from multiple single cells (including cells 1-6).
[0108] Depend on Figure 1 It can be seen that the reconnected fragment consists of two spatially interacting DNA fragments and two sequencing adapters between them, and the two are connected by reverse complementation through the CG sequences of the sequencing adapters they carry.
[0109] The sequences involved in this embodiment are shown in Table 1.
[0110] Table 1:
[0111]
[0112] The cell barcode sequences used in this example are shown in Table 2.
[0113] Table 2:
[0114]
[0115] (II) Library sequencing and data analysis
[0116] The library (library 1) obtained by the above method was sequenced and data analyzed. The library was sequenced using Miseq PE150 double-end sequencing. The sequencing volume was 20M read pairs. In the steps of detecting read 1 and read 2, customized sequencing primers such as the oligonucleotide sequence shown in SEQ ID NO: 8 were used. Figure 1 It can be seen that in this embodiment, the sequencing primer sequence between the reconnected fragments is used as the sequencing starting point to perform two-step sequencing of read1 and read2 towards both ends.
[0117] Note that only cell barcodes in this library were detected by the Index 2 sequencing step, and the sequencer can only use Illumina sequencers using the Forward Strand Workflow when using Index 2 detection (such as NovaSeq 6000 with v1.0 reagent kits, MiniSeq with rapid reagent kits, MiSeq, HiSeq 2500, or HiSeq2000).
[0118] After the data was split, each index2 data point represented a single-cell dataset. Residual Illumina library adapter sequences and sequencing primer sequences were trimmed, and then the sequencing data from read1 and read2 were reverse-complemented. The data were analyzed using the method described in reference [1]. Data filtering statistics for the entire library (library 1) were generated by combining the data from the six single cells.
[0119] Comparative Example 1
[0120] One million E14 mouse macrophage RAW264.7 cells (sample 2) were processed according to the treatment method described in reference [1], and six single cells (cells 7-12) were selected for single-cell Hi-C library (library 2) construction and sequencing analysis.
[0121] Statistics and analysis of sequencing results of examples and comparative examples
[0122] The statistics of the logarithm of DNA interactions captured by each single cell in the Example and Comparative Example libraries are shown in Table 3.
[0123] Table 3:
[0124]
[0125] The data filtering statistics of the embodiments and comparative examples are shown in Table 4.
[0126] Table 4:
[0127]
[0128] The above statistical data show that the average percentage of valid data in the sequencing data obtained using the method of Example 1 exceeds 55%. In contrast, the average percentage of valid data in the sequencing data obtained using the method of Comparative Example 1 is only approximately 27%. This demonstrates that the single-cell Hi-C library construction method of the present invention can significantly increase the percentage of valid data (the ratio of valid read pairs to total read pairs).
[0129] Furthermore, the average sequencing load in Example 1 was 18M, while the average sequencing load in the comparative example was 20.4M. However, the number of DNA captured per cell and the number of interaction pairs in Example 1 were both higher than those in the comparative example. Therefore, using the method in Example 1, a higher amount of valid data can be obtained with a sequencing load equivalent to or less than that in the comparative example. This shows that, when a comparable amount of valid data is required, the library construction method of the present invention requires a lower sequencing load.
[0130] At the same time, by comparing the methods of reference [1] and Example 1, it can be seen that the method of the present invention has fewer steps in library construction, takes less time, and is simpler to operate.
[0131] References[1]: Nimrod Rappoport, Elad Chomsky, Takashi Nagano, CharlieSeibert, Yaniv Lubling, Yael Baran, Aviezer Lifshitz, Wing Leung, ZoharMukamel, Ron Shamir, Peter Fraser & Amos Tanay. Single cell Hi-C identifies plastic chromosome conformations underlying the gastrulation enhancerlandscape. Nature Communications volume 14, Article number: 3844 (2023)
[0132] According to the present invention, other methods for studying DNA, RNA, and proteins can be used in combination to study cell characteristics or functions, as well as chromatin conformation, DNA, RNA, and protein functions. These other methods for studying DNA, RNA, and proteins include, but are not limited to, single-cell sequencing, gene chips, QPCR, first-generation sequencing, second-generation sequencing, third-generation sequencing, fourth-generation sequencing, gene sequencing, genome sequencing, metagenomic sequencing, exon sequencing, intron sequencing, target gene capture sequencing, RNA sequencing, expression profiling sequencing, transcriptome sequencing, small RNA transcriptome sequencing, microRNA sequencing, macrotranscriptome sequencing, LncRNA sequencing, tumor gene sequencing, tumor genome sequencing, Bisulfite methylation sequencing, ChIP-DNA sequencing, MeDIP sequencing, RRBS sequencing, Target-BS sequencing, and hmC sequencing.
[0133] It should also be noted that, provided that it is feasible and does not obviously violate the main purpose of the present invention, any technical feature or combination of technical features described in this specification as a component of a technical solution can also be applied to other technical solutions; and, provided that it is feasible and does not obviously violate the main purpose of the present invention, the technical features described as components of different technical solutions can also be combined in any manner to form other technical solutions. The present invention also includes technical solutions obtained by combination in the above circumstances, and these technical solutions are equivalent to those described in this specification.
[0134] The foregoing description shows and describes preferred embodiments of the present invention. As previously mentioned, it should be understood that the present invention is not limited to the forms disclosed herein and should not be construed as excluding other embodiments, but is applicable to various other combinations, modifications, and environments, and is capable of modification within the scope of the inventive concept described herein, through the above teachings or through technology or knowledge in the relevant field. Modifications and variations made by those skilled in the art without departing from the spirit and scope of the present invention are intended to be protected by the appended claims.
Claims
1. A method for constructing a single-cell Hi-C library, comprising the following steps: Cross-linking: cross-linking the DNA and proteins in the sample cells to obtain a cross-linked product; Enzyme digestion: Enzyme digestion of DNA in the cross-linked body to obtain cross-linked DNA fragments; Adding sequencing adapters: adding sequencing adapters to the cross-linked DNA fragments to obtain DNA fragments with sequencing adapters, wherein the sequencing adapters include at least one sequencing primer sequence; Reconnect DNA fragments: Reconnect the DNA fragments with sequencing adapters to obtain cross-linked reconnected DNA fragments; Isolation of single cells; Decrosslinking: Decrosslinking the cross-linked reconnected DNA fragments from proteins to obtain reconnected DNA fragments; adding a library linker, wherein the library linker comprises a first library structure sequence and a second library structure sequence; Library amplification: Use primers with the first library structure sequence and primers with the second library structure sequence to amplify the DNA fragments with library adapters to obtain a single-cell Hi-C library. The sequencing adapter includes at least one sequencing primer sequence, and the DNA fragment with the sequencing adapter added is reconnected through the added sequencing adapter, and the reconnected DNA fragment includes two sequencing adapters. The sequencing adapter consists of two DNA single strands, and the sequencing primer sequence is set on the single strand connected to the cross-linked DNA fragment at the 3' end. The 5' end of the single strand with the sequencing primer sequence set is a protruding nucleotide sequence, or the 3' end of the single strand without the sequencing primer sequence set is a protruding nucleotide sequence, and the protruding nucleotide sequence is as follows: 5'-N 1… N m N' 1… N' m -3', Wherein, N is any one of A, T, C, G deoxyribonucleotides, N 1… N m With N' m… N'1 is the reverse complementary sequence, m is 1-4, in, The oligonucleotide sequence of the first library structure sequence primer is: 5'-AATGATACGGCGACCACCGAGATCTACAC-3', The oligonucleotide sequence of the second library structural sequence primer is: 5'-CAAGCAGAAGACGGCATACGAGAT-3'.
2. The single-cell Hi-C library construction method according to claim 1, wherein the length of the sequencing adapter is 10-40 bp.
3. The method for constructing a single-cell Hi-C library according to claim 1, wherein: After the enzyme digestion step, the cross-linked DNA fragments are end-repaired and A-added.
4. The method for constructing a single-cell Hi-C library according to claim 3, wherein: Set the sequencing primer sequence so that the 3' end of the single strand has an overhanging T.
5. The method for constructing a single-cell Hi-C library according to claim 1, wherein: The library adapter includes a cellular barcode.
6. A single-cell Hi-C library obtained by the single-cell Hi-C library construction method according to any one of claims 1 to 5, wherein the reconnected fragments comprise two sequencing adapters.
7. A method for detecting a single-cell Hi-C library, comprising sequencing the single-cell Hi-C library according to claim 6, wherein the sequencing primer sequence within the reconnected fragment is used as the sequencing starting point.
Citation Information
Patent Citations
DLO Hi-C (Digestion-Ligation-Only Hi-C) chromosome conformation capture method
CN106480178A
Single cell library construction sequencing method and application thereof
CN117089597A