Oligonucleotide adapters and kits for single-cell chromatin accessibility sequencing

CN122563952APending Publication Date: 2026-08-14HUAZHONG AGRI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-30
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

1)通量有限,难以同时处理大量细胞;

Benefits of technology

1.高通量与低成本:单次实验能够检测数千至上万级别的单细胞,在保持高灵敏度的同时大幅降低检测成本;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

This invention discloses an oligonucleotide adapter and kit for single-cell chromatin accessibility sequencing, belonging to the field of single-cell omics technology. This invention achieves high-throughput detection of tens of thousands of single cells in a single batch by preparing Tn5 transposase complexes with different barcodes and combining them with a two-round cell sorting strategy. The experimental procedure includes cell nucleus isolation and sorting, multiplex transposition reaction, library construction, and high-throughput sequencing and data analysis. The accompanying kit contains Tn5 transposase complexes and various primer combinations. Compared with existing single-cell ATAC-seq methods, this invention simplifies the process, reduces costs, and increases throughput, making it suitable for single-cell chromatin accessibility sequencing of complex tissues. The kit and method described in this invention can also be applied to single-cell genotyping and quantitative trait locus localization in germ cells, and can be used in epigenetics, developmental biology, and crop functional genomics research.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of single-cell omics technology, specifically to an oligonucleotide adapter and kit for single-cell chromatin accessibility sequencing. Background Technology

[0002] Chromatin accessibility is a crucial feature of epigenetic regulation, characterizing the binding regions of regulatory elements and transcription factors in the genome. In recent years, single-cell sequencing technology has developed rapidly, leading to the development of single-cell chromatin accessibility sequencing (scATAC). ATAC seq has become an important tool for studying cellular heterogeneity and chromatin state. Currently, single-cell ATAC seq mostly uses commercial microfluidic systems (such as 10X Genomics Chromium), but its high experimental cost and strong equipment dependence have limited its research and promotion.

[0003] Although index-based scATAC-seq methods have been developed in recent years, these methods have the following limitations: 1) Limited throughput, making it difficult to process a large number of cells simultaneously; 2) High reagent consumption and high experimental costs; 3) The operation process is complex and requires high experimental conditions; 4) Data quality is unstable and reproducibility is poor. Especially in plant cell research, due to the presence of cell walls and tissue specificity, existing techniques are difficult to apply effectively to the study of chromatin accessibility in some special tissues.

[0004] Quantitative trait locus (QTL) mapping is a core tool in genetics and breeding. Traditional QTL studies typically require the construction of large-scale populations and systematic phenotyping, resulting in long experimental cycles, high costs, and difficulty in revealing cell-type-specific regulation. The rapid development of single-cell omics technology has provided new possibilities for obtaining multi-dimensional molecular phenotypes and elucidating the genetic mechanisms of quantitative traits at single-cell resolution.

[0005] However, QTL mapping at the single-cell level still faces many challenges, including the sparsity of single-cell data, the complexity of genotype inference, and the difficulty of applying traditional QTL analysis methods directly to single-cell omics data.

[0006] Therefore, there is an urgent need to develop an efficient and low-cost single-cell chromatin accessibility sequencing technology and apply it to QTL mapping of single-cell molecular phenotypes. This technology is expected to provide a new entry point for revealing the cell type-specific regulatory mechanisms of complex traits and promote in-depth research in plant functional genomics. Summary of the Invention

[0007] The purpose of this invention is to overcome the shortcomings of existing technologies and provide an oligonucleotide adapter and kit for single-cell chromatin accessibility sequencing. This invention utilizes an improved adapter oligonucleotide design and a two-round cell sorting combined with a multiple barcoding strategy to achieve high-throughput, low-cost chromatin accessibility detection of cells such as pollen at single-cell resolution, and further achieves single-cell level quantitative trait locus (QTL) localization. This method can achieve high-throughput chromatin accessibility detection at single-cell resolution and is widely applicable to epigenetics, developmental biology, and functional genomics research.

[0008] To achieve the above objectives, the technical solution designed by the present invention is as follows: This invention provides an oligonucleotide adapter for single-cell chromatin accessibility sequencing, the oligonucleotide adapter comprising a first strand and a second strand: The first chain includes, from the 5' end to the 3' end, a universal connector sequence, a barcode sequence, and an ME sequence in sequence; The second chain is a complementary chain containing the ME sequence, and its length does not exceed that of the first chain; The universal connector sequence has a length of 10-50 bp and a GC content of 25%-75%. The barcode sequence (used to identify the origin of the cell nucleus in the transposation reaction) has a GC content of 25% to 75% and a barcode sequence length of 4 to 20 bp.

[0009] Furthermore, the universal connector sequence is GACCTCGTGCACCTGGATC; The ME sequence is AGATGTGTATAAGAGACAG, or AGATGTGTATAAGAGACA, or its functionally equivalent sequence.

[0010] Furthermore, the expression for the first chain is: GACCTCGTGCACCTGGATC<>AGATGTGTATAAGAGACAG; Among them, GACCTCGTGCACCTGGATC is a universal connector sequence; <> represents a barcode sequence, which is composed of n N's, where n is an integer from 4 to 20, and each N is independently selected from A, T, C, and G; AGATGTGTATAAGAGACAG represents an ME sequence.

[0011] Depending on the actual situation, in order to improve the base balance of sequencing data, a spacer sequence of 0-20 bp in length can be introduced between the universal adapter sequence and the barcode sequence and / or between the barcode sequence and the ME sequence. The spacer sequence can be a random or balanced combination of nucleotides to maintain the base composition balance.

[0012] The adapter sequences described above are compatible with the Illumina PE150 sequencing platform and can bind smoothly with standard P5 / P7 sequencing adapters and sequencing primers, thereby supporting library amplification and sequencing on existing commercial sequencing platforms.

[0013] The present invention also provides a Tn5 transposase complex, characterized in that: the complex is formed by assembling the above-mentioned oligonucleotide linker with Tn5 transposase.

[0014] The present invention also provides a Tn5 transposase complex composition comprising two or more of the above-described Tn5 transposase complexes, wherein the barcode sequences of the oligonucleotide linkers contained in each complex are different from each other.

[0015] The above composition is used to distinguish different reaction batches or cell nuclear sources in multiple transposable reactions, thereby enabling high-throughput and multi-sample parallel analysis.

[0016] This invention also provides a single-cell barcoded PCR primer set, including a first-round PCR amplification primer set, which includes Ad1 series primers and Ad2 series primers, used together as a pair of forward / reverse primers in the first round of PCR amplification: (1) The general formula of the Ad1 series primers is: CCCTACACGACCGTCTCTCCGATCT(N)mGACCTCGTGCACCTGGATC; Among them, CCCTACACGACGCTCTTCCGATCT is the P5 end connector sequence. (N)m is a barcode sequence, where m is an integer from 0 to 50. GACCTCGTGCACCTGGATC is a universal connector sequence; (2) Ad2 series primers, their general formula is: CAAGCAGAAGACGGCATACGAGAT[i7]GTCTCGTGGGCTCGGAGATGT; Among them, GTCTCGTGGGCTCGGAGATGT is the universal connector sequence 1. [i7] is a tag sequence of length 8 bp.

[0017] Furthermore, it also includes primer pairs for the second round of amplification, said primer pairs being: Forward primer: AATGATACGGCGACCACCGAGATCTACAC[i5]ACACTCTTTCCCTACACGACCGCTTCCGATCT Where [i5] is a tag sequence of length 8 bp; The reverse primers are Illumina universal P7 primers.

[0018] This invention also provides a kit for single-cell chromatin accessibility sequencing, comprising: (1) The above-mentioned Tn5 transposase complex composition; (2) Transposition reaction buffer; (3) The above-mentioned single-cell barcoded PCR primer set; (4) PCR amplification system; (5) DNA purification reagents.

[0019] The transposable reaction buffer contains 50 mM Tris-HCl (pH 7.5), 25 mM MgCl2 and 50% (v / v) DMF.

[0020] The PCR amplification system described above includes high-fidelity PCR enzyme, buffer, dNTPs, and components to improve GC preference; the DNA purification reagent is a magnetic bead purification system.

[0021] The kit also includes 96-well or 384-well plates, and / or codes or documentation describing methods for barcode splitting, sequence alignment, genotype determination, quantitative analysis of chromatin accessibility, and identification of quantitative trait loci.

[0022] This invention also provides a multi-index-based single-cell chromatin accessibility sequencing method (MPS-ATAC-seq) using the above-described kit, comprising the following steps: (1) Collect target samples and separate cell nuclei, and distribute cell nuclei to multiple reaction tubes through the first round of sorting; (2) Add the Tn5 transposase complex composition from the above kit to each reaction tube to carry out the transposition reaction, thereby labeling the cell nuclei in each tube and thus carrying the first layer of barcode; (3) The transposable cell nuclei are sorted in the second round and assigned to multi-well plates with different primer combinations, so that the cell nuclei in each well obtain a second layer of barcode, realizing dual labeling at the single cell level; (4) The sorted cell nuclei were lysed and the first round of PCR amplification was performed using the kit described above to introduce secondary index sequences; (5) Merge and purify the first-round amplification products, and use the above-mentioned kit to perform a second round of PCR amplification to introduce standard sequencing adapters and obtain a high-throughput sequencing library; (6) Perform high-throughput sequencing on the obtained library, and perform two rounds of splitting, comparison and quality control on the sequencing data according to the barcode information to obtain chromatin accessibility data at single-cell resolution.

[0023] The sorting in steps (1) and (3) above is performed using a flow cytometer; the first round of sorting distributes cell nuclei to two or more reaction tubes, and the multi-well plate is a 96-well plate or a 384-well plate.

[0024] The target samples mentioned above include pollen, leaves, root tips, stem tips or floral organs of plants, or cells or nuclei of other eukaryotes.

[0025] The above step (6) involves two rounds of splitting based on barcode information: the first round is splitting based on the sample index, and the second round is splitting based on the barcode in the tube; the alignment is to align to two parental reference genomes and compare the alignment scores; the quality control includes filtering low alignment quality reads, PCR duplicate reads and reads aligned to organelle genomes, and removing cells that do not meet the standards for Tn5 cut point number, PCR duplication rate or FRiP value.

[0026] As an application of the above-mentioned kits or methods, single-cell genotyping can also be performed. Using the methods described above, chromatin accessibility data of single-cell pollen is obtained, and a genotypic recombination map is constructed according to the following steps: a. Based on the barcode information, the raw sequencing data is split into data corresponding to individual cell nuclei in two rounds; b. Align the sequencing reads to the two parental reference genomes respectively; c. Screen for valid sequences with differential alignment scores between the two parental reference genomes; d. Divide the reference genome into fixed-size windows, count the number of valid sequences supporting different parental sources in each window, and determine the window genotype based on the proportion of sequences supporting each parental source in the window; e. Use Hidden Markov Models (HMMs) in conjunction with recombination rate, physical distance, and sequencing error rate to correct and fill in window genotypes; f. Obtain single-cell genotype maps, and then construct recombination maps based on the genotype information of individual single cells.

[0027] Preferably, the window size in step (d) is 1 kb to 1 Mb; A 10 kb window is preferred for initial genotype determination; A 50 kb window is preferred for recombination analysis.

[0028] Preferably: When the proportion of sequences originating from the same parent and having a high alignment score within the window is greater than or equal to 80%, it is determined to be the corresponding parent genotype.

[0029] As a further application of the above kits or methods, quantitative trait locus localization at the single-cell level can also be performed, including the following steps: a. Obtain single-cell chromatin accessibility sequencing data, and split the raw sequencing data into data corresponding to individual cell nuclei in two rounds according to barcode information; b. Compare, deduplicate, and perform quality control on the split data, and statistically analyze the chromatin accessibility data of each single cell in a fixed window as the molecular phenotype; c. Divide the reference genome into multiple contiguous windows; d. Within each window, cells are dynamically divided into different genotype groups based on their individual genotypes; e. Compare the molecular phenotypic differences among different genotype groups; f. Screen genomic regions that are significantly associated with molecular phenotypes to obtain quantitative trait loci (QTLs). Preferably, the dynamic partitioning in step (d) refers to: for each window to be detected, the comparison population is reconstructed according to the single-cell genotype corresponding to the window, and different cell grouping methods are used for different windows.

[0030] Preferably, the molecular phenotype includes: chromatin accessibility level; transposon accessibility level; gene expression level, etc.

[0031] Preferably, in step (e), statistical analysis is performed using DESeq2, Fisher's exact test, or t-test.

[0032] The above-mentioned kits are used in the detection of chromatin accessibility and / or the localization of quantitative trait loci at the single-cell level.

[0033] The above applications are used in epigenetics, developmental biology or crop functional genomics research to locate cell type-specific chromatin regulatory regions or to identify cell type-specific chromatin accessibility quantitative trait loci.

[0034] The beneficial effects of this invention are: 1. High throughput and low cost: A single experiment can detect thousands to tens of thousands of single cells, significantly reducing detection costs while maintaining high sensitivity; 2. Simple and standardized operation: The process is modularized and standardized through the accompanying reagent kit, reducing the dependence on special equipment; 3. Cell type specificity: It can accurately identify different cell types inside the sample and achieve cell type-specific chromatin regulation region localization; 4. Wide range of applications: Applicable to a variety of organisms, especially special samples such as plant tissues such as pollen and embryos, and can also be extended to animal samples; 5. QTL Localization Innovation: Combining the Dynamic-BSA method, accurate identification of quantitative trait loci can be achieved at the single-cell level without the need for large-scale construction of F2 populations, significantly shortening the research cycle.

[0035] Furthermore, by introducing a two-round sorting strategy in the cell sorting and library construction stages, this invention solves the problems of pollen cell rupture and low nuclear envelope acquisition efficiency, thereby improving the capture efficiency of real single-cell signals and facilitating the differentiation of individual cells. Attached Figure Description

[0036] Figure 1 This is a schematic diagram of a single-cell chromatin accessibility sequencing workflow based on multiple indexes. Figure 2 A schematic diagram of a library based on multi-indexed single-cell chromatin accessibility sequencing technology; In the diagram, different colors distinguish the various functional elements, including ME, Ad1 / Ad2 / Ad3, BC1 / BC2, gDNA, and i5 / P5 / P7 / i7.

[0037] Figure 3 A schematic diagram of MPS-ATAC-seq data quality control; a is a distribution diagram of MPS-ATAC-seq fragment lengths for different cell types in this embodiment; b is a schematic diagram of the TSS enrichment fold for different cell types in this embodiment; c is a schematic diagram illustrating the correlation between chromatin accessibility and FRiP values; d is a schematic diagram illustrating the correlation between chromatin accessibility and PCR repeatability; e represents the correlation of chromatin accessibility signals between Bulk ATAC-seq and MPS-ATAC-seq data (n=53,619, Spearman's ρ=0.73). f represents the single-cell quality control result.

[0038] Figure 4 A schematic diagram of genotype map construction; In the diagram, 'a' is a flowchart of the genotype map construction process; b is a schematic diagram of genotype determination for randomly selected single cells (showing the entire process from initial genotype, HMM correction, to genotype imputation). The blue area represents the ZS97 genotype, and the red area represents the MH63 genotype.

[0039] Figure 5 A diagram illustrating the genetic basis of chromatin accessibility variation in dynamic BSA analysis; In the figure, 'a' is the Dynamic-BSA analysis flowchart; b is a heatmap of caQTLs identified in all cells and in each cell type. Rows represent caQTL sites, and columns correspond to different cell types used in Dynamic-BSA analysis. Numerical values ​​represent -log10 (P-value) obtained from DESeq2 analysis. Hierarchical clustering was performed using the hcluster method. c is a Manhattan diagram showing the caQTL localization of the DTM1 site in all cells and cell types. d represents the chromatin accessibility profile of the DTM1 promoter region (Chr07:26,911–26,913 kb) in all cells, pseudo-sc-bulk, and various cell types. Grouped by MH63 (blue) and ZS97 (red) genotypes; e is a UMAP diagram showing the DTM1 gene activity scores in MH63 and ZS97 genotype cells. Detailed Implementation

[0040] The present invention will now be described in further detail with reference to specific embodiments, so that those skilled in the art can understand it.

[0041] Example 1 The kit for single-cell chromatin accessibility sequencing was constructed as follows: 1. Oligonucleotide linkers for single-cell chromatin accessibility sequencing The oligonucleotide linker consists of a first chain (Tn5Me-Ax, where x is 1, 2, 3...) and a second chain (Tn5ME-B). Table 1 Oligonucleotide Linker Sequences Note: The universal connector sequence is GACCTCGTGCACCTGGATC; the ME sequence is AGATGTGTATAAGAGACAG. The sequence between the universal connector sequence and the ME sequence is the BC sequence.

[0042] 2. Tn5 transposase complex composition The Tn5 transposase complex composition comprises two or more Tn5 transposase complexes. The Tn5 transposase complex is formed by the assembly of an oligonucleotide linker and the Tn5 transposase.

[0043] The Tn5me-A sequence anneals with the Tn5MErev sequence to form the A-series adapter, and the Tn5ME-B sequence anneals with the Tn5MErev sequence to form the B-series adapter. Then, the different A-series adapters and B-series adapters are mixed and assembled in a 1:1 ratio to form a Tn5 transposase complex containing BC1.

[0044] 3. Single-cell barcoded PCR primer set The single-cell barcoded PCR primer set includes primer pairs for the first round of PCR amplification and primer pairs for the second round of amplification. The primer combination for the first round of PCR amplification was: Forward primer (Ad1) Fx, where x is 1, 2, 3, ...): CCCTACACGACCGTCTCTCCGATCT(N)mGACCTCGTGCACCTGGATC, Reverse primer (Ad2) Rx, where x is 1, 2, 3, ...): CAAGCAGAAGACGGCATACGAGAT[i7]GTCTCGTGGGCTCGGAGATGT; Where (N)m is the sample barcode sequence, m is an integer from 0 to 50, and [i7] is an 8 bp sample index sequence.

[0045] The primer combination for the second round of PCR amplification is as follows: Forward primer (Ad3) Fx, where x is 1, 2, 3, ...): AATGATACGGCGACCACCGAGATCTACAC[i5]ACACTCTTTCCCTACACGACGCTCTCTCCGATCT; Where [i5] is a tag sequence of length 8 bp; The reverse primers are Illumina universal P7 primers.

[0046] Specifically as follows: 4. Based on the above-mentioned Tn5 transposase complex composition and PCR primer set, a kit for single-cell chromatin accessibility sequencing was obtained, which includes... (1) The above-mentioned Tn5 transposase complex composition; (2) Transposition reaction buffer; (3) The above single-cell barcoded PCR primer set; (4) PCR amplification system; (5) DNA purification reagents.

[0047] The transposable reaction buffer contains 50 mM Tris-HCl (pH 7.5), 25 mM MgCl2 and 50% (v / v) DMF.

[0048] PCR amplification system: NEBNext® High Fidelity 2X PCR Master Mix or equivalent.

[0049] The above-described multi-index-based single-cell chromatin accessibility sequencing method (MPS-ATAC-seq) using the above-described kit includes the following steps: (1) Collect target samples and separate cell nuclei, and distribute cell nuclei to multiple reaction tubes through the first round of sorting; (2) Add the Tn5 transposase complex composition from the above kit to each reaction tube to carry out the transposition reaction, thereby labeling the cell nuclei in each tube and thus carrying the first layer of barcode; (3) The transposable cell nuclei are sorted in the second round and assigned to multi-well plates with different primer combinations, so that the cell nuclei in each well obtain a second layer of barcode, realizing dual labeling at the single cell level; (4) The sorted cell nuclei were lysed and the first round of PCR amplification was performed using the kit described above to introduce secondary index sequences; (5) Merge and purify the first-round amplification products, and use the above-mentioned kit to perform a second round of PCR amplification to introduce standard sequencing adapters and obtain a high-throughput sequencing library; (6) Perform high-throughput sequencing on the obtained library, and perform two rounds of splitting, comparison and quality control on the sequencing data according to the barcode information to obtain chromatin accessibility data at single-cell resolution.

[0050] The sorting in steps (1) and (3) above is performed using a flow cytometer; the first round of sorting distributes cell nuclei to two or more reaction tubes, and the multi-well plate is a 96-well plate or a 384-well plate.

[0051] The target samples mentioned above include pollen, leaves, root tips, stem tips or floral organs of plants, or cells or nuclei of other eukaryotes.

[0052] The above step (6) involves two rounds of splitting based on barcode information: the first round is splitting based on the sample index, and the second round is splitting based on the barcode in the tube; the alignment is to align to two parental reference genomes and compare the alignment scores; the quality control includes filtering low alignment quality reads, PCR duplicate reads and reads aligned to organelle genomes, and removing cells that do not meet the standards for Tn5 cut point number, PCR repeat rate or FRiP value.

[0053] Example 2 A method for single-cell chromatin accessibility sequencing of rice pollen using the above-mentioned kit. like Figure 1The method described first involves isolating cell nuclei from plant tissue and then performing a first-round sorting (1st FACS) using flow cytometry to distribute the nuclei into multiple reaction tubes (24 in this example), with approximately 30,000 nuclei per tube. Subsequently, Tn5 transposases with different barcodes are used to transpose the nuclei in each tube, ensuring that each nucleus carries a specific Tn5 barcode. After transposition, the nuclei are sorted a second time using flow cytometry and individually assigned to 384-well plates pre-filled with different primers, ensuring that each well contains multiple cells (24 in this example) with a unique primer combination. Finally, a single-cell library is constructed and high-throughput sequencing is performed, enabling the detection of chromatin accessibility for thousands (9216 in this example) of single cells in a single experiment. The specific steps are as follows: 1. Sample Rice pollen (or young leaves of Arabidopsis thaliana, or mammalian cell nuclei) were used as starting materials.

[0054] 2. Isolation of cell nuclei Fresh pollen was collected, washed with PBS, and lysed using nuclear lysis buffer (10 mM Tris). HCl, 10 mM NaCl, 3 mMMgCl2, 0.1% NP 40) Process, filter to remove debris, and obtain cell nucleus suspension.

[0055] 3. First-round sorting and transposition reaction The prepared cell nuclei were distributed into multiple 1.5 mL centrifuge tubes, and Tn5 with different barcodes was added to each tube. Add 2 μL of the complex to the transposition buffer and incubate at 37 °C for 30 min. Immediately after incubation, remove excess Tn5 enzyme using a DNA purification kit.

[0056] 4. Two-round sorting and single-cell barcoding The second round of sorting was performed using flow cytometry. The processed cell nuclei were sequentially distributed into 384-well plates, each well pre-filled with different combinations of PCR primers mentioned above. Each single cell was ensured to have a double barcode (tube barcode + well barcode).

[0057] 5. Lysis and first round of PCR amplification After lysis, single cells in the plate were directly subjected to the first round of PCR amplification. The amplification system was as follows: a. 5 μL of lysis products b. 2X PCR Mix 25 μL c.Forward Primer 1 μL (10 μM) d.Reverse Primer 1 μL (10 μM) Add e.ddH2O to a final volume of 50 μL. PCR program settings: 72 ℃ for 3 min; 98 ℃ for 30 s; 98 ℃ for 10 s, 63 ℃ for 30 s, 72 ℃ for 1 min, for a total of 11 cycles; 72 ℃ for 5 min, then incubate at 4 ℃. After amplification, purify the product and combine the samples from each well.

[0058] 6. Second round of PCR amplification and adapter introduction A second round of amplification was performed using primers containing Illumina P5 / P7 adapters: the procedure was the same as above, with 5 to 8 cycles; after amplification, the sample was purified with magnetic beads to obtain the final library.

[0059] 7. Sequencing and Data Structure: The final library was sequenced using PE150 sequencing on an Illumina NovaSeq or HiSeq platform. The library structure is as follows: P5--index5--Universal connector A--barcode1--ME--DNA insert--ME--Universal connector B--index7 --P7; Where: P5 and P7 are standard Illumina sequencing adapter sequences; index5 and index7 are replaceable 6-8 bp sample differentiation sequences; barcode1 is the first-round transposon barcode; ME is the Tn5 recognition region; DNA insert is the real genome fragment (…). Figure 2 ).

[0060] First, using Tn5 transposase, sequences containing adapters (Ad1 / Ad2) and barcode 1 (BC1) are inserted into both ends of a genomic DNA (gDNA) fragment, forming a library fragment with ME sequences. Then, a first round of PCR amplification introduces secondary adapters (Ad3) and barcode 2 (BC2). Next, a second round of PCR further introduces sequencing adapters (P5, i5, P7, i7) and barcode 2, completing the construction of the final sequencing library.

[0061] 8. Experimental Results 1. A single experiment can yield results on the order of 10³~10⁻⁶. 4 Single-cell library data ( Figure 1 ); 2. After quality control, adapter removal, and alignment, the raw data, after removing low-quality and duplicate reads, showed a clear nucleosome distribution and TSS enrichment, with high FRiP values. The PCR repeatability indicated sequencing saturation. Figure 3 ad).

[0062] Example 3 Using the kit for single-cell chromatin accessibility sequencing from Example 1, MPS-based sequencing was performed. ATAC The quantitative trait locus localization method of seq includes the following steps: S1. Experimental Materials 1. Sample Source The hybrid rice MH63 and ZS97, SY63, was selected as the experimental material, and pollen from its F2 generation was collected as the experimental subject. Because pollen is a reproductive cell, it has genotypic recombination diversity.

[0063] 2. Sequencing Platform Illumina NovaSeq or HiSeq PE150.

[0064] S2. Library Construction and Sequencing 1. Nucleus preparation: Cell nuclei were isolated from pollen, and debris was removed by filtration to obtain a single-nucleus suspension; Perform barcode-based Tn5 transposation reaction: Cell nuclei were distributed to multiple reaction tubes using a flow cytometer, and different Tn5 enzyme complexes were added to each tube. The tubes were then incubated at 37 °C for 30 min.

[0065] 2. Second sorting and single-core allocation: Using a flow cytometer, transposable single cells were distributed into 384-well plates, with each well pre-filled with PCR primers containing different barcode combinations.

[0066] 3. PCR amplification and library construction: First round of PCR: Introduce secondary index sequences; merge and purify amplification products; Second round of PCR: Introduction of P5 / P7 sequencing adapters; The final library was obtained by purifying with magnetic beads.

[0067] 4. Sequencing data structure: P5--index5--Universal connector A--barcode1—ME--DNA insert--ME--Universal connector B--index7--P7.

[0068] S3. Bioinformatics Analysis Workflow 1. Data preprocessing, i.e., barcode parsing and data splitting: a. Sequence identification is key to distinguishing single-cell data. The raw data from single-cell experiments consisted of 48 FASTQ files (24 READ1 files and 24 READ2 files). b. For each i7 index file, the combination of X_(i5) and Y_(Tn5 library) generates an X×Y barcode sequence (16×24 in this embodiment); c. Since only READ1 contains the barcode sequence, the original data needs to be split in two rounds. The first 35 bases of READ1 are used as sequence identifiers to split the data; d. A two-stage splitting strategy was adopted: First, the fastx_barcode_splitter.pl tool was used to split the READ1 file into 384 independent FASTQ files, each FASTQ file corresponding to a unique index, with the parameter set to "--mismatches 1" to allow one mismatch; then, the repair.sh tool of UMItools was used to extract the corresponding READ2 sequence from each FASTQ file generated in the previous step, combined with 24 original READ2 files, finally obtaining 24×384 (9216) READ1 and 24×384 (9216) READ2 FASTQ files, thus splitting the sequencing data of a single cell; 2. Data comparison and quality control: a. Use the BWA-MEM algorithm to align the data of each split cell to the MH63RS3 and ZS97RS3 genomes respectively, and obtain the sequence score (AS:i value in the BAM file) for each sequence aligned to MH63 and ZS97 respectively. b. Use SAMTools to filter out alignment reads, PCR replicates, and reads aligned to mitochondria and chloroplasts with a MAPQ quality score below 30. c. Quality control: Remove cells with fewer than 1,000 Tn5 cut sites, PCR repeatability of less than 50%, and FRiP value of less than 0.5; 3. Cell clustering and annotation: a. Use the ArchR package for dimensionality reduction and unsupervised clustering; b. Annotate different cell types based on chromatin accessibility patterns of known marker genes; c. Use the getMarkerFeatures function of the ArchR package to further identify cell type-specific open chromatin regions; 4. Reconstruction of the reconstructed map: a. Divide the MH63 genome into 10kb windows. For each window, calculate the alignment score of all sequences within that window that align to MH63 and ZS97. Based on the score, divide the cells in that region into two groups (MH63 group and ZS97 group). Specifically, determine the group by calculating the proportion of sequences with higher alignment scores originating from MH63 among all valid sequences within the window. If this proportion is greater than or equal to 80%, the window is classified as MH63 type; if the proportion of sequences with higher alignment scores originating from ZS97 is greater than or equal to 80%, and if neither reaches 80%, the window is classified as heterozygous or indeterminate. b. Based on the above preliminary determination results, the parental origin status of each single-cell whole genome was corrected using a Hidden Markov Model (HMM) to remove stray and noise signals, obtain the genotype information of each single cell, and finally construct a recombination map. Figure 4 ); 5. Dynamic Grouping BSA Analysis (Dynamic BSA) a. For MH63 and ZS97 cells, chromatin accessibility signals were statistically analyzed within each 10kb window for both groups of cells. b. Use the statistical test method DESeq2 (such as Fisher's exact test) to compare the differences in chromatin accessibility between the two genotype groups and calculate the significance of the differences in each window.

[0069] c. Screen windows with significant differences and identify them as QTL segments that are significantly associated with the molecular phenotype.

[0070] S4. Results Analysis In this embodiment, the Dynamic-BSA method was used to analyze 4,887 pollen cell nuclei, and a total of 15,376 significant quantitative trait loci of chromatin accessibility were identified. Figure 5 ab). A chromatin accessibility region (Chr07:26,912–26,913 kb) located in the DTM1 promoter was found to contain a significant QTL in VC. Figure 5 c). The DTM1 gene plays an important role in the early development of rice pollen; mutations in it reduce the expression of the pectinase gene Osg1, leading to abnormal pollen development. VC cell nuclei were grouped according to the genotype at this locus (MH63 or ZS97).

[0071] The results showed that the chromatin accessibility of the ZS97 genotype at the DTM1 promoter was significantly higher than that of the MH63 genotype. Figure 5 d). UMAP results showed that in the nuclei of VC cells with the ZS97 genotype, the activity level of DTM1 was higher than that of the MH63 genotype ( Figure 5 e).

[0072] To verify the above regulatory relationship, this embodiment further utilized four import lines (L1–L4) with MH63 as the background and carrying different chromosome 8 segments of the ZS97 genome for analysis. qRT-PCR results showed that the expression level of DTM1 in the L1 line (carrying the 2.28–9.65 Mb region on chromosome 8 of ZS97, covering the identified caQTL region) was significantly higher than that in the MH63 background and the other three import lines (L1–L4). Figure 5 h). This demonstrates that the regulatory relationships identified by Dynamic-BSA analysis can effectively identify cell type-specific chromatin accessibility genetic regulatory relationships at the single-cell level, proving that this method can overcome the technical limitations caused by the sparsity of single-cell sequencing data, thereby enabling the analysis of the relationship between genetic variation and cell type-specific chromatin states. Compared with traditional population QTL analysis, this method can complete the localization without large-scale population construction, significantly shortening the experimental cycle.

[0073] Furthermore, this embodiment utilizes single-cell ATAC sequencing data to simultaneously achieve single-cell genotype inference and recombination map construction. By aligning the sequencing sequence of each pollen cell nucleus to two parental reference genomes, and combining sliding window statistics and hidden Markov model correction, parental origin information and recombination breakpoint locations can be accurately obtained at the single-cell level. The results show that even with low single-cell sequencing coverage, this invention can still stably recover the genotype composition and recombination events of haploid pollen, providing a reliable genetic basis for subsequent quantitative trait locus localization.

[0074] This embodiment demonstrates that the present invention can not only achieve chromatin accessibility detection at the single-cell level, but also simultaneously obtain single-cell genotype information, realizing the joint analysis of chromatin accessibility, genotype and cell type information, thereby significantly improving the resolution and efficiency of quantitative trait locus localization.

[0075] Example 4 The method for single-cell level quantitative trait locus localization in rice leaves or pollen using the kit from Example 1 includes the following steps: 1. Nuclear isolation and transposition reaction Take target tissues (e.g., rice leaves or pollen) and prepare mononuclear suspensions; Add 2 μL of transposase complex to every 10,000-50,000 cell nuclei and react at 37°C for 30 min in a transposase buffer system. The free enzyme was purified to obtain the barcode-labeled transposon product.

[0076] 2. Single-cell sorting and PCR labeling Cell nuclei were assigned to 384-well plates using flow cytometry or an automated sorting system. Pre-placed PCR primers in the wells assign a second layer of barcode to each single cell.

[0077] 3. PCR amplification First round of PCR: Introduction of secondary index sequences; Second round of PCR: further amplification and introduction of sequencing adapters.

[0078] 4. Library sequencing The amplified products were purified with magnetic beads and then sequenced using the Illumina HiSeq / NovaSeq platform for PE150 sequencing. The library structure is as follows: P5--index5--Universal connector A--barcode1--ME--DNA insert--ME--Universal connector B--index7--P7; This kit allows for the convenient construction of MPS ATAC seq libraries; it can obtain thousands of single-cell ATAC libraries in a single experiment, with a mapping rate consistently above 90%. By combining the genotype inference and dynamic BSA data analysis methods provided in this invention, single-cell chromatin accessibility information and single-cell genotype information can be obtained simultaneously, and a high-resolution recombination map can be constructed to achieve single-cell-level quantitative trait locus localization. Therefore, this invention realizes an integrated analysis workflow for single-cell chromatin accessibility determination, single-cell genotyping, recombination map construction, and quantitative trait locus localization, with advantages such as high throughput, high resolution, and wide applicability.

[0079] All other parts not described in detail are existing technologies. Although the above embodiments have provided a detailed description of the present invention, they are only some embodiments of the present invention, not all embodiments. People can obtain other embodiments based on these embodiments without creative effort, and these embodiments all fall within the protection scope of the present invention.

Claims

1. An oligonucleotide adapter for single-cell chromatin accessibility sequencing, characterized in that: The oligonucleotide linker comprises a first strand and a second strand: The first chain includes, from the 5' end to the 3' end, a universal connector sequence, a barcode sequence, and an ME sequence in sequence; The second chain is a complementary chain containing the ME sequence, and its length does not exceed that of the first chain; The universal connector sequence has a length of 10-50 bp and a GC content of 25%-75%. The GC content in the barcode sequence is 25%~75%, and the barcode sequence length is 4~20bp.

2. The oligonucleotide linker according to claim 1, characterized in that: The universal connector sequence is GACCTCGTGCACCTGGATC; The ME sequence is AGATGTGTATAAGAGACAG, or AGATGTGTATAAGAGACA, or its functionally equivalent sequence.

3. The oligonucleotide linker according to claim 1, characterized in that: The expression for the first chain is: GACCTCGTGCACCTGGATC<>AGATGTGTATAAGAGACAG; Among them, GACCTCGTGCACCTGGATC is a universal connector sequence; <> represents a barcode sequence, which is composed of n N's, where n is an integer from 4 to 20, and each N is independently selected from A, T, C, and G; AGATGTGTATAAGAGACAG represents an ME sequence.

4. A Tn5 transposase complex, characterized in that: The complex is formed by assembling an oligonucleotide linker with Tn5 transposase as described in any one of claims 1 to 3.

5. A Tn5 transposase complex composition, characterized in that, The composition comprises two or more Tn5 transposase complexes as described in claim 4, wherein the barcode sequences of the oligonucleotide linkers contained in each complex are different from each other.

6. A single-cell barcoded PCR primer set, characterized in that, This includes the primer combination for the first round of PCR amplification, which includes Ad1 series primers and Ad2 series primers: (1) The general formula of the Ad1 series primers is: CCCTACACGACCGTCTCTCCGATCT(N)mGACCTCGTGCACCTGGATC; Among them, CCCTACACGACGCTCTTCCGATCT is the P5 end connector sequence. (N)m is a barcode sequence, where m is an integer from 0 to 50. GACCTCGTGCACCTGGATC is a universal connector sequence; (2) Ad2 series primers, their general formula is: CAAGCAGAAGACGGCATACGAGAT[i7]GTCTCGTGGGCTCGGAGATGT; Among them, GTCTCGTGGGCTCGGAGATGT is the universal connector sequence 1. [i7] is a tag sequence of length 8 bp.

7. The primer set according to claim 6, characterized in that: It also includes primer pairs for the second round of amplification, the primer pairs being: Forward primer: AATGATACGGCGACCACCGAGATCTACAC[i5]ACACTCTTTCCCTACACGACCGCTTCCGATCT Where [i5] is a tag sequence of length 8 bp; The reverse primers are Illumina universal P7 primers.

8. A kit for single-cell chromatin accessibility sequencing, characterized in that, include: (1) The Tn5 transposase complex composition according to claim 5; (2) Transposition reaction buffer; (3) The single-cell barcoded PCR primer set as described in claim 6 or 7; (4) PCR amplification system; (5) DNA purification reagents.

9. A method for constructing a single-cell chromatin accessibility sequencing library, characterized in that, The method is constructed using a kit and includes the following steps: (1) Collect target samples and separate cell nuclei, and distribute cell nuclei to multiple reaction tubes through the first round of sorting; (2) The Tn5 transposase complex composition of claim 5 is added to each reaction tube to carry out a transposition reaction, thereby labeling the cell nuclei in each tube and thus carrying the first layer of barcode; (3) The transposable cell nuclei are sorted in a second round and assigned to a multi-well plate with a pre-set single-cell barcoded PCR primer set as described in claim 6 or 7, so that the cell nuclei in each well obtain a second layer of barcode and achieve dual labeling at the single-cell level. (4) The sorted cell nuclei were lysed and subjected to the first round of PCR amplification to introduce secondary index sequences; (5) Merge and purify the first round of amplification products, perform a second round of PCR amplification to introduce P5 / P7 standard sequencing adapters, and obtain a high-throughput sequencing library.

10. The use of the Tn5 transposase complex of claim 4, the composition of claim 5, the single-cell barcoded PCR primer set of claim 6 or 7, or the kit of claim 8 in the detection of chromatin accessibility at the single-cell level.