An ultra-low input single cell nucleus chromatin accessibility sequencing method and application thereof

The ULI-snATAC-seq method solves the problem of inefficient analysis of preimplantation embryonic samples in traditional methods, realizes high-resolution chromatin accessibility sequencing at the single-cell level, constructs a panoramic map of mouse preimplantation development, and reveals the dynamic changes of transcription factor networks and transposable DNA sequence elements.

CN122235285APending Publication Date: 2026-06-19FUJIAN FUYAO UNIVERSITY OF SCIENCE & TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
FUJIAN FUYAO UNIVERSITY OF SCIENCE & TECHNOLOGY
Filing Date
2026-05-08
Publication Date
2026-06-19

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

This invention discloses a method for sequencing the accessibility of single-cell nuclear chromatin at ultra-low starting amounts and its applications. The method includes: transferring single cells or single-cell nuclei into microdroplets of a labeling mixture for Tn5 transposon labeling; covering the microdroplets with mineral oil to prevent evaporation; terminating the reaction and lysing the cells after labeling; adding a PCR premix for library fragment amplification after quenching; obtaining a sequencing library after purification and quantification; and performing high-throughput sequencing. This invention, through a micromanipulation-based oil-phase-covered microdroplet reaction system, minimizes sample loss and enables the analysis of chromatin accessibility at the single-cell nuclear level for rare, minute samples. This invention also discloses the application of this method in the construction of a single-nuclear chromatin accessibility map for mouse preimplantation embryonic development.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the technical field of single-cell sequencing methods, specifically relating to a single-cell nuclear chromatin accessibility sequencing method with ultra-low starting amount and its application. Background Technology

[0002] The life journey of mammals begins with a remarkable transformation: the terminally differentiated gamete transforms into a totipotent embryo capable of developing into a fully formed individual. This fundamental stage of preimplantation development is regulated by a series of meticulously coordinated epigenetic and transcriptional reprogramming events. A key milestone in this process is zygotic genome activation (ZGA), in which transcriptional control is transferred from maternally deposited factors to the newly formed embryonic genome. The timing and regulation of ZGA are species-specific; in mice, a small number of ZGA events occur in the single-cell stage, but the major wave of transcriptional activation occurs in the two-cell stage. This event is not merely a switch, but a comprehensive reorganization of the embryonic epigenome that lays the foundation for subsequent cell fate determination, including the initial differentiation of cells into the inner cell mass (ICM) and trophectoderm (TE) during the blastocyst stage.

[0003] Activation of the zygotic genome is regulated by the interaction of trans-regulatory factors, such as transcription factors (TFs), with cis-DNA sequence elements in the genome. While some transcription factors, including the Zscan 4 family, Obox family, Dux, and Nr5a2, have been identified as key regulators of zygotic genome activation (ZGA), they are likely just the tip of the iceberg. It is hypothesized that a broader, potentially redundant, network of transcription factors works synergistically to ensure the robustness of this process. Equally important are the cis-regulatory sequences to which these transcription factors bind. Notably, transposable DNA sequence elements (TEs), long considered "genomic junk," are now recognized as a significant source of evolutionary innovation. Transposable DNA sequence elements (TEs), such as endogenous retroviruses (ERVs) and short, scattered nuclear DNA sequence elements (SINEs), can contain multiple transcription factor (TF) binding sites and function as species-specific enhancers or promoters, potentially playing a central role in regulating large-scale transcriptional transitions during preimplantation genetic development. Therefore, systematically cataloging all zygotic genome activation (ZGA)-specific transcription factors and clarifying the role of specific repetitive DNA sequence elements are key steps in deciphering the molecular blueprint of the origin of life.

[0004] High-throughput sequencing technology is undoubtedly the most effective method for comprehensively and unbiasedly analyzing this regulatory network. However, its application is severely limited due to the extremely limited number of cells in preimplantation embryos. To overcome this limitation, previous studies have optimized low-starting-volume chromatin accessibility analysis methods (e.g., liDNase-seq and miniATAC-seq) for use in preimplantation embryos. These advances have made it possible to sequence dozens of blastomeres, overcoming the need for tens of thousands of cells and elucidating the dynamic changes in chromatin during early development. Among these, chromatin accessibility sequencing (ATAC-seq) has become the preferred tool due to its sensitivity and simplicity, enabling genome-wide localization of regulatory DNA sequence elements and transcription factor binding sites (TFBS). In recent years, single-cell ATAC sequencing (scATAC-seq) technology has made significant progress in resolving cellular heterogeneity and identifying cell type-specific transcription factor binding. However, current scATAC-seq methods, like batch analyses, still require hundreds of thousands of cells, which limits their application in preimplantation embryos. Therefore, developing strategies to analyze the chromatin accessibility of individual blastomeres is crucial for elucidating transposable DNA sequence elements and transcription factor networks during preimplantation development. Summary of the Invention

[0005] The purpose of this invention is to provide a method for sequencing the accessibility of single-cell nuclear chromatin with ultra-low starting amounts and its applications. The term "single-cell nucleus" is used broadly, applicable to both intact single cells and isolated single-cell nuclei—two rare and minute sample types.

[0006] To achieve the above objectives, the present invention adopts the following technical solution:

[0007] A method for sequencing the accessibility of single-cell nuclear chromatin with ultra-low starting amounts includes the following steps:

[0008] (1) The single cell or single cell nucleus sample to be tested is transferred to a labeling mixture droplet for Tn5 transposition labeling reaction, wherein the labeling mixture droplet is covered with mineral oil to prevent evaporation;

[0009] (2) After the labeling reaction is complete, add stop buffer to terminate the reaction and place on ice;

[0010] (3) Pick a single cell or cell nucleus and transfer it to a PCR tube containing lysis buffer and sequencing index primers for lysis to release the bound Tn5 transposase;

[0011] (4) Add SDS to the Tween-20 quenching lysis buffer;

[0012] (5) Add PCR premix to amplify the library fragments;

[0013] (6) Purify and quantify the PCR amplification products to obtain sequencing libraries;

[0014] (7) Perform high-throughput sequencing on the sequencing library to obtain chromatin accessibility sequencing data of the sample to be tested.

[0015] Further, in step (1), the formulation of the labeled mixture droplets is: 33mM Tris-acetic acid with a pH of 7.8, 66mM potassium acetate, 10mM magnesium acetate, 16vol% dimethylformamide, 0.01wt% digitalis saponin; 25vol% Tn5 transposase.

[0016] Further, in step (1), the labeling reaction conditions are as follows: the labeling reaction is carried out in a 37°C incubator for 30 minutes, and then the sample is transferred to a new labeling mixture droplet and the reaction is continued at 37°C for 30 minutes; wherein each 0.1 μL droplet contains no more than 20 cells or cell nuclei for the reaction.

[0017] Furthermore, in step (2), the formulation of the termination buffer is: 10mM Tris-HCl with a pH of 8.0 and 20mM EDTA; the conditions for terminating the reaction are: placing on ice for 10 minutes.

[0018] Furthermore, in step (3), the lysis buffer formulation is as follows: 100mM Tris-HCl, pH 8.0; 100mM NaCl; 40μg / mL proteinase K; 0.4wt% SDS; the lysis conditions are: incubation at 65℃ for 15 minutes.

[0019] Furthermore, in step (5), the PCR amplification conditions are: 72℃ for 5 minutes; 98℃ for 5 minutes; followed by 18 cycles of amplification, each cycle being 98℃ for 10 seconds, 63℃ for 30 seconds, and 72℃ for 20 seconds; after amplification, the temperature is maintained at 4℃.

[0020] Furthermore, in step (7), the high-throughput sequencing uses the Illumina HiSeq X sequencing platform.

[0021] Furthermore, the single-cell or single-nucleus sample is derived from mouse preimplantation embryos, mouse oocytes, or mouse embryonic stem cells.

[0022] Furthermore, the single-cell nucleus sample is prepared by staining oocytes, zygotes, 2-cell embryos, or 4-cell embryos at the MI and MII stages with Hoechst 33342, and then using a micromanipulator to aspirate and separate the nuclei.

[0023] The above-mentioned method for sequencing single-cell nuclear chromatin accessibility with ultra-low starting amounts is applied in the construction of single-cell chromatin accessibility maps.

[0024] The beneficial effects of this invention are as follows:

[0025] The ultra-low starting amount single-cell nuclear chromatin accessibility sequencing method provided by this invention overcomes the technical deficiency of traditional methods that cannot perform single-nucleus level chromatin accessibility analysis on rare and trace samples. For the first time, it realizes independent chromatin accessibility analysis of single maternal and paternal pronuclei and single blastomeres during mouse preimplantation development, and successfully constructs a high-resolution single-nucleus chromatin accessibility map covering all developmental stages from oocyte to blastocyst. Attached Figure Description

[0026] Figure 1 a) is a schematic diagram of the experimental procedure for the ultra-low starting amount single-cell nuclear chromatin accessibility sequencing method (ULI-snATAC-seq) of this invention; b) is a graph showing the distribution of insert fragment size in the single-cell library constructed by ULI-snATAC-seq; c) is a graph showing the distribution of enriched signals around the transcription start site (TSS) in ULI-snATAC-seq sequencing data; d) is a heatmap showing the correlation analysis between ULI-snATAC-seq aggregated pseudo-batch data and traditional batch ATAC-seq data.

[0027] Figure 2 a) Comparison of the number of open chromatin peaks detected by ULI-snATAC-seq and traditional batch ATAC-seq; b) Visualization comparison of chromatin accessibility signals at the Nanog and Slc2a3 loci of pluripotency genes by ULI-snATAC-seq and traditional batch ATAC-seq.

[0028] Figure 3 : Schematic diagram of the experimental design for constructing a mononuclear chromatin accessibility map of the entire preimplantation embryonic development of mice using the ULI-snATAC-seq method of this invention.

[0029] Figure 4 Figure a shows the results of UMAP dimensionality reduction clustering analysis of chromatin accessibility data of mouse preimplantation embryos at various stages of embryonic development; Figure b shows the results of pseudo-time trajectory reconstruction of mouse preimplantation embryonic development based on chromatin accessibility data.

[0030] Figure 5 Heatmap of chromatin accessibility signals of specific marker genes at different stages of mouse preimplantation embryonic development.

[0031] Figure 6 : Peak diagram of dynamic changes in chromatin accessibility of key developmental genes during mouse preimplantation embryonic development.

[0032] Figure 7 Correlation analysis of the accessibility of key transcription factor motifs and the expression levels of corresponding genes during preimplantation embryonic development in mice.

[0033] Figure 8 Figure a shows the statistical results of the proportion of mitochondrial-derived reads in a single-cell library of ULI-snATAC-seq; Figures b-d show the results of the correlation analysis of sequencing data between repetitive samples using ULI-snATAC-seq technology; Figure e shows the statistical results of the overlap ratio of chromatin open peaks detected by ULI-snATAC-seq and traditional batch ATAC-seq.

[0034] Figure 9 a~d are quality control results of the mouse preimplantation embryonic development mononuclear chromatin accessibility dataset; e is a UMAP clustering stage labeling diagram of cell nuclei at each stage of mouse preimplantation embryonic development; f is a diagram showing the differentiation results of maternal pronuclei (MPN) and paternal pronuclei (PPN) in zygotes based on sex chromosome reads.

[0035] Figure 10 a) UMAP clustering results distinguishing the inner cell mass (ICM) and trophectoderm (TE) cells during the blastocyst stage; b) Signal distribution heatmap of specific chromatin open peaks at various stages of mouse preimplantation embryonic development.

[0036] Figure 11 Figures a to b show the correlation analysis results between the ULI-snATAC-seq aggregated pseudo-batch data and the published batch ATAC-seq and DNase-seq datasets. Detailed Implementation

[0037] The present invention will be further described in detail below through specific implementation examples and accompanying drawings. It should be understood that these embodiments are only used to illustrate the present invention and are not intended to limit the scope of protection of the present invention. After reading the present invention, any modifications of the present invention in various equivalent forms by those skilled in the art fall within the scope of the appended claims.

[0038] Unless otherwise specified, all raw materials and reagents used in this invention are from the conventional market.

[0039] Example 1: Establishment of an Ultra-Low-Input Single-Nucleus ATAC-sequencing (ULI-snATAC-seq) method for accessibility sequencing of single-nuclei chromatin.

[0040] 1.1 Laboratory Animals and Ethical Statement

[0041] Female C57BL / 6N mice (4–6 weeks old) and male DBA / 2N mice (8 weeks old) were purchased from Beijing Vital River Laboratory Animal Technology Co., Ltd. All animal experiments were conducted in accordance with guidelines approved by the Laboratory Animal Ethics Committee (IACUC) of the Guangzhou Institutes of Biomedicine and Health, Chinese Academy of Sciences.

[0042] 1.2 Preparation of Experimental Samples

[0043] 1.2.1 Preparation of mouse preimplantation embryo and oocyte samples

[0044] (1) Superovulation induction: Superovulation was induced in 4-6 week old female C57BL / 6N mice by intraperitoneal injection of 7.5 IU pregnant mare serum gonadotropin (PMSG, 3SBio), followed by intraperitoneal injection of 7.5 IU human chorionic gonadotropin (hCG, 3SBio) 44-48 hours later.

[0045] (2) MII stage oocyte collection: 13-14 hours after hCG injection, female C57BL / 6N mice induced by superovulation were sacrificed, and mature MII stage oocytes were collected from the ampulla of their oviducts.

[0046] (3) Preimplantation embryo collection at each stage: Female C57BL / 6N mice that have undergone superovulation induction treatment were mated with 8-week-old DBA / 2N male mice. After injection of hCG, embryos at the corresponding stages were collected at the following time intervals: zygote (18-20 hours), early 2-cell stage (29-30 hours), late 2-cell stage (39-43 hours), 4-cell stage (54-56 hours), 8-cell stage (68-70 hours), morula (76-80 hours), and blastocyst (92-95 hours).

[0047] (4) GV phase oocyte collection: 42 hours after PMSG injection, female C57BL / 6N mice induced by superovulation were sacrificed, and fully mature germinal follicle (GV) oocytes were collected by puncturing the ovaries with a 30-gauge needle. They were then stored in M2 medium (Millipore) containing 0.2 mM3-isobutyl-1-methylxanthine (IBMX, Sigma-Aldrich) to prevent germinal follicle rupture (GVBD).

[0048] (5) Preparation of MI stage oocytes: GV stage oocytes were transferred to M16 medium (Sigma-Aldrich) without IBMX and cultured at 37℃ and 5% CO2 for 2 hours to obtain MI stage oocytes.

[0049] 1.2.2 Preparation of single-cell / single-nucleus samples

[0050] (1) Nuclear separation: Oocytes, zygotes, 2-cell and 4-cell embryos at the MI and MII stages were stained with 1 μg / mL Hoechst 33342 at 37℃ for 10-15 minutes, washed with 1×PBS (pH 7.2-7.4), and the nuclei were separated by micromanipulation. The male and female pronuclei of the zygote were distinguished according to their size and distance from the polar body.

[0051] (2) GV stage oocyte treatment: GV stage oocytes were treated with acidic Tyrode's solution (pH 2.2~2.5, Sigma-Aldrich, catalog number T1788) for 30~90 seconds to remove the zona pellucida. After washing with 1×PBS (pH 7.2~7.4), the whole oocyte was used for subsequent reactions.

[0052] (3) Sample processing for 8-cell stage and morula stage: The zona pellucida was removed by treatment with acidic Tyrode's solution (pH 2.2~2.5, Sigma-Aldrich, catalog number T1788) for 30~90 seconds. After washing with 1×PBS (pH 7.2~7.4), the blastomeres were mechanically separated by gentle blowing and aspiration to obtain single-cell samples.

[0053] (4) Preparation of single cells during the blastocyst stage: The zona pellucida was removed by treatment with acidic Tyrode's solution (pH 2.2~2.5, Sigma-Aldrich, catalog number T1788) for 30~90 seconds. After washing with 1×PBS (pH 7.2~7.4), the cells were digested with TrypLE stock solution (Gibco, catalog number 12604013) at 37°C for 3~5 minutes. The single cell suspension was obtained by gentle aspiration.

[0054] (5) Preparation of mouse embryonic stem cell (mESC) samples: mESCs prepared in our laboratory were cultured on a feeder layer consisting of mouse embryonic fibroblasts (MEFs) inactivated by mitomycin C or gamma rays, at a seeding density of 8 × 10⁻⁶. 5 ~1.5×10 6 Cells per 100mm culture dish; mESC medium contained 83% high-glucose DMEM (Hyclone), 15% fetal bovine serum (FBS, Gibco), GlutaMAX (100×, Gibco), non-essential amino acids (NEAA, 100×, Gibco), 0.1mM β-mercaptoethanol (Sigma-Aldrich), 3μM CHIR99021 (Stemgent), 1μM PD0325901 (Stemgent), and 1000U / mL leukemia inhibitory factor (LIF, Millipore); single-cell suspensions were prepared after cell collection for subsequent experiments.

[0055] 1.3 Preparation of ULI-snATAC-seq libraries

[0056] Preparation of ULI-snATAC-seq libraries for mESCs, oocytes, and early mouse embryos ( Figure 1 a) The specific steps are as follows:

[0057] (1) Preparation of labeled droplets: Prepare a labeling mixture with the following components: 33 mM Tris-acetic acid (pH 7.8), 66 mM potassium acetate, 10 mM magnesium acetate, 16 vol% dimethylformamide, 0.01 wt% digitalis saponin, and 25 vol% pre-assembled Tn5 transposons (commercial Tn5 transposase containing Tn5 protein and transposon DNA linker, from Vazyme TruePrep DNA Library Preparation Kit V2 for Illumina, catalog number TD501-02); prepare the labeling mixture into droplets, place them in a 60 mm culture dish, and cover them with embryo culture grade mineral oil (Sigma-Aldrich, catalog number M8410) to prevent evaporation.

[0058] (2) Tn5 transposable labeling reaction: The prepared cell or nucleus sample was transferred to the above microdroplet using a micromanipulator and the labeling reaction was carried out in a 37°C incubator for 30 minutes. Then the sample was transferred to a new labeling mixture microdroplet and the reaction was continued at 37°C for 30 minutes. Each 0.1 μL microdroplet can hold about 20 cells / nuclei for the reaction.

[0059] (3) Reaction termination: Add an equal volume of termination buffer (10mM Tris-HCl, pH 8.0; 20mM EDTA) to the reaction system to terminate the reaction and place on ice for 10 minutes.

[0060] (4) Single sample sorting and lysis: Use a hand pipette to pick up a single cell or cell nucleus and transfer it to a PCR tube. The tube is pre-filled with 2 μL of 2× lysis buffer (100 mM Tris-HCl, pH 8.0; 100 mM NaCl; 40 μg / mL proteinase K; 0.4 wt% SDS), 1.2 μL of enzyme-free water (Takara, catalog number 9012), and 0.8 μL of 10 μM Nextera Index primer mixture (0.4 μL each of forward and reverse primers, Illumina, catalog number FC-131-1001 / 1002, containing i5 Index Primer and i7 Index Primer). Incubate the PCR tube at 65°C for 15 minutes to release the bound Tn5 transposase.

[0061] (5) SDS quenching: Add 4 μL of 10 vol% Tween-20 to the above system and mix thoroughly to quench SDS.

[0062] (6) PCR amplification and enrichment: Mix the above mixture thoroughly with 10 μL of PCR premix (NEB, catalog number M0541L) and 2 μL of enzyme-free water to amplify the library fragments. The PCR reaction conditions are: 72℃ for 5 minutes; 98℃ for 5 minutes; followed by 18 cycles of amplification (98℃ for 10 seconds, 63℃ for 30 seconds, 72℃ for 20 seconds); after amplification, incubate at 4℃.

[0063] (7) Library purification and quantification: PCR amplification products were purified using VAHTS DNA Clean Beads (catalog number N411-02), and the purified library was quantified using a quantum dot quantitative kit (Vazyme Biotech, catalog number EQ111-01).

[0064] 1.4 Sequencing and Data Processing

[0065] 1.4.1 Sequencing

[0066] Equal volumes of quantitatively qualified sample libraries were mixed and sequenced using the Illumina HiSeq X sequencing platform (Annord Gene Technology Co., Ltd.).

[0067] 1.4.2 ULI-snATAC-seq Data Processing

[0068] (1) Raw data preprocessing: The raw sequencing reads were preprocessed using fastp (version 0.20.1) to filter low-quality reads and remove sequencing adapters.

[0069] (2) Genome alignment: High-quality clean reads were aligned to the mm10 reference genome using HISAT2 (version 2.2.1). The alignment parameters were set as follows: -X 2000 --very-sensitive --no-spliced-alignment --no-mixed --no-discordant.

[0070] (3) Data filtering: Remove all unaligned reads, non-unique aligned reads, mitochondrial-derived reads, and PCR repetitive sequences.

[0071] (4) Peak identification: MACS2 (version 2.2.6) was used to identify open regions (peaks) of chromatin with the parameters set to: --nomodel --shift -100 --extsize 200, while excluding blacklisted regions of the genome; the fragments command (sinto fragments -b input.bam -f fragments.tsv) of sinto (version 0.8.0) was used to adjust the BAM alignment interval of sequencing read pairs to obtain the BED interval of the fragments.

[0072] (5) Quality control: Generate quality indicators, including alignment rate, number of unique fragments, proportion of peak region reads, number of mitochondrial genome reads and number of sex chromosome reads, through custom Python scripts to complete sample quality control.

[0073] (6) Pseudobulk sample construction: By merging the bam files of the same sample group, Pseudobulk samples of each developmental stage or cell type are obtained; peak identification is performed using MACS2 (version 2.2.6, parameters are the same as in step (4)) to exclude blacklisted areas.

[0074] (7) Signal normalization and trajectory generation: Using the bamCoverage function in deeptools (version 3.5.4), the number of reads per thousand bases per million reads (RPKM) is calculated with a bin size of 100 bp, the read count is normalized, and a signal trajectory file is generated.

[0075] (8) Counting matrix construction: Use the coverageBed function in bedtools (version 2.26.0) to count the number of reads from a single cell that overlap with the union peaks and generate a counting matrix of the peak union.

[0076] (9) Analysis of repetitive DNA sequence elements: alignment reads at repetitive sites were quantified using featureCounts (version 2.0.1), with the repeatMasker-defined trajectories downloaded from the UCSC Table Browser as GTF files; read counts were normalized by calculating RPKM values.

[0077] 1.4.3 Unified Processing Flow for Accompanying Sequencing Data

[0078] For batch ATAC-seq, DNase-seq, and ChIP-seq data, the same processing flow as steps (1) to (4) and (7) to (8) in Section 1.4.2 is used. For genome alignment, HISAT2 (version 2.2.1) is used with the same parameters: -X2000 --very-sensitive --no-spliced-alignment --no-mixed --no-discordant; for peak identification, MACS2 (version 2.2.6, parameters: --nomodel --shift -100 --extsize 200) is used and blacklisted regions are excluded.

[0079] The processing workflow for batch and single-cell RNA sequencing data is as follows:

[0080] (1) Raw data preprocessing: The raw sequencing reads were preprocessed using fastp (version 0.20.1) to filter low-quality sequences and remove adapters.

[0081] (2) Genome alignment and gene expression quantification: cleanreads were aligned to the mm10 genome using STAR (version 2.7.10b) with default parameters; alignment reads on annotated genes (ncbiRefSeq.gtf) were quantified using FeatureCounts (version 2.0.1) from the Subread package.

[0082] (3) Quantification of repetitive DNA sequence elements: High-quality reads were aligned to the mm10 genome using STAR (version 2.7.10b) with the following alignment parameters: --twopassMode Basic --outFilterType BySJout --outFilterMultimapNmax 1 --winAnchorMultimapNmax 50 --chimSegmentMin 12 --chimJunctionOverhangMin 8 --alignSJoverhangMin 8 --alignSJDBoverhangMin 10 --outFilterMismatchNmax 999 --outFilterMismatchNoverReadLmax 0.04 --alignSJstitchMismatchNmax 5 -1 5 5 --outSAMattrRGline ID:GRPundef --alignIntronMin 20 --alignIntronMax 1000000 --alignMatesGapMax 1000000; The aligned reads at the repetitive sites were quantified using featureCounts. RepeatMasker-defined orbitals were extracted from the UCSC Table in GTF format. Browser download.

[0083] (4) Data normalization: The read count is normalized by calculating the RPKM value for downstream analysis.

[0084] 1.4.4 Downstream Bioinformatics Analysis Methods

[0085] (1) ATAC-seq insert size and TSS enrichment analysis: The insert size was extracted from the 9th column of the bam file, and the read distribution for each unique length was statistically analyzed; In the TSS enrichment analysis, the makeTagDirectory function in homer (version 2.26.0) was first used to convert the bam file into a tag directory, and then the annotatePeaks.pl function in homer was used to align and annotate the transcription start sites (TSS) defined in the mm10 genome with the parameters -size -1500,1500 -hist 10 -norm 1e6 -fragLength 1; The matrices generated by both analyses were processed and visualized using a custom R script.

[0086] (2) Peak annotation: The peak coordinates were annotated based on mm10 genome information using the annotatePeaks.pl program in the Homer package (version 2.26.0); peaks less than 2.5kb from TSS were defined as promoter peaks, and the rest were defined as distal peaks.

[0087] (3) Dimensionality reduction and cluster analysis of single-cell level data: merge the bam files of all single cells and generate reference peak files based on peak calls; score the insertion events in the reference peaks of each cell to generate a binary matrix of single cell × reference peak; based on this matrix and fragment files, use the R packages Signac (version 1.12.0) and ArchR (version 1.0.3) for subsequent analysis; use term frequency-inverse document frequency (TF-IDF) transformation to reduce the dimension of the binary matrix, select the top 25% of peaks to perform singular value decomposition (SVD) analysis on the TF-IDF normalized matrix to obtain a 50-dimensional matrix; remove the first principal component affected by sequencing depth, and use the 2nd to 25th principal components for downstream uniform manifold approximation and projection (UMAP) and cluster analysis.

[0088] (4) Transcription factor motif enrichment analysis (batch ATAC-seq data): For batch ATAC-seq data, motif enrichment analysis was performed using the findMotifsGenome.pl function in the homer suite (version 2.26.0).

[0089] Example 2: Performance Testing of the ULI-snATAC-seq Method

[0090] This embodiment uses mouse embryonic stem cells (mESCs) to validate the performance of the ULI-snATAC-seq method established in Example 1. The minimum number of input cell nuclei is only 20. A traditional batch ATAC-seq method is also used as a control. Details are as follows:

[0091] 2.1 Preparation of batch ATAC-seq libraries

[0092] The preparation process for mESC batch ATAC libraries is as follows:

[0093] (1) Increase the volume of the standard labeling reaction system from 0.1 μL to 50 μL (that is, increase the proportion of each component of the labeling mixture described in Example 1 by 500 times) and resuspend approximately 30,000 to 50,000 mESC cell nuclei;

[0094] (2) After reacting at 37℃ for 30 minutes, add 50 μL of stop buffer (10 mM Tris-HCl, pH 8.0; 20 mM EDTA) and place on ice for 10 minutes;

[0095] (3) Add 100 μL of reverse cross-linking buffer (100 mM Tris-HCl, 100 mM NaCl, 0.4 wt% SDS, 40 μg / mL proteinase K) to the mixture and react at 65 °C for 30 minutes.

[0096] (4) Add 200 μL of Tween-20 to neutralize SDS, then add 100 μL of enzyme-free water, and purify the template using VAHTS DNA CleanBeads (catalog number N411-02);

[0097] (5) Using the purified product as a template, PCR enrichment was performed using the TruePrep DNA Library Preparation Kit V2 for Illumina (Vazyme Biotech, catalog number TD501-02). The reaction system and reaction conditions were performed according to the kit instructions.

[0098] (6) After PCR library purification, sequencing was performed. The sequencing and data processing procedures were the same as in Example 1.

[0099] 2.2 Performance Verification Results

[0100] (1) Library quality: The ULI-snATAC-seq method generated high-quality libraries for each cell nucleus. The library characteristics were: the total alignment rate of the vast majority of cells was higher than 90%, and there were a reasonable number of unique alignment fragments; the size distribution of the inserted fragments showed a significant nucleosome periodicity. Figure 1 b) There is significant enrichment of open chromatin signals around the transcription start site (TSS). Figure 1 c), and the proportion of mitochondrial reads is generally below 10% (most cells are in the 3%~10% range). Figure 8 It meets the quality control standards for high-quality ATAC-seq libraries.

[0101] (2) Data Consistency: The pseudo-batch sequencing results generated by aggregating single-cell data from ULI-snATAC-seq showed high consistency with traditional batch ATAC-seq data: the Pearson correlation coefficients between the pseudo-batch data generated from aggregating different numbers of single cells and the batch results of this study were 0.84 (aggregates of 20 cells) to 0.92 (aggregates of all single cells). Figure 1 d、 Figure 8 b); The correlation coefficients between the pseudo-batch data and multiple publicly available batch ATAC-seq data in this study are 0.85~0.89 ( Figure 8 d); The correlation between pseudo-batch data aggregated with different numbers of single cells (10-200 cells) significantly increased with increasing cell number, with the correlation coefficient rising from 0.79 (10 cells) to 0.99 (200 cells). Figure 8(c) This confirms that stable data consistency can be maintained even under low cell count conditions. These results collectively demonstrate that the ULI-snATAC-seq method possesses excellent reproducibility and data fidelity.

[0102] (3) Detection sensitivity: Under the condition of an ultra-low input volume of only 1 / 1000 of that of traditional batch ATAC-seq, the ULI-snATAC-seq method exhibits excellent detection sensitivity: the peak overlap rate of its pseudo-batch results with that of this study and several publicly available batch ATAC-seq datasets can reach 78.13%~94.23%, of which the overlap rate with the PRJNA525408 dataset is as high as 94.23%, successfully capturing the vast majority of accessible chromatin regions detected by traditional batch methods ( Figure 8 e); The dynamics of chromatin opening at each stage from oocyte to blastocyst were clearly elucidated, including site changes in key lineage-specific genes (such as the Obox family, Pou5f1, Nanog, Cdx2, etc.), validating the method's efficient detection capability for functionally open chromatin in low-starting-volume samples. Figure 6 In summary, ULI-snATAC-seq exhibits significantly higher detection sensitivity than traditional batch ATAC-seq methods, making it particularly suitable for research scenarios involving rare embryo samples.

[0103] (4) Single-molecule resolution: Visualization analysis of representative pluripotency gene loci such as Nanog and Slc2a3 showed that the chromatin accessibility spectra of the pseudo-batch data of ULI-snATAC-seq were highly consistent with those of the traditional batch ATAC-seq data, and the positions and intensity distributions of the open peaks almost completely overlapped; at the same time, the distribution of reads at the single-cell level could clearly show the variation of chromatin open state among different cells, realizing the high-resolution cellular heterogeneity analysis that traditional batch methods could not provide. Figure 2 b).

[0104] The above results confirm that the ULI-snATAC-seq method established in this invention can achieve high-resolution epigenomics studies at the single-cell / single-nucleus level and is suitable for rare and minute samples that cannot be processed by traditional methods.

[0105] Example 3: Application of ULI-snATAC-seq in the construction of mononuclear chromatin accessibility maps in mouse preimplantation embryonic development

[0106] This embodiment uses the ULI-snATAC-seq method established in Example 1 to construct a mononuclear chromatin accessibility map covering all key stages of mouse preimplantation development (from oocyte to blastocyst), as detailed below:

[0107] 3.1 Experimental Design

[0108] Samples were collected from mice at all stages of preimplantation development, including: GV stage oocytes, MI stage oocytes, MII stage oocytes, zygotes (containing paternal pronuclear PPN and maternal pronuclear MPN), early 2-cell stage, late 2-cell stage, 4-cell stage, 8-cell stage, morula, and blastocysts (inner cell mass ICM and trophectoderm TE). A ULI-snATAC-seq library was prepared and sequenced using the method described in Example 1. A total of 3185 high-quality cell nuclei were analyzed. Figure 3 ).

[0109] 3.2 Map Construction and Analysis Results

[0110] All samples underwent rigorous quality control, and were validated using multiple indicators including peak distribution, sequencing alignment rate, number of unique fragments, FRIP value, insert periodicity, and TSS enrichment. This confirmed the reliability of the ULI-snATAC-seq library and ensured the technical stability and biological consistency of the dataset. Figure 9 a- Figure 9 d).

[0111] (1) Developmental stage-specific clustering: UMAP was used to perform dimensionality reduction and clustering analysis on open chromatin data of single cells ( Figure 4 a) The results showed that the cell nuclei were significantly separated according to developmental stages, and the cell clusters were arranged in an orderly manner along the embryonic development timeline. The pseudo-time trajectory clearly reflected the gradual remodeling process of chromatin in the open state. Figure 4 b、 Figure 9 e); In addition to the paternal and maternal pronuclei of the zygote, the nuclei of the 2-cell, 4-cell, 8-cell, morula, and blastocyst stages all exhibit stage-specific chromatin opening patterns. Among them, there is obvious clustering and separation between the inner cell mass (ICM) and trophectoderm (TE) cells of the blastocyst, and they respectively show the opening characteristics of pluripotency genes and lineage-specific genes. Figure 4 b、 Figure 10 a).

[0112] (2) Parental genome-specific analysis: ULI-snATAC-seq can distinguish between maternal pronuclear MPN and paternal pronuclear PPN in zygotes based on differences in sex chromosome read content. Figure 9 f): Maternal pronuclei carry only the maternal X chromosome, exhibiting high X chromosome reads signal and no Y chromosome reads; paternal pronuclei (especially those carrying the Y chromosome in male zygotes) exhibit specific Y chromosome reads signal. Figure 9 f), thereby enabling independent epigenomic analysis of each parent's genome within a single fertilized egg.

[0113] (3) Marker gene verification: The identity of each cell cluster was determined by stage-specific marker peaks ( Figure 10 b) and dual verification of the accessibility characteristics of typical marker genes ( Figure 5 The markers include oocyte-specific marker Zp1, 2-cell stage-specific marker Zscan4, 2-8 cell stage marker Klf17, 8-cell stage marker Trim43a, ICM-specific marker Sox2, and TE-specific marker Cdx2, all of which have opening patterns that are highly consistent with the biological characteristics of the corresponding developmental stages.

[0114] (4) Data reliability verification: Single-cell data from each stage were aggregated into a pseudo-batch spectrum, which was highly correlated with the published batch ATAC-seq and DNase-seq datasets. Figure 11 a- Figure 11 (b) verified the technical reproducibility and biological fidelity of ULI-snATAC-seq.

[0115] (5) Developmental trajectory reconstruction: The pseudo-time trajectory reconstruction arranges all cell nuclei in an orderly manner on a continuous developmental axis according to the dynamic changes of chromatin open state. The trajectory completely covers all preimplantation developmental stages from oocyte to blastocyst. The color gradient is highly consistent with the in vivo embryonic development process, faithfully reproducing the complete epigenetic dynamics of mouse preimplantation embryonic development.

[0116] (6) Peak-gene association analysis: Integrating peak-gene association analysis, this study identified over 50,000 significant cis-regulatory peak-gene associations, visualized in a heatmap. Figure 6 The results showed that the peaks and gene accessibility signals exhibited clear dynamic changes during embryonic development, revealing the dynamic coordination between chromatin openness and gene regulation. Specifically, the cis-regulatory peaks near Obox and Zscan4 family genes showed highly active openness at the 2-cell stage, consistent with the stage characteristics of embryonic genome activation. Meanwhile, the accessibility peaks of the pluripotency core regulatory gene Pou5f1 and the trophoblast lineage marker gene Cdx2 were specifically activated at the inner cell mass (ICM) and trophoblast (TE) stages, respectively, corresponding to the lineage differentiation process during the blastoblast stage.

[0117] (7) Dynamic analysis of transcriptional regulatory factors: A series of transcription factor (TF) motifs were identified, and their chromatin accessibility was positively correlated with the expression of the corresponding TF genes distributed along the developmental trajectory. Figure 7 These motifs are core candidate regulators driving the continuous chromatin transitions during early embryogenesis.

[0118] The above results confirm that the present invention has successfully constructed a high-resolution single-nuclear chromatin accessibility map of mouse preimplantation development. This dataset has both excellent technical quality and biological depth, and fully captures global and parent-specific epigenomic transitions during embryonic development, providing a basic framework for elucidating the molecular mechanism by which chromatin accessibility dynamically regulates the initiation of mammalian life.

[0119] Data availability

[0120] All sequencing data related to this invention have been stored in the GEO database, accession number GSE282209. The accession numbers of the published data used in this study are as follows: batch ATAC-seq of mouse embryonic stem cells, accession numbers GSE127790, GSE164130, GSE159468, GSE67298, GSE66581, and GSE175632; batch ATAC-seq of oocytes and early embryos, accession numbers GSE116854, GSE207222, GSE169632, and GSE66581; DNase-seq of oocytes and early embryos, accession numbers GSE76642 and GSE92605; scRNA-seq of oocytes and early embryos, accession numbers GSE45719 and GSE38495.

Claims

1. An ultra-low input single cell nucleus chromatin accessibility sequencing method, characterized by: The method comprises the following steps: (1) transferring the single cell or single nucleus sample to be tested into a microdroplet of labeling mixture for Tn5 transpositional labeling reaction, and covering the microdroplet with mineral oil to prevent evaporation; (2) after the labeling reaction is completed, adding a termination buffer to terminate the reaction and placing on ice; (3) picking up a single cell or nucleus and transferring it into a PCR tube containing a lysis buffer and a sequencing index primer for lysis to release the bound Tn5 transposase; (4) adding Tween-20 to quench the SDS in the lysis buffer; (5) adding a PCR premix to amplify the library fragments; (6) purifying and quantifying the PCR amplification product to obtain a sequencing library; (7) performing high-throughput sequencing on the sequencing library to obtain chromatin accessibility sequencing data of the sample to be tested.

2. The method of claim 1, wherein: In step (1), the formula of the microdroplet of labeling mixture is: 33 mM Tris-acetate with a pH value of 7.8, 66 mM potassium acetate, 10 mM magnesium acetate, 16 vol% dimethylformamide, and 0.01 wt% digitonin; and 25 vol% Tn5 transposase.

3. The method of claim 1, wherein: In step (1), the conditions of the labeling reaction are: performing the labeling reaction in a 37℃ incubator for 30 minutes, then transferring the sample to a new microdroplet of labeling mixture and continuing the reaction at 37℃ for 30 minutes; wherein no more than 20 cells or nuclei are contained in each 0.1 μL microdroplet for the reaction.

4. The method of claim 1, wherein: In step (2), the formula of the termination buffer is: 10 mM Tris-HCl with a pH value of 8.0, 20 mM EDTA; and the conditions of the termination reaction are: placing on ice for 10 minutes.

5. The method of claim 1, wherein: In step (3), the formula of the lysis buffer is: 100 mM Tris-HCl with a pH value of 8.0, 100 mM NaCl, 40 μg / mL proteinase K, and 0.4 wt% SDS; and the conditions of the lysis are: incubating at 65℃ for 15 minutes.

6. The method of claim 1, wherein: In step (5), the conditions of the PCR amplification are: 72℃ for 5 minutes, 98℃ for 5 minutes, then performing 18 cycles of amplification, each cycle being 98℃ for 10 seconds, 63℃ for 30 seconds, and 72℃ for 20 seconds, and then incubating at 4℃ after the amplification is completed.

7. The method of claim 1, wherein: In step (7), the high-throughput sequencing is performed by using an Illumina HiSeq X sequencing platform.

8. The method of claim 1, characterized by: The single cell or single nucleus sample is derived from a mouse pre-implantation embryo, a mouse oocyte, or a mouse embryonic stem cell.

9. The method of claim 8, wherein: The single nucleus sample is prepared by the following method: after MI and MII oocytes, zygotes, 2-cell stage embryos, or 4-cell stage embryos are stained with Hoechst 33342, the cell nuclei are separated by using a micromanipulator.

10. Use of the method of any one of claims 1-9 in constructing a single cell chromatin accessibility map.