ScATAC-seq data-based ecDNA identification method

Through the ecDNA recognition method based on single-cell ATAC-seq data, combined with pretreatment and verification technology, single-cell analysis of ecDNA is realized, solving the limitations of the existing methods and providing the detailed structural and functional analysis capabilities of ecDNA.

CN120290730AActive Publication Date: 2025-07-11KUNMING MEDICAL UNIVERSITY

Patent Information

Application Number
CN202510712909.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-07-11
Estimated Expiration
2045-05-30

AI Technical Summary

Technical Problem

The existing single-cell analysis methods of ecDNA are limited to hundreds of cells, and cannot deeply explore ecDNA heterogeneity and clonal dynamics, and there is a lack of single-cell research on the clonal trajectory of ecDNA.

Method used

A method of ecDNA recognition based on single-cell ATAC-seq data was developed to visualize and verify ecDNA through pretreatment, cell type annotation, amplified fragments and breakpoint recognition, combined with fluorescence in situ hybridization and immunohistochemistry verification.

Benefits of technology

It can accurately detect copy number variation, breakpoints and characteristic distribution of ecDNA at the single-cell level, reveal the functional role of ecDNA in tumors, fill the gaps in existing methods, and support the analysis of enhancer hijacking, transacting regulation, tumor heterogeneity and uneven distribution of ecDNA.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120290730A_ABST
    Figure CN120290730A_ABST
Patent Text Reader

Abstract

The invention discloses an ecDNA (ectodeoxyribonucleic acid) identification method based on scATAC-seq (scATAC-Seq) data. The ecDNA recognition method provided by the invention can be used for accurately detecting copy number amplification, breakpoint quantity and ecDNA feature quantity on the whole cell level. According to the method provided by the invention, the specific carried annular oncogene in the tumor can be identified, the distribution characteristics of the annular oncogene in the genome can be analyzed, and the overall composition of all annular DNAs in a single cell and the specific composition of the single annular DNA can be disclosed. In addition, according to the method, the existence of the circular DNA can be verified through a dyeing experiment, and the reliability of a detection result is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of bioinformatics, and particularly to a method for identifying ecDNA based on scATAC-seq data. Background Art

[0002] Recent studies have shown that ecDNA usually contains specific oncogenes, drug resistance genes, immune regulatory genes, and other regulatory elements, making it an important factor in cancer pathogenesis and poor prognosis in multiple cancer types. However, many aspects of ecDNA are still unclear, including its role in abnormal gene expression, cancer evolution, biogenesis, maintenance, and clearance. Given the high tumor gene copy number and enhanced chromatin accessibility of ecDNA, ATAC-seq is a powerful tool for studying the functional roles of ecDNA in enhancer hijacking, trans-acting regulation, tumor heterogeneity, and uneven distribution.

[0003] Current analysis tools can provide detailed insights into the structure of ecDNA. By using whole-genome sequencing (WGS), ATAC-seq, Circle-seq, and long-read sequencing data, they can identify amplified fragments and breakpoints, revealing the circular structure of ecDNA, including complex breakpoint edges. However, the single-cell analysis of ecDNA is still limited. Although single-cell extrachromosomal circular DNA and transcriptome sequencing (scEC&T-seq) can isolate and sequence mRNA and circular DNA at the single-cell level, its application is usually limited to a few hundred cells. This limitation hinders the in-depth exploration of ecDNA heterogeneity and clonal dynamics. In addition, single-cell studies on the clonal trajectory of ecDNA are still lacking. Summary of the Invention

[0004] To meet the need for single-cell analysis of ecDNA, the present invention has developed a method for identifying ecDNA based on single-cell ATAC-seq data. Using this method, ecDNA in single-cell data can be identified and visualized, and subsequently, the presence of the identified ecDNA is verified by fluorescence in situ hybridization (FISH) and immunohistochemistry. In summary, this method has the ability to powerfully analyze the copy number variation (CNV), breakpoints, and characteristic distribution of ecDNA in single-cell ATAc-seq data.

[0005] To solve the above technical problems, the technical solution of the present invention is as follows: A method for identifying ecDNA based on scATAC-seq data, and the method operates as follows: Step (1): After preprocessing single-cell chromatin accessibility sequencing (scATAC-seq) data, quantify gene activity by calculating the number of chromatin-open DNA fragments in gene bodies and promoter regions, and then perform cell type annotation based on the expression of typical genes in various cell types in the tumor microenvironment; Step (2): Identify amplified fragments in single-cell sequencing data, and the specific operations are as follows: Remove data contained in the blacklist region of the hg38 reference genome; Filter single-cell data based on multiple identification conditions; Identify fragments (gene fragments) with the number of reads (DNA short fragment sequences) greater than 2; Use AMULET (ATAC-seq MULtipletEstimation Tool) to identify intervals with CNV greater than 2 in each fragment; The AMULET used is a python toolkit for detecting amplified gene fragments based on single-nucleus ATAC-seq data; The multiple identification conditions are: the mapping quality score (MAPQ) is greater than 5, the fragment length is less than or equal to 300MB, the reads are paired, the reads can fully correspond to the hg38 reference genome, and the reads are not PCR duplicates; Step (3): Identify fragments with breakpoints in single-cell sequencing data, and the specific operations are as follows: Perform quality control on the data obtained by single-cell sequencing; Identify breakpoints based on three criteria; The three criteria are: consider reads with a length greater than 500bp to contain breakpoints, consider reads aligned to different chromosomes to contain breakpoints, and consider reads with the same alignment direction to contain breakpoints; Step (4): Identify ecDNA in single-cell data; Step (5): Visualize the structure of ecDNA in single cells.

[0006] As a further description of the above solution: The detailed process of Step (4) to identify ecDNA in single-cell data is as follows: Extend 2000bp upstream and downstream of the amplified fragments; Find the intersection of the amplified fragment region and the reads containing breakpoints; Correspond the cells containing breakpoints to the single-cell data according to the cell barcode (barcode); Annotate the genes contained in the ecDNA fragment; Find the ecDNA characteristic genes.

[0007] As a further description of the above solution: The selected cellular ecDNAs were visualized using the R packages OmicCircos and circlize.

[0008] The OmicCircos used is an R package focused on multi-omics visualization, capable of displaying different types of data on multiple tracks. Circlize is an R package for creating circular plots, especially suitable for visualizing complex relationships and data hierarchies. By using these two toolkits in combination, the circular conformation and biological characteristics of ecDNA can be clearly and elaborately presented.

[0009] As a further description of the above solution: The single-cell sequencing data is derived from single-cell sequencing of tissue samples, and the specific process is as follows: Normal tissue samples, pre-cancerous tissue samples, and cancer tissue samples paired from multiple patients were collected; The tissue samples were separately stored in MACS tissue storage solution; the tissue samples were enzymatically digested and mechanically dissociated using a dissociation kit and a dissociator, and a cell suspension was obtained by filtration; The apoptotic and dead cells were removed from the filtered cell suspension using a Dead Cell Removal kit; then a single-cell kit was used for library preparation; the prepared library was evaluated for library quality; the integrity of cDNA was detected, and single-cell sequencing was performed.

[0010] As a further description of the above solution: The preprocessing process of the single-cell sequencing data is as follows: Single-cell data normalization to eliminate technical biases and make gene expression levels comparable between cells; Feature selection analysis to screen for highly variable genes, reduce noise, and retain biological differences; Single-cell data dimensionality reduction to compress the data dimension, retain key biological variations, and facilitate visualization and clustering; Clustering analysis to divide cells into biologically meaningful groups based on the dimensionality-reduced data.

[0011] As a further description of the above solution: Experimental methods were used to verify the existence of the speculated ecDNA, including: fluorescence in situ hybridization to verify the presence of the top marker gene and immunohistochemistry to localize the protein.

[0012] The ecDNA identification method provided by the present invention can accurately detect copy number amplification, the number of breakpoints, and the number of ecDNA features at the whole cell level. This method can identify specific circular oncogenes carried in tumors, analyze their distribution characteristics in the genome, and reveal the overall composition of all circular DNAs within a single cell and the specific composition of a single circular DNA. In addition, this method can also verify the existence of circular DNA through staining experiments to ensure the reliability of the detection results.

[0013] Compared with the prior art, the present invention has the following beneficial effects: The present invention identifies ecDNA based on large-scale single-cell ATAC-seq data, filling the gap in existing ecDNA research methods; the present invention can be used to analyze the roles of ecDNA in enhancer hijacking, trans-acting regulation, tumor heterogeneity, and uneven distribution, and can more comprehensively characterize the functional characteristics of ecDNA. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 It is a schematic diagram for identifying ecDNA and its structure in single-cell ATAC-seq data; a shows how to predict the ecDNA structure using single-cell ATAC-seq data from skin squamous cell carcinoma, precancerous actinic keratosis, and normal skin tissue samples; b shows the distribution of 27,792 cells in the UMAP plot according to cell type and sample type after dimensionality reduction, clustering, and cell type annotation in the single-cell ATAC-seq dataset; c shows the distribution of copy number amplification related to ecDNA in cells; d shows the distribution of breakpoints related to ecDNA in cells; e shows the distribution of genes constituting ecDNA in different cells. Figure 2 It is a schematic diagram of ecDNA genes, cell types, and chromatin distribution; a shows the total number of amplified fragments (CNV>2) and breakpoints identified in different cell types; b shows the calculated fragment amplification and breakpoint numbers of known ecDNA-related genes; c shows the proportion of high-quality fragments at the transcription start site (TSS) in single-cell ATAC-seq data; d shows the distribution of ecDNA in different genomic regions, classified according to the proportion of various genomic elements. Figure 3 It is a schematic diagram of KLF7 ecDNA structure in single cells; a shows the distribution of various ecDNA features within a representative cell; b shows the KLF7 ecDNA structure in the same cell; c shows that the FISH experiment confirmed the existence of KLF7 ecDNA in frozen and formalin-fixed paraffin-embedded (FFPE) skin squamous cell carcinoma tissue samples. DETAILED DESCRIPTION OF THE INVENTION

[0015] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are some, but not all, of the embodiments of this application. In the following description, specific details such as specific configurations and components are provided only to assist in a comprehensive understanding of the embodiments of this application. Therefore, those skilled in the art should clearly understand that various changes and modifications can be made to the embodiments described here without departing from the scope and spirit of this application. Additionally, descriptions of known functions and structures are omitted for clarity and conciseness in the embodiments.

[0016] It should be understood that the phrase "one embodiment" or "this embodiment" mentioned throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, the phrase "one embodiment" or "this embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in one or more embodiments in any suitable manner.

[0017] In addition, this application may repeat reference numerals and / or letters in different instances. This repetition is for the purpose of simplicity and clarity and does not in itself indicate the relationship between the various embodiments and / or arrangements discussed.

[0018] The term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, B exists alone, and both A and B exist simultaneously. The term " / and" in this document describes another association relationship between associated objects, indicating that there can be two relationships. For example, A / and B can represent: A exists alone, and both A and B exist. Additionally, the character " / " in this document generally indicates that the associated objects before and after are in an "or" relationship.

[0019] The term "at least one" in this document is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, at least one of A and B can represent: A exists alone, both A and B exist simultaneously, and B exists alone.

[0020] It should also be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprises", "comprising", or any other variation thereof is intended to cover a non-exclusive inclusion.

[0021] Embodiment 1 An ecDNA recognition method based on scATAC-seq data, including identifying cells containing ecDNA based on copy number, an ecDNA breakpoint recognition rule, and ecDNA visualization, includes: S101. Collect skin tissue samples; S102. Perform single-cell sequencing on the tissue samples to obtain single-cell sequencing data of the tissue samples; S103. Preprocess the single-cell sequencing data of the tissue samples; S104. Perform cell annotation on the preprocessed single-cell sequencing data; S105. Identify amplified fragments in the single-cell sequencing data; S106. Identify fragments with breakpoints in the single-cell sequencing data; S107. Identify ecDNA in the single-cell data; S108. Visualize the structure of ecDNA in single cells.

[0022] Preferably, the specific operation of collecting skin tissue samples in step S101 is as follows: Collect paired normal skin tissue samples, pre-cancerous actinic keratosis tissue samples, and cutaneous squamous cell carcinoma tissue samples from three patients.

[0023] Preferably, the performing single-cell sequencing on the tissue samples to obtain single-cell sequencing data of the tissue samples in step S102 specifically includes the following steps: S301. Store the skin tissue samples in MACS tissue storage solution; S302. Use a dissociation kit and dissociator to perform enzymatic digestion and mechanical dissociation on the tissue samples to obtain a cell suspension; S303. Filter the cell suspension of the tissue samples using a 70 µm cell filter to obtain a filtered cell suspension; S304. Use a Dead Cell Removal kit to remove apoptotic and dead cells from the filtered cell suspension; S305. Use a single-cell kit to prepare a library for the cell suspension; S306. Evaluate the quality of the prepared library; S307. Detect the integrity of cDNA; S308. Perform single-cell sequencing.

[0024] Preferably, the preprocessing of the single-cell sequencing data of the tissue samples in step S103 specifically includes the following steps: S401. Cell Ranger ATAC data upstream processing process; This application uses the Cell Ranger ATAC pipeline of 10× Genomics to process the freshly sequenced scATAC-seq data. This pipeline includes aligning FASTQ files to the reference genome, removing PCR duplicates, identifying valid cells, generating open chromatin regions (peaks), and outputting files such as fragments. The finally generated files such as fragments.tsv.gz and peaks.bed can be used for downstream analysis.

[0025] S402: Construct a Seurat object using the result data of Cell Ranger; 1. Data reading First, according to the process of processing scATAC data in the Signac R package, read the peak file (peaks.bed) of each sample and convert it into a GenomicRanges object for subsequent genomic range operations. Then, read the single-cell metadata file (singlecell.csv) of each sample, which contains the filtering information of each cell (such as passed_filters). By filtering out low-quality cells (passed_filters > 500), it is ensured that the cells for subsequent analysis have sufficient sequencing depth.

[0026] 2. Peak merging and filtering The peak data of all samples are merged into a unified peak set (combined.peaks), and overlapping peak regions are removed through the reduce function to ensure that each peak is unique on the genome. The merged peak set is further filtered to remove peaks that are too short (<20 bp) or too long (>10,000 bp), and high-quality peak regions are retained for subsequent analysis.

[0027] 3. Creation of fragment object and generation of count matrix For each sample, a fragment object (FragmentObject) is created, which contains the fragment information of each cell. This fragment information is used to generate the count matrix (FeatureMatrix) of each sample, where each row represents a peak, each column represents a cell, and the value in the matrix represents the fragment count of each cell on each peak. The generated count matrix is encapsulated into a ChromatinAssay object and further encapsulated into a Seurat object, and metadata information (such as cell filtering information) is added to the Seurat object.

[0028] 4. Merge all samples into one Seurat object Finally, all the Seurat objects of the samples are merged into a unified Seurat object (combined_atac), and the cells of different samples are distinguished by adding the sample source information (dataset). Gene annotations of the human genome are added to the seurat object. The merged object contains the peak count matrix and metadata information of all samples, which is convenient for subsequent joint analysis.

[0029] S403, perform quality control on Seurat objects; 1. Nucleosome signaling Calculate the nucleosome signal for each cell. The nucleosome signal reflects the length of the open region of DNA, which is usually within the range of 1 to 2 nucleosome lengths (147bp). Cells with too high nucleosome signals (such as nucleosome_signal>4) may contain more nucleosome contamination and are therefore usually filtered out.

[0030] 2. TSS enrichment score Calculate the TSS enrichment score for each cell. The TSS enrichment score reflects the enrichment of open chromatin near the transcription start site (TSS) of the gene in the sequencing data. A higher TSS enrichment score generally indicates that there are more active gene regions in the cell, while cells with lower scores (TSS < 3) will be filtered out.

[0031] 3. Ratio of fragments within the peak The proportion of fragments falling within the open chromatin region (peak) in each cell is calculated. A higher proportion of fragments within the peak usually indicates that there are more open chromatin regions in the cell, which may be related to the active state of the gene. Cells with a peak fragment ratio greater than 15% are retained.

[0032] 4. Proportion of blacklisted areas The proportion of fragments falling within the blacklist region of the genome is calculated for each cell. Blacklist regions often contain repetitive sequences or other difficult-to-map regions. A higher blacklist ratio may indicate poor data quality. Cells with a blacklist region ratio less than 0.05 are retained.

[0033] Preferably, the step S104 of annotating the pretreated cells specifically includes the following steps: S405, data dimensionality reduction; The signacR package was used to perform TF-IDF normalization on the peaks of the cells to correct for differences in sequencing depth between cells and the sparsity between peaks. Peaks with a total count greater than 10 in all cells were selected as highly expressed features, and the peak matrix of the standardized cells was subjected to singular value decomposition to reduce the cells to low latitudes.

[0034] S406, Cell clustering; Construct a K-nearest neighbor graph of cells based on the dimensionality reduction results, and use the Louvain algorithm to cluster the cells, dividing the cells into different clusters.

[0035] S407, Gene activity matrix calculation; Quantify gene activity by calculating the number of chromatin-open DNA fragments in the gene body and promoter regions, and perform LogNormalize normalization on the gene activity matrix of the cells.

[0036] S408, Cell annotation; Analyze the gene activity characteristics of each cluster by calculating the expression differences of typical marker genes of various cell types (such as immune cells, stromal cells, tumor cells, etc.) in the tumor microenvironment in different clusters, and then annotate the clusters as specific cell types based on these characteristics.

[0037] Preferably, the specific steps for identifying the amplified fragments in the single-cell sequencing data in step S105 are as follows: S501, Remove the data contained in the blacklist region of the hg38 reference genome; S502, Filter the single-cell data based on five conditions; S503, Identify the fragments with more than 2 reads; S504, Use AMULET to identify the intervals with CNV greater than 2 in each fragment.

[0038] Preferably, the five conditions are: the mapping quality score (MAPQ) is greater than 5; the fragment length is less than or equal to 300MB; the reads are paired; the reads can be fully mapped to the hg38 reference genome; the reads are not PCR duplicates.

[0039] Preferably, the specific steps for identifying the fragments with breakpoints in the single-cell sequencing data in step S106 are as follows: S601, Perform quality control on the data obtained by single-cell sequencing; S602, Identify breakpoints based on three criteria.

[0040] Preferably, the three criteria are: consider the reads with a length greater than 500bp to contain breakpoints; consider the reads mapped to different chromosomes to contain breakpoints; consider the reads with the same alignment direction to contain breakpoints.

[0041] Preferably, the identification of ecDNA in the single-cell data in step S107 specifically includes the following steps: S701. Extend 2000 bp upstream and downstream of the amplified fragments; S702. Find the intersection of the amplified fragment region and the reads containing breakpoints; S703. Correlate the cells containing breakpoints with the single-cell data according to the cell barcode; S704. Annotate the genes contained in the ecDNA fragments; S705. Find the ecDNA characteristic genes.

[0042] Preferably, the visualization of the ecDNA structure in single cells in step S108 specifically includes the following steps: S801. Visualize the selected cell ecDNA using the R packages OmicCircos and circlize.

[0043] Example 2 This example is based on the above Example 1, and the same parts as those in the above Example 1 will not be elaborated.

[0044] This example introduces a method for verifying the existence of the speculated ecDNA using experimental methods. The specific steps are as follows: S901. Verify the existence of the top marker gene by fluorescence in situ hybridization; In this experiment, fluorescence in situ hybridization (FISH) technology was used to detect and analyze the target DNA in paraffin-embedded tissue sections of human cutaneous squamous cell carcinoma (CSCC). First, the paraffin sections were subjected to routine dewaxing and hydrated to distilled water for standby. In the antigen retrieval step, the sections were placed in boiling distilled water or EDTA solution at pH 9.0 for 25 - 30 minutes for retrieval. Meanwhile, a staining jar was preheated in a 37°C water bath, 40 mL of distilled water was added, and after equilibration for 10 - 20 minutes, 1 mL of 1 M hydrochloric acid and 1 mL of 10% pepsin solution were added successively as the digestion solution. After the retrieval was completed, the sections were taken out, washed with PBS buffer for 5 minutes, and then placed in the preheated digestion solution at 37°C for digestion for 15 - 30 minutes, and the specific digestion time was appropriately adjusted according to the tissue density and section thickness of the sample. Subsequently, the sections were washed twice with PBS buffer, 5 minutes each time, and the sections were dehydrated with 70%, 95%, and 100% ethanol for 2 minutes each, and air-dried naturally at room temperature. After dehydration and air-drying, 10 μL of probe mixture (the probe was diluted with hybridization buffer at a ratio of 1:9) was added to the center of the section, covered with a 22 mm × 22 mm coverslip, and sealed with FISH mounting glue around to prevent evaporation. Before the hybridization step, two incubators were prepared in advance, set at 88°C and 37°C respectively, and the light-shielded humid box was preheated in the 88°C incubator for at least 10 minutes. Subsequently, the sectioned slides together with the humid box were placed in the 88°C incubator for denaturation for 10 minutes, and then transferred to the 37°C incubator for overnight hybridization for 16 - 20 hours. Before the hybridization was completed, two tanks of 2×SSC (containing 0.1% NP-40) washing solution were prepared, one tank was preheated in a 60±1°C water bath, and the other tank was kept at room temperature. After hybridization was completed, the mounting glue and coverslip were removed, and the sections were first washed in the room temperature washing solution for 5 minutes, and then transferred to the 60°C washing solution for washing for 2 minutes. After washing, the excess liquid was shaken off, 20 μL of DAPI anti-quenching mounting solution was added dropwise to the sections, covered with a coverslip, incubated in the dark for 20 minutes, and then observed and photographed under a fluorescence microscope. The probes used in this experiment included the target gene probes KLF7 (FAM-labeled), MYC (Cy5-labeled), NRP2 (FAM-labeled), and PTMA (FAM-labeled) synthesized by Beijing Tsingke Biotechnology Co., Ltd.; the centromere probes provided by EXONBIO included the centromere probe for human chromosome 8 (FITC-labeled) and the centromere probe for human chromosome 2 (Cy3-labeled).

[0045] S902, immunohistochemistry experiment locates proteins; In this experiment, immunohistochemical staining (IHC) was used to detect and analyze the expression location of KLF7 protein in paraffin-embedded tissue sections of cutaneous squamous cell carcinoma (CSCC). The sections used in the experiment had a thickness of 4–5 μm and were baked at 75 °C for 40 minutes after being attached to glass slides. The sections were dewaxed twice with xylene for 10 minutes each time, and then dehydrated through a gradient of ethanol at 100%, 95%, 85%, and 75% for 3 minutes each step to complete the rehydration process. Subsequently, the sections were placed in a citrate buffer at 10 mM and pH 6.0, and antigen retrieval was performed using microwave heating for 10–15 minutes. After natural cooling to room temperature, the sections were washed three times with PBS buffer for 5 minutes each time. Endogenous peroxidase activity was blocked by treating with 3% hydrogen peroxide for 10 minutes, and then non-specific antigen sites were blocked with 10% goat serum and incubated at room temperature for 30 minutes. Subsequently, the primary antibody against KLF7 (rabbit anti-KLF7 polyclonal antibody, HUABIO, catalog number ER1911-99, dilution ratio 1:200) was added and incubated overnight at 4 °C. After washing three times with PBS the next day, a biotinylated secondary antibody (goat anti-rabbit IgG, HUABIO, catalog number HA1001) was added and incubated at room temperature for 30 minutes, and then a streptavidin-horseradish peroxidase (streptavidin-HRP) complex was added and incubated for another 30 minutes. DAB chromogenic solution was used for color development, and the reaction time was controlled until a satisfactory brown-yellow signal appeared under the microscope (about 2–5 minutes). After terminating the reaction, the sections were washed with distilled water, counterstained with hematoxylin for 1–2 minutes, blued in Scott's tap water substitute solution, and then dehydrated through a gradient of ethanol at 75%, 85%, 95%, and 100% and cleared in xylene. Finally, neutral gum (Leica DPX mounting medium, catalog number 3808600E) was added for mounting. After air-drying at room temperature, observations and photographic records were made under a bright-field microscope.

[0046] Example 3 This example is based on Example 1 above, and the same parts as in Example 1 will not be elaborated. This example introduces the identification and visualization of ecDNA.

[0047] In this example, paired normal skin tissue samples, precancerous actinic keratosis tissue samples, and cutaneous squamous cell carcinoma tissue samples from three patients were collected, and single-cell ATAC sequencing was performed.

[0048] Step 1: Identify ecDNA and its structure in single-cell ATAC-seq data. The results are as shown in the appendix Figure 1 as follows: Figure 1 in: a. Schematic diagram of the workflow of this method, showing how to predict ecDNA structures using single-cell ATAC-seq data from skin squamous cell carcinoma, precancerous actinic keratosis, and normal skin tissue samples. The process includes steps such as doublet removal, identifying DNA amplification fragments with copy number variation (CNV)>2, and reconstructing ecDNA circular structures through ecDNA breakpoints of the above three types.

[0049] b. In the single-cell ATAC-seq dataset, the distribution of 27,792 cells in the UMAP plot is shown according to cell type and sample type (skin squamous cell carcinoma, precancerous actinic keratosis, and normal skin tissue) after dimensionality reduction, clustering, and cell type annotation.

[0050] c. The figure shows the distribution of copy number amplifications related to ecDNA in cells.

[0051] d. The figure shows the distribution of breakpoints related to ecDNA in cells.

[0052] e. The figure shows the distribution of genes constituting ecDNA in different cells.

[0053] Step 2: ecDNA genes, cell types, and chromatin distribution, and the results are as shown in the appendix Figure 2 as follows Figure 2 in: a. The figure shows the total number of amplification fragments (CNV>2) and breakpoints identified in different cell types.

[0054] b. The figure shows the number of fragment amplifications and breakpoints of known ecDNA-related genes calculated in this experiment.

[0055] c. The figure shows the proportion of high-quality fragments at transcription start sites (TSSs) in single-cell ATAC-seq data.

[0056] d. The figure shows the distribution of ecDNA in different genomic regions, classified according to the proportion of various genomic elements (such as promoters, exons, introns, intergenic regions).

[0057] Step 3: Construction of ecDNA molecules in single cells and the ecDNA structure of a single gene (KLF7) in specific cells, and the results are as shown in the appendix Figure 3 as follows

[0058] a. The figure shows the distribution of various ecDNA features within a representative cell (ID: TACCTCGGTAAGCCGA-1). The outer ring shows gene names and their corresponding chromosomes, the middle ring shows copy number variations (CNVs), and the inner ring indicates the positions of detected breakpoints.

[0059] b. The figure shows the KLF7 ecDNA structure in the same cell (ID: TACCTCGGTAAGCCGA-1). The outer ring shows the genomic start and end positions of the CNV associated with KLF7, the middle ring highlights the CNV values, and the inner ring marks the start and end positions of the breakpoints of KLF7.

[0060] c. FISH experiments confirmed the presence of KLF7 ecDNA in frozen and formalin-fixed paraffin-embedded (FFPE) skin squamous cell carcinoma tissue samples.

[0061] It should be understood that the above specific embodiments of the present invention are only for illustrative or explanatory purposes of the principles of the present invention, and do not constitute a limitation on the present invention. Therefore, any modifications, equivalent substitutions, improvements, etc. made without departing from the spirit and scope of the present invention shall be included within the protection scope of the present invention. In addition, the appended claims of the present invention are intended to cover all variations and modifications that fall within the scope and boundaries of the appended claims, or equivalent forms of such scope and boundaries.

Claims

1. An ecDNA recognition method based on scATAC-seq data, characterized in that The method operates as follows: Step (1): After preprocessing the single-cell chromatin accessibility sequencing (scATAC-seq) data, quantify gene activity by calculating the number of chromatin-open DNA fragments in the gene body and promoter regions, and then perform cell type annotation based on the expression of typical genes of various cell types in the tumor microenvironment; Step (2): Identify the amplified fragments in the single-cell sequencing data, and the specific operations are as follows: Remove the data contained in the blacklist region of the hg38 reference genome; Filter the single-cell data based on multiple identification conditions; Identify the fragments with the number of reads greater than 2; Use AMULET to identify the intervals with CNV greater than 2 in each fragment; The AMULET used is a Python toolkit for detecting amplified gene fragments based on single-nucleus ATAC-seq data; The multiple identification conditions are: the mapping quality score is greater than 5, the fragment length is less than or equal to 300MB, the reads are paired, the reads can be fully mapped to the hg38 reference genome, and the reads are not PCR duplicates; Step (3): Identify the fragments with breakpoints in the single-cell sequencing data, and the specific operations are as follows: Perform quality control on the data obtained by single-cell sequencing; Identify breakpoints based on three criteria; The three criteria are: consider the reads with a length greater than 500bp to contain breakpoints, consider the reads mapped to different chromosomes to contain breakpoints, and consider the reads with the same alignment direction to contain breakpoints; Step (4): Identify ecDNA in the single-cell data; Step (5): Visualize the structure of ecDNA in single cells.

2. The ecDNA recognition method based on scATAC-seq data according to claim 1, wherein The detailed process of Step (4) for identifying ecDNA in the single-cell data is as follows: Extend 2000bp upstream and downstream of the amplified fragments; Find the intersection of the amplified fragment region and the reads containing breakpoints; Correspond the cells containing breakpoints to the single-cell data according to the cell barcode; Annotate the genes contained in the ecDNA fragments; Find the ecDNA characteristic genes.

3. The ecDNA identification method based on scATAC-seq data according to claim 1, wherein, Use the R packages OmicCircos and circlize to visualize the selected cell ecDNA.

4. The ecDNA identification method based on scATAC-seq data according to claim 1, wherein The single-cell sequencing data is derived from the single-cell sequencing of tissue samples, and the specific process is as follows: Collect paired normal tissue samples, pre-cancerous tissue samples, and cancer tissue samples from multiple patients; Store the tissue samples in MACS tissue storage solution respectively; Use a dissociation kit and a dissociator to perform enzymatic digestion and mechanical dissociation on the tissue samples, and filter to obtain a cell suspension; Use the Dead Cell Removal kit to remove apoptotic and dead cells from the filtered cell suspension; Then use a single-cell kit to prepare a library; Evaluate the quality of the prepared library; Detect the integrity of cDNA and perform single-cell sequencing.

5. The ecDNA recognition method based on scATAC-seq data according to claim 1, wherein The preprocessing process of the single-cell sequencing data is as follows: Single-cell data normalization to eliminate technical biases and make gene expression levels comparable between cells; Feature selection analysis to screen for highly variable genes, reduce noise, and retain biological differences; Single-cell data dimensionality reduction to compress the data dimensions, retain key biological variations, and facilitate visualization and clustering; Clustering analysis to divide cells into biologically meaningful groups based on the dimensionality-reduced data.

6. The ecDNA identification method based on scATAC-seq data according to claim 1, wherein Use experimental methods to verify the existence of putative ecDNA, including: fluorescence in situ hybridization to verify the presence of top marker genes and immunohistochemistry experiments to localize proteins.

Citation Information

Patent Citations

  • Method for identifying chromosome outer circular DNA composition gene of tumor cell

    CN113724788A

  • Extrachromosomal DNA identification and methods of use

    CN114040982A

  • Single-cell multi-omics parallel sequencing method based on third-generation sequencing platform and application of single-cell multi-omics parallel sequencing method

    CN117106873A

  • METHODS AND COMPOSITIONS FOR DETECTING ecDNA

    US20220364182A1

  • Methods for targeted purification and profiling of human extra chromosomal DNA

    WO2023091825A1

Cited By

  • Method for identifying single cell ecDNA cloning trajectory

    CN120656545A

  • A method for identifying single-cell ecDNA clonal trajectories

    CN120656545B