An ecDNA identification method based on scat ac-seq data

By using an ecDNA identification method based on single-cell ATAC-seq data, the limitations of single-cell ecDNA analysis have been overcome, enabling efficient identification and visualization of ecDNA, revealing its functional characteristics in tumors, and making up for the deficiencies of existing technologies.

CN120290730BActive Publication Date: 2026-03-24KUNMING MEDICAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing technologies are insufficient for efficient single-cell analysis and research on the heterogeneity and clonal dynamics of ecDNA. They also lack in-depth exploration of ecDNA clonal trajectories. The application of single-cell extracellular chromosomal circular DNA and transcriptome sequencing is limited to a few hundred cells, which cannot fully characterize the functional features of ecDNA.

Method used

A method for identifying ecDNA based on single-cell ATAC-seq data was developed. By preprocessing single-cell chromatin accessibility sequencing data, gene activity was quantified, amplified fragments and breakpoints were identified, CNVs were detected using the AMULET tool, and visualization was performed using OmicCircos and circlize to verify the presence of ecDNA.

Benefits of technology

This method enables accurate detection of copy number variations, breakpoints, and characteristic distributions of ecDNA at the single-cell level, revealing the functional role of ecDNA in tumors and filling the gaps in existing methods. It can also analyze enhancer hijacking, cross-functional regulation, and tumor heterogeneity of ecDNA.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120290730B_ABST
    Figure CN120290730B_ABST
Patent Text Reader

Abstract

The application discloses an ecDNA recognition method based on scATAC-seq data. The ecDNA recognition method provided by the application can accurately detect copy number amplification, breakpoint quantity and ecDNA feature quantity on the whole cell level. The method provided by the application can recognize specific circular oncogenes carried in tumors, analyze the distribution characteristics of the circular oncogenes in the genome, and reveal the overall composition of all circular DNAs in a single cell and the specific composition of a single circular DNA. In addition, the method can also verify the existence of the circular DNA through a staining experiment, so that the reliability of the detection result is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of biological information, and in particular to an ecDNA identification method based on scATAC-seq data. BACKGROUND

[0002] Recent studies have shown that ecDNA often contains specific oncogenes, drug resistance genes, immune regulatory genes and other regulatory elements, making it an important factor in the pathogenesis and poor prognosis of various cancer types. However, many aspects of ecDNA are still unclear, including its role in abnormal gene expression, cancer evolution, biogenesis, maintenance and clearance. Given that ecDNA has high tumor gene copy number and enhanced chromatin accessibility, ATAC-seq is a powerful tool for studying the functional roles of ecDNA in enhancer hijacking, trans-acting regulation, tumor heterogeneity and uneven distribution.

[0003] Current analysis tools can provide detailed insights into the structure of ecDNA. They can identify amplified fragments and breakpoints and reveal the circular structure of ecDNA, including complex breakpoint edges, by using whole genome sequencing (WGS), ATAC-seq, Circle-seq and long read sequencing data. However, single-cell analysis of ecDNA is still limited. Although single-cell extrachromosomal circular DNA and transcriptome sequencing (scEC&T-seq) can isolate and sequence mRNA and circular DNA at the single-cell level, its application is usually limited to a few hundred cells. This limitation hinders in-depth exploration of ecDNA heterogeneity and clonal dynamics. In addition, single-cell studies on ecDNA clonal trajectories are still lacking. SUMMARY

[0004] To meet the demand for ecDNA single-cell analysis, the present application develops an ecDNA identification method based on single-cell ATAC-seq data, which can identify ecDNA in single-cell data and visualize it, and subsequently verifies the existence of the identified ecDNA through fluorescence in situ hybridization (FISH) and immunohistochemistry. In summary, the present method has strong capabilities in analyzing copy number variation (CNV), breakpoints and feature distribution of ecDNA in single-cell ATAc-seq data.

[0005] To solve the above technical problems, the technical solution of the present application is as follows: an ecDNA identification method based on scATAC-seq data, the method operates as follows:

[0006] Step (1) After preprocessing the single-cell chromatin accessibility sequencing (scATAC-seq) data, gene activity is quantified by calculating the number of DNA fragments with open chromatin in the gene body and promoter regions, and then cell type annotation is performed by the expression of typical genes of various cell types in the tumor microenvironment.

[0007] Step (2) Identify amplified fragments in single-cell sequencing data. The specific operations are as follows: remove data contained in the blacklist region of the hg38 reference genome; filter single-cell data based on multiple identification conditions; identify fragments (gene fragments) with more than 2 reads (short DNA fragment sequences); use AMULET (ATAC-seq MULtipletEstimation Tool) to identify intervals with more than 2 CNVs in each fragment;

[0008] The AMULET used is a Python toolkit for detecting amplified gene fragments based on single-core ATAC-seq data;

[0009] The multiple identification conditions are: Mapping Quality Score (MAPQ) greater than 5, fragment length less than or equal to 300MB, reads are paired, reads can completely correspond to the hg38 reference genome, and reads are not PCR repetitive sequences.

[0010] Step (3) Identify fragments with breakpoints in single-cell sequencing data. The specific operations are as follows: perform quality control on the data obtained from single-cell sequencing; identify breakpoints based on three criteria; the three criteria are: reads with a length greater than 500bp are considered to contain breakpoints, reads aligned to different chromosomes are considered to contain breakpoints, and reads aligned in the same direction are considered to contain breakpoints.

[0011] Step (4) Identify ecDNA in single-cell data;

[0012] Step (5) Visualize the structure of ecDNA in a single cell.

[0013] As a further description of the above scheme: the detailed process of step (4) identifying ecDNA in single-cell data is as follows:

[0014] The amplified fragments were extended by 2000 bp upstream and downstream;

[0015] Find the intersection between the amplified fragments region and the reads containing the breakpoints;

[0016] Based on the cell barcode, the cells containing the breakpoints are mapped to single-cell data;

[0017] Annotate the genes contained in the ecDNA fragment;

[0018] Identify the characteristic genes of ecDNA.

[0019] As a further description of the above scheme: the selected cellular ecDNA is visualized using the R packages OmicCircos and circlize.

[0020] OmicCircos, an R package focused on multi-omics visualization, is used to display different types of data across multiple tracks. irclize, an R package for creating pie charts, is particularly suitable for visualizing complex relationships and hierarchical data structures. Using these two packages together allows for a clear and detailed representation of the circular conformation of ecDNA and its biological characteristics.

[0021] As a further description of the above scheme: the single-cell sequencing data comes from single-cell sequencing of tissue samples, and the specific process is as follows:

[0022] Collected paired normal tissue samples, precancerous lesion tissue samples, and cancerous tissue samples from multiple patients;

[0023] Tissue samples were stored in MACS tissue storage solution; the tissue samples were subjected to enzymatic and mechanical dissociation using a dissociation kit and dissociation instrument, and the cell suspension was obtained by filtration.

[0024] The filtered cell suspension was treated with a Dead Cell Removal kit to remove apoptotic and dead cells; then, a single-cell kit was used for library preparation; the prepared library was evaluated for quality; cDNA integrity was checked, and single-cell sequencing was performed.

[0025] As a further description of the above scheme, the preprocessing procedure for the single-cell sequencing data is as follows:

[0026] Single-cell data normalization eliminates technical biases and makes gene expression levels comparable between cells;

[0027] Feature selection analysis screens highly variable genes, reducing noise and preserving biological differences;

[0028] Single-cell data dimensionality reduction compresses data dimensions, preserves key biological variations, and facilitates visualization and clustering;

[0029] Cluster analysis, based on dimensionality-reduced data, divides cells into biologically significant groups.

[0030] As a further description of the above scheme: the existence of the hypothesized ecDNA was verified using experimental methods, including: fluorescence in situ hybridization to verify the presence of the top marker gene and immunohistochemical experiments to locate the protein.

[0031] The ecDNA identification method provided by this invention can accurately detect copy number amplification, breakpoint number, and ecDNA feature quantity at the whole-cell level. This method can identify specific circular oncogenes carried in tumors, analyze their distribution characteristics in the genome, and reveal the overall composition of all circular DNA within a single cell and the specific composition of individual circular DNA. Furthermore, this method can also verify the presence of circular DNA through staining experiments, ensuring the reliability of the detection results.

[0032] Compared with existing technologies, the present invention has the following advantages: The present invention identifies ecDNA based on large-scale single-cell ATAC-seq data, filling the gap in existing ecDNA research methods; the present invention can be used to analyze the role of ecDNA in enhancer hijacking, cross-functional regulation, tumor heterogeneity and uneven distribution, and can more comprehensively characterize the functional features of ecDNA. Attached Figure Description

[0033] Figure 1 This document provides a schematic diagram for identifying ecDNA and its structure in single-cell ATAC-seq data. a) It demonstrates how to predict ecDNA structure using single-cell ATAC-seq data from squamous cell carcinoma, precancerous actinic keratosis, and normal skin tissue samples. b) It shows the distribution of 27,792 cells in the single-cell ATAC-seq dataset by cell type and sample type in a UMAP plot after dimensionality reduction clustering and cell type annotation. c) It shows the distribution of ecDNA-related copy number amplifications in cells. d) It shows the distribution of ecDNA-related breakpoints in cells. e) It shows the distribution of genes constituting ecDNA in different cells.

[0034] Figure 2 This diagram illustrates the distribution of ecDNA genes, cell types, and chromatin. a) shows the total number of amplified fragments (CNV>2) and breakpoints identified in different cell types; b) shows the calculated number of amplified fragments and breakpoints for known ecDNA-related genes; c) shows the proportion of high-quality fragments at transcription start sites (TSS) in single-cell ATAC-seq data; d) shows the distribution of ecDNA in different genomic regions, categorized by the proportion of various genomic elements.

[0035] Figure 3A schematic diagram of the KLF7 ecDNA structure in a single cell; a shows the distribution of various ecDNA features in a representative cell; b shows the KLF7 ecDNA structure in the same cell; c shows the presence of KLF7 ecDNA confirmed by FISH experiments in frozen and formalin-fixed paraffin-embedded (FFPE) skin squamous cell carcinoma tissue samples. Detailed Implementation

[0036] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. In the following description, specific details such as specific configurations and components are provided merely to help fully understand the embodiments of this application. Therefore, those skilled in the art should understand that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this application. In addition, for clarity and brevity, descriptions of known functions and structures are omitted in the embodiments.

[0037] It should be understood that the phrase "an embodiment" or "this embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "an embodiment" or "this embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments.

[0038] Furthermore, reference numerals and / or letters may be repeated in different examples within this application. Such repetition is for the purpose of simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or settings discussed.

[0039] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, B exists alone, and A and B exist simultaneously. The term " / and" in this article describes another type of relationship between related objects, indicating that two relationships can exist. For example, A / and B can mean: A exists alone, and A and B exist alone. In addition, the character " / " in this article generally indicates that the related objects before and after it are in an "or" relationship.

[0040] In this article, the term "at least one" is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, "at least one of A and B" can mean: A exists alone, A and B exist simultaneously, or B exists alone.

[0041] It should also be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion.

[0042] Example 1

[0043] A method for identifying ecDNA based on scATAC-seq data includes copy number-based identification of cells containing ecDNA, ecDNA breakpoint identification rules, and ecDNA visualization, comprising:

[0044] S101. Collect skin tissue samples;

[0045] S102. Perform single-cell sequencing on the tissue sample to obtain single-cell sequencing data of the tissue sample;

[0046] S103. Preprocess the single-cell sequencing data of the tissue sample;

[0047] S104. Perform cell annotation on the preprocessed single-cell sequencing data;

[0048] S105, Identify fragments amplified in single-cell sequencing data;

[0049] S106. Identify fragments with breakpoints in single-cell sequencing data;

[0050] S107. Identify ecDNA in single-cell data;

[0051] S108, Visualization of the structure of ecDNA in single cells.

[0052] Preferably, the specific operation of collecting skin tissue samples in step S101 is as follows: collecting normal skin tissue samples, precancerous actinic keratosis tissue samples, and skin squamous cell carcinoma tissue samples from three patients.

[0053] Preferably, step S102, which involves performing single-cell sequencing on the tissue sample to obtain single-cell sequencing data, specifically includes the following steps:

[0054] S301. Skin tissue samples are stored in MACS tissue storage solution;

[0055] S302. The tissue sample is subjected to enzymatic digestion and mechanical dissociation using a dissociation kit and a dissociation instrument to obtain a cell suspension;

[0056] S303. Filter the cell suspension of the tissue sample using a 70µm cell filter to obtain the filtered cell suspension.

[0057] S304. Use the Dead Cell Removal kit to remove apoptotic and dead cells from the filtered cell suspension;

[0058] S305. Prepare a library from the cell suspension using a single-cell kit;

[0059] S306. Perform library quality assessment on the prepared library;

[0060] S307. Detect cDNA integrity;

[0061] S308, single-cell sequencing.

[0062] Preferably, the preprocessing of the single-cell sequencing data of the tissue sample in step S103 specifically includes the following steps:

[0063] S401, Cell Ranger ATAC data upstream processing flow;

[0064] This application uses the 10×Genomics Cell Ranger ATAC workflow to process freshly sequenced scATAC-seq data. This workflow includes aligning FASTQ files to a reference genome, removing PCR duplicates, identifying valid cells, generating open chromatin regions (peaks), and outputting fragments files. The final generated fragments.tsv.gz and peaks.bed files can be used for downstream analysis.

[0065] S402. Construct a Seurat object using the result data from the Cell Ranger;

[0066] 1. Data Reading

[0067] First, following the workflow of the SignacR package for processing scATAC data, the peak file (peaks.bed) for each sample was read and converted into GenomicRanges objects for subsequent genome-wide operations. Next, the single-cell metadata file (singlecell.csv) for each sample was read. These files contain filtering information for each cell (such as passed_filters). By filtering out low-quality cells (passed_filters>500), sufficient sequencing depth was ensured for cells used in subsequent analyses.

[0068] 2. Peak Merging and Filtering

[0069] Peak data from all samples were merged into a single peak set (combined.peaks), and overlapping peak regions were removed using the reduce function to ensure that each peak is genomically unique. The merged peak set was further filtered to remove peaks that were too short (<20 bp) or too long (>10000 bp), retaining high-quality peak regions for subsequent analysis.

[0070] 3. Fragment object creation and counting matrix generation

[0071] For each sample, a FragmentObject is created containing fragment information for each cell. This fragment information is used to generate a FeatureMatrix for each sample, where each row represents a peak, each column represents a cell, and the values ​​in the matrix represent the fragment count for each cell at each peak. The generated count matrix is ​​encapsulated into a ChromatinAssay object, and further encapsulated into a Seurat object, while metadata information (such as cell filtering information) is added to the Seurat object.

[0072] 4. Merge all samples into a single seurat object.

[0073] Finally, the Seurat objects from all samples are merged into a single unified Seurat object (combined_atac), and cell origin information (dataset) is added to differentiate between different samples. Human genome gene annotations are added to the Seurat object. The merged object contains peak count matrices and metadata information for all samples, facilitating subsequent joint analysis.

[0074] S403. Perform quality control on Seurat objects;

[0075] 1. Nucleosome signaling

[0076] Calculate the nucleosome signal for each cell. The nucleosome signal reflects the length of open regions of DNA, which are typically 1 to 2 nucleosome lengths (147 bp). Cells with excessively high nucleosome signals (e.g., nucleosome_signal > 4) may contain more nucleosome contamination and are therefore usually filtered out.

[0077] 2. TSS enrichment score

[0078] Calculate the TSS enrichment score for each cell. The TSS enrichment score reflects the degree of open chromatin enrichment near the transcription start site (TSS) in the sequencing data. A higher TSS enrichment score generally indicates that there are more active gene regions in the cell, while cells with lower scores (TSS < 3) will be filtered out.

[0079] 3. Proportion of intrapeak segments

[0080] Calculate the proportion of fragments falling within open chromatin regions (peaks) in each cell. A higher proportion of fragments within peaks generally indicates more open chromatin regions in the cell, which may be associated with gene activity. Cells with a proportion of fragments within peaks greater than 15% are retained.

[0081] 4. Proportion of blacklisted areas

[0082] Calculate the proportion of fragments in each cell that fall within a genome blacklist region. Blacklist regions typically contain repetitive sequences or other difficult-to-map areas; a higher blacklist proportion may indicate poor data quality, and cells with a blacklist region proportion less than 0.05 are retained.

[0083] Preferably, the cell annotation of the pretreated cells in step S104 specifically includes the following steps:

[0084] S405, Data Dimensionality Reduction;

[0085] The signacR package was used to perform TF-IDF normalization on the cell peaks to correct for sequencing depth differences between cells and sparsity among peaks. Peaks with a total count greater than 10 in all cells were selected as high-expression features. Singular value decomposition was then performed on the peak matrix of the normalized cells to reduce the cells to a lower dimension.

[0086] S406, Cell clustering;

[0087] Based on the dimensionality reduction results, a K-nearest neighbor graph of the cells is constructed, and the Louvain algorithm is used to cluster the cells into different clusters.

[0088] S407, Calculation of gene activity matrix;

[0089] Gene activity is quantified by calculating the number of DNA fragments with open chromatin in the gene body and promoter regions, and the gene activity matrix of the cell is normalized by LogNormalize.

[0090] S408, Cell Annotation;

[0091] By calculating the expression differences of typical marker genes of various cells (such as immune cells, stromal cells, and tumor cells) in different clusters in the tumor microenvironment, the gene activity characteristics of each cluster are analyzed, and then the cluster is annotated as a specific cell type based on these characteristics.

[0092] Preferably, the identification of amplified fragments in single-cell sequencing data in step S105 specifically includes the following steps:

[0093] S501. Remove data contained in the blacklisted region of the hg38 reference genome;

[0094] S502, Single-cell data filtering based on 5 conditions;

[0095] S503, Identify fragments with a read count greater than 2;

[0096] S504. Use AMULET to identify intervals in each fragment where CNV is greater than 2.

[0097] Preferably, the five conditions are: Mapping Quality Score (MAPQ) greater than 5; fragment length less than or equal to 300 MB; reads are paired; reads can be completely mapped to the hg38 reference genome; and reads are not PCR repetitive sequences.

[0098] Preferably, the identification of fragments with breakpoints in single-cell sequencing data in step S106 specifically includes the following steps:

[0099] S601. Perform quality control on data obtained from single-cell sequencing;

[0100] S602. Breakpoint identification based on three criteria.

[0101] Preferably, the three criteria are: reads longer than 500 bp are considered to contain breakpoints; reads aligned to different chromosomes are considered to contain breakpoints; and reads aligned in the same direction are considered to contain breakpoints.

[0102] Preferably, the identification of ecDNA in single-cell data in step S107 specifically includes the following steps:

[0103] S701. Extend the amplified fragments by 2000bp upstream and downstream;

[0104] S702, Find the intersection of the amplified fragments region and the reads containing the breakpoints;

[0105] S703. Match cells containing breakpoints with single-cell data based on cell barcodes;

[0106] S704. Annotate the genes contained in the ecDNA fragment;

[0107] S705, Identify the characteristic genes of ecDNA.

[0108] Preferably, the visualization of the ecDNA structure in a single cell described in step S108 specifically includes the following steps:

[0109] S801. Use the R packages OmicCircos and circlize to visualize the selected cell ecDNA.

[0110] Example 2

[0111] This embodiment is based on the above embodiment 1, and the similarities with the above embodiment 1 will not be repeated.

[0112] This embodiment describes a method for experimentally verifying the presence of hypothesized ecDNA. The specific steps are as follows:

[0113] S901, fluorescence in situ hybridization confirmed the presence of the top marker gene;

[0114] This experiment used fluorescence in situ hybridization (FISH) to detect and analyze target DNA in paraffin-embedded tissue sections of human squamous cell carcinoma of the skin (CSCC). First, the paraffin sections underwent routine dewaxing and were hydrated to distilled water for later use. For antigen retrieval, the sections were placed in boiling distilled water or pH 9.0 EDTA solution for 25-30 minutes. Simultaneously, a staining tank was preheated in a 37°C water bath, 40 mL of distilled water was added, and after equilibration for 10-20 minutes, 1 mL of 1 M hydrochloric acid and 1 mL of 10% pepsin solution were added sequentially as a digestion solution. After retrieval, the sections were removed, washed with PBS buffer for 5 minutes, and then placed in the preheated 37°C digestion solution for 15-30 minutes. The specific digestion time was adjusted according to the sample tissue density and section thickness. Subsequently, the sections were washed twice with PBS buffer for 5 minutes each time, and then dehydrated sequentially with 70%, 95%, and 100% ethanol for 2 minutes each, and then air-dried at room temperature. After dehydration and air drying, add 10 μL of probe mixture (probe and hybridization buffer diluted at a ratio of 1:9) to the center of the slide, cover with a 22 mm × 22 mm coverslip, and seal the edges with FISH mounting adhesive to prevent evaporation. Before the hybridization step, prepare two incubators, set to 88℃ and 37℃ respectively. Preheat the light-protected humidified chamber in the 88℃ incubator for at least 10 minutes. Then, place the mounted slides along with the humidified chamber in the 88℃ incubator for denaturation for 10 minutes, and then transfer them to the 37℃ incubator for overnight hybridization for 16-20 hours. Before the end of hybridization, prepare two jars of 2×SSC (containing 0.1% NP-40) washing buffer. Preheat one jar in a 60±1℃ water bath, and keep the other at room temperature. After hybridization, remove the mounting adhesive and coverslip, wash the slides in room temperature washing buffer for 5 minutes, and then wash them in 60℃ washing buffer for 2 minutes. After washing, remove excess liquid, add 20 μL of DAPI anti-quenching mounting solution to the slide, cover with a coverslip, incubate in the dark for 20 minutes, and then observe and photograph under a fluorescence microscope. The probes used in this experiment included the gene-targeting probes KLF7 (FAM-labeled), MYC (Cy5-labeled), NRP2 (FAM-labeled), and PTMA (FAM-labeled) synthesized by Beijing Qingke Biotechnology Co., Ltd.; the centromere probes provided by EXONBIO included the human chromosome 8 centromere probe (FITC-labeled) and the human chromosome 2 centromere probe (Cy3-labeled).

[0115] S902, immunohistochemical localization of the protein;

[0116] This experiment used immunohistochemical staining (IHC) to detect and analyze the expression location of KLF7 protein in paraffin-embedded tissue sections of cutaneous squamous cell carcinoma (CSCC). The sections used were 4–5 μm thick, mounted on glass slides, and baked at 75°C for 40 minutes. The sections were dewaxed twice with xylene for 10 minutes each time, followed by dehydration with a gradient of 100%, 95%, 85%, and 75% ethanol for 3 minutes each step to complete the rehydration process. Subsequently, the sections were placed in 10 mM, pH 6.0 citrate buffer and microwaved for 10–15 minutes to repair the antigen. After natural cooling to room temperature, they were washed three times with PBS buffer for 5 minutes each. Endogenous peroxidase activity was blocked by 3% hydrogen peroxide for 10 minutes, followed by blocking of non-specific antigen sites with 10% goat serum and incubation at room temperature for 30 minutes. KLF7 primary antibody (rabbit anti-KLF7 polyclonal antibody, HUABIO, catalog number ER1911-99, dilution 1:200) was then added, and the slides were incubated overnight at 4°C. The next day, after washing three times with PBS, biotin-labeled secondary antibody (goat anti-rabbit IgG, HUABIO, catalog number HA1001) was added, and the slides were incubated at room temperature for 30 minutes. Then, a streptavidin-HRP complex was added, and the slides were incubated for another 30 minutes. DAB chromogenic solution was used for development, and the reaction time was controlled by observation under a microscope until a satisfactory brownish-yellow signal appeared (approximately 2–5 minutes). After terminating the reaction, the slides were washed with distilled water, counterstained with hematoxylin for 1–2 minutes, and then blued again in Scott's tap water substitute solution. The slides were then dehydrated sequentially with 75%, 85%, 95%, and 100% graded ethanol solutions and cleared in xylene. Finally, add neutral resin (Leica DPX mounting solution, catalog number 3808600E) to mount the slide, let it air dry at room temperature, and then observe and photograph it under a bright field microscope.

[0117] Example 3

[0118] This embodiment is based on Embodiment 1 described above, and the similarities with Embodiment 1 will not be repeated. This embodiment introduces the identification and visualization of ecDNA.

[0119] In this embodiment, normal skin tissue samples, precancerous actinic keratosis tissue samples, and squamous cell carcinoma tissue samples from three patients were collected and single-cell ATAC sequencing was performed.

[0120] Step 1: Identify ecDNA and its structure from single-cell ATAC-seq data. Results are attached. Figure 1 As shown, Figure 1 middle:

[0121] a. A schematic diagram of the workflow of this method illustrates how to predict ecDNA structure using single-cell ATAC-seq data from squamous cell carcinoma, precancerous actinic keratosis, and normal skin tissue samples. The process includes steps such as double cell removal, identification of DNA amplification fragments with copy number variation (CNV) > 2, and reconstruction of the ecDNA circular structure using the three types of ecDNA breakpoints mentioned above.

[0122] b. In the single-cell ATAC-seq dataset, the distribution of 27,792 cells in the UMAP plot after dimensionality reduction clustering cell type annotation is shown by cell type and sample type (squamous cell carcinoma of the skin, precancerous actinic keratosis, and normal skin tissue).

[0123] c. The figure shows the distribution of ecDNA-related copy number amplification in cells.

[0124] d. The figure shows the distribution of ecDNA-related breakpoints in the cell.

[0125] e. The figure shows the distribution of genes that make up ecDNA in different cells.

[0126] Step 2: ecDNA gene, cell type, and chromatin distribution; results are attached. Figure 2 As shown, Figure 2 middle:

[0127] a. The figure shows the total number of amplified fragments (CNV>2) and breakpoints identified in different cell types.

[0128] b. The figure shows the number of fragment amplifications and breakpoints of known ecDNA-related genes calculated in this experiment.

[0129] c. The figure shows the proportion of high-quality fragments at transcription start sites (TSS) in single-cell ATAC-seq data.

[0130] d. The figure shows the distribution of ecDNA in different genomic regions, classified according to the proportion of various genomic elements (e.g., promoters, exons, introns, intergenic regions).

[0131] Step 3: Construction of ecDNA molecules in single cells and the ecDNA structure of a single gene (KLF7) in specific cells. The results are shown in the attached figure. Figure 3 As shown.

[0132] a. The figure shows the distribution of various ecDNA features in a representative cell (ID: TACCTCGGTAAGCCGA-1). The outer loop shows the gene name and its corresponding chromosome, the middle loop shows copy number variation (CNV), and the inner loop indicates the location of detected breakpoints.

[0133] b. The figure shows the KLF7 ecDNA structure in the same cell (ID: TACCTCGGTAAGCCGA-1). The outer loop shows the genomic start and end positions of CNVs associated with KLF7, the middle loop highlights the CNV values, and the inner loop marks the breakpoint start and end positions of KLF7.

[0134] c. FISH assays confirmed the presence of KLF7 ecDNA in frozen and formalin-fixed paraffin-embedded (FFPE) squamous cell carcinoma tissue samples of the skin.

[0135] It should be understood that the specific embodiments described above are merely illustrative or explanatory of the principles of the invention and do not constitute a limitation thereof. Therefore, any modifications, equivalent substitutions, improvements, etc., made without departing from the spirit and scope of the invention should be included within the protection scope of the invention. Furthermore, the appended claims are intended to cover all variations and modifications falling within the scope and boundaries of the appended claims, or equivalent forms of such scope and boundaries.

Claims

1. A non-disease diagnostic method for identifying ecDNA based on scATAC-seq data, characterized in that, The method is operated as follows: Step (1) After preprocessing the scATAC-seq data of single-cell chromatin accessibility sequencing, gene activity is quantified by calculating the number of DNA fragments with open chromatin in the gene body and promoter regions, and then cell type annotation is performed by the expression of typical genes of various cell types in the tumor microenvironment. Step (2) Identify amplified fragments in single-cell sequencing data. The specific operations are as follows: remove data contained in the blacklist region of the hg38 reference genome; filter single-cell data based on multiple identification conditions; identify fragments with more than 2 reads. Use AMULET to identify intervals in each fragment where CNV is greater than 2; The AMULET used is a Python toolkit for detecting amplified gene fragments based on single-core ATAC-seq data; The multiple identification conditions are: mapping quality score greater than 5, fragment length less than or equal to 300MB, reads are paired, reads can completely correspond to the hg38 reference genome, and reads are not PCR repetitive sequences. Step (3) Identify fragments with breakpoints in single-cell sequencing data. The specific operations are as follows: perform quality control on the data obtained from single-cell sequencing; identify breakpoints based on three criteria; the three criteria are: reads with a length greater than 500bp are considered to contain breakpoints, reads aligned to different chromosomes are considered to contain breakpoints, and reads aligned in the same direction are considered to contain breakpoints. Step (4) Identify ecDNA in single-cell data; Step (5) Visualize the structure of ecDNA in a single cell; The detailed process of step (4) identifying ecDNA in single-cell data is as follows: The amplified fragments were extended by 2000 bp upstream and downstream; Find the intersection between the amplified fragments region and the reads containing the breakpoints; Based on the cell barcode, associate the cells containing breakpoints with single-cell data; Annotate the genes contained in the ecDNA fragment; Identify the characteristic genes of ecDNA.

2. The ecDNA identification method based on scATAC-seq data for non-disease diagnostic purposes according to claim 1, characterized in that, Use the R packages OmicCircos and circlize to visualize the selected cell ecDNA.

3. The method for identifying ecDNA based on scATAC-seq data for non-disease diagnostic purposes according to claim 1, characterized in that, The single-cell sequencing data is derived from single-cell sequencing of tissue samples, and the specific process is as follows: Collected paired normal tissue samples, precancerous lesion tissue samples, and cancerous tissue samples from multiple patients; Tissue samples were stored separately in MACS tissue storage solution; The tissue samples were subjected to enzymatic and mechanical dissociation using a dissociation kit and dissociation instrument, and the resulting cell suspension was obtained by filtration. The filtered cell suspension was treated with the Dead Cell Removal kit to remove apoptotic and dead cells; Then, a single-cell kit was used to prepare the library; the prepared library was evaluated for quality; cDNA integrity was checked, and single-cell sequencing was performed.

4. The method for identifying ecDNA based on scATAC-seq data for non-disease diagnostic purposes according to claim 1, characterized in that, The preprocessing procedure for the single-cell sequencing data is as follows: Single-cell data normalization eliminates technical biases and makes gene expression levels comparable between cells; Feature selection analysis screens highly variable genes, reducing noise and preserving biological differences; Single-cell data dimensionality reduction compresses data dimensions, preserves key biological variations, and facilitates visualization and clustering; Cluster analysis, based on dimensionality-reduced data, divides cells into biologically significant groups.

5. The method for identifying ecDNA based on scATAC-seq data for non-disease diagnostic purposes according to claim 1, characterized in that, The existence of the hypothesized ecDNA was verified using experimental methods, including fluorescence in situ hybridization to verify the presence of the top marker gene and immunohistochemical experiments to locate the protein.

Citation Information

Patent Citations

  • Single-cell multi-omics parallel sequencing method based on third-generation sequencing platform and application of single-cell multi-omics parallel sequencing method

    CN117106873A

  • METHODS AND COMPOSITIONS FOR DETECTING ecDNA

    US20220364182A1