Single-cell multi-body full-length sequencing analysis method adopting DNA fragment multi-combination assembly reaction

Through the multi-combination assembly reaction of DNA fragments and CRISPR/Cas nuclease treatment, the low quality and cellular heterogeneity of single-cell multiple full-length sequencing are solved, and efficient and accurate single-cell multi-group full-length sequencing is achieved, which is suitable for genotype analysis and cancer-specific multi-group analysis.

CN120380168APending Publication Date: 2025-07-25EYEONCELL CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380084726.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-08
Filing Date
2023-12-08
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The prior art has limitations in the study of low quality and cellular heterogeneity in single-cell multiple full-length sequencing, resulting in low sequencing efficiency and high error rate, making it difficult to accurately analyze complex genomic structures and cellular heterogeneity.

Method used

The DNA fragment multi-combination assembly reaction is adopted, including connecting the multi-link assembly linker of the DNA fragment to a single-cell library, selectively cyclizing DNA molecules through the multiple ligation assembly reaction, and using exonuclease to remove fragments other than circular DNA, and selective linearization treatment is performed in combination with CRISPR/Cas nuclease.

Benefits of technology

It significantly improves the efficiency and accuracy of full-length sequencing, and can perform efficient multi-group full-length sequencing at the single-cell level, especially selective sequencing of targeted DNA molecules, and is suitable for genotype analysis, mutation detection and cancer-specific multi-group analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120380168A_ABST
    Figure CN120380168A_ABST
Patent Text Reader

Abstract

The invention relates to a single-cell multi-body full-length sequencing analysis technology. The present invention is a developmental study that overcomes the limitations of cell heterogeneity studies and low quality of single-cell multiple full-length sequencing, in which an assembly method that allows a combination of a plurality of DNA fragments is used to greatly improve the efficiency and error rate of full-length sequencing and complete single-cell multi-body full-length sequencing, therefore, the method can be greatly used for gene subtype expression level analysis, mutation detection or cancer cell specific multi-body analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a method for single-cell multi-body full-length sequencing analysis using DNA fragment multi-combination assembly reaction. Background Art

[0002] The ability to accurately and rapidly sequence genomes has enabled innovative biology and medicine. Research on complex genomes, especially on the genetic basis of human diseases, involves large-scale genetic analysis. Genome-wide genetic analysis is not only costly but also time-consuming and laborious. These costs increase when using protocols that involve analyzing individual single DNA samples. Sequencing polymorphic regions in disease-related genomes greatly contributes to understanding diseases such as cancer and drug development and leads to the completion of pharmacogenomics tasks aimed at identifying genes and functional polymorphisms associated with variability in drug response.

[0003] After the completion of the Human Genome Project, next-generation sequencing (NGS) technologies have made great progress in the past decade. Second-generation NGS technologies (such as Illumina or IonTorrent) are very powerful, but they can only sequence short reads with lengths of 100 to 500 bp, so the length of analyzable molecules is limited. Therefore, second-generation technologies have limitations in understanding complex genomic structures such as repetitive sequences and copy number variations (CNVs). To improve this, third-generation long-read sequencing methods have been developed. However, there are limitations in preparing long-read single-molecule sequencing libraries for third-generation methods, especially in long-read sequencing, with problems such as high costs and limitations in the resolution of cell heterogeneity studies. Therefore, there is an increasing demand for more efficient and improved sequencing.

[0004] Therefore, the present invention is a research and development aimed at overcoming the low quality and high costs of multi-omics full-length sequencing including single-cell transcriptomes, as well as the limitations of cell heterogeneity studies. By using an assembly method that allows the combination of multiple DNA fragments, the efficiency and error rate of full-length sequencing have been significantly improved, and multi-group full-length sequencing of single cells has been achieved. It is expected to be widely used for expression analysis of gene subtypes, mutation detection, or cancer cell-specific multi-group analysis. Summary of the Invention

[0005]

Technical Problem

[0006] The inventors of the present invention made efforts to research to overcome the low quality of single-cell multi-group full-length sequencing and the limitations of cell heterogeneity. The results showed that by using an assembly method that allows the combination of multiple DNA fragments, they significantly improved the efficiency and error rate of full-length sequencing and completed multi-group full-length sequencing of single cells, thus completing the present invention.

[0007] Therefore, an object of the present invention is to provide a method for single-cell full-length multi-omics sequencing analysis, which provides a sequencing analysis method including the following:

[0008] (a) A step of ligating a multi-ligation assembly adapter of a DNA fragment to a multi-omics library prepared from a single cell;

[0009] (b) A step of selectively circularizing DNA molecules in the library through a multi-ligation assembly reaction of DNA fragments to screen DNA molecules within the library; and

[0010] (c) A step of using an exonuclease to remove DNA fragments other than circular DNA.

[0011] However, the object to be achieved by the present disclosure is not limited to the above object, and other unmentioned objects can be clearly understood by those skilled in the art from the following description.

[0012]

Technical Solution

[0013] Hereinafter, various embodiments described herein will be described in conjunction with the accompanying drawings. In the following description, many specific details are listed, such as specific configurations, compositions, and processes, etc., in order to provide a thorough understanding of the present disclosure. However, some embodiments can be implemented without one or more of these specific details, or in combination with other known methods and configurations. In other cases, known processes and preparation techniques are not described in particular detail so as not to unnecessarily obscure the present disclosure. The phrase "in one embodiment" or "embodiment" mentioned throughout this specification means that the specific features, configurations, compositions, or characteristics described in connection with the embodiment are included in at least one implementation of the present disclosure. Therefore, the phrases "in one embodiment" or "embodiment" that appear in various places in this specification do not necessarily refer to the same embodiment of the present disclosure. In addition, in one or more embodiments, the specific features, configurations, compositions, or characteristics can be combined in any suitable manner.

[0014] Unless otherwise specified in the specification, all scientific and technical terms used in the specification have the same meaning as commonly understood by those skilled in the art to which the present disclosure pertains.

[0015] According to one aspect of the present invention, there is provided a method for single-cell full-length multi-omics sequencing analysis, including the following steps: (a) ligating a multi-ligation assembly adapter of a DNA fragment to a multi-omics library prepared from a single cell; (b) selectively circularizing DNA molecules in the library through a multi-ligation assembly reaction of DNA fragments to select DNA molecules within the library; and (c) using an exonuclease to remove DNA fragments other than circular DNA.

[0016] The method further includes, after step (c), (d) linearizing the circular DNA in the library.

[0017] In this method, the multiple groups may be a transcriptome, spatial transcriptome, genome, spatial genome, proteome, epigenome, or spatial epigenome, and the multiple-group library may contain different adapters (forward and reverse adapters) respectively linked to the 3' and 5' ends. Step (a) may be a process in which two different adapters are assembled to form circular DNA.

[0018] When step (d) is included in the method, the linearization in step (d) may be carried out using a CRISPR / Cas nuclease, a restriction endonuclease, or a transcription activator-like effector nuclease (TALEN), and the CRISPR / Cas nuclease is selected from the group consisting of: CRISPR / Cas9, CRISPR / Cas12, CRISPR-Cas13, or CRISPR / Cas14.

[0019] The inventors of the present invention have made efforts to overcome the limitations of low quality in single-cell multi-group full-length sequencing and the challenges in cell heterogeneity research. Therefore, by utilizing the multiple ligation assembly reaction of DNA fragments, the efficiency and error rate of full-length sequencing have been significantly improved, achieving single-cell multi-group full-length sequencing, especially achieving selective sequencing of only the desired target DNA molecules.

[0020] According to the present invention, "Ouroboros" refers to an ancient symbol in the form of a snake eating its own tail, representing a technique for selectively removing or enriching various DNA targets using an assembly method that allows the ligation of multiple DNA fragments. By utilizing the assembly reaction that allows the ligation of multiple DNA fragments, this technique can selectively target DNA with cell barcodes. Using this selection technique, target enrichment can be performed to detect cancer gene mutations, cDNA derived from the mitochondrial genome, and specific cell populations. In addition, depletion techniques can be used to remove ribosomes and cDNA derived from the mitochondrial genome.

[0021] In this specification, the term "biological sample" refers to any substance, biological fluid, tissue, or cell obtained or derived from an organism. Examples include but are not limited to whole blood, white blood cells, peripheral blood mononuclear cells, buffy coat, plasma, serum, sputum, tears, mucus, nasal wash, nasal aspirate, breath, urine, semen, saliva, peritoneal lavage fluid, pelvic fluid, cystic fluid, cerebrospinal fluid, amniotic fluid, glandular fluid, pancreatic fluid, lymph fluid, pleural fluid, nipple aspirate, bronchial aspirate, synovial fluid, joint aspirate, organ secretions, cells, cell extracts, or cerebrospinal fluid.

[0022] In this specification, the term "transcriptome" refers to the total collection of all expressed RNAs. Related terms include genome, epigenome, spatial transcriptome, proteome, connectome, and multi-omics, which encompass all of these.

[0023] In this specification, the term "library" refers to a collection of DNA molecules in a sufficient quantity to include all individual genes present in a particular organism.

[0024] According to the present invention, the multi-fragment assembly reaction of DNA fragments, also known as multi-DNA fragment assembly, refers to a molecular assembly method that allows the ligation of multiple DNA fragments in a single isothermal reaction. Specifically, it can be a Gibson assembly reaction, but is not limited thereto.

[0025] In this specification, the term "exonuclease" refers to an enzyme that acts by cleaving one nucleotide at a time from the end of a polynucleotide. A hydrolysis reaction occurs, breaking the phosphodiester bond at the 3' or 5' end.

[0026] According to a specific embodiment of the present invention, the multi-omics can be a transcriptome, genome, spatial transcriptome, epigenome, or proteome.

[0027] According to a specific embodiment of the present invention, single cells are used to construct single-cell transcriptome and genome libraries based on one or more techniques selected from the group consisting of: 10X Genomics, Drop Seq, SPLiT-Seq, Sci-Seq, Microwell Seq, SMART Seq, Fluidigm C1 single-cell technology, BGIDNBelab C4 single-cell technology, CITE-Seq, CEL-Seq, scifi RNA-Seq, Sci-Seq, Visium, Slide-Seq, Seq-Scope, and Stereo-Seq technologies.

[0028] In this specification, the term "10X Genomics single-cell sequencing" refers to a single-cell RNA sequencing (RNA-seq) technology. It uses microfluidic dispensing to capture single cells and prepare barcoded next-generation sequencing (NGS) cDNA libraries, thereby allowing transcriptome analysis on a per-cell basis.

[0029] According to a specific embodiment of the present invention, a transcriptome library is constructed such that the Read1 (R1) primer and the cell barcode are ligated to the 3' end of the RNA sequence, and the template-switching oligonucleotide (TSO) primer is ligated to the 5' end.

[0030] According to the present invention, a template-switching oligonucleotide (TSO) refers to an oligonucleotide that hybridizes with the untemplated C nucleotides added by reverse transcriptase during reverse transcription. The TSO adds a common 5' sequence to the full-length cDNA for downstream cDNA amplification.

[0031] According to a specific embodiment of the present invention, multiple sets of libraries are constructed such that the forward adapter sequence is ligated to the 3' end and the reverse adapter sequence is ligated to the 5' end.

[0032] The "primer" of the present invention is a fragment that recognizes the target gene sequence and includes a pair of forward and reverse primers. Preferably, the primer pair provides analytical results with specificity and sensitivity. High specificity can be obtained when the nucleotide sequence of the primer does not match the non-target sequence present in the sample, and the primer only amplifies the target sequence containing the complementary primer binding site without causing non-specific amplification.

[0033] According to a specific implementation manner of the present invention, step (b) involves forming circular DNA by assembling the forward adapter and the reverse adapter into an assembled adapter.

[0034] According to a specific implementation scheme of the present invention, the DNA to be linearized in this step is the full-length multiple sets of libraries of a single cell.

[0035] According to a specific implementation scheme of the present invention, the linearization step involves linearizing the DNA using a CRISPR / Cas nuclease, a restriction endonuclease, or a TALEN (transcription activator-like effector nuclease).

[0036] According to a specific implementation scheme of the present invention, the CRISPR / Cas nuclease is selected from the group consisting of CRISPR / Cas9, CRISPR / Cas12, CRISPR / Cas13, or CRISPR / Cas14.

[0037] In this specification, the term "CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats)-Cas" refers to a gene editing technology originating from the microbial adaptive immune system. It enables simpler and more efficient gene manipulation by precisely cleaving target DNA and then allowing the DNA to repair itself naturally. This system consists of guide RNA (guide RNA, gRNA) and a Cas nuclease, and is also known as RNA-guided engineered nuclease (RGEN) technology. The gRNA is composed of a crRNA that binds complementarily to the target sequence and a tracrRNA required for Cas binding. The gRNA is used to guide Cas to the target region and bind complementarily to the target DNA sequence. The target DNA sequence must contain a PAM (Protospacer Adjacent Motif) sequence at the 3'-end. When the gRNA forms a complex with the Cas nuclease and specifically binds to the target sequence, the Cas nuclease class recognizes the PAM sequence and induces a double-strand break (DSB) in the 3-bp region upstream of the PAM sequence.

[0038] According to the present invention, the Cas is selected from the group consisting of: Cas9, Cas12, Cas13 or Cas14, which are components of the CRISPR / Cas complex. Specifically, it may be Cas9, but is not limited thereto.

[0039] In this specification, the term "Cas9 (CRISPR-associated protein 9)" refers to a 160 kDa protein that plays a crucial role in the immune defense against DNA viruses and plasmids in certain bacteria. It is widely used in genetic engineering applications, and its main function is to cleave DNA, thereby modifying the genome of cells.

[0040] In this specification, the term "Cas12" refers to an RNA-guided endonuclease that forms part of the CRISPR system in certain bacteria. In addition to cleaving the target gene based on the recognition of the protospacer, Cas12 also induces collateral cleavage of non-target single-stranded DNA (ssDNA). Cas12 is used to develop various portable diagnostic technologies by utilizing its collateral cleavage function. When the Cas12a-crRNA complex binds to the target double-stranded DNA, it causes approximately 1250 collateral cleavages per second in the single-stranded DNA, making Cas12a an enzyme capable of simultaneous nucleic acid detection and signal amplification. Cas12 specifically recognizes sequences with a 5'-(T)TTN-3' PAM sequence in the double-stranded structure. Commonly used Cas12 enzymes are Cas12a and Cas12b. In the case where the target is single-stranded DNA, the PAM sequence does not need to be 100% matched, but it can still induce catalytic cleavage even with variations.

[0041] In this specification, the term "Cas13" refers to an RNA-guided RNA endonuclease that does not cleave DNA but only single-stranded RNA. The Cas13 complex forms a ribonucleoprotein complex with gRNA, and the gRNA consists of a 28-30bp spacer region capable of recognizing the target gene. The Cas13 protein is an RNA-degrading enzyme, different from Cas12 which cleaves DNA. Cas13 induces collateral RNA cleavage rather than DNA cleavage. Among the Cas13 group, Cas13a (formerly known as C2C2) was the first to be used for nucleic acid detection. Cas13a cleaves single-stranded RNA targets through a protospacer matching the crRNA. Once bound to single-stranded RNA, the Cas13a protein acts as a non-specific RNA-degrading enzyme and degrades the surrounding RNA even if it is not the target RNA. It is reported that one target recognition can induce at least 10^4 collateral cleavages. Cas13 cleaves the target gene and non-specific surrounding RNA sequences. Since Cas13 cleaves around the complementary protospacer rather than directly on the protospacer, the protospacer remains intact, which allows for multiple cleavages and may lead to an amplification effect, making it useful for sensitive detection. Cas13 does not require a PAM sequence but requires a protospacer flanking site (PFS), where the base guanine is adjacent to the protospacer. The most widely used Cas13 proteins in biotechnology and diagnostics are Cas13a and Cas13b, and Cas13b is known to be more stable for cell gene manipulation.

[0042] According to another aspect of the present invention, there is provided a method for sequencing analysis after selectively removing DNA molecules from a full-length multi-group library of single cells, comprising the following steps: (a) ligating a multi-ligation assembly linker of a DNA fragment to the multi-group library prepared from a single cell; (b) generating circular DNA in the library through a multi-ligation assembly reaction of the DNA fragment; (c) selectively linearizing the circular DNA containing the target sequence to be excised using a CRISPR / Cas nuclease; and (d) removing linear DNA that is not circular DNA using an exonuclease.

[0043] The method may further include (e) linearizing the circular DNA in the library after step (d).

[0044] In the method, the CRISPR / Cas nuclease may be selected from the group consisting of: CRISPR / Cas9, CRISPR / Cas12, CRISPR / Cas13 or CRISPR / Cas14, and the CRISPR / Cas nuclease can only be applied to circular DNA. The linear DNA in step (d) may be a molecule generated by the degradation of the selected circular DNA recognized by the CRISPR / Cas nuclease.

[0045] In this specification, the term "CRISPR / Cas9" refers to a powerful third-generation genome editing tool that is more economical and efficient than the first-generation ZFN (zinc finger nuclease) and the second-generation TALEN (transcription activator-like effector nuclease). The CRISPR-Cas9 system consists of a guide RNA that accurately recognizes and binds to the target DNA and a Cas9 nuclease that cleaves the target DNA, inducing a double-strand break. The double-strand break is repaired, allowing gene manipulation.

[0046] In a specific embodiment of the present invention, the CRISPR / Cas nuclease is selected from the group consisting of: CRISPR / Cas9, CRISPR / Cas12, CRISPR / Cas13, or CRISPR / Cas14.

[0047] According to the present invention, the cleavage of a DNA molecule refers to a reaction that cleaves one of the covalent sugar-phosphate bonds between nucleotides that make up the DNA sugar-phosphate backbone. This reaction is catalyzed by an enzyme, a chemical substance, or radiation.

[0048] In a specific embodiment of the present invention, the CRISPR / Cas nuclease is only applicable to circular DNA.

[0049] In another specific embodiment of the present invention, the linear DNA in the step of removing linear DNA is a selected DNA molecule recognized by the CRISPR / Cas nuclease.

[0050] In a further specific embodiment of the present invention, the DNA in step (e) is a full-length multi-group library of single cells.

[0051] According to another aspect of the present invention, there is provided a sequencing analysis method after screening a DNA molecule from a full-length multi-group library of a single cell, comprising the following steps: (a) preparing a multi-group library from a single cell and ligating a multi-ligation assembly adapter of a DNA fragment to the library; (b) generating circular DNA in the library through a multi-ligation assembly reaction of the DNA fragment; (c) using an exonuclease to remove DNA fragments that are not circular DNA; and (d) selectively converting the circular DNA containing the target sequence to be selected into linear DNA using a CRISPR / Cas nuclease.

[0052] The method may further include (e) amplifying the linear DNA in the library after step (d).

[0053] In the method, the CRISPR / Cas nuclease can be selected from the group consisting of: CRISPR / Cas9, CRISPR / Cas12, CRISPR / Cas13, or CRISPR / Cas14, and the linear DNA in step (d) can be a molecule generated by degradation of the selected circular DNA recognized by the CRISPR / Cas nuclease.

[0054] When step (e) is included in the method, this step can further include a step of ligating a Y-junction or a hairpin junction.

[0055] According to the present invention, Y-junction ligation refers to a method of ligating synthetic oligonucleotides with known nucleotide sequences to both ends or termini of a DNA fragment.

[0056] In a specific embodiment of the present invention, the plurality of groups can be a full-length cDNA transcriptome, a spatial transcriptome, a genome, a spatial genome, a proteome, an epigenome, or a spatial epigenome library derived from a single cell.

[0057] According to the present invention, single-cell sequencing analysis technology can analyze the proteome, transcriptome, genome, spatial transcriptome, or epigenome of a single cell. It allows for the simultaneous analysis of various cells present in a tissue and provides accurate information about the type, lineage, disease, and mutations of individual cells. Single-cell sequencing technology can be effectively applied to fields such as cancer tumor heterogeneity research, immunology, or developmental biology. Specifically, in immunology research, it can comprehensively analyze the immune system, characterize immune cell populations (clonal analysis), and reconstruct immune cell developmental pathways. In addition, through cancer tumor heterogeneity research, it can analyze cancer stem cells or circulating tumor cells (CTCs).

[0058] According to the present invention, single-cell analysis refers to techniques for analyzing genomics, transcriptomics, spatial transcriptomics, epigenomics, or proteomics at the single-cell level, rather than methods for analyzing multiple cells simultaneously. A single cell refers to a cell isolated from a cell population. Due to cell heterogeneity even among cells of the same origin, research at the single-cell level is crucial for accurate analysis. This is particularly emphasized in diseases that require analysis of inherently heterogeneous cell populations, such as stem cell research or cancer, where molecular markers must be studied at the single-cell level rather than the average level of the cell population. With the advancement of single-cell analysis techniques, there is an increasing potential to construct a systematic reference map of all human cells. Some countries are carrying out projects to analyze various diseased tissues using single-cell analysis techniques, and these studies are making positive contributions to diagnosis, healthcare, and drug development. Traditional bulk analysis methods aggregate data from multiple cells and play an important role in accumulating information related to specific omics fields. However, they only provide the average value of each single omics, making it difficult to accurately identify signals directly related to diseases. In contrast, single-cell multi-omics analysis can achieve a more comprehensive understanding of each cell by independently analyzing omics data in different dimensions, thereby more precisely elucidating the complex interaction network system that regulates cell functions. Even cells derived from the same parent can give rise to different cell types and tissues composed of these heterogeneous cell populations. Therefore, identifying and characterizing the specificity of individual cells through single-cell multi-omics analysis is crucial for accurate disease diagnosis and treatment. In addition, single-cell multi-omics analysis integrates and analyzes multi-layer omics information of specific cells, enabling rapid extraction of meaningful insights.

[0059] According to the present invention, multi-omics analysis is an advanced research technique that can measure the comprehensive changes of various biomolecules in cells and integrate and analyze this data. It is particularly useful in basic life sciences and clinical research, and can be used to explain gene functions, understand disease mechanisms, etc. With the introduction of next-generation sequencing (NGS) technology, it has become possible to generate large-scale multi-omics data at low cost, enabling the wide application of multi-omics big data.

[0060] According to the present invention, multi-omics includes genomics, epigenomics, transcriptomics, spatial transcriptomics, or proteomics, which are related to information transmission and regulation in organisms. It also includes metabolites such as metabolomics and lipidomics, which are cell metabolites.

[0061] According to the present invention, genomics research refers to the omics field of studying the genetic composition of an individual using next-generation sequencing (NGS) technology. It includes whole-genome sequencing (WGS) for measuring the entire genome, whole-exome sequencing (WES) for analyzing protein-coding regions, and single nucleotide polymorphism (SNP) genotyping studies. The genomic information generated by these technologies is used in genome-wide association studies (GWAS) to identify disease-causing variants. Based on this, genomics research in genomic medicine is actively exploring the prediction of disease genetic risks and drug responses. In addition, this field is also applied to various fields of life sciences, such as population genetics, evolutionary biology, and synthetic biology.

[0062] According to the present invention, epigenomics research refers to the field of studying epigenetic modifications of the genome, which involves interpreting phenotypic traits affected by environmental factors. It mainly studies changes in DNA methylation and histone modifications, as well as their relationship with genetic regulation. This research allows for the interpretation of phenotypic diversity (phenotypic plasticity) that cannot be explained solely by genetic composition. This field uses experimental techniques such as bisulfite sequencing for measuring DNA methylation, chromatin immunoprecipitation (ChIP) sequencing for identifying histone modification-marked regions, and DNase / ATAC-seq for analyzing open / closed chromatin regions. Recently, epitranscriptomics research focusing on transcriptome modifications such as m6A has been actively conducted, revealing the causes of various biological phenomena such as cell differentiation, development, and tumor formation.

[0063] According to the present invention, transcriptomics research refers to the omics field of comprehensively studying gene expression, including not only coding genes involved in protein synthesis but also non-coding RNAs (such as microRNA) that do not participate in coding. With the introduction of microarray technology, it received great attention in the early 2000s and has since been largely replaced by RNA sequencing (RNA-seq) after the emergence of next-generation sequencing (NGS) technology in 2010. Recently, with the introduction of single-cell transcriptomics, it has become possible to analyze gene expression in different cell types within a tissue, making it a valuable tool for studying disease mechanisms. Transcriptomics research is widely used in life sciences, especially in fields such as cell differentiation, stem cells, tumor formation, gene regulation, and biomarker discovery. It is also applied to expression profiling, alternative splicing analysis, RNA editing analysis, etc.

[0064] According to the present invention, spatial transcriptomics research refers to a technique that simultaneously observes the positional information of each cell and its gene expression pattern in a tissue section, enabling the study of cell distribution and cell-cell interactions within an actual tissue. High-resolution techniques have been developed that can distinguish compartments within dozens of cells down to single cells, and this method is particularly useful for understanding the microenvironment of cancer tissues.

[0065] According to the present invention, proteomics research refers to the omics field that analyzes the protein expression or modification directly mediating physiological processes in an organism. It is a key aspect of functional genomics and has received great attention since the Human Genome Project. This field allows the analysis of post-translational modifications and protein isoforms that cannot be identified by genomic or transcriptomic analysis, as well as the discovery of various biomarkers used in clinical settings. Due to the inherent complexity of proteins, proteomics requires a variety of analytical techniques, including antibody-based affinity proteomics and mass spectrometry-based shotgun proteomics.

[0066] According to the present invention, bead-based single-cell transcriptomics technology is one of the most commonly used technologies in single-cell sequencing. This droplet-based technology revolves around beads containing barcodes. The principle of bead-based single-cell transcriptomics technology is as follows: First, a single cell is paired with a bead. Second, the barcode extracted from the bead is attached to the DNA or RNA extracted from the cell. Third, after amplification and sequencing, the barcode is identified, enabling the analysis of the characteristics of a single cell. An example of using single-cell transcriptomics technology is the Human Cell Atlas (Tabula Sapiens), in which the transcriptomes of approximately one million cells from 24 tissues and organs were analyzed.

[0067] According to the present invention, the multi-ligation assembly reaction of DNA fragments is different from the prior art. Although this process still uses traditional circular DNA ligation, different from the traditional method, the present invention adopts a more complex cyclization process, using uracil-containing primers in semi-nested gene-specific PCR, USERII enzyme and T4 DNA ligase. This design aims to solve the problem of reduced PCR reaction efficiency of circular DNA, including steps: adding a step of linearizing circular DNA using CRISPR / Cas9 endonuclease before the PCR process. In contrast, the existing method does not include this linearization step but directly performs PCR. Specifically, in the existing single-cell RNA sequencing library, biotin-labeled R1 primers are used to select molecules containing cell barcodes. Then these molecules are captured using streptavidin beads that bind to the biotin tag. However, this biotin-based selection method is greatly affected by the molecular size, and when the DNA length exceeds 1 or 2 kb (the average length of human mRNA is about 2.2 kb), the selection efficiency will decrease sharply. In addition, molecules with two cell barcodes (i.e., molecules with barcodes at both ends) are also selected, which may cause potential problems because these molecules are smaller and contain two biotin tags, but they are captured more effectively. In contrast, the method of using multi-ligation of DNA fragments to select molecules containing cell barcodes in the present invention has significant advantages. It is less affected by the DNA molecular size, can effectively select molecules with a single barcode at one end, and eliminates the problem of selecting molecules with two barcodes at both ends.

[0068] In addition, the traditional method of selecting cell barcode-containing molecules in a single-cell RNA sequencing library uses asymmetric PCR to linearly amplify molecules with cell barcodes, so that these molecules can be selectively identified. However, a major problem has emerged with this method: only molecules with a cell barcode at one end will be linearly amplified, while atypical molecules with barcodes at both ends will be exponentially amplified. Therefore, the number of these atypical molecules increases disproportionately, causing the problem of an excess of these unwanted molecules. In contrast, the key advantage of the method used in the present invention is that this method uses multi-ligation of DNA fragments to select molecules containing cell barcodes. It selectively identifies molecules with a single cell barcode at one end and effectively eliminates atypical molecules with barcodes at both ends or without barcodes. This method solves the problem of unnecessary amplification of abnormal molecules and ensures more accurate selection of molecules with expected characteristics.

[0069] According to the present invention, the multi-ligation assembly reaction for selectively removing DNA molecules is different from the prior art. Specifically, conventional techniques for negative enrichment in single-cell RNA-seq libraries use the CRISPR / Cas9 system to remove unwanted DNA (such as ribosomal RNA-derived cDNA) from the library. This method involves using 58 sgRNAs targeting specific sequences, and any linear DNA molecule containing one of these target sequences is recognized and cleaved by the CRISPR / Cas9 system. The resulting fragmented linear DNA molecules cannot bind correctly to the primers required for PCR amplification, and therefore, they will not be amplified in subsequent PCR steps. In contrast, the method for selectively removing DNA molecules in the present invention applies the CRISPR / Cas9 system to circular DNA rather than linear DNA. When circular DNA is recognized and cleaved by the CRISPR / Cas9 system, it is converted into linear DNA, which is then completely removed from the library by exonuclease treatment. This method avoids the potential problem in the traditional method that unwanted molecular fragments (such as cleaved linear DNA) may remain in the PCR process, causing interference. By completely removing unwanted DNA molecules by exonuclease treatment before PCR amplification, the present invention ensures that the library amplification process is more accurate and free from such interference, having significant technical advantages compared to traditional methods.

[0070] According to the present invention, the multi-ligation assembly reaction for selective cDNA target screening is different from the prior art. Specifically, the targeting selection technology using the targeted gene expression Panel provided by 10X Genomics employs a method called hybridization capture. In this method, first, the single-cell RNA-Seq library is separated into single-stranded molecules, and then hybridized with biotin-labeled RNA / DNA probes. Biotin is attached to the target DNA molecules, and streptavidin magnetic beads are used to selectively capture the target DNA molecules. The targeting selection strategy based on hybridization capture relies on the binding of target DNA molecules to RNA / DNA probes, which requires complex experimental conditions to ensure that the target DNA molecules are separated into single strands and can effectively bind to the probes. Optimizing these conditions is challenging because the multiple secondary structures of single-stranded target DNA may hinder their binding to RNA / DNA probes. In addition, a complex capture probe design algorithm is required to predict and prevent the formation of secondary structures that may be caused by the interaction between thousands or tens of thousands of RNA / DNA probes. Another limitation of the hybridization capture method is that as the length of the selected molecules increases, the efficiency of the selection process decreases, making it difficult to select full-length cDNA. In contrast, the targeting selection strategy of the present invention uses the CRISPR / Cas9 system, which can precisely identify the target sequences of DNA molecules. Even with thousands or tens of thousands of target DNA sequences, the sgRNAs that recognize these sequences will form stable complexes with the CRISPR / Cas9 protein, thus forming an RNP (ribonucleoprotein) structure. This physical separation enables each sgRNA to independently search for its respective target DNA sequence without interference from other sgRNAs. Different from the hybridization capture method that requires processing the RNA-Seq library into a single-stranded form under complex conditions, the method in the present invention allows the selection process to be carried out on double-stranded (dsDNA) molecules, making the selection process simpler and more efficient.

[0071] In addition, the prior art of using the CRISPR / Cas9 system to select long DNA molecules relied on the principle that the DNA molecules cleaved by CRISPR / Cas9 were labeled with 5'-phosphates. In this method, the 5'-phosphate groups at the ends of linear DNA molecules were first removed. Then, the CRISPR / Cas9 system was used to cleave the regions around the target DNA, and the cleaved target DNA regions were labeled with 5'-phosphates. Since DNA ligation requires 5'-phosphates, this feature allowed sequencing adapters to be ligated only to the DNA molecules cleaved by CRISPR / Cas9. Therefore, only the DNA molecules cleaved by CRISPR / Cas9 were analyzed using long-read sequencing techniques. However, this prior art encountered a problem that DNA inevitably underwent physical fragmentation (mechanical shearing) during sample processing, and the resulting fragments were randomly labeled with 5'-phosphates at their cleavage sites. This random 5'-phosphate labeling led to inefficient selection of the target DNA. In contrast, the selective DNA target selection method in the present invention employs a multi-ligation assembly reaction, in which the unselected molecules are retained in their circular DNA form, and only the selected target DNA molecules are converted into linear DNA for amplification. Therefore, there is no need to remove or label 5'-phosphate groups. In addition, when purifying circular DNA, exonuclease is used to degrade linear DNA, preventing any adapter from accidentally ligating to the linear DNA to be removed. If necessary, alkaline phosphatase (similar to the previous technique) can be used to remove unwanted 5'-phosphates, ensuring that the adapter is ligated only to the target DNA cleaved by CRISPR / Cas9. This improvement reduces the need for complex procedures while increasing the efficiency of target DNA selection.

[0072] The prior art of using the CRISPR / Cas9 system to select long DNA molecules also relied on the principle that the CRISPR / Cas9 RNP (protein + RNA complex) strongly binds to the target DNA. In this method, the DNA bound to the CRISPR / Cas9 RNP was protected from degradation by exonuclease, similar to being "covered by a cap". As a result, the CRISPR / Cas9 system cleaved the ends of the target DNA region, and only the DNA molecules protected by the CRISPR / Cas9 RNP could be selected through exonuclease treatment. In contrast, the selective DNA target screening method in the present invention uses a multi-ligation assembly reaction, in which the unselected molecules are left in their circular DNA form, and only the selected target DNA is converted into linear DNA for amplification. Therefore, there is no need to use exonuclease to select the molecules protected by the CRISPR / Cas9 RNP. This difference highlights the technique used in the present invention, which, although also utilizing CRISPR / Cas9, is fundamentally different from the prior art.

[0073] Therefore, by effectively performing sequencing analysis of the full-length sequences of multiple genomes in single cells according to the present invention, not only can gene expression be analyzed, but gene subtypes can also be effectively identified and their expression levels analyzed. In addition, gene mutations can be effectively detected. By accurately distinguishing gene subtypes at the single-cell level, it is possible to propose new biomarkers for various types of cancers that were previously difficult to accurately diagnose. Examples of gene subtypes differentially expressed according to cancer types are well summarized in the article "Splice Variations as Cancer Biomarkers, Clinical Biochemistry, 2004". Currently, the method for selecting targeted cancer therapies is to first detect genes highly expressed or mutated in cancer tissues and then select therapies targeting these genes. However, cancer tissues consist of various cancer cell clones, each containing different mutations. Therefore, the conventional method of analyzing the entire cancer tissue to find target genes may be difficult to detect a small number of cancer cell clones with different mutations. Therefore, even with targeted therapy for mutations, a small number of cancer cells with different mutations will remain, leading to cancer recurrence. Therefore, by analyzing cancer tissues at the single-cell level using the present invention, it will be possible to detect a small number of cancer cell clones with different mutations. With this information, targeted cancer therapy can be carried out according to combinations targeting different mutations, thereby potentially reducing the risk of recurrence.

[0074] According to another aspect of the present invention, the present invention provides a single-cell full-length multi-group sequencing analysis kit, which uses the above method to perform full-length multi-group sequencing analysis on a single cell.

[0075] According to another aspect of the present invention, the present invention provides a gene subtype diagnostic marker analysis kit at the single-cell level, which uses the above method to perform full-length multi-group sequencing analysis on a single cell.

[0076] According to another aspect of the present invention, the present invention provides a cancer targeted therapy candidate analysis kit based on single-cell mutations, which uses the above method to perform full-length multi-group sequencing analysis on a single cell.

[0077]

Advantages and Effects

[0078] The features and advantages of the present invention can be summarized as follows:

[0079] (a) The present invention provides a multi-group full-length sequencing analysis technology for single cells.

[0080] (b) The present invention is a development research aimed at overcoming the limitations of low-quality in single-cell multi-group full-length sequencing and cell heterogeneity research. By using an assembly method that allows the combination of multiple DNA fragments, the efficiency and error rate of full-length sequencing have been significantly improved. Through the completion of single-cell multi-group full-length sequencing, it can be widely applied to gene expression analysis, gene subtype analysis, mutation detection, or multi-group analysis specific to cancer cells. In addition, it can also be used for disease diagnosis based on gene subtypes and targeted cancer therapy based on mutations. BRIEF DESCRIPTION OF THE DRAWINGS

[0081] Figure 1 Shows a schematic diagram of the multi-combination assembly and replication method of DNA fragments.

[0082] Figure 2 Shows the working principle of the DNA screening method using the multi-combination assembly and replication method of DNA fragments.

[0083] Figure 3 Shows the result of screening DNA molecules with cell barcodes using the multi-combination assembly reaction of DNA fragments.

[0084] Figure 4 Shows the result of screening DNA molecules with cell barcodes using the multi-combination assembly reaction of single-cell DNA fragments from mouse heart.

[0085] Figure 5 Shows the working principle of selectively removing DNA molecules using the multi-combination assembly reaction of DNA fragments.

[0086] Figure 6 Shows an example of a method for selectively removing DNA molecules using the multi-combination assembly reaction of DNA fragments.

[0087] Figure 7 Shows the working principle of full-length cDNA target screening using the multi-combination assembly reaction of DNA fragments.

[0088] Figure 8 Shows an example of a method for full-length cDNA target screening using the multi-combination assembly reaction of DNA fragments.

[0089] Figure 9 Shows an example of applying the method for selecting full-length cDNA targets using the multi-combination assembly reaction of DNA fragments to a peripheral blood mononuclear cell (PBMC) sample from a patient with relapsed acute myeloid leukemia (AML).

[0090] Figure 10 Shows the result of screening DNA molecules with cell barcodes using the multi-combination assembly reaction of DNA fragments.

[0091] Figure 11 Shows an example of applying the full-length cDNA target selection method.

[0092] Figure 12 Shows the results of generating single-cell RNA-seq libraries using SPLiT-Seq technology.

[0093] Figure 13 Shows a comparison of the results with or without applying the Ouroboros cell barcode selection technology of the present invention.

[0094]

Best Mode

[0095] Single-cell RNA-seq libraries were produced from two mouse cell lines and one human cell line using SPLiT-Seq technology. Thus, before applying the Ouroboros cell barcode screening technology of the present invention, the proportion of molecules with cell barcodes was 8.3%, but after applying the Ouroboros cell barcode screening technology, this proportion increased significantly to 85.1%, more than 10 times higher.

[0096]

Mode of Invention

[0097] Hereinafter, the present disclosure will be described in detail with reference to the following examples. However, the following examples are only for illustrative purposes of the present disclosure, and the content of the present disclosure is not limited by the following examples. Detailed Description of the Invention

[0098] [Example 1]

[0099] 1. Multi-combinatorial assembly reaction scheme of DNA fragments

[0100] Step 1 - Preparation and purification of single-cell cDNA libraries

[0101] Use the 10X Genomics Single Cell 3' Expression Kit or other single-cell RNA-seq library preparation methods.

[0102] Step 2 - Screening of DNA molecules with cell barcodes

[0103] (1) Perform PCR using primers containing a part of R1 plus an additional 22 bp and a part of TSO plus an additional 22 bp, so as to add a 22-bp multi-combinatorial assembly adapter required for self-cyclization assembly reaction to the ends of molecules containing R1 and TSO sequences using the following program.

[0104] [Table 1]

[0105]

[0106] (2) Assemble the self - ligation reaction mixture according to the following table and procedure. Before setting up the reaction, transfer the NEBuilder HiFi DNA Assembly Master Mix to ice and gently tap the tube with your finger to mix it before use.

[0107] [Table 2]

[0108] Component Amount (final concentration) 2X NEBuilder HiFi DNA Assembly Premix 5 μL Deoxyribonucleic acid 10 - 100 ng (1 - 10 ng / μL) Nuclease-free water To 10 μL Total reaction volume 10 μL

[0109] [Table 3]

[0110] Lid temperature Reaction volume Running time 105℃ 10 μL 60 minutes Step Temperature Time 1 50℃ 01:00:00 2 4℃ Hold

[0111] (3) On ice, add 1 μL of Exonuclease III, 1 μL of Lambda Exonuclease, and 1 μL of Exonuclease I to the reaction mixture and use the following procedure to remove linear DNA molecules.

[0112] [Table 4]

[0113] Lid temperature Reaction volume Running time 105℃ 12 μL 30 minutes Step Temperature Time 1 37℃ 00:20:00 2 70℃ 00:10:00 3 4℃ Permanent

[0114] (4) After the degradation process, add 14 μL of SPRI magnetic beads to purify the circular DNA and elute the purified circular DNA with 10 μL of elution buffer. To confirm the amount of circular DNA molecules, analyze using an Invitrogen Qubit 4 fluorometer (using 1 μL).

[0115] (5) In a new tube, add reagents in the following order to assemble the Cas9 - gRNA ribonucleoprotein targeting the self - ligation linker and incubate at room temperature for 10 minutes.

[0116] [Table 5]

[0117]

[0118] (6) Set up the following reaction at room temperature and use the following procedure to linearize all circular DNA molecules.

[0119] [Table 6]

[0120] Component Amount (final concentration) Purified circular DNA 9 μL NEBuffer r3.1 (10x) 1 μL Assembled Cas9-gRNA ribonucleoprotein 20 μL Total reaction volume 30 μL

[0121] [Table 7]

[0122]

[0123]

[0124] (7) Add 1 μL of Proteinase K to the reaction mixture, gently mix, and use a centrifuge to collect the solution at the bottom of the tube.

[0125] (8) Incubate at room temperature for 10 minutes.

[0126] (9) Add 37 μl of SPRI beads for DNA purification and elute the purified DNA with 11 μl of elution buffer. Analyze the amount of circular DNA molecules using Qubit (using 1 μl).

[0127] (10) Use Q5 DNA polymerase or KAPA HiFi HotStart ReadyMix and primers with partial R1 and partial TSO sequences to perform PCR in a 50 μL volume and amplify full-length cDNA using the program shown below.

[0128] [Table 8]

[0129]

[0130] (11) Use SPRI magnetic beads for DNA purification and analyze the sample using Qubit or a bioanalyzer to check the quantity and quality of the DNA contained in the sample.

[0131] This is the process of amplifying linearized full-length molecules (molecules with both a "head" and a "tail") through a PCR reaction.

[0132] Step 3 - Long-read sequencing

[0133] Analyze full-length molecules from single cells using Nanopore / PacBio long-read sequencing technology.

[0134] Analysis of the DNA fragment multi-combination assembly reaction scheme

[0135] This is the process of the DNA fragment multi-combination reaction. Only molecules with both a "head" (R1 and cell barcode, representing the 3' end of mRNA) and a "tail" (TSO, representing the 5' end of mRNA) assemble into a circle, while the remaining molecules do not assemble. After the multi-combination reaction of DNA fragments is completed, the molecules that do not assemble into a circle are removed. In this process, linear molecules are removed by exonucleases. Technically, using a combination of various exonucleases is effective. Once the removal of linear molecules that do not assemble into a circle is completed, the circular molecules are re-linearized using CRISPR / Cas9 endonuclease, which can precisely recognize the multi-combination linker sequence of DNA fragments and linearize them. Other endonucleases, such as restriction endonucleases, can also be used.

[0136] 2. Working principle of the selection method using multi-combinatorial assembly reaction of DNA fragments

[0137] The multi-combinatorial assembly reaction of DNA fragments is based on the characteristic of self-cyclization. Its working principle starts with self-cycling, where the DNA form contains a cellular barcode added to the 3' end of mRNA and a Read1 (R1) primer, as well as a template-switching oligonucleotide (TSO) primer added to the 5' end. Next, circular DNA molecules are selected, while those that have not been circularized are removed by exonuclease. Then, a linearization step is carried out to select the single-cell full-length cDNA library.

[0138] 3. Results of multi-combinatorial assembly reaction of DNA fragments

[0139] Using the multi-combinatorial assembly reaction of DNA fragments, DNA molecules with cellular barcodes are selected. The multi-combinatorial assembly reaction of DNA fragments is used to isolate cells from the mouse heart. First, single cells are isolated from the mouse heart. Then, a 10X single-cell transcriptome library is created, and the multi-combinatorial assembly reaction of DNA fragments is used to select DNA molecules with cellular barcodes, ultimately leading to the selection of target molecules. During the selection of mouse heart single cells, it was confirmed that molecules with correctly ligated cellular barcodes were effectively screened out during the multi-combinatorial assembly reaction ( Figure 4 ).

[0140] As a result of the screening, the proportion of molecules with cellular barcodes in various genes increased from less than 50% to more than 90% ( Figure 3 ). The same effect was also confirmed in 5 other tissues of the mouse, excluding the heart.

[0141] 4. Technical verification

[0142] The results of applying the SPLiT-Seq technology to the single-cell RNA-Seq library created using this technology are as follows. Single-cell RNA-seq libraries were created from two mouse cell lines and one human cell line using the SPLiT-Seq technology. Therefore, before applying the Ouroboros cell barcode screening technology of the present invention, the ratio of molecules to cell barcodes was 8.3%, but after applying the Ouroboros cell barcode screening technology, this ratio increased significantly to 85.1%, more than 10 times higher ( Figure 12 ).

[0143] Using the 10X Genomics technology (3' Gene Expression Kit, v3), a single-cell RNA-seq library was constructed from cells isolated from mouse liver tissue. Before applying the Ouroboros cell barcode screening technology of the present invention, the proportion of molecules with cell barcodes was 19.1%, but after applying the Ouroboros cell barcode screening technology, this proportion increased to 91.1%, approximately 4.8 times higher ( Figure 13 ).

[0144] [Example 2]

[0145] 1. Multi-combinatorial assembly reaction scheme of DNA fragments (selective DNA molecule removal)

[0146] Step 1 – Preparation and Purification of Single-Cell cDNA Libraries

[0147] The 10X Genomics Single Cell 3' Expression Kit or other single-cell RNA-seq library preparation methods can be used.

[0148] Step 2 – Screening of DNA Molecules with Cell Barcodes and Selective Removal of DNA Molecules

[0149] (1) Perform PCR using primers containing partial R1 plus an additional 22 bp and partial TSO plus an additional 22 bp to add the 22-bp multi-combination assembly adapter required for the self-cycling assembly reaction to the ends of molecules containing R1 and TSO sequences using the following program.

[0150] [Table 9]

[0151]

[0152] (2) Assemble the self-ligation reaction mixture according to the following table and program. Before setting up the reaction, transfer the NEBuilder HiFi DNA Assembly premix to ice and gently tap the tube with your finger to mix before use.

[0153] [Table 10]

[0154] Component Amount (final concentration) 2X NEBuilder HiFi DNA Assembly Premix 5 μL Deoxyribonucleic acid 10 - 100 ng (1 - 10 ng / μL) Nuclease-free water To 10 μL Total reaction volume 10 μL

[0155] [Table 11]

[0156] Lid temperature Reaction volume Running time 105℃ 10 μL 60 minutes Step Temperature Time 1 50℃ 01:00:00 2 4℃ Hold

[0157] (3) In a new tube, add reagents in the following order to assemble the Cas9-gRNA ribonucleoprotein and incubate at room temperature for 10 minutes.

[0158] [Table 12]

[0159]

[0160]

[0161] (4) Set up the following reaction at room temperature and linearize the unwanted circular DNA molecules using the following program.

[0162] [Table 13]

[0163] Component Amount (final concentration) Gibson reaction mixture (DNA) 10 μL Assembled Cas9-gRNA ribonucleoprotein 20 μL Total reaction volume 30 μL

[0164] [Table 14]

[0165] Lid temperature Reaction volume Running time 105℃ 30 μL 30 minutes Step Temperature Time 1 37℃ 00:30:00 2 4℃ Hold

[0166] (5) Add 1 μl of heat-labile proteinase K to the reaction mixture. Mix gently and pulse-spin in a centrifuge. Incubate the reaction mixture using the following program.

[0167] [Table 15]

[0168] Lid temperature Reaction volume Running time 105℃ 31 μL 25 minutes Step Temperature Time 1 37℃ 00:15:00 2 55℃ 00:10:00 3 4℃ Hold

[0169] (6) Add 1 μl of exonuclease III, 1 μl of Lambda exonuclease, and 1 μl of exonuclease I to the reaction mixture on ice. Use the following program to remove linear DNA molecules.

[0170] [Table 16]

[0171]

[0172]

[0173] (7) After the linear DNA degradation reaction is complete, add 14 μl of SPRI magnetic beads to purify the circular DNA, and elute the purified DNA with 10 μl of elution buffer. Analyze the amount of circular DNA molecules using Qubit (using 1 μl).

[0174] (8) In a new tube, add reagents in the following order to assemble the Cas9-gRNA ribonucleoprotein targeting the self-ligating adapter, and incubate at room temperature for 10 minutes.

[0175] [Table 17]

[0176]

[0177] (9) Construct the following reaction at room temperature and use the following program to linearize the circular DNA molecules.

[0178] [Table 18]

[0179] Component Amount (final concentration) Purified circular DNA 9 μL NEBuffer r3.1 (10x) 1 μL Assembled Cas9-gRNA ribonucleoprotein 20 μL Total reaction volume 30 μL

[0180] [Table 19]

[0181] Lid temperature Reaction volume Running time 105℃ 30 μL 25 minutes Step Temperature Time 1 37℃ 00:25:00 2 4℃ Hold

[0182] (10) Add 1 μL of proteinase K to the reaction mixture, mix gently, and pulse-spin using a centrifuge.

[0183] (11) Incubate at room temperature for 10 minutes.

[0184] (12) After the digestion reaction, 37 μl of SPRI beads were added for DNA purification, and the purified DNA was eluted with 11 μl of elution buffer. The amount of circular DNA molecules was analyzed using Qubit (using 1 μl).

[0185] (13) Using Q5 DNA polymerase or KAPA HiFi HotStart ReadyMix, primers were used for PCR in a 50 μL volume for a part of the R1 sequence and a part of the TSO sequence, and the full-length cDNA was amplified using the program shown below.

[0186] [Table 20]

[0187]

[0188] (14) DNA purification was performed using SPRI beads, and the samples were analyzed using Qubit or a bioanalyzer to check the quantity and quality of the DNA contained in the samples.

[0189] The above (13) to (14) are the processes for amplifying linearized full-length molecules (molecules with both a "head" and a "tail") by PCR reaction.

[0190] Step 3 - Long-read sequencing (Nanopore / PacBio)

[0191] The full-length molecules from single cells were analyzed using Nanopore / PacBio long-read sequencing technology.

[0192] Analysis of the DNA fragment multi-combination assembly reaction scheme (selective removal of DNA molecules)

[0193] This is the process of the DNA fragment multi-combination reaction. Only the molecules with both a "head" (R1 and cell barcode, representing the 3' end of mRNA) and a "tail" (TSO, representing the 5' end of mRNA) are assembled into a circle, while the remaining molecules are not assembled. The process of using CRISPR / Cas9 to remove unwanted DNA is very accurate because the CRISPR / Cas9 system allows only the precise removal of target molecules. One of the advantages of this method is the ability to easily remove hundreds of targets simultaneously.

[0194] After the multi-ligation assembly reaction of DNA fragments, a process occurs to remove unassembled linear molecules and linearized DNA molecules in the initially circular molecules. Exonucleases remove the linear molecules. Technically, the use of a combination of different exonucleases is effective. In addition, after removing the unassembled linear molecules, a process of re-converting the circular molecules into a linear form is carried out using the CRISPR / Cas9 endonuclease, which precisely recognizes the multi-ligation junction sequence of the DNA fragments and continues to linearize. In addition, other endonucleases, such as restriction endonucleases, can also be used.

[0195] 2. Mechanism and application of multi-ligation assembly reaction of DNA fragments (selective removal of DNA molecules)

[0196] After the multi-ligation assembly reaction of DNA fragments, the process of selective removal of DNA molecules starts from the full-length DNA library. The process of selective removal of DNA molecules is as follows:

[0197] 1) Full-length DNA library self-assembled into circular DNA

[0198] 2) Use CRISPR / Cas9 nuclease to target DNA molecules containing specific DNA sequences (from 1 to thousands) (selectively convert unwanted DNA targets into linear DNA)

[0199] 3) Circular DNA purification, screening of full-length DNA molecules removing unwanted DNA molecules

[0200] Through these steps, DNA molecules are selectively removed, and the remaining circular DNA undergoes a linearization process. Subsequently, the single-cell full-length DNA library is analyzed.

[0201] CRISPR / Cas9 recognizes the target DNA sequences in unwanted DNA molecules, converts them into linear DNA, and finally purifies the remaining circular DNA. After that, linearization is achieved, and the linear molecules are amplified by PCR for further analysis. The results of the process of selective removal of DNA molecules are as Figure 6 shown. In this case, 26 sgRNAs targeting cDNA from ribosomal and mitochondrial genomes were used, and increasing the number of sgRNAs will enable more effective removal. Applying the method of selective removal of DNA molecules to mouse tissues confirmed that the sequencing efficiency of the required molecules was doubled.

[0202] 3. Technical verification

[0203] Using 26 target sequences targeting mitochondrial and ribosomal RNA-derived cDNA, mitochondrial and ribosomal RNA-derived cDNA was removed from single-cell RNA-Seq libraries. Single-cell RNA-seq libraries were prepared for analysis from single cells isolated from various mouse tissues using 10X Genomics technology (3' Gene Expression Kit, v3). Thus, before using the Ouroboros cell barcode selection and molecular target removal technology, the proportion of mitochondrial and ribosomal RNA-derived cDNA molecules in the libraries generated from kidney tissue was 56%. However, after using the Ouroboros cell barcode selection and molecular target removal technology, the proportion of mitochondrial and ribosomal RNA-derived cDNA molecules in the kidney tissue library decreased to 17%. The sequencing efficiency of the desired molecules (non-mitochondrial and non-ribosomal RNA-derived cDNA molecules) increased approximately two-fold, from 44% to 83%( Figure 6 ).

[0204] [Example 3]

[0205] 1. Multi-combinatorial assembly reaction scheme of DNA fragments (DNA Target Selection)

[0206] Step 1 – Preparation and purification of single-cell cDNA libraries

[0207] The 10X Genomics Single Cell 3' Expression Kit or other single-cell RNA-seq library preparation methods can be used.

[0208] Step 2 – Screening of DNA molecules with cell barcodes

[0209] (1) PCR was performed using primers containing partial R1 plus an additional 22 bp and partial TSO plus an additional 22 bp in order to add a 22-bp multi-combinatorial assembly adapter required for the self-cycling assembly reaction to the ends of the molecules containing the R1 and TSO sequences using the following program.

[0210] [Table 21]

[0211]

[0212]

[0213] (2) Assemble the self-ligation reaction mixture according to the following table and program. Before setting up the reaction, transfer the NEBuilder HiFi DNA Assembly premix to ice and gently tap the tube with your finger to mix before use.

[0214] [Table 22]

[0215] Component Amount (final concentration) 2X NEBuilder HiFi DNA Assembly Premix 5 μL Deoxyribonucleic acid 10 - 100 ng (1 - 10 ng / μL) Nuclease-free water To 10 μL Total reaction volume 10 μL

[0216] [Table 23]

[0217] Lid temperature Reaction volume Running time 105℃ 10 μL 60 minutes Step Temperature Time 1 50℃ 01:00:00 2 4℃ Hold

[0218] (3) On ice, add 1 μL of exonuclease III, 1 μL of Lambda exonuclease, and 1 μL of exonuclease I to the reaction mixture, and use the following program to remove linear DNA molecules.

[0219] [Table 24]

[0220] Lid temperature Reaction volume Running time 105℃ 12 μL 30 minutes Step Temperature Time 1 37℃ 00:20:00 2 70℃ 00:10:00 3 4℃ Hold

[0221] (4) After the degradation process, add 14 μL of SPRI beads for purification, and elute the purified circular DNA with 10 μL of elution buffer. To confirm the amount of circular DNA molecules, analyze using an Invitrogen Qubit 4 fluorometer (using 1 μL).

[0222] (5) In a new tube, add reagents in the following order to assemble the Cas9-gRNA ribonucleoprotein targeting the self-ligating adapter, and incubate at room temperature for 10 minutes.

[0223] [Table 25]

[0224]

[0225] (6) Set up the following reaction at room temperature and use the following program to linearize all circular DNA molecules.

[0226] [Table 26]

[0227] Component Amount (final concentration) Purified circular DNA 9 μL NEBuffer r3.1 (10x) 1 μL Assembled Cas9-gRNA ribonucleoprotein 20 μL Total reaction volume 30 μL

[0228] [Table 27]

[0229] Lid temperature Reaction volume Running time 105℃ 30 μL 25 minutes Step Temperature Time 1 37℃ 00:25:00 2 4℃ Maintain

[0230] (7) Add 1 μL of proteinase K to the reaction mixture, mix gently, and pulse centrifuge using a centrifuge.

[0231] (8) Incubate at room temperature for 10 minutes.

[0232] (9) After the digestion reaction, add 37 μl of SPRI beads for DNA purification, and elute the purified DNA with 11 μl of elution buffer. Analyze the amount of circular DNA molecules using Qubit (using 1 μl).

[0233] (10) Set up the following dA tailing (End-Prep) reaction at room temperature, incubate at 20 °C for 20 minutes, and then incubate at 65 °C for 10 minutes.

[0234] [Table 28]

[0235] Component Amount (Final Concentration) Cas9-digested Target DNA 10 μl Ultra II End-Prep Reaction Buffer 1.75 μl Ultra II End-prep Enzyme Mix 0.75 μl Nuclease-Free Water 2.5 μl (to 15 μl) Total Reaction Volume 15 μl

[0236] (11) For clean purification, add 20 μl of SPRI beads and elute the purified DNA with 11 μl of elution buffer.

[0237] (12) Set up the following reaction at room temperature and incubate for 30 minutes at room temperature.

[0238] [Table 29]

[0239]

[0240] (13) For clean purification, add 20 μl of SPRI magnetic beads and elute the purified DNA with 11 μl of elution buffer.

[0241] The above (12) to (13) is the process of ligating the Y adapter to the linear target DNA (end-prepared DNA).

[0242] (14) Using the forward and reverse primers of the Y adapter, perform 50 μl of PCR with Q5 DNA polymerase or KAPA HiFi HotStart ReadyMix according to the procedure outlined below to selectively amplify the full-length cDNA target.

[0243] [Table 30]

[0244]

[0245] (15) Clean up the SPRI magnetic beads and analyze the sample using Qubit or a bioanalyzer to check the quantity and quality of the DNA contained in the sample.

[0246] The above (14) to (15) is the process of amplifying the linear DNA target with the appropriate primers of the Y adapter used in the ligation. During this amplification process, unselected circular DNA is removed.

[0247] Step 3 – Long-read sequencing (Nanopore / PacBio)

[0248] Analyze full-length molecules from single cells using Nanopore / PacBio long-read sequencing technology.

[0249] DNA fragment multi-combination assembly reaction scheme analysis (DNA target selection)

[0250] This is the process of the multi - combination reaction of DNA fragments. Only the molecules that have both a "head" (R1 and cell barcodes, representing the 3' end of mRNA) and a "tail" (TSO, representing the 5' end of mRNA) assemble into a ring, while the remaining molecules do not assemble. After the multi - combination reaction of DNA fragments, the unassembled molecules are removed. In this process, linear molecules are removed by exonucleases. Technically, it is important to use a combination of various exonucleases. Once the removal of unassembled linear molecules is completed, the circular molecules are re - linearized using the CRISPR / Cas9 endonuclease, which can precisely recognize the multi - combination junction sequence of DNA fragments and linearize them. Other endonucleases, such as restriction endonucleases, can also be used.

[0251] The protocol for the multi - fragment assembly reaction of DNA targets is a process of selecting the desired DNA targets using CRISPR / Cas9 without losing information. By utilizing the CRISPR / Cas9 system, only the desired molecules can be accurately selected, and it has the advantage of being able to effectively select hundreds of targets simultaneously. When the number of targets is assumed to be the same, this method is simpler and more economical compared to traditional techniques such as RNA hybridization capture. Technically, before performing the Y - junction ligation reaction, it is necessary to add an A - tail to the target DNA that has been converted into a linear form using the CRISPR / Cas9 system. This step is crucial for reducing the formation of adapter dimers and improving the reaction efficiency.

[0252] 2. Mechanism and Application of DNA Fragment Multi-Ligation Assembly Reaction (Selective Removal of DNA Molecules)

[0253] After the multi - ligation assembly reaction of DNA fragments, the process of selective removal of DNA molecules starts from the full - length DNA library. The process of selective removal of DNA molecules is as follows:

[0254] 1) The full - length DNA library that self - assembles into circular DNA

[0255] 2) Use CRISPR / Cas9 nuclease to target DNA molecules containing specific DNA sequences (from 1 to thousands) (selectively convert unwanted DNA targets into linear DNA)

[0256] 3) Y - adapters only ligate to linear DNA, and the targets are screened through an amplification process

[0257] The results of cDNA target screening are as Figure 8As shown. This process was carried out using mouse tissue samples with the lowest cDNA content from mitochondrial genomes. Targeting cDNA from 11 genes in the mitochondrial genome, after target screening, 72% of the sequenced bases were identified as belonging to the target cDNA. The target selection efficiency was found to be 15 times higher than that of the original library. In addition, single-cell isolation was performed on peripheral blood mononuclear cell (PBMC) samples from cancer patients, and full-length cDNA target screening was carried out. A 10X single-cell transcriptome library was generated from the isolated single cells, and target screening was performed using 22 sgRNA-targeted regions with single nucleotide variants (SNVs). The results showed that the target selection efficiency was increased by 17.5 times ( Figure 9 ).

[0258] 3. Technical Verification

[0259] The results of selectively targeting cDNA from mitochondrial genomes in single-cell RNA-Seq libraries using 11 types of cDNA from mitochondrial genomes are as follows. A single-cell RNA-seq library was generated using mouse thymus tissue and isolated using the 10X Genomics 3' gene expression kit (v3). Before applying the Ouroboros cell barcode selection and molecular target selection technology of the present invention, the proportion of cDNA molecules from mitochondrial genomes in the library was 4.2%. After applying the Ouroboros technology, this proportion increased to 72%, indicating that the sequencing efficiency of the required molecules (cDNA from mitochondrial genomes) was increased by approximately 17 times ( Figure 8 ).

[0260] The following are the results of selectively isolating cDNA molecules derived only from 99 fibroblasts, which were identified in a single-cell library containing 6,500 single-cell cDNAs. A single-cell RNA-seq library was prepared using the kidney tissue of the mother mouse and processed using the 10X Genomics 3' gene expression kit (v3). Through sequencing and bioinformatics analysis, the cell barcodes of 99 fibroblasts were identified. The specific DNA oligonucleotides for these 99 cell barcodes were designed using the oPool service of IDT. Then, CRISPR / Cas9 sgRNAs targeting the 99 cell barcodes were generated using these DNA oligonucleotides. The application of the Ouroboros cell barcode selection and molecular target selection technology confirmed that only the cDNA from 99 fibroblasts was accurately screened ( Figure 11 ).

[0261] Although the present disclosure has been described in detail with reference to specific features, it is obvious to those skilled in the art that this description is only its preferred embodiment and does not limit the scope of the present disclosure. Therefore, the substantial scope of the present disclosure will be defined by the appended claims and their equivalents.

[0262]

Industrial Applicability

[0263] The present invention is a development study aimed at overcoming the low quality and high cost of multi-omics full-length sequencing, including single-cell transcriptomes, and the limitations of cell heterogeneity research. By using an assembly method that allows the combination of multiple DNA fragments, it significantly improves the efficiency and error rate of full-length sequencing and enables multi-group full-length sequencing of single cells. It is expected to be widely used for the expression analysis of gene subtypes, mutation detection, or cancer cell-specific multi-group analysis.

Claims

1. A method for full-length multi-omics sequencing analysis of single cells, comprising: (a) ligating a multi-ligation assembly linker of DNA fragments to a multi-omics library prepared from a single cell; (b) selectively circularizing DNA molecules in the library through a multi-ligation assembly reaction of DNA fragments to screen DNA molecules within the library; and (c) using an exonuclease to remove DNA fragments other than circular DNA.

2. The method according to claim 1, wherein After step (c), it further includes: (d) linearizing the circular DNA in the library.

3. The method according to claim 1, wherein The multi-omics are transcriptome, spatial transcriptome, genome, spatial genome, proteome, epigenome or spatial epigenome.

4. The method according to claim 1, wherein The multi-omics library contains different linkers (forward and reverse linkers) ligated to the 3' and 5' ends respectively.

5. The method according to claim 1, characterized in that, In step (a), two different linkers are assembled to form circular DNA.

6. The method according to claim 2, wherein The linearization in step (d) is carried out using a CRISPR / Cas nuclease, a restriction endonuclease or a transcription activator-like effector nuclease (TALEN).

7. The method according to claim 6, wherein The CRISPR / Cas nuclease is selected from the group consisting of: CRISPR / Cas9, CRISPR / Cas12, CRISPR / Cas13 or CRISPR / Cas14.

8. A sequencing analysis method after selectively removing DNA molecules from a full-length multi-omics library of single cells, comprising: (a) ligating a multi-ligation assembly linker of DNA fragments to a multi-omics library prepared from a single cell; (b) generating circular DNA within the library through a multi-ligation assembly reaction of DNA fragments; (c) using a CRISPR / Cas nuclease to selectively linearize circular DNA containing a target sequence to be excised; and (d) using an exonuclease to remove linear DNA that is not circular DNA.

9. The method according to claim 8, wherein After step (d), it further includes: (e) linearizing the circular DNA in the library.

10. The method according to claim 8, wherein The CRISPR / Cas nuclease is selected from the group consisting of: CRISPR / Cas9, CRISPR / Cas12, CRISPR / Cas13 or CRISPR / Cas14.

11. The method according to claim 8, wherein The CRISPR / Cas nuclease is only applied to circular DNA.

12. The method according to claim 8, wherein The linear DNA in step (d) is a molecule generated by the degradation of selected circular DNA recognized by the CRISPR / Cas nuclease.

13. A sequencing analysis method after screening DNA molecules from a full-length multi-omics library of single cells, comprising: (a) preparing a multi-omics library from a single cell and ligating a multi-ligation assembly linker of DNA fragments to the library; (b) generating circular DNA within the library through a multi-ligation assembly reaction of DNA fragments; (c) using an exonuclease to remove DNA fragments that are not circular DNA; and (d) using a CRISPR / Cas nuclease to selectively convert circular DNA containing a target sequence to be selected into linear DNA.

14. The method according to claim 13, wherein After step (d), it further includes: (e) amplifying the linear DNA in the library.

15. The method according to claim 13, wherein The CRISPR / Cas nuclease is selected from the group consisting of: CRISPR / Cas9, CRISPR / Cas12, CRISPR / Cas13 and CRISPR / Cas14.

16. The method according to claim 13, wherein, The linear DNA in step (d) is a molecule generated by degradation of the selected circular DNA recognized by the CRISPR / Cas nuclease.

17. The method according to claim 14, characterized in that, Step (e) also includes a step of ligating a Y-junction or a hairpin junction.

18. A single-cell full-length multi-omics sequencing analysis kit using the method according to claim 1.

19. A single-cell level gene subtype diagnostic marker analysis kit using the method according to claim 1.

20. A cancer targeted therapy candidate analysis kit based on single-cell mutations using the method according to claim 1.