Nucleic acid editing system, and nucleic acid editing method and use

By utilizing single-strand cleavage before homologous recombination and ORF2p-EN endonuclease in a nucleic acid editing system, the safety and efficiency issues of existing gene editing technologies have been resolved, enabling efficient and precise gene editing in humans and in vitro, and applying it to the treatment of various diseases.

WO2026061540A1PCT designated stage Publication Date: 2026-03-26SUI YUNPENG
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-23
Publication Date
2026-03-26

AI Technical Summary

Technical Problem

Existing gene editing technologies have risks such as double-strand breaks, low homologous recombination efficiency, off-target risks, and uncertainties in gene editing, which pose difficulties, especially in human applications.

Method used

A nucleic acid editing system is provided, including a framework structure and a cutting structure. Single-strand cutting is performed before homologous recombination, and precise nucleic acid editing is performed by using the ORF2p-EN endonuclease in the human body in combination with the framework structure. This eliminates the need for retrotransposon sequences and reverse transcription, thereby improving safety and efficiency.

Benefits of technology

It enables safe and efficient gene editing in humans, which can be applied to the treatment of a variety of diseases, including genetic diseases and cancer. Furthermore, the in vitro nucleic acid editing process is simple, reducing the uncertainty and off-target risk of gene editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025123302_26032026_PF_FP_ABST
    Figure CN2025123302_26032026_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present invention are a nucleic acid editing system, and a nucleic acid editing method and the use. The nucleic acid editing system comprises a framework structure and a cleavage structure. The framework structure comprises an upstream sequence of a target site, a sequence to be inserted and a downstream sequence of the target site; and the cleavage structure comprises an ORF2p protein and a derivative protein thereof. Further provided in the present invention are a nucleic acid editing method using the editing system; and the use of the editing system in the preparation of a drug for preventing and / or treating cancer, gene-related diseases or neurodegenerative diseases.
Need to check novelty before this filing date? Find Prior Art

Description

A nucleic acid editing system, nucleic acid editing method and application TECHNICAL FIELD

[0001] The present application belongs to the field of molecular biology, and relates to a nucleic acid editing system, nucleic acid editing method and application. BACKGROUND

[0002] Currently, the main gene editing technologies mainly include CRISPR and ZFN, and the way of inserting an exogenous sequence in these technologies mainly relies on first cutting a double strand of a genome and then introducing the exogenous sequence to the cut part through homologous recombination. Such a way will cause a dangerous double strand break, and the efficiency of introducing the exogenous sequence through homologous recombination is low. In the case where homologous recombination does not occur, a random sequence will be introduced to the double strand break through the NHEJ pathway, causing uncertainty of gene editing and even harm. Meanwhile, CRISPR has the risk of off-target, and may cut the genome at a non-target site. Therefore, the human application of CRISPR and other gene editing technologies has been full of difficulties. SUMMARY

[0003] In order to solve the above problems, the present application aims to provide a nucleic acid editing system.

[0004] Another object of the present application is to provide a vector system.

[0005] A third object of the present application is to provide a gene sequence editing method.

[0006] A fourth object of the present application is to provide the application of the nucleic acid editing system or the vector system.

[0007] In order to achieve the above objects, the present application provides a nucleic acid editing system, comprising a frame structure and a cleavage structure.

[0008] The frame structure is one or more of a single-stranded DNA frame, a double-stranded DNA frame, a single-stranded DNA derivative frame, a double-stranded DNA derivative frame, a DNA-RNA hybrid frame, a DNA / RNA hybrid derivative frame, a single-stranded RNA frame, a double-stranded RNA frame, a single-stranded RNA derivative frame and a double-stranded RNA derivative frame; and the cleavage structure is a protein or a polypeptide.

[0009] The single-stranded DNA framework, double-stranded DNA framework, single-stranded DNA derivative framework, and double-stranded DNA derivative framework are composed of a target site upstream sequence, a sequence to be inserted, and a target site downstream sequence in the 5'→3' direction; the target site upstream sequence of the single-stranded DNA framework, double-stranded DNA framework, single-stranded DNA derivative framework, and double-stranded DNA derivative framework or the complement sequence of the target site upstream sequence is used for hybridization with the target site upstream sequence or the complement sequence of the target site upstream sequence in a target site in a nucleic acid sequence to be cleaved or cleaved and edited; the target site downstream sequence or the complement sequence of the target site downstream sequence on the single-stranded DNA framework, double-stranded DNA framework, single-stranded DNA derivative framework, and double-stranded DNA derivative framework is used for hybridization with the target site downstream sequence or the complement sequence of the target site downstream sequence in a target site in a nucleic acid sequence to be cleaved or cleaved and edited; the target site upstream sequence and the target site downstream sequence on the single-stranded DNA framework, double-stranded DNA framework, single-stranded DNA derivative framework, and double-stranded DNA derivative framework are directly connected to the corresponding sequences in a nucleic acid sequence to be cleaved or cleaved and edited; and the target site upstream sequence and the target site downstream sequence of the nucleic acid sequence to be cleaved or cleaved and edited or the complement sequence thereof are connected by a target site;

[0010] The DNA / RNA hybridization framework or DNA / RNA hybridization derivative framework includes a target site upstream sequence, a sequence to be inserted, and a target site downstream sequence in the 5'→3' direction of the single-stranded RNA or single-stranded DNA in the DNA / RNA hybridization framework or DNA / RNA hybridization derivative framework;

[0011] The target site upstream sequence or the complement sequence of the target site upstream sequence on the DNA / RNA hybridization framework or DNA / RNA hybridization derivative framework is used for hybridization with the target site upstream sequence or the complement sequence of the target site upstream sequence in a nucleic acid sequence to be cleaved or cleaved and edited; the target site downstream sequence or the complement sequence of the target site downstream sequence on the DNA / RNA hybridization framework or DNA / RNA hybridization derivative framework is used for hybridization with the target site downstream sequence or the complement sequence of the target site downstream sequence in a target site in a nucleic acid sequence to be cleaved or cleaved and edited; the target site upstream sequence and the target site downstream sequence on the DNA / RNA hybridization framework or DNA / RNA hybridization derivative framework are directly connected to the corresponding sequences in a nucleic acid sequence to be cleaved or cleaved and edited; and the target site upstream sequence and the target site downstream sequence of the nucleic acid sequence to be cleaved or cleaved and edited or the complement sequence thereof are connected by a target site;

[0012] The single-stranded RNA framework, double-stranded RNA framework, single-stranded RNA derivative framework, and double-stranded RNA derivative framework include a target site upstream sequence, a sequence to be inserted, and a target site downstream sequence in the 5'→3' direction;

[0013] The target site upstream sequence on the single-stranded RNA framework, double-stranded RNA framework, single-stranded RNA derived framework, double-stranded RNA derived framework or the complement sequence of the target site upstream sequence is used for hybridizing with the upstream sequence of the target site in the nucleic acid sequence to be cleaved or edited after cleavage, or the complement sequence of the target site upstream sequence; the target site downstream sequence on the single-stranded RNA framework, double-stranded RNA framework, single-stranded RNA derived framework, double-stranded RNA derived framework or the complement sequence of the target site downstream sequence is used for hybridizing with the downstream sequence of the target site in the nucleic acid sequence to be cleaved or edited after cleavage; the target site upstream sequence and the target site downstream sequence on the single-stranded RNA framework, double-stranded RNA framework, single-stranded RNA derived framework, double-stranded RNA derived framework are directly connected in the corresponding sequence in the nucleic acid sequence to be cleaved or edited after cleavage; the target site upstream sequence and the target site downstream sequence of the nucleic acid sequence to be cleaved or edited after cleavage or the complement sequence thereof are target sites;

[0014] The cleavage structure is a protein or a polypeptide or DNA or RNA encoding the protein or the polypeptide, wherein the protein is ORF2p-EN, ORF2p, an ORF2p-EN homologous protein, an ORF2p derived protein, an ORF2p-EN derived protein, an ORF2p-EN homologous protein derived protein; the polypeptide is a length part of 50%-99% of the protein.

[0015] The cleaved or edited nucleic acid sequence is single-stranded DNA, double-stranded DNA, single-stranded RNA, double-stranded RNA or DNA / RNA hybrid chain.

[0016] Further, the framework structure is completely matched with the 4bp immediately upstream of the target site and the 4bp immediately downstream of the target site of the cleaved or edited nucleic acid sequence or the complement sequence thereof; the remaining sequence is complementary to 0%-100%.

[0017] Further, the cleaved or edited nucleic acid is single-stranded DNA, double-stranded DNA, single-stranded RNA, double-stranded RNA or DNA / RNA hybrid chain in vivo or in vitro; wherein the double-stranded DNA or single-stranded DNA comprises the genome of eukaryotes, prokaryotes or viruses; the double-stranded RNA or single-stranded RNA comprises the genome of viruses, mRNA, rRNA, miRNA or siRNA of eukaryotes, prokaryotes or viruses.

[0018] Further, the framework structure is linear or circular.

[0019] The application also provides a vector system comprising the framework structure and the cleavage structure with the above nucleic acid editing system, wherein the framework structure and the cleavage structure are located in the same vector or different vectors.

[0020] The present application also provides the use of the nucleic acid editing system as described above or the vector system as described above in the sequence insertion, sequence deletion, sequence replacement, site insertion, site deletion, site replacement, sequence inversion, and / or sequence inversion correction in any region of the genome.

[0021] The present application also provides a method for editing a genetic sequence, comprising the following steps:

[0022] 1) selecting the insertion site of the nucleic acid to be cleaved or edited, determining the upstream sequence of the target site and the downstream sequence of the target site on both sides of the insertion site;

[0023] 2) preparing the vector system as described above;

[0024] 3a) adding or transfecting the vector system into a solution, cell, tissue, organ or organism containing the nucleic acid to be cleaved or edited;

[0025] or 3b) combining the vector system with homologous recombination proteins and / or reverse transcriptase first, and then adding or transfecting into a solution, cell, tissue, organ or organism containing the nucleic acid to be cleaved or edited;

[0026] or 3c) adding or transfecting the vector system and homologous recombination proteins and / or reverse transcriptase into a solution, cell, tissue, organ or organism containing the nucleic acid to be cleaved or edited at the same time;

[0027] or 3d) adding or transfecting the vector system into a solution, cell, tissue, organ or organism containing the nucleic acid to be cleaved or edited and then adding homologous recombination proteins and / or reverse transcriptase;

[0028] or 3e) adding or transfecting homologous recombination proteins and / or reverse transcriptase into a solution, cell, tissue, organ or organism containing the nucleic acid to be cleaved or edited, and then adding the vector system.

[0029] Further, the primer is added at the same time as the vector is added in step 3a), 3b), 3c), 3d) or 3e).

[0030] Further, the DNA polymerase or RNA polymerase is added at the same time as the vector is added in step 3a), 3b), 3c), 3d) or 3e).

[0031] Further, the compound RS-1, SCR7, NU7441, Vanillin, Chlorpromazine, Resveratrol, Benomyl, Inositol for improving the efficiency of homologous recombination is added at the same time as the vector is added in step 3a), 3b), 3c), 3d) or 3e).

[0032] Further, the homologous recombination protein is Rec / Rad51 family, RAD51, RecA, Uvs X, RAD52, RAD51AP1, RAD51AP2, RAD54L, Swi5-Sfr1 complex, BRCA1, BRCA2, PALB2, RAD53, RAD54, RAD55, RAD56, RAD57, BARD1, RecA, RPA, Mre11, NBS, DMC1, SCR7, SRS2, PSF, PARP1, RBM6, RAD58, RAD59, RAD50, TID1 / RDH54, MSH3, Mlh1, XRS2, NBS1, SSBP, PHF1, SMC1, DHX9, RECQ helicase such as RECQ1, P53, Mus81, SWI2 / SNF2 family protein, Hop2 or Mnd1.

[0033] The present application also provides the use of the nucleic acid editing system or the vector system as described above in the preparation of a medicament for preventing and / or treating cancer, a genetic-related disease or a neurodegenerative disease.

[0034] Further, the cancer is glioma, breast cancer, cervical cancer, lung cancer, gastric cancer, colorectal cancer, duodenal cancer, leukemia, prostate cancer, endometrial cancer, thyroid cancer, lymphoma, pancreatic cancer, liver cancer, melanoma, skin cancer, pituitary tumor, germ cell tumor, meningioma, meningeal cancer, glioblastoma, various astrocytoma, various oligodendroglioma, astrocytic oligodendroglioma, various ependymoma, choroid plexus papilloma, choroid plexus carcinoma, chordoma, various ganglioneuroma, olfactory neuroblastoma, sympathetic nervous system neuroblastoma, pineal cell tumor, pinealoblastoma, medulloblastoma, retinoblastoma, trigeminal schwannoma, facial acoustic neuroma, jugular bulb tumor, angiomatous reticulocyte, craniopharyngioma or granular cell tumor.

[0035] Still further, the disease associated with the gene is Huntington's disease, fragile X syndrome, phenylketonuria, Duchenne's progressive muscular dystrophy, Duchenne's muscular dystrophy, mitochondrial encephalomyopathy, mucopolysaccharidosis type I, mucopolysaccharidosis type II, mucopolysaccharidosis type IIIA, mucopolysaccharidosis type IIIB, mucopolysaccharidosis type III C, mucopolysaccharidosis type IIID, mucopolysaccharidosis type IVA, mucopolysaccharidosis type IVB, mucopolysaccharidosis type VI, mucopolysaccharidosis type VII, mucopolysaccharidosis type IX, spinal muscular atrophy, Parkinson-plus syndrome, albinism, color blindness, achondroplasia, alkaptonuria, congenital deaf mutism, thalassemia, sickle cell anemia, hemophilia, epilepsy associated with genetic alterations, myoclonus, dystonia, stroke and schizophrenia, antivitamin D rickets, familial colonic polyposis, 21-hydroxylase deficiency, arginase deficiency, Alport syndrome, Angelman syndrome, Tay-Sachs disease, atypical hemolytic uremic syndrome, autoimmune encephalitis, autoimmune hypophysitis, autoimmune insulin receptor disease, beta-ketothiolase deficiency, biotinidase deficiency, cardiac ion channelopathy, primary carnitine deficiency, Castleman disease, Charcot-Marie-Tooth disease, citrullinemia, congenital adrenal hypoplasia, congenital hyperinsulinemic hypoglycemia, congenital myasthenic syndrome, non-dystrophic myotonia syndrome, congenital scoliosis, coronary artery ectasia, congenital pure red cell aplasia, Erdheim-Chester disease, Fabry disease, familial Mediterranean fever, Fanconi anemia, galactosemia, Gaucher disease, generalized myasthenia gravis, Gitelman syndrome, glutaric aciduria type I, glycogen storage disease (type I, type II), hemophilia, hepatolenticular degeneration, hereditary angioedema, hereditary epidermolysis bullosa, hereditary fructose intolerance, hereditary hypomagnesemia, hereditary multi-infarct dementia, hereditary spastic paraplegia, holocarboxylase synthetase deficiency, homocysteinuria, homozygous familial hypercholesterolemia, HHH syndrome, hyperphenylalaninemia, hypophosphatasia, hypophosphatemic rickets, idiopathic cardiomyopathy, idiopathic hypogonadotropic hypogonadism, idiopathic pulmonary arterial hypertension, idiopathic pulmonary fibrosis, IgG4-related disease, inborn error of bile acid synthesis, isovaleric acidemia, Kallmann syndrome, Langerhans cell histiocytosis, Leigh syndrome, Leber hereditary optic neuropathy, long-chain 3-hydroxyacyl-CoA dehydrogenase deficiency, lymphangioleiomyomatosis, lysinuric protein intolerance, lysosomal acid lipase deficiency, maple syrup urine disease, Marfan syndrome, McCune-Albright syndrome, medium-chain acyl-CoA dehydrogenase deficiency, methylmalonic acidemia, multifocal motor neuropathy, multiple acyl-CoA dehydrogenase deficiency, multiple sclerosis, myotonic dystrophy,N-acetylglutamate synthetase deficiency, neonatal diabetes, neuromyelitis optica, Niemann-Pick disease, non-syndromic deafness, Noonan syndrome, ornithine transcarbamylase deficiency, osteogenesis imperfecta, young-onset Parkinson's disease, early-onset Parkinson's disease, paroxysmal nocturnal hemoglobinuria, Peutz-Jeghers syndrome, POEMS syndrome, porphyria, Prader-Willi syndrome, primary combined immunodeficiency, primary hereditary muscular atonia, primary light chain amyloidosis, progressive familial intrahepatic cholestasis, progressive muscular dystrophy, propionic acidemia, pulmonary alveolar proteinosis, pulmonary cystic fibrosis, retinitis pigmentosa, severe congenital granulocytopenia, infantile severe myoclonic epilepsy, Dravet syndrome, Silver-Russell syndrome, sitosterolemia, spinal and bulbar muscular atrophy, spinal muscular atrophy, spinocerebellar ataxia, systemic sclerosis, tetrahydrobiopterin deficiency, tuberous sclerosis, primary tyrosinemia, very long chain acyl-CoA dehydrogenase deficiency, Williams syndrome, X-linked agammaglobulinemia, X-linked adrenoleukodystrophy, X-linked lymphoproliferative syndrome, arteriosclerotic cerebral small vessel disease, cerebral amyloid angiopathy, cerebral arteriopathy with subcortical infarcts and leukoaraiosis, cerebral arteriopathy with subcortical infarcts and leukoaraiosis, cysteine string protein-associated arteriopathy with stroke and leukoaraiosis, pyridoxine-dependent epilepsy, AADC enzyme deficiency of serotonin metabolism, AADC deficiency, or hereditary nephritis.

[0036] Further, the neurodegenerative disease is Parkinson's disease, Alzheimer's disease, Huntington's disease, amyotrophic lateral sclerosis, spinocerebellar ataxia, multiple system atrophy, primary lateral sclerosis, Pick's disease, frontotemporal dementia, Lewy body dementia, or progressive supranuclear palsy.

[0037] The present application also provides use of the ORF2p-EN, ORF2p, ORF2p-EN homologous protein, ORF2p-derived protein, ORF2p-EN-derived protein, or ORF2p-EN homologous protein-derived protein in the preparation of a substance of a nucleic acid editing system.

[0038] Further, a nuclear localization signal is further added in the N terminus, N part, C terminus, C part, or sequence middle of the ORF2p-EN, ORF2p, ORF2p-EN homologous protein, ORF2p-derived protein, ORF2p-EN-derived protein, or ORF2p-EN homologous protein-derived protein.

[0039] Still further, a tag is further added to the N-terminus, N-part, C-terminus, C-part or sequence middle of the ORF2p-EN, ORF2p, ORF2p-EN homologous protein, ORF2p derivative protein, ORF2p-EN derivative protein or ORF2p-EN homologous protein derivative protein; the tag is selected from HIS, Flag, GST, Myc, eGFP, eCFP, eYFP, mCherry eGFP, HA, SUMO, MBP, Strep-II Tag, Avi Tag or SNAP-Tag.

[0040] The application also provides a nucleic acid editing method, comprising the following steps:

[0041] 1) selecting an insertion site of a nucleic acid to be edited, determining the upstream sequence of the target site and the downstream sequence of the target site on both sides of the insertion site;

[0042] 2) preparing a nucleic acid editing system as described above or a vector system as described above;

[0043] 3a) when either or both of the frame structure in the nucleic acid to be edited or the nucleic acid editing system is double-stranded, mixing the nucleic acid to be edited and the frame structure to denature and then anneal, so that the single strand of the frame structure is combined with the single strand of the nucleic acid to be edited;

[0044] or 3b) when one of the frame structure in the nucleic acid to be edited or the nucleic acid editing system is single-stranded and the other is double-stranded, the nucleic acid to be edited, the frame structure and the homologous recombination protein can be mixed to perform strand displacement, so that the single strand of the frame structure is combined with the single strand of the nucleic acid to be edited;

[0045] or 3c) when either or both of the frame structure in the nucleic acid to be edited or the nucleic acid editing system is double-stranded, the nucleic acid to be edited and the frame structure are mixed and an exonuclease is added, so that the double-stranded is converted into single-stranded, and then the single strand of the frame structure is combined with the single strand of the nucleic acid to be edited;

[0046] or 3d) when both of the frame structure in the nucleic acid to be edited or the nucleic acid editing system are single-stranded, the two can be directly mixed, and the single strand of the frame structure is combined with the single strand of the nucleic acid to be edited;

[0047] 4) adding a cleavage structure to the hybridization product of the single strand of the frame structure and the single strand of the nucleic acid to be edited, so as to cut the single strand of the nucleic acid to be edited;

[0048] 5) continuing to add a DNA polymerase or an RNA polymerase to complete the editing of the nucleic acid to be edited.

[0049] The application has the following beneficial effects:

[0050] The application provides a nucleic acid editing system, a nucleic acid editing method and application, compared with the first cleavage of double strands and then homologous recombination of other gene editing technologies, the nucleic acid editing system first performs homologous recombination, and then performs single strand cleavage, is safer, and can use an endonuclease ORF2p-EN which is more human-friendly and exists in a normal human body, has clinical application and other application prospects, and can be applied to treatment of various diseases related to genes, such as genetic diseases or cancers, and the like. Meanwhile, the application omits the retrotransposon sequence and the reverse transcription process of the retrotransposon, so that the whole gene editing process is more controllable. The application can also be applied to in vitro nucleic acid editing, sequence insertion into various nucleic acids in vitro, and preparation of an in vitro nucleic acid editing kit which is simpler than other methods at present. The application can also be used for destroying various RNAs in the body and applied to treatment of various diseases such as cancers. Compared with the current related field which focuses on modifying the structure of the endonuclease for cleaving DNA, the application promotes cleavage by changing the structure of genomic DNA, which is a brand-new idea. BRIEF DESCRIPTION OF DRAWINGS

[0051] Fig. 1 is a schematic diagram of the basic architecture of the application.

[0052] Fig. 2 is a schematic diagram of a DNA-derived framework and / or a DNA / RNA hybrid-derived framework and / or an RNA-derived framework.

[0053] Fig. 3 is a schematic diagram of the relationship and interaction between a primer and a framework structure.

[0054] Fig. 4 is a schematic diagram of the principle of nucleic acid editing by a framework structure.

[0055] Fig. 5 is a schematic diagram of the principle of editing of in vitro double-stranded DNA by a protein with strand displacement function and ORF2p-EN.

[0056] Fig. 6 is a schematic diagram of the principle that ORF2p-EN can directly and effectively bind to nucleic acid after removing steric hindrance.

[0057] Fig. 7 is a schematic diagram of the structure cleaved by ORF2p and the related cleavage mechanism.

[0058] Fig. 8 is a schematic diagram of further adding NLS or a tag to the cleaved structure.

[0059] Fig. 9 is a schematic diagram of further adding NLS or a tag to the cleaved structure (fusion protein).

[0060] Fig. 10 is a schematic diagram of connection between a framework structure and a cleaved structure.

[0061] Fig. 11 is a schematic diagram of the principle of combination of a framework structure and a homologous recombination protein.

[0062] Figure 12 is a schematic diagram of two target sites at the same locus, on two complementary strands to each other.

[0063] Figure 13 is a schematic diagram of the basic framework structure of the present application.

[0064] Figure 14 is a schematic diagram of the principle of the target site being on a hypothetical strand when the nucleic acid to be edited is single-stranded.

[0065] Figure 15 is a schematic diagram of the engineered ORF2p-EN of Example 1 with a nuclear localization signal.

[0066] Figure 16 is a schematic diagram of the engineered ORF2p-EN of Example 2 with a 6xHis tag.

[0067] Figure 17 is a graph of the results of the in vitro implementation of the ORF2p-EN cleaving DNA and inserting a sequence of Example 4.

[0068] Figure 18 is a graph of the results of the double-stranded DNA framework insertion into the genome of a eukaryote of Example 5.

[0069] Figure 19 is a graph of the results of the single-stranded DNA framework insertion into the genome of a eukaryote of Example 6.

[0070] Figure 20 is a graph of the results of the DNA / RNA hybrid framework insertion into the genome of a eukaryote of Example 7.

[0071] Figure 21 is a graph of the results of the GFP sequence being inserted into the genome of Example 8. DETAILED DESCRIPTION

[0072] The embodiments of the present application will be described in detail below, so that the advantages and features of the present application can be more easily understood by those skilled in the art, and the scope of protection of the present application can be more clearly defined.

[0073] The present application provides a nucleic acid editing system, comprising a framework structure and a cleavage structure, as shown in Figure 1;

[0074] wherein the framework structure is one or more of a single-stranded DNA framework, a double-stranded DNA framework, a single-stranded DNA derivative framework, a double-stranded DNA derivative framework, a DNA-RNA hybrid framework, a DNA / RNA hybrid derivative framework, a single-stranded RNA framework, a double-stranded RNA framework, a single-stranded RNA derivative framework, a double-stranded RNA derivative framework; and the cleavage structure is a protein or a polypeptide.

[0075] The single-stranded DNA framework, double-stranded DNA framework, single-stranded DNA derivative framework, and double-stranded DNA derivative framework are composed of a target site upstream sequence, a sequence to be inserted, and a target site downstream sequence in the 5'→3' direction; the target site upstream sequence of the single-stranded DNA framework, double-stranded DNA framework, single-stranded DNA derivative framework, and double-stranded DNA derivative framework or the complement sequence of the target site upstream sequence is used for hybridization with the target site upstream sequence or the complement sequence of the target site upstream sequence in a target site in a nucleic acid sequence to be cleaved or cleaved and edited; the target site downstream sequence or the complement sequence of the target site downstream sequence on the single-stranded DNA framework, double-stranded DNA framework, single-stranded DNA derivative framework, and double-stranded DNA derivative framework is used for hybridization with the target site downstream sequence or the complement sequence of the target site downstream sequence in a target site in a nucleic acid sequence to be cleaved or cleaved and edited; the target site upstream sequence and the target site downstream sequence on the single-stranded DNA framework, double-stranded DNA framework, single-stranded DNA derivative framework, and double-stranded DNA derivative framework are directly connected to the corresponding sequences in a nucleic acid sequence to be cleaved or cleaved and edited; and the target site upstream sequence and the target site downstream sequence of the nucleic acid sequence to be cleaved or cleaved and edited are connected by a target site;

[0076] The DNA / RNA hybridization framework or DNA / RNA hybridization derivative framework includes a target site upstream sequence, a sequence to be inserted, and a target site downstream sequence in the 5'→3' direction of the single-stranded RNA or single-stranded DNA in the DNA / RNA hybridization framework or DNA / RNA hybridization derivative framework;

[0077] The target site upstream sequence or the complement sequence of the target site upstream sequence on the DNA / RNA hybridization framework or DNA / RNA hybridization derivative framework is used for hybridization with the target site upstream sequence or the complement sequence of the target site upstream sequence in a nucleic acid sequence to be cleaved or cleaved and edited; the target site downstream sequence or the complement sequence of the target site downstream sequence on the DNA / RNA hybridization framework or DNA / RNA hybridization derivative framework is used for hybridization with the target site downstream sequence or the complement sequence of the target site downstream sequence in a nucleic acid sequence to be cleaved or cleaved and edited; the target site upstream sequence and the target site downstream sequence on the DNA / RNA hybridization framework or DNA / RNA hybridization derivative framework are directly connected to the corresponding sequences in a nucleic acid sequence to be cleaved or cleaved and edited; and the target site upstream sequence and the target site downstream sequence of the nucleic acid sequence to be cleaved or cleaved and edited are connected by a target site;

[0078] The single-stranded RNA framework, double-stranded RNA framework, single-stranded RNA derivative framework, and double-stranded RNA derivative framework include a target site upstream sequence, a sequence to be inserted, and a target site downstream sequence in the 5'→3' direction;

[0079] The target site upstream sequence on the single-stranded RNA framework, the double-stranded RNA framework, the single-stranded RNA derivative framework, the double-stranded RNA derivative framework, or the complement sequence of the target site upstream sequence is used for hybridizing with the upstream sequence of the target site or the complement sequence of the target site in the nucleic acid sequence to be cleaved or edited after cleavage, and the target site downstream sequence on the single-stranded RNA framework, the double-stranded RNA framework, the single-stranded RNA derivative framework, the double-stranded RNA derivative framework, or the complement sequence of the target site downstream sequence is used for hybridizing with the downstream sequence of the target site or the complement sequence of the target site in the nucleic acid sequence to be cleaved or edited after cleavage; the target site upstream sequence and the target site downstream sequence on the single-stranded RNA framework, the double-stranded RNA framework, the single-stranded RNA derivative framework, the double-stranded RNA derivative framework are directly connected in the corresponding sequences in the nucleic acid sequence to be cleaved or edited after cleavage; the target site upstream sequence and the target site downstream sequence of the nucleic acid sequence to be cleaved or edited after cleavage are connected by a target site;

[0080] The cleavage structure is a protein or a polypeptide or DNA or RNA encoding the protein or the polypeptide, wherein the protein is ORF2p-EN, ORF2p, an ORF2p-EN homologous protein, an ORF2p derivative protein, an ORF2p-EN derivative protein, an ORF2p-EN homologous protein derivative protein; the polypeptide is a length of 50%-99% of the protein.

[0081] The cleaved or edited nucleic acid sequence after cleavage is single-stranded DNA, double-stranded DNA, single-stranded RNA, double-stranded RNA, or DNA / RNA hybrid chain.

[0082] Long interspersed element (LINE) is a transposon widely distributed in various types of organisms, mainly including various types of LINE-1 (L1) such as L1 (L1RE1 (L1.2, LRE1) and LRE2) in human body, various types of LINE-2 (L2) and various types of LINE-3 (L3), Ta element (such as Ta-1d) and R2, RandI, L1, RTE, I and Jockey six types of LINE, and other LINE types such as LINE-1 in mice, LINE UnaL2 in eels, LINE R2 in insects, LINE ZfL2-1 and ZfL2-2 in zebrafish, L1 in algae, LINE SART1 in silkworms, L1 in monocotyledons, Tad1 in fungi, L2 in fish, and RTE in some mammals.

[0083] The open reading frame 2 protein (ORF2p, or L1 endonuclease) is a reverse transcriptase and endonuclease encoded by the retrotransposon LINE, and the endonuclease part in the ORF2p is ORF2p-EN. The ORF2p and ORF2p-EN in the present application refer to the ORF2p and ORF2p-EN in all LINE species.

[0084] Before the in vitro nucleic acid or genome is cleaved, the frame structure needs to be first subjected to strand exchange with the in vitro nucleic acid or genome to form a bulge structure in the shape of Ω, and then the single strand is cleaved by the ORF2p-EN. Therefore, the proteins and RNAs related to homologous recombination or strand displacement (strand exchange), such as RAD51, RAD52, PALB2, BRCA2, and DNA polymerases with strand displacement function, can affect the effect of the present application.

[0085] RAD51 can promote the strand exchange between the exogenous single-stranded DNA and the double-stranded DNA on the in vitro nucleic acid or genome, so that the exogenous DNA binds to the complementary DNA single strand on the in vitro nucleic acid or genome and displaces another in vitro nucleic acid or genome DNA single strand. When the exogenous single-stranded DNA does not match the corresponding in vitro nucleic acid or genome DNA single strand, the strand exchange stops. Therefore, only when the exogenous single-stranded DNA has sufficient homologous sequences with the corresponding in vitro nucleic acid or genome DNA single strand, the strand exchange can be effectively initiated, so that the exogenous single-stranded DNA binds to the complementary single strand on the in vitro nucleic acid or genome and forms a bulge structure in the shape of Ω at the target site, and then the DNA single strand at the target site is cleaved under the action of the ORF2p-EN.

[0086] RAD52 and PALB2 can promote the strand exchange between the single-stranded DNA and the in vitro nucleic acid or genome, or even the strand exchange between the DNA analog such as single-stranded RNA and the genome, just like RAD51.

[0087] RAD52 and PALB2 can also perform reverse strand exchange, that is, bind to double-stranded DNA and promote the strand exchange between the double-stranded DNA and single-stranded DNA or single-stranded RNA.

[0088] Meanwhile, homologous recombination-related proteins such as BRCA2 can also unwind DNA / RNA hybrid double strands, DNA double strands, and RNA double strands, and promote the binding and strand exchange between various frames and the editing objects.

[0089] The DNA polymerases with strand displacement function include Bst DNA polymerase of family A and phi29 DNA polymerase from family B. DNA polymerases, such as bacteriophage, etc., including but not limited to Bst polymerase, phi29 DNA polymerase and its mutants such as EquiPhi29 DNA polymerase, Bst X DNA polymerase, Sau DNA polymerase, Bsu DNA polymerase.

[0090] The DNA framework and / or DNA / RNA hybrid framework for nucleic acid editing in the present application can be reverse transcribed by reverse transcriptase, generated from RNA in vitro or in vivo, and play a nucleic acid editing role.

[0091] In order to achieve the purpose of reverse transcription, other sequences can be further added outside the framework structure for nucleic acid editing, referred to as DNA derived framework and / or DNA / RNA hybrid derived framework and / or RNA derived framework, for example, transposon sequences, retrotransposon sequences or part of the sequences (especially the part related to the binding of the corresponding reverse transcriptase, generally the retrotransposon terminal sequence or the sequence close to the terminal, such as the LTR sequence on both sides of the transposon, the UTR sequence, the ME (Mosaic End) sequence, or the polyA sequence bound by ORF2p) can be added downstream of the sequence downstream of the target site and / or upstream of the sequence upstream of the target site to initiate reverse transcription, as shown in Figure 2.

[0092] The transposon as described above includes various class I and class II transposons. Class I transposons include LTR retrotransposons and non-LTR retrotransposons (Non-LTR), LTR retrotransposons include five superfamilies of Copia, Gypsy, BEL, DIRS and endogenous retroviruses (ERV). ERV includes ERV1, ERV2, ERV3, ERV4 and endogenous lentivirus (ELV), ERV1 includes Gammaretrovirus and Epsilonretrovirus, ERV2 includes Alpharetrovirus and Betaretrovirus, LTR retrotransposons also include nucleic acids of hepatitis B virus and nucleic acids of cruciferous virus.

[0093] The non-LTR retrotransposons include various types of CRE, NeSL, R4, R2, Hero, RandI / Dualen, L1, Proto1, Tx1, Proto2, RTE, RTEX, RTETP, I, Nimb, Ingi, Vingi, Tad1, Loa, R1, Outcast, Jockey, CR1, L2, L2A, L2B, Kiri, Rex1, Crack, Daphne, Ambal, Penelope transposons; the non-LTR retrotransposons also include non-autonomous SINE (short interspersed element) and autonomous LINE (long interspersed element), wherein the SINE further includes SINE1 representing 7SL RNA, SINE2 representing tRNA, SINE3 representing 5S rRNA, SINEU representing U1 or U2 snRNA, SINE4 representing unknown origin, and the SINE can further include various types of CORE-SINE, V-SINE, Deu-SINE (or Nin-SINE), Ceph-SINE and Meta-SINE.

[0094] The class I transposons further include DIRS retrotransposon (also known as tyrosine recombinase encoding retrotransposon, YR retrotransposon) and Penelope-like retrotransposon (Penelope-like element, PLE).

[0095] The retrotransposons as described above include various types of long terminal repeat (LTR) transposons and non-LTR transposons; the LTR transposons include Ty1-copia type, Ty3-gypsy type and BEL-Pao type, such as DIRS retrotransposon (tyrosine recombinase retrotransposon), mammalian apparent LTR-retrotransposons (MaLR), endogenous retroviruses (ERV), human endogenous retroviruses (ERV), miniature terminal repeat retrotransposon (TRIM); the non-LTR transposons include various types of SINE, various types of LINE, such as LINE-1, Alu element, SVA element, R2 element, class II intron, etc.

[0096] The retrotransposon sequence as described above includes various types of short interspersed elements (SINEs), long interspersed elements (LINEs), Alu elements in SINEs, SVA elements, etc.; the main feature of SINEs is that they are relatively short transposons distributed in the genome, containing internal RNA polymerase III promoters and ending with A or T-rich tails or short simple repeats, and are reverse transcribed by means of LINEs, and the right half of the transcription product contains reverse transcription functional structure; the feature of LINEs is that they are transposons widely distributed in the genome containing reverse transcriptase coding sequences. Short interspersed elements mainly include Alu elements (such as Alu Jo element, Alu Jb element, Alu Sq element, Alu Sx element, Alu Sp element, Alu Sc element, Alu Sg element, Alu S element, Alu Y element, Alu Yb8 element, Alu Ya5 element, Alu Ya8 element, Alu J element, etc.) and SVA elements in primates (including humans), various types of mammalian-wide interspersed repeat elements (MIRs) commonly found in mammals such as MIR and MIR3, Mon-1 in Monodelphis, B1 and B2 elements in mice, C-element in rabbits, HE1 family in zebrafish, SINE SmaI in salmon, Anolis SINE2 and Sauria SINE in reptiles, IdioSINE1, IdioSINE2, SepiaSINE, Sepioth-SINE1, Sepioth-SINE2A, Sepioth-SINE2B and OegopSINE in invertebrates such as cuttlefish, and p-SINE1 in plants such as rice, etc. SINEs and LINEs are widely present in various animals and plants and are scattered throughout the genome, and each organism has its specific SINE and corresponding LINE.

[0097] Other transposons also include Tol1 transposon, type II transposon such as Tn3, Tn5, Tn10, Mariner transposon, Tol2 transposon, Tc1 / mariner transposon superfamily such as Sleeping Beauty (SB) transposon, PiggyBac (PB) transposon, Helitrons transposon, hAT transposon family, etc.

[0098] Some of the sequences can include a short interspersed element (SINE), a long interspersed element (LINE), or a part of the sequence of the short interspersed element or the long interspersed element. For example, the sequence of the Alu element is the remaining part of the sequence of the scAlu element (the polyA sequence and the sequence after the polyA sequence in the Alu element, which can further include the ACT sequence before the polyA sequence, the part of the sequence after the 118th position of the Alu sequence), the polyA sequence, and the 3' UTR of the R2 element. The right half of the sequence of the complete SINE transcription product after the cleavage at the middle site is a part of the short interspersed element, and the cleavage sites of different short interspersed elements in different species are different. The natural cleavage site of the short interspersed element is generally located in the middle of the full-length sequence or before the middle. For the short interspersed element with a full length of about 100-400 nt, the natural cleavage site is generally located at the 100th-250th position. For example, the cleavage site (the scAlu cleavage site or the natural cleavage site, which is generally located before the polyA sequence in the middle of the Alu transcription product, and the actual situation can be floating) of the Alu element with a full length of about 300 bp is located at the 118th position. The cleavage product includes the right monomer of the Alu element, and can also include the middle adenine repeat sequence of the Alu element transcription product after the right monomer, as well as the 2-3 bases upstream of the sequence and the 3' polyA repeat sequence after the right monomer, which can be referred to as a part of the Alu element. For the MIR with a full length of about 260 nt, the cleavage site can be observed in the range of the 100th-150th position, and the part after the cleavage site is also a part of the MIR.

[0099] Further, in order to achieve the purpose of reverse transcription, the primer, the primase, or the vector expressing the primase can be simultaneously given to the target system to start the reverse transcription (the double-stranded DNA framework can also be synthesized by the DNA polymerase, or the double-stranded RNA can be synthesized by the RNA polymerase), as shown in FIG. 3.

[0100] The primer can be in the same vector as the framework structure for nucleic acid editing or in different vectors, such as the primer can be separated from the framework structure and administered separately, such as in different vectors, or connected to the framework structure and administered simultaneously, such as in the same vector; when the primer is in the same vector as the framework structure for nucleic acid editing, it can be upstream of the upstream sequence of the target site and complementary to the upstream sequence of the target site, or downstream of the downstream sequence of the target site and complementary to the downstream sequence of the target site, or downstream of the added transposon sequence, retrotransposon sequence or partial sequence thereof and complementary to the transposon sequence, retrotransposon sequence or partial sequence thereof, or at the 3' end of the vector containing the framework structure or any other position; regardless of whether the primer is in the same vector or different vectors as the single-stranded DNA framework, DNA / RNA hybrid framework and / or single-stranded RNA framework, its 3' end must be free, or generate a 3' free end in the reaction system, such as in solution, nucleus or cytoplasm, tissue, organ, organism.

[0101] The primer can also be designed to be complementary to any part of the nucleic acid to be edited to facilitate the conversion of the nucleic acid to be edited into double-stranded DNA, DNA / RNA hybrid chain or double-stranded RNA, so as to facilitate the subsequent implementation of nucleic acid editing.

[0102] The above-mentioned primer is a DNA primer or an RNA primer.

[0103] The reverse transcriptase in the present application includes various reverse transcriptases in nature, partial regions of corresponding reverse transcriptases, especially regions responsible for reverse transcription in corresponding reverse transcriptases, and combinations of each part of the same reverse transcriptase or combinations of multiple parts of multiple reverse transcriptases, including AMV, M-MLV, ORF2p, reverse transcriptase function related parts in ORF2p, such as domains, type II intron reverse transcriptase Induro, R2 element protein, reverse transcriptase function related parts in R2 element protein, such as domains, R2Bm element protein, reverse transcriptase function related parts in R2Bm element protein, such as domains, etc.

[0104] The present application can also synthesize double-stranded DNA framework or DNA / RNA hybrid framework with DNA single strand or RNA single strand as substrate with the aid of reverse transcriptase, transposase (such as DD[E / D]-transposase (Tpase)), integrase (such as Ginger1 and Ginger2 superfamily), tyrosine recombinase (such as Crypton), and Rep / Helicase1 (such as Helitron), and further used for nucleic acid cleavage or editing.

[0105] The present application can be used for in vitro nucleic acid editing, such as editing double-stranded DNA or single-stranded DNA in vitro, for example, inserting other DNA sequences in double-stranded DNA or single-stranded DNA. The present application can be used for in vitro nucleic acid destruction, such as cleaving and destroying double-stranded DNA or single-stranded DNA in vitro.

[0106] The present application can be used for in vitro nucleic acid editing, such as editing double-stranded RNA or single-stranded RNA in vitro, for example, inserting other sequences in double-stranded RNA or single-stranded RNA. The present application can be used for in vitro nucleic acid destruction, such as cleaving and destroying double-stranded RNA or single-stranded RNA in vitro.

[0107] The above-mentioned in vitro nucleic acid editing also requires the additional addition of DNA polymerase or RNA polymerase or primase, DNA polymerase including various DNA polymerases in prokaryotes, eukaryotes and other organisms, such as DNA polymerase I (Pol I), Klenow fragment, DNA polymerase II (Pol II), DNA polymerase III (Pol III), DNA polymerase IV (Pol IV), DNA polymerase V (Pol V), family D DNA polymerase (family D) in prokaryotes, Pol a, Pol b, Pol g, Pol d, Pol e, Pol z in eukaryotes, RNA polymerase including bacterial RNA polymerase, archaeal RNA polymerase, viral RNA polymerase such as bacteriophage T7, eukaryotic RNA polymerase such as RNA polymerase I, RNA polymerase II, RNA polymerase III, other mitochondrial and chloroplast RNA polymerase species. Primase includes DnaG, PRIM1, PRIM2.

[0108] The DNA polymerase used in the present application also includes group A such as T7 DNA polymerase, Pol I, and DNA Polymerase g, group B such as Pol II, Pol B, Pol z, Pol a, d, and e; group C such as Pol III, group D, group X such as Pol b, Pol s, Pol l, Pol m, and Terminal deoxynucleotidyl transferase, group Y such as Pol iota (iota), Pol kappa (kappa), Pol eta (eta), Pol IV, Pol V, RT group such as telomerase, Hepatitis B virus.

[0109] The DNA polymerase used in the present application includes any DNA polymerase, and also includes E. coli DNA polymerase I, T4 DNA polymerase, 9°N polymerase, phi29 DNA polymerase, taq DNA polymerase, Bst polymerase, high-fidelity DNA polymerases such as pfu DNA polymerase and KOD, Pfx DNA polymerase, Phusion high-fidelity enzyme, phi29 DNA polymerase, and mutants thereof such as EquiPhi29 DNA polymerase, Bst X DNA polymerase, Sau DNA polymerase, Bsu DNA polymerase, and the like.

[0110] The DNA polymerase used in the present application includes DNA polymerases with strand displacement function, including Bst DNA polymerase of family A and phi29 DNA polymerase from family B, and the like, including but not limited to Bst polymerase, phi29 DNA polymerase, and mutants thereof such as EquiPhi29 DNA polymerase (and other phi 29 DNA polymerases that have been mutated to remove exonuclease activity), Bst X DNA polymerase, Sau DNA polymerase, Bsu DNA polymerase.

[0111] The DNA polymerase used in the present application includes mutants or modified proteins of the above-mentioned DNA polymerases; the RNA polymerase used in the present application includes mutants or modified proteins of the above-mentioned RNA polymerases.

[0112] The recombinant protein or DNA polymerase with strand displacement function in the present application can be replaced by other proteins with strand displacement function.

[0113] The present application can be used for nucleic acid editing in prokaryotic organisms or eukaryotic organisms, such as editing double-stranded DNA, single-stranded DNA, or genomes in prokaryotic organisms or eukaryotic organisms, for example, inserting other DNA sequences into double-stranded DNA, single-stranded DNA, or genomes. The present application can be used for destroying nucleic acids in prokaryotic organisms or eukaryotic organisms, such as cleaving and destroying double-stranded DNA, single-stranded DNA, or genomes in prokaryotic organisms or eukaryotic organisms.

[0114] The present application can be used for nucleic acid editing in prokaryotic organisms or eukaryotic organisms, such as editing double-stranded RNA or single-stranded RNA in prokaryotic organisms or eukaryotic organisms, for example, inserting other sequences into double-stranded RNA or single-stranded RNA. The present application can be used for destroying nucleic acids in prokaryotic organisms or eukaryotic organisms, such as cleaving and destroying double-stranded RNA or single-stranded RNA in prokaryotic organisms or eukaryotic organisms.

[0115] The RNA can be any RNA such as mRNA, pre-mRNA, rRNA, tRNA, lncRNA, miRNA, snRNA, circRNA, Alu-RNA, LINE-RNA, and can be RNA related to various diseases such as cancer.

[0116] The application can be applied to treat various diseases such as cancer, neurodegenerative diseases, etc.

[0117] The prokaryote includes various prokaryotes such as bacteria. The eukaryote includes various eukaryotes such as various animals, plants, or fungi, etc. The mammal includes primates such as humans, apes, monkeys, etc., rodents such as mice, etc., cats such as lions, cats, tigers, etc., canids such as wolves, dogs, etc., reptiles, birds, fish, etc., and various invertebrates, etc. Other objects of the application can also include viruses, bacteriophages, and other life forms.

[0118] The editing object of the application includes various nucleic acids in cells, tissues, organs, or organisms of the above-mentioned organisms.

[0119] The editing object of the application can be the genome of various organisms, such as the genome of a mammal such as a human. In particular, the editing object can be SINE (such as Alu element), LINE-1, and CNV of various genes and the CNV end in the human genome to hinder further changes in the genome caused by them.

[0120] The editing object of the application can be nucleic acid in an in vitro solution.

[0121] For the site actually cut, the 3' adjacent nucleic acid or deoxy nucleic acid sequence of several (not more than 4) contains more purines, and / or the 5' adjacent nucleic acid or deoxy nucleic acid sequence of several (not more than 4) contains more pyrimidines, or improves the nucleic acid cutting or editing efficiency to a certain extent, such as 5'-TTTT / AA-3', 5'-TCTT / AG-3', etc.

[0122] The framework structure of the application refers to one or more of double-stranded DNA framework, single-stranded DNA framework, DNA / RNA hybrid framework, single-stranded RNA framework, double-stranded RNA framework, double-stranded DNA derivative framework, single-stranded DNA derivative framework, DNA / RNA hybrid derivative framework, single-stranded RNA derivative framework, and double-stranded RNA derivative framework.

[0123] The destruction of the application refers to the destruction of the continuity of nucleic acid, which can affect its biological function.

[0124] As shown in Figure 4 is the principle of the frame structure for nucleic acid editing in the present application, the gene editing technology can achieve "in vitro editing or cleavage of double-stranded DNA, single-stranded DNA, double-stranded RNA, single-stranded RNA", "in vivo cleavage or editing of the genome of eukaryotes or prokaryotes", "in vivo cleavage or editing of various types of RNA in eukaryotes or prokaryotes", "genome sequence replacement technology in eukaryotes or prokaryotes (including sequence replacement, site deletion, site addition, sequence addition, sequence deletion and site replacement)", "in vitro DNA sequence replacement technology (including sequence replacement, site deletion, site addition, sequence addition, sequence deletion and site replacement)", "preventing genomic changes caused by transposons, stabilizing the genome and CNVs thereon" and other technologies, which will be explained one by one.

[0125] I. In vitro editing or cleavage of double-stranded DNA, single-stranded DNA, double-stranded RNA, single-stranded RNA, as shown in Figure 5:

[0126] 1. In vitro generation of purified cleavage structure and editing or cleavage of double-stranded DNA, single-stranded DNA, double-stranded RNA, single-stranded RNA:

[0127] (1) Determine the nucleic acid object and its sequence that needs to be edited or cleaved in vitro and the corresponding site that needs to be cleaved or edited, and design the frame structure accordingly. Here the nucleic acid object refers to double-stranded DNA, single-stranded DNA, double-stranded RNA or single-stranded RNA.

[0128] (2) Express and purify the cleavage structure (cleavage protein or polypeptide) through a prokaryotic system or a eukaryotic system, and put it together with the frame structure or a carrier containing the frame structure into a system with the nucleic acid to be cleaved or edited for reaction.

[0129] (3) If the nucleic acid object to be cleaved or edited and the frame structure are both single-stranded DNA or single-stranded RNA, the frame structure and the single-stranded DNA or single-stranded RNA to be edited can directly bind to form an Ω-shaped bulge structure and be cleaved by the cleavage structure.

[0130] (4) If the nucleic acid object to be cleaved or edited and one side of the frame structure are double-stranded DNA, double-stranded RNA or DNA / RNA hybrid, and the other side is single-stranded DNA or single-stranded RNA, it is necessary to add homologous recombination-related proteins such as Rec / Rad51 family such as RAD51, RecA, Uvs X, BRCA2, RAD52, etc. to the reaction system to promote the melting of the double-stranded DNA, double-stranded RNA or DNA / RNA hybrid of the nucleic acid object to be cleaved or edited and the strand exchange with the frame structure, forming a bulge structure. If further improvement of efficiency is required, the nucleic acid object to be cleaved or edited and the frame structure, or the nucleic acid to be edited can be subjected to chromosome assembly. In addition, if further improvement of efficiency is required, compounds that improve homologous recombination efficiency such as RS-1, SCR7, NU7441, Vanillin, Chlorpromazine, Resveratrol, Benomyl, Inositol can be added. In addition, if further improvement of efficiency is required, single-stranded binding proteins such as gp32 (such as T4), Redβ, RecT, SSB, HSSB1, SSAPs, RPA, etc. can be added to bind to the displaced single strand to stabilize the Ω-shaped bulge structure formed by the binding of the single-stranded DNA or single-stranded RNA to be edited to the single-stranded DNA or single-stranded RNA of the frame structure. In addition, if further improvement of efficiency is required, auxiliary factors such as T4 uvs Y or crowding agents such as peg, glycerol, Carbowax20M, etc. can be added. After the frame structure binds to the single-stranded DNA or single-stranded RNA to be edited to form an Ω-shaped bulge structure, it is cleaved by the cleavage structure.

[0131] (5) If the nucleic acid object to be cleaved or edited and one side of the frame structure are double-stranded DNA, double-stranded RNA or DNA / RNA hybrid, and the other side is single-stranded DNA or single-stranded RNA, it is necessary to add recombination enzymes or exoI, exoIII, exoV, RecBCD, exoVII, RecJ, RecE, λexo, Redα, Xrn1, TREX1, TREX2, Pol γ, Pol δ, Pol ε, T7 gene 6 exonuclease, MRN-CtIP, DNA2, EXO1, WRN protein, P53, MRE11, RAD1, RAD9, APE1, VDJP protein, etc. to the reaction system. Nucleic acid exonuclease or protein with exonuclease function to make the double-stranded structure into a single-stranded structure, so that the frame structure and the single-stranded DNA or single-stranded RNA to be edited can be directly combined to form an Ω-shaped bulge structure, and then cleaved by the cleavage structure.

[0132] (6) The frame structure exchanges strands with the double-stranded DNA, double-stranded RNA or DNA / RNA hybrid chain to be edited, and is cut by a cleavage protein or cleavage polypeptide after the two are combined and form an Ω-shaped bulge structure.

[0133] (7) After the nucleic acid object to be edited is cut, a DNA polymerase or RNA polymerase can be added to the reaction system according to requirements to further edit the corresponding nucleic acid object, and insert the DNA or RNA sequence of the sequence to be inserted into the insertion site in the nucleic acid object to be edited.

[0134] (8) The frame structure, or the carrier containing the frame structure, can be one or more, and multiple sites can be edited; multiple single-stranded DNA frames, single-stranded RNA frames, single-stranded DNA derivative frames, and single-stranded RNA derivative frames can edit the complementary sites of the same position of the double-stranded DNA, double-stranded RNA, and DNA / RNA hybrid chain to be edited to increase efficiency.

[0135] 2. In vitro production of purified cleavage structure and editing or cleavage of double-stranded DNA, single-stranded DNA, double-stranded RNA, and single-stranded RNA:

[0136] (1) Determine the nucleic acid object and its sequence that needs to be edited or cleaved in vitro and the corresponding site that needs to be cut or edited, and design the frame structure (referred to as frame) accordingly.

[0137] (2) When one or both of the nucleic acid to be edited or cleaved or the frame described above is double-stranded (it is recommended to use a single-stranded frame with a double-stranded nucleic acid to be edited, and to increase the concentration of the single-stranded frame to increase efficiency), the nucleic acid to be edited and the frame can be mixed and annealed after denaturation, such as PCR, to open the double strands and combine the single-stranded frame with the single-stranded nucleic acid to be edited to form a hybrid chain, and the two are combined and form an Ω-shaped bulge structure. Recover the hybrid chain and place it in the reaction system;

[0138] (3) The cleavage structure is expressed and purified by a prokaryotic system or a eukaryotic system, and is added to the reaction system containing the hybrid product of the single-stranded frame and the single-stranded nucleic acid to be edited, and cuts the single-stranded nucleic acid to be edited. In addition, if further efficiency improvement is required, single-stranded binding proteins such as gp32 (such as T4), Redβ, RecT, SSB, HSSB1, SSAPs, RPA, etc. can be added to bind with free single strands to stabilize the Ω-shaped bulge structure formed by the combination of the single-stranded frame and the single-stranded nucleic acid to be edited. In addition, if further efficiency improvement is required, auxiliary factors such as T4 uvs Y can also be added.

[0139] (4) When the nucleic acid object to be edited is cut, DNA polymerase or RNA polymerase can be added to the reaction system according to the requirement (if a heat-resistant DNA polymerase such as Taq enzyme is added, it can be added at any stage of the whole process) to further edit the corresponding nucleic acid object, and the DNA or RNA sequence of the sequence to be inserted is inserted into the insertion site in the nucleic acid object to be edited.

[0140] (5) The framework structure or the carrier containing the framework structure can be one or more, and multiple sites can be edited; multiple single-stranded DNA frameworks, single-stranded RNA frameworks, single-stranded DNA derivative frameworks, and single-stranded RNA derivative frameworks can edit the complementary sites of the same position of the double-stranded DNA, double-stranded RNA, and DNA / RNA hybrid chain to be edited to increase the efficiency.

[0141] 3. In vitro production of purified cleavage structure and editing or cleavage of double-stranded DNA, single-stranded DNA, double-stranded RNA, and single-stranded RNA:

[0142] (1) Determine the nucleic acid object and its sequence that needs to be edited or cleaved in vitro and the corresponding site that needs to be cut or edited, and design the framework structure (referred to as framework) accordingly.

[0143] (2) When one or both of the nucleic acid to be edited or the above framework is double-stranded, the nucleic acid to be edited and the framework can be mixed in the reaction system, and then an appropriate amount of DNA polymerase with strand displacement function can be added, and a cleavage structure expressed and purified in a prokaryotic or eukaryotic system can also be added at the same time, to open the double-stranded and combine the single-stranded framework with the single-stranded nucleic acid to be edited to form a hybrid chain, and then the two are combined to form an Ω-shaped bulge structure, and the single-stranded nucleic acid to be edited is cut.

[0144] (3) When the nucleic acid object to be edited is cut, DNA polymerase with strand exchange function synthesizes DNA, and the DNA or RNA sequence of the sequence to be inserted is inserted into the insertion site in the nucleic acid object to be edited.

[0145] (4) Other DNA polymerases or RNA polymerases can be added to the reaction system according to the requirement, or PCR means can be used to edit the corresponding nucleic acid object.

[0146] (5) The framework structure or the carrier containing the framework structure can be one or more, and multiple sites can be edited; multiple single-stranded DNA frameworks, single-stranded RNA frameworks, single-stranded DNA derivative frameworks, and single-stranded RNA derivative frameworks can edit the complementary sites of the same position of the double-stranded DNA, double-stranded RNA, and DNA / RNA hybrid chain to be edited to increase the efficiency.

[0147] 4. By extracting the cytosol of the cell expressing the cleavage structure, the reaction is carried out in the liquid:

[0148] (1) The vector expressing the cleavage structure is introduced into prokaryotic cells or eukaryotic cells to produce the cleavage structure; or the cell stably expressing the cleavage structure is obtained by biological engineering.

[0149] (2) The cell producing the cleavage structure is taken and the cytoplasm is extracted, and the frame structure, or the vector containing the frame structure, is put into the reaction system together with the nucleic acid object to be cut or edited and reacted.

[0150] (3) If the nucleic acid object to be cut or edited is single-stranded DNA or single-stranded RNA, the frame structure can be directly combined with the single-stranded DNA or single-stranded RNA to be edited and cut by the cleavage structure.

[0151] (4) If the nucleic acid object to be cut or edited is double-stranded DNA, double-stranded RNA or DNA / RNA hybrid chain, it is necessary to add homologous recombination related proteins such as Rec / Rad51 family such as RAD51, RecA, Uvs X, RAD51, BRCA2, RAD52, etc. to the reaction system to promote the melting of the nucleic acid object to be cut or edited as double-stranded DNA, double-stranded RNA or DNA / RNA hybrid chain and the strand exchange of the frame structure. If further improvement of efficiency is required, the nucleic acid object to be cut or edited and / or the frame structure, or the nucleic acid to be edited can be subjected to chromosome assembly. In addition, if further improvement of efficiency is required, compounds such as RS-1, SCR7, NU7441, Vanillin, Chlorpromazine, Resveratrol, Benomyl, Inositol can be added to improve the efficiency of homologous recombination. In addition, if further improvement of efficiency is required, single-stranded binding proteins such as gp32 (such as T4), Redβ, RecT, SSB, HSSB1, SSAPs, RPA, etc. can be added to bind to the displaced single strand to stabilize the Ω-shaped bulge structure formed by the single-stranded binding of the frame structure and the single-stranded DNA or single-stranded RNA to be edited. In addition, if further improvement of efficiency is required, auxiliary factors such as T4 uvs Y or crowding agents such as peg, glycerol, Carbowax20M, etc. can be added. After the frame structure and the single-stranded DNA or single-stranded RNA to be edited form the Ω-shaped bulge structure, they are cut by the cleavage structure.

[0152] (5) If the nucleic acid object to be cleaved or edited is double-stranded DNA, double-stranded RNA or DNA / RNA hybrid on both sides of the frame structure, or one is double-stranded DNA, double-stranded RNA or DNA / RNA hybrid and the other is single-stranded DNA or single-stranded RNA, a recombinase or exoI, exoIII, exoV, RecBCD, exoVII, RecJ, RecE, lambda exo, Red alpha, Xrn1, TREX1, TREX2, Pol gamma, Pol delta, Pol epsilon, T7 gene 6 exonuclease, MRN-CtIP, DNA2, EXO1, WRN protein, P53, MRE11, RAD1, RAD9, APE1, VDJP protein and other exonucleases or proteins with exonuclease function are added to the reaction system to make the double-stranded structure into a single-stranded structure, and then the frame structure and the single-stranded DNA or single-stranded RNA to be edited can be directly combined to form an omega-shaped bulge structure and be cleaved by the cleavage structure.

[0153] (6) The frame structure and the double-stranded DNA, double-stranded RNA, DNA / RNA hybrid to be edited perform strand exchange, and when they are combined to form an omega-shaped bulge structure, they are cleaved by the cleavage structure.

[0154] (7) After the nucleic acid object to be edited is cleaved, a DNA polymerase or an RNA polymerase can be added to the reaction system as needed to further edit the corresponding nucleic acid object and insert the DNA or RNA sequence of the sequence to be inserted into the insertion site in the nucleic acid object to be edited.

[0155] (8) The frame structure or the carrier containing the frame structure can be one or more, and multiple sites can be edited; multiple single-stranded DNA frames, single-stranded RNA frames, single-stranded DNA derivative frames, and single-stranded RNA derivative frames can edit the complementary sites at the same position of the double-stranded DNA, double-stranded RNA, and DNA / RNA hybrid to be edited to increase efficiency.

[0156] II. Cleavage or editing of the genome in eukaryotic or prokaryotic organisms:

[0157] (1) Determine the genomic sequence that needs to be edited or cleaved and the corresponding site that needs to be cleaved or edited, and design the frame structure accordingly.

[0158] (2) Introduce the frame structure or the carrier containing the frame structure into the cell to be edited to express the cleavage structure,

[0159] or if the system to be edited does not express the cleavage structure itself, the frame structure or the carrier containing the frame structure is introduced into the system to be edited together or sequentially with the carrier (DNA or RNA) expressing the cleavage structure or the protein of the cleavage structure.

[0160] (3) can be selected to express or improve the expression of homologous recombination related proteins in the system to be edited to improve efficiency.

[0161] (4) can be selected to assemble the chromosome of the framework structure or the vector containing the framework structure, or the nucleic acid to be edited, to improve efficiency.

[0162] (5) can be selected to add compounds RS-1, SCR7, NU7441, Vanillin, Chlorpromazine, Resveratrol, Benomyl, Inositol to improve the efficiency of homologous recombination.

[0163] (6) Since the homologous recombination mechanism required for the auxiliary strand exchange in the present application mainly occurs in the S phase of the cell, the cell cycle has a certain auxiliary effect on the present application, which can make the cell divide or remain in the corresponding cell cycle such as S phase after being given the vector system of the present application to promote the effect of the present application. However, some cell lines express homologous recombination related proteins throughout the cell cycle, so the influence is relatively small.

[0164] (7) The framework structure and the corresponding target sequence of the genome to be edited perform strand exchange, and when the two are combined and form a bulge structure in the shape of Ω, the target site is cut by the scissors structure.

[0165] (8) After the target site to be edited is cut, the genome repair system will repair it, and the sequence to be inserted will be inserted into the corresponding site of the genome.

[0166] (9) The framework structure or the vector containing the framework structure can be one or more, and multiple sites can be edited; multiple single-stranded DNA frameworks, single-stranded RNA frameworks, single-stranded DNA derivative frameworks, and single-stranded RNA derivative frameworks can edit the complementary sites of the same position of the genome to be edited to increase efficiency.

[0167] III. Cutting or editing various types of RNA in eukaryotic or prokaryotic organisms:

[0168] (1) Determine the genomic sequence that needs to be edited or cut and the corresponding site that needs to be cut or edited (insertion site, target site), and design the framework structure accordingly.

[0169] (2) Introduce the framework structure or the vector containing the framework structure into the edited cell expressing the scissors structure,

[0170] Or if the system to be edited itself does not express the production of the cleavage structure, the framework structure or the carrier containing the framework structure is introduced into the system to be edited together or successively with the DNA or RNA carrier expressing the cleavage structure, or the protein / polypeptide of the cleavage structure.

[0171] (3) The expression of the homologous recombination related protein expressed in the system to be edited can be selected to increase the efficiency, and this step can be omitted as the RNA is mostly single-stranded.

[0172] (4) The chromosome assembly of the framework structure or the carrier containing the framework structure can be selected to improve the efficiency. Since the RNA is mostly single-stranded, the single-stranded DNA framework, the single-stranded RNA framework, the single-stranded DNA derived framework, and the single-stranded RNA derived framework can be directly combined with the RNA, and this step can be omitted in this case.

[0173] (5) Since the homologous recombination mechanism required for the auxiliary strand exchange in the present application mainly occurs in the S phase of the cell, the cell cycle has a certain auxiliary effect on the present application, which can make the cell divide or remain in the corresponding cell cycle such as the S phase after the carrier system of the present application is given to promote the effect of the present application. However, some cell lines express homologous recombination related proteins throughout the cell cycle, so the effect is relatively small. Since the RNA is mostly single-stranded, the single-stranded DNA framework, the single-stranded RNA framework, the single-stranded DNA derived framework, and the single-stranded RNA derived framework can be directly combined with the RNA, and this step can be omitted in this case; this step can also be omitted as appropriate in other cases.

[0174] (6) The framework structure and the corresponding target sequence of the RNA to be edited perform strand exchange, and when they are combined and form an Ω-shaped bulge structure, the target site is cut by the cleavage structure, thereby destroying the RNA to be edited.

[0175] (7) After the target site to be edited is cut, a repair system such as a DNA polymerase or an RNA polymerase is used to repair it, and the insertion sequence is inserted into the corresponding site of the genome.

[0176] (8) The framework structure or the carrier containing the framework structure can be one or more, and multiple sites can be edited; multiple single-stranded DNA frameworks, single-stranded RNA frameworks, single-stranded DNA derived frameworks, and single-stranded RNA derived frameworks can edit the complementary sites of the same position of the RNA to be edited to increase the efficiency.

[0177] Four, genomic sequence replacement technology in eukaryotic or prokaryotic organisms, such as sequence replacement, site deletion, site addition, sequence addition, sequence deletion, or site replacement:

[0178] The sequence to be inserted in the framework designed in the above insertion technology is replaced by a replacement sequence and the surrounding sequence of the sequence to be replaced on the genome, i.e. the DNA sequence of the replacement sequence to be inserted and the sequence on the genome that will be deleted after homologous recombination. When constructing the vector, the 3' or 5' of the replacement sequence depends on whether the insertion point is upstream or downstream of the sequence to be replaced on the genome. The DNA sequence of the replacement sequence should be homologous to the sequence to be replaced on the genome. If the inserted sequence homologously recombines with the upstream sequence, any sequence can be inserted between the corresponding sequence of the sequence to be inserted on the genome and the corresponding sequence upstream of the target site, which does not affect the results or can promote subsequent homologous recombination and / or the effects it produces.

[0179] If the inserted sequence homologously recombines with the downstream sequence, any sequence can be inserted between the sequence to be inserted on the genome and the downstream sequence of the target site, which does not affect the results or can promote subsequent homologous recombination and / or the effects it produces. By the above gene editing insertion method, the replacement sequence and the surrounding sequence of the sequence to be replaced on the genome are inserted upstream or downstream of the sequence to be replaced on the genome. When the inserted replacement sequence homologously recombines with the sequence to be replaced on the genome, the sequence to be replaced on the genome is replaced by the inserted replacement sequence homologous to it. At the same time, the part of the surrounding sequence of the sequence to be replaced that is deleted due to homologous recombination is reinserted together with the replacement sequence at the time of insertion.

[0180] The replacement on the genome includes sequence replacement and site replacement. Sequence replacement means that the replacement sequence to be inserted has some sequence, such as one or several, that is inconsistent with the corresponding sequence on the genome. Site replacement means that the replacement sequence to be inserted has some site, such as one or several, that is inconsistent with the corresponding sequence on the genome. Site deletion means that the replacement sequence to be inserted has some site, such as one or several, that is missing compared to the corresponding sequence on the genome. Site addition means that the replacement sequence to be inserted has some site, such as one or several, that is added compared to the corresponding sequence on the genome. Sequence addition means that the replacement sequence to be inserted has some sequence, such as one or several, that is added compared to the corresponding sequence on the genome. Sequence deletion means that the replacement sequence to be inserted has some sequence, such as one or several, that is missing compared to the corresponding sequence on the genome.

[0181] The smaller the difference between the replacement sequence to be inserted and the corresponding homologous sequence on the genome, the higher the efficiency. The inconsistent part of the replacement sequence to be inserted and the corresponding homologous sequence on the genome is preferably located at or near the two ends or sides of the replacement sequence to be inserted to improve efficiency. If directional transfer is required, modifications can be made on the wrapping outside the vector.

[0182] Five, in vitro DNA sequence replacement technology, such as sequence replacement, site deletion, site addition, sequence addition, sequence deletion and site replacement:

[0183] The sequence to be inserted in the frame structure designed in the above insertion technology is changed to a replacement sequence, and the surrounding sequence of the sequence to be replaced in vitro DNA, i.e. the DNA sequence of the replacement sequence to be inserted and the sequence to be deleted after homologous recombination of the sequence on the in vitro DNA, is located 3' or 5' of the replacement sequence in the construction of the vector, which depends on whether the insertion point is upstream or downstream of the sequence to be replaced on the in vitro DNA. The DNA sequence of the replacement sequence should be homologous to the sequence to be replaced on the in vitro DNA. If the inserted sequence homologously recombines with the upstream sequence, any sequence can be inserted between the corresponding sequence of the sequence to be inserted on the in vitro DNA and the corresponding sequence upstream of the target site, which does not affect the results or can promote subsequent homologous recombination and / or the effects produced thereby. If the inserted sequence homologously recombines with the downstream sequence, any sequence can be inserted between the sequence to be inserted on the in vitro DNA and the downstream sequence of the target site, which does not affect the results or can promote subsequent homologous recombination and / or the effects produced thereby.

[0184] The replacement sequence and the surrounding sequence of the sequence to be replaced on the in vitro DNA are inserted upstream or downstream of the sequence to be replaced on the in vitro DNA by the above editing insertion method. When the inserted replacement sequence homologously recombines with the sequence to be replaced on the in vitro DNA, the sequence to be replaced on the in vitro DNA is replaced by the inserted replacement sequence homologous thereto, and the surrounding sequence of the sequence to be replaced that is deleted due to homologous recombination is reinserted together with the replacement sequence at the time of insertion.

[0185] The replacement on the in vitro DNA includes sequence replacement and site replacement. The sequence replacement means that the replacement sequence to be inserted has part of the sequence, such as one or several, inconsistent with the corresponding sequence on the in vitro DNA. The site replacement means that the replacement sequence to be inserted has part of the site, such as one or several, inconsistent with the corresponding sequence on the in vitro DNA. The site deletion means that the replacement sequence to be inserted has part of the site, such as one or several, missing compared to the corresponding sequence on the in vitro DNA. The site addition means that the replacement sequence to be inserted has part of the site, such as one or several, added compared to the corresponding sequence on the in vitro DNA. The sequence addition means that the replacement sequence to be inserted has part of the sequence, such as one or several, added compared to the corresponding sequence on the in vitro DNA. The sequence deletion means that the replacement sequence to be inserted has part of the sequence, such as one or several, missing compared to the corresponding sequence on the in vitro DNA.

[0186] The smaller the difference between the replacement sequence to be inserted and the corresponding homologous sequence on the in vitro DNA, the higher the efficiency. The inconsistent part of the replacement sequence to be inserted and the corresponding homologous sequence on the in vitro DNA is preferably located away from or not close to the two ends or sides of the replacement sequence to be inserted to improve the efficiency. Note that the homologous recombination protein, exonuclease and / or the frame structure or the vector containing the frame structure, or the edited nucleic acid are subjected to chromosome assembly to promote the occurrence of homologous recombination.

[0187] Six, impede transposon caused genomic changes, stable genome and CNVs thereon:

[0188] By the above technical means, the SINE such as Alu element or SVA element, LINE on the genome of various organisms, especially the human genome, and the CNV and the corresponding CNV end of each gene can be edited, such as sequence insertion, sequence replacement, site deletion, site addition, sequence addition, sequence deletion and site replacement, so that the corresponding transposon is inactivated, or the CNV no longer extends, thereby stabilizing the genome, reducing its reconfiguration and changes, and thereby making the phenotype of cells, tissues, organs, organisms no longer change.

[0189] The present application uses the reverse transcriptase and endonuclease ORF2p expressed by the long interspersed element LINE, which only cuts single-stranded DNA or RNA, and is safer than Cas protein which produces DSB by cutting DNA double strands. As the other part of ORF2p except EN has a steric hindrance effect on ORF2p-EN, as shown in Figure 6, the cleavage activity of ORF2p-EN and the binding property to the edited nucleic acid will be greatly reduced, therefore the present application expresses ORF2p-EN separately, and further adds a nuclear localization signal (NLS) to facilitate its application in eukaryotes. The other part of ORF2p except EN refers to the other part of the ORF2p amino acid sequence except 1-239.

[0190] The various frameworks used in the present application will first form a bulge structure in the shape of Ω with the nucleic acid object to be edited, which is generated by strand exchange if the nucleic acid object is double-stranded. Previous studies have shown that when the base stacking force of the cleavage site is reduced, ORF2p-EN can efficiently cut double-stranded DNA. The Ω-shaped bulge structure used in the present application can exactly reduce the base stacking force of the cleavage site while maintaining the normal structure of double-stranded DNA, and can also make the double-stranded DNA produce any bending angle and twist at the cleavage site, as shown in Figure 7. The principle is that ORF2p-EN basically only binds to the nucleic acid on the side of the cleavage site, while basically not contacting the bases therein. Therefore, ORF2p-EN can efficiently cut the cleavage site such as the target site or the site opposite to the strand under the guidance of the Ω-shaped bulge structure, as shown in Figure 4. In addition, the structure of ORF2p-EN shows that the length between its binding site and the nucleic acid is significantly smaller than the normal length of nucleic acid, so that the general nucleic acid, if not extremely twisted or has the Ω-shaped bulge structure as described in the present application, is difficult to bind or further cleave with ORF2p-EN, thus ensuring the safety of ORF2p-EN application and having the potential for future human and clinical applications.

[0191] In addition, in order to promote the binding of the cleavage structure to the frame structure, improve the gene editing efficiency, increase the targeting and safety (so that the cleavage structure is enriched in the action site), the ORF2p-EN, ORF2p, ORF2p-EN homologous protein can also be combined with a protein, polypeptide or domain thereof having DNA, RNA binding function, strand displacement function or other functions to form a fusion protein, the N terminal or C terminal of the ORF2p-EN, ORF2p, ORF2p-EN homologous protein is directly or indirectly connected with the N terminal or C terminal of the fused protein, which is called ORF2p-EN fusion protein, ORF2p fusion protein, ORF2p-EN homologous protein fusion protein or ORF2p-EN derived protein, ORF2p derived protein, ORF2p-EN homologous protein derived protein.

[0192] Fusion proteins of ORF2p-EN and ORF2p-C terminal sequence (CTS) and / or RNA binding protein (RBD), fusion proteins of ORF2p-EN and the rest of ORF2p except for EN, fusion proteins of ORF2p-EN and Rec / Rad51 family, fusion proteins of ORF2p-EN and RecA, fusion proteins of ORF2p-EN and Uvs X, fusion proteins of ORF2p-EN and RAD51, fusion proteins of ORF2p-EN and RAD52, fusion proteins of ORF2p-EN and RAD51AP1, fusion proteins of ORF2p-EN and RAD51AP2, fusion proteins of ORF2p-EN and RAD54L, fusion proteins of ORF2p-EN and Swi5, fusion proteins of ORF2p-EN and Sfr1, fusion proteins of ORF2p-EN and Swi5 and Sfr1, fusion proteins of ORF2p-EN and RAD56, fusion proteins of ORF2p-EN and RAD57, fusion proteins of ORF2p-EN and BARD1, fusion proteins of ORF2p-EN and RecA, fusion proteins of ORF2p-EN and Mre11, fusion proteins of ORF2p-EN and NBS, fusion proteins of ORF2p-EN and DMC1, fusion proteins of ORF2p-EN and SCR7, fusion proteins of ORF2p-EN and ORF2p, fusion proteins of ORF2p-EN and BRCA2, fusion proteins of ORF2p-EN and BRCA1, fusion proteins of ORF2p-EN and PALB2, fusion proteins of ORF2p-EN and RAD53, fusion proteins of ORF2p-EN and RAD54, fusion proteins of ORF2p-EN and RAD55, fusion proteins of ORF2p-EN and RPA, fusion proteins of ORF2p-EN and SSBP;fusion protein of ORF2p-EN and SRS2, a fusion protein of ORF2p-EN and PSF, a fusion protein of ORF2p-EN and PARP1, a fusion protein of ORF2p-EN and RBM6, a fusion protein of ORF2p-EN and RAD58, a fusion protein of ORF2p-EN and RAD59, a fusion protein of ORF2p-EN and RAD50, a fusion protein of ORF2p-EN and TID1 (RDH54), a fusion protein of ORF2p-EN and MSH3, a fusion protein of ORF2p-EN and Mlh1, a fusion protein of ORF2p-EN and XRS2, a fusion protein of ORF2p-EN and NBS1, a fusion protein of ORF2p-EN and PHF1, a fusion protein of ORF2p-EN and SMC1, a fusion protein of ORF2p-EN and DHX9, a fusion protein of ORF2p-EN and various RECQ helicases, a fusion protein of ORF2p-EN and RECQl, a fusion protein of ORF2p-EN and P53, a fusion protein of ORF2p-EN and Mus81, a fusion protein of ORF2p-EN and Swi2, a fusion protein of ORF2p-EN and Snf2, a fusion protein of ORF2p-EN and Sth1, a fusion protein of ORF2p-EN and ISWI, a fusion protein of ORF2p-EN and Ino80, a fusion protein of ORF2p-EN and Mi-2, a fusion protein of ORF2p-EN and CHD3, a fusion protein of ORF2p-EN and CHD4, a fusion protein of ORF2p-EN and SWI1, a fusion protein of ORF2p-EN and SWI3, a fusion protein of ORF2p-EN and SWI5, a fusion protein of ORF2p-EN and SWI6, a fusion protein of ORF2p-EN and other various SWI2 / SNF2 family proteins, a fusion protein of ORF2p-EN and Hop2, a fusion protein of ORF2p-EN and Mnd1, a fusion protein of ORF2p-EN and Bst polymerase, a fusion protein of ORF2p-EN and phi29 DNA polymerase, a fusion protein of ORF2p-EN and EquiPhi29 DNA polymerase, a fusion protein of ORF2p-EN and Bst X DNA polymerase, a fusion protein of ORF2p-EN and Sau DNA polymerase, a fusion protein of ORF2p-EN and Bsu DNA polymerase; these fusion proteins belong to the ORF2p-EN derived proteins.

[0193] It is worth noting that the fusion protein of ORF2p-EN and the rest of ORF2p except EN, and the indirect connection between ORF2p-EN and the rest of ORF2p except EN, that is, any amino acid sequence of any length can be inserted in the middle, so as to reduce the steric hindrance effect of the rest of ORF2p except EN on the cleavage of ORF2p-EN. This fusion protein can increase the cleavage ability of ORF2p-EN while retaining the affinity of ORF2p to the human body and the reverse transcription function, and at the same time cooperate with the retrotransposon sequence such as Alu sequence or partial Alu sequence added downstream of the sequence downstream of the target site. Reverse transcription + high-efficiency cleavage makes the nucleic acid cleavage function more efficient and flexible.

[0194] In addition, some polypeptide fragments of different lengths cut from ORF2p can also be understood as fusion proteins of ORF2p-EN. For example, ORF2p amino acid sequence 1-440 is a fusion protein of ORF2p-EN and ORF2p-tower, ORF2p amino acid sequence 1-347 is a fusion protein of ORF2p-EN and ORF2p-cryptic, and ORF2p amino acid sequence 1-239+270-274 is a fusion protein of ORF2p-EN and ORF2p-aa270-274. These fusion proteins of ORF2p-EN and other ORF2p parts can improve the nuclear localization of ORF2p to a certain extent and reduce the cytotoxicity of ORF2p-EN;

[0195] The ORF2p-EN (L1-EN) homologous proteins include AP-like ENs, APE1 and proteins with 50% or more of its nucleotide sequence or amino acid sequence, DNaseI and proteins with 50% or more of its nucleotide sequence or amino acid sequence, ExoIII and proteins with 50% or more of its nucleotide sequence or amino acid sequence, IP5P and proteins with 50% or more of its nucleotide sequence or amino acid sequence, TRAS1-EN and proteins with 50% or more of its nucleotide sequence or amino acid sequence.

[0196] The ORF2p-EN homologous proteins can form fusion proteins with the above-mentioned proteins involved in fusion in the manner of the above-mentioned ORF2p-EN fusion proteins, and also belong to ORF2p-EN homologous protein derived proteins.

[0197] In the above-mentioned fusion proteins, the Rec / Rad51 family such as RAD51, RecA, Uvs X, RAD52, ORF2p, BRCA2, BRCA1, PALB2, RAD53, RAD54, RAD55, RPA, RAD51AP1, RAD51AP2, RAD54L, Swi5-Sfr1 complex, RAD56, RAD57, BARD1, RecA, Mre11, NBS, DMC1, SCR7, SRS2, PSF, PARP1, RBM6, RAD58, RAD59, RAD50, TID1 (RDH54), MSH3, Mlh1, XRS2, NBS1, SSBP, HSSB1, SSAPs, PHF1, SMC1, DHX9, RECQ helicase such as RECQ1, P53, Mus81, SWI2 / SNF2 family protein, Hop2, Mnd1 can be replaced by a protein similar to the corresponding protein in nucleotide sequence or amino acid sequence by 50% or more, and such a protein also belongs to the ORF2p-EN derived protein, ORF2p derived protein, or ORF2p-EN homologous protein derived protein;

[0198] In the above-mentioned fusion proteins, the ORF2p-EN can be replaced by ORF2p or ORF2p-EN homologous protein, and the resulting fusion protein belongs to ORF2p derived protein or ORF2p-EN homologous protein derived protein, respectively;

[0199] The above-mentioned proteins involved in fusion with ORF2p-EN, ORF2p or ORF2p-EN homologous protein can be one or more of the same protein fused with ORF2p-EN, ORF2p or ORF2p-EN homologous protein; the above-mentioned proteins involved in fusion with ORF2p-EN, ORF2p or ORF2p-EN homologous protein can be one or more proteins fused with ORF2p-EN, ORF2p or ORF2p-EN homologous protein; the above-mentioned ORF2p-EN, ORF2p or ORF2p-EN homologous protein involved in fusion can also be one or more. Multiple ORF2p-EN, ORF2p or ORF2p-EN homologous proteins and multiple fusion proteins can be arranged in any order. These fusion proteins also belong to ORF2p-EN homologous protein derived protein, ORF2p derived protein or ORF2p-EN derived protein, respectively.

[0200] ORF2p also contains a protein similar to ORF2p in nucleotide sequence or amino acid sequence by 50% or more; ORF2p-EN also contains a protein similar to ORF2p-EN in nucleotide sequence or amino acid sequence by 50% or more.

[0201] Lig4, DNA-PK, XRCC6 can be specifically inhibited by sgRNA, ASO, siRNA or specific antibodies to promote the progress of DNA homologous recombination, thereby improving the efficiency of gene editing in the present application.

[0202] To increase efficiency, one or more nuclear localization sequences (NLS) can be further added to the amino acid sequence of the above-described cleavage structure (which can be added to the N-terminus or C-terminus of the cleavage structure, or located in the middle region of the cleavage structure, such as between two fusion proteins), as shown in FIGS. 8 and 9, to facilitate the nuclear entry of the corresponding vector system. NLS consists of 4-8 amino acids, containing Pro, Lys and Arg, such as the NLS in the T antigen of SV40 (Pro-Lys-Lys-Lys-Arg-Lys-Val); NLS can be divided into classic NLS and non-classic NLS, classic NLS consists of a group of lysine and arginine-rich amino acids, such as KRKR, KKKR, KKKRK. The structure of non-classic NLS is more diverse, which can be a proline or serine-rich sequence, or a sequence containing lysine, arginine or phenylalanine, or a lysine or other basic amino acid-rich sequence, or it can contain a short peptide sequence or domain.

[0203] To facilitate application, production, industrialization or commercialization, a protein tag can be further added to the C-terminus or N-terminus of the coding sequence of the protein / polypeptide of the cleavage structure or the cleavage structure with added NLS described in the present application, as shown in FIGS. 8 and 9, such as 6xHIS, Flag, GST, Myc, eGFP / eCFP / eYFP / mCherry eGFP, HA, SUMO, MBP, Strep-II Tag, Avi Tag, SNAP-Tag, etc., to facilitate the expression, detection, tracking and purification of the corresponding protein, etc. The corresponding protein can be further synthesized and purified by conventional methods and further applied to the present application.

[0204] It is worth noting that the cleavage structure with added NLS and / or tag still belongs to the cleavage structure, such as ORF2p-EN with added NLS, which still belongs to ORF2p-EN.

[0205] The cleavage structure added with NLS and connected with ORF2p, ORF2p-EN, ORF2p-EN homologous protein, ORF2p derived protein, ORF2p-EN derived protein and / or ORF2p-EN homologous protein derived protein can be directly or indirectly connected with DNA and / or RNA, such as oligonucleotide or longer nucleic acid or deoxy nucleic acid at the N terminal and / or C terminal, and can be connected or indirectly connected at the 5' terminal, 3' terminal or any position of DNA or RNA, as shown in Figure 10.

[0206] The DNA or RNA connected with the cleavage structure added with NLS and connected with ORF2p, ORF2p-EN, ORF2p-EN homologous protein, ORF2p derived protein, ORF2p-EN derived protein and / or ORF2p-EN homologous protein derived protein can be complementary to any sequence in the frame structure, which can be close to or on the sequence to be inserted upstream or downstream of the target site.

[0207] The DNA or RNA connected with the cleavage structure added with NLS and connected with ORF2p, ORF2p-EN, ORF2p-EN homologous protein, ORF2p derived protein, ORF2p-EN derived protein and / or ORF2p-EN homologous protein derived protein can be the primer for gene editing described in the application, which is directly or indirectly connected with the cleavage structure at the 5' terminal of the primer.

[0208] The DNA or RNA connected with the cleavage structure added with NLS and connected with ORF2p, ORF2p-EN, ORF2p-EN homologous protein, ORF2p derived protein, ORF2p-EN derived protein and / or ORF2p-EN homologous protein derived protein can be the frame structure for gene editing described in the application.

[0209] The protein directly or indirectly connected with the above-mentioned various DNA and / or RNA can also be a homologous recombination related protein.

[0210] The cleavage structure described in the application can exert nucleic acid editing function before or after being administered to the target system, with or without any nucleic acid (double-stranded DNA, single-stranded DNA, double-stranded RNA, single-stranded RNA or DNA / RNA hybrid chain) being combined.

[0211] The nucleic acid combined with the cleavage structure can be a frame structure or a nucleic acid to be cut or edited.

[0212] The sequence of the nucleic acid combined with the cleavage structure can be the sequence upstream of the target site (or its complement), the sequence upstream of the target site and its surrounding sequence (or its complement), the sequence downstream of the target site (or its complement), the sequence downstream of the target site and its surrounding sequence (or its complement), the sequence upstream of the target site and the sequence downstream of the target site (or its complement), the sequence upstream of the target site and the sequence to be inserted and the sequence downstream of the target site (or its complement), the sequence containing the sequence upstream of the target site and / or the sequence downstream of the target site and containing any other sequence (or its complement), the transposon sequence, the reverse transposon sequence or a partial sequence thereof connected with the frame structure, and any sequence.

[0213] The sequence to be inserted in the frame structure provided by the present application can be an exogenous sequence or an endogenous sequence, and the length of the sequence to be inserted can be 1 bp-10 bp or 11 bp-100000 bp. When inserted multiple times, the genomic insertion of a DNA sequence of any length can be achieved. The length of the sequence upstream of the target site can be 1 bp-100000 bp, and the length of the sequence downstream of the target site can be 1 bp-100000 bp. The efficiency of strand exchange or homologous recombination can be improved by extending or shortening the length of the sequence upstream of the target site and / or the sequence downstream of the target site.

[0214] The sequence to be inserted needs a certain length, and too low length or can increase the base stacking force at the target site, thereby reducing the efficiency of cleavage and editing of the nucleic acid. When the cleavage and editing efficiency is desired to be improved, the length of the sequence to be inserted can be considered to be increased, for example, designed as a sequence greater than 10 bp.

[0215] The frame structure in the present application can be combined with the cleavage structure to form an RNP before being introduced into the target system, and the RNP is introduced into the target system to play a role. This method is performed in the presence or absence of ATP according to conventional methods.

[0216] The framework structure can be bound to a homologous recombination associated protein such as a Rec / Rad51 family such as RAD51, RecA, Uvs X, RAD52, ORF2p, BRCA2, BRCA1, PALB2, RAD53, RAD54, RAD55, RPA, RAD51AP1, RAD51AP2, RAD54L, Swi5-Sfr1 complex, RAD56, RAD57, BARD1, RecA, Mre11, NBS, DMC1, SCR7, SRS2, PSF, PARP1, RBM6, RAD58, RAD59, RAD50, TID1 (RDH54), MSH3, Mlh1, XRS2, NBS1, SSBP, HSSB1, SSAPs, PHF1, SMC1, DHX9, RECQ helicases such as RECQ1, P53, Mus81, SWI2 / SNF2 family proteins, Hop2, or Mnd1 prior to being switched into the editing system (target system) to promote the occurrence of strand exchange or homologous recombination and to protect the framework structure from degradation to some extent, as shown in Figure 11.

[0217] The introduction target (target system to be edited) can also be caused to express, increase, and / or inhibit expression of a homologous recombination associated protein such as a Rec / Rad51 family such as RAD51, RecA, Uvs X, RAD52, RAD51AP1, RAD51AP2, RAD54L, Swi5-Sfr1 complex, BRCA1, BRCA2, PALB2, RAD53, RAD54, RAD55, RAD56, RAD57, BARD1, RecA, RPA, Mre11, NBS, DMC1, SCR7, SRS2, PSF, PARP1, RBM6, RAD58, RAD59, RAD50, TID1 (RDH54), MSH3, Mlh1, XRS2, NBS1, SSBP, HSSB1, SSAPs, PHF1, SMC1, DHX9, RECQ helicases such as RECQ1, P53, Mus81, SWI2 / SNF2 family proteins, Hop2, or Mnd1 before, at the same time as, and / or after the introduction of the vector system into a solution, cell, tissue, organ, and / or organism.

[0218] A compound that increases the efficiency of homologous recombination, RS-1, SCR7, NU7441, Vanillin, Chlorpromazine, ResVeratrol, Benomyl, Inositol, can be added to the target system to be edited.

[0219] The framework structure can be combined with a related protein having strand displacement function, such as a DNA polymerase having strand displacement function, such as Bst polymerase, phi29 DNA polymerase, and mutants thereof, such as EquiPhi29 DNA polymerase, Bst X DNA polymerase, Sau DNA polymerase, Bsu DNA polymerase, before being introduced into the system to be edited, to promote the occurrence of strand exchange and to protect the framework structure to a certain extent from degradation.

[0220] The DNA framework or DNA / RNA hybrid framework described in the present application can be subjected to chromotin assembly before being introduced into the target system, before or after being combined with the cleavage structure: the DNA framework or DNA / RNA hybrid framework can be incubated with core histones H2A, H2B, H3 and / or H4, or histones H1, H2A, H2B, H3 and / or H4 according to conventional methods and combined with each other; or can be further incubated with ACF and / or NAP-1 in the presence of ATP according to conventional methods and combined therewith.

[0221] All proteins mentioned in the present application, including ORF2p, ORF2p-EN, Rec / Rad51 family such as RAD51, RecA, Uvs X, RAD52, RAD51AP1, RAD51AP2, RAD54L, Swi5-Sfr1 complex, BRCA1, BRCA2, PALB2, RAD53, RAD54, RAD55, RAD56, RAD57, BARD1, RPA, RecA, Mre11, NBS, DMC1, SCR7, SRS2, PSF, PARP1, RBM6, RAD58, RAD59, RAD50, TID1 (RDH54), MSH3, Mlh1, XRS2, NBS1, SSBP, HSSB1, SSAPs, PHF1, SMC1, DHX9, RECQ helicase such as RECQ1, P53, Mus81, SWI2 / SNF2 family proteins, Hop2, Mnd1, Dmc1, all include the homologous proteins of the corresponding protein in each species, including all eukaryotes and prokaryotes, including humans, mammals such as mice, birds, fish, reptiles, invertebrates, fungi, bacteria and the like, for example, RAD51 in humans hRAD51, and the corresponding RecA in bacteria; for example, human hRAD51 or mouse mRAD51; or for example, Saccharomyces cerevisiae strand exchange protein 1 (Sep1; also referred to as Xrn1, Kem1, Rar5, or Stpp) in Saccharomyces cerevisiae; or fungal RecA homologs (S. cerevisiae RAD51, RAD55, RAD57, DMC1; Schizosaccharomyces pombe rad51; Neurospora crassa mei3).

[0222] The RNA framework or RNA-derived framework of the framework structure of the present application can be linear RNA, mRNA, circular RNA. The RNA framework can be obtained by in vivo or in vitro prokaryotic system transcription, in vivo or in vitro eukaryotic system transcription or chemical synthesis; wherein the eukaryotic transcription is transcription by eukaryotic RNA polymerase I, eukaryotic RNA polymerase II or eukaryotic RNA polymerase III. For example, the RNA framework can be produced by in vitro T7 promoter or SP6 promoter transcription, or chemical synthesis.

[0223] The DNA framework or DNA derived framework of the framework structure of the present application can be obtained by prokaryotic DNA polymerization, eukaryotic DNA polymerization, prokaryotic DNA amplification (similar to plasmid), eukaryotic DNA amplification (similar to plasmid), PCR amplification, in vitro DNA polymerase synthesis or chemical synthesis.

[0224] The DNA / RNA hybrid framework or DNA / RNA hybrid derived framework of the framework structure of the present application can be generated after denaturation and annealing after obtaining the DNA framework or RNA framework, or generated by reverse transcription of the RNA framework, or generated by transcription of the DNA framework.

[0225] The spliced structure RNA of the present application can be linear RNA, circular RNA, mRNA.

[0226] The spliced structure RNA or DNA of the present application, the RNA form is obtained by in vivo or in vitro prokaryotic system transcription, in vivo or in vitro eukaryotic system transcription or chemical synthesis; the DNA can be subjected to in vivo or in vitro prokaryotic system transcription, in vivo or in vitro eukaryotic system transcription. The prokaryotic transcription is transcription by the RNA polymerase of prokaryotes; the eukaryotic transcription is transcription by the RNA polymerase I of eukaryotes, the RNA polymerase II of eukaryotes or the RNA polymerase III of eukaryotes.

[0227] The present application can, on the basis of inserting the desired sequence into the genomic target, carry out more accurate genomic fragment sequence deletion, fragment sequence replacement and individual site replacement, etc. through the mechanism such as homologous recombination or genome repair of the recipient editing system such as prokaryotes or eukaryotes itself. At the same time, the present application can continue to design a vector for insertion through the new site formed after the insertion of the to-be-inserted sequence, and the progressive insertion makes the length of the inserted sequence into the genome theoretically unlimited in length, and can complete various and various forms of sequence insertion, deletion, replacement and site replacement, etc. for the purpose of gene editing, and the usage is flexible. In addition, a single or multiple CNV and its end on the genome can also be edited by the present application to make it stable, unchanged, lengthened, shortened or change its expression sequence, so as to achieve the purpose of changing or stabilizing the gene expression and the state of cells, tissues, organs or organisms.

[0228] The nucleic acid editing system or vector system provided by the present application is used for preventing and / or treating cancer, gene-related diseases or neurodegenerative diseases.

[0229] Cancer is glioma, breast cancer, cervical cancer, lung cancer, gastric cancer, colorectal cancer, duodenal cancer, leukemia, prostate cancer, endometrial cancer, thyroid cancer, lymphoma, pancreatic cancer, liver cancer, melanoma, skin cancer, pituitary tumor, germ cell tumor, meningioma, meningeal cancer, glioblastoma, various astrocytoma, various oligodendroglioma, astrocytic oligodendroglioma, various ependymoma, choroid plexus papilloma, choroid plexus carcinoma, chordoma, various ganglioneuroma, olfactory neuroblastoma, sympathetic nervous system neuroblastoma, pineal cell tumor, pinealoblastoma, medulloblastoma, retinoblastoma, trigeminal schwannoma, facial acoustic neuroma, jugular bulb tumor, hemangioblastoma, craniopharyngioma, or granular cell tumor.

[0230] The present application can be used to edit or disrupt the following genes or related genes: 2B4, 4-1BB, 4-1BBL, A33, Adenosine A2a receptor, Akt, Androgen receptor, Aurora A, Aurora B, B7-H3, B7-H4, Bcl-2, Bcr-Ab1, BRAF, BTK, BTLA, BTN2A1, CAIX, CCR4, CD155, CD160, CD19, CD20, CD200, CD200R, CD25, CD27, CD28, CD30, CD33, CD36, CD38, CD40, CD40L, CD47, CD48, CD52, CD70, CD80, CD86, CD96, CDK4, CEA, CEACAM1, ChK1, ChK2, c-KIT, c-Met / HGFR, COX2, CSF-1R, CTLA-4, DDR2, DNAM-1, DR5, EGFR, EpCAM, EPHA3, ERK1, ERK2 / p38 MAPK, FAP, FGFR1, FGFR2, FGFR3, FGFR4, Flt-3, Gal-9, GITR, GITRL, HDAC1, HDAC2, HER2, HER3, HER4 / ERBB4, HGF, HHLA2, HVEM, ICOS, ICOS ligand, IDO, IGF1R, KRAS, LAG-3, LIGHT, MDM2, MEK1, MEK2, mTOR, Mucin 1, NRAS, NTRK1, NTRK2, NTRK3, OX40, OX40L, p53n, PARP1, PD1, PDGFR-alpha, PDGFR-beta, PD-L1, PD-L2, PI3K alpha, PI3K beta, PI3K gamma, PI3K delta, PSMA, PTEN, RAF-1, RANKL, RET, SIRP alpha, SLAMF7, SYK, TDO, TIGIT, TIM-3, TMIGD2, TRAIL TRAILR1, VEGF, VEGFR-1, VEGFR-2, VEGFR-3, VISTA. Editing or disruption of the above genes or related genes can be used to treat tumors.

[0231] Treatment of ovarian cancer is achieved by editing or disrupting genes or related genes expressing Ang-1, Ang-2, CA-125, c-KIT, c-Met / HGFR, CTLA-4, DLL4, EGFR, Flt-3, HER2, HER3, HER4 / ERBB4, HSP90, IGF1R, IL-6, MEK1, MEK2, mTOR, NKG2A, PARP, PARP-1, PD1, PDGFR-a, PDGFR-b, PD-L1, PD-L2, PI3K a, PI3K b, PI3K g, PI3K d, Src, Tie-2, TLR8, TNF-a, VEGF-A, VEGFR-1, VEGFR-2, or VEGFR-3.

[0232] Treatment of lung cancer is achieved by editing or disrupting genes or related genes expressing ALK, BRAF, CDK4, CTLA-4, EGFR, HER2, HER3, K-Ras, MEK1, MEK2, MET, PARP1, PD-1, PDGFR-a, PDGFR-b, PD-L1, RET, VEGF-A, VEGFR-1, VEGFR-2, VEGFR-3.

[0233] Treatment of breast cancer is achieved by editing or disrupting genes or related genes expressing Akt, B7-H3, CDK4, CDK6, COX-2, CTLA4, EGFR, ERBB3, HER2, LAG-3, mTOR, NF-Kb, OX40, p53, PARP-1, PD1, PD-L1, K-Ras, TRAIL, VEGF, VEGFR-1, VEGFR-2, VEGFR-3.

[0234] Treatment of bladder cancer is achieved by editing or disrupting genes or related genes expressing Adenosine A2a receptor, Androgen receptor, B7-H3, CD24, CD27, CDK4, CTLA-4, EGFR, EphB4, FGFR3, GITR, HER2, HIF-1a, IGF1R, mTOR, PD1, PD-L1, STAT3, VEGF, VEGF-C, VEGFR-2, VEGFR-3.

[0235] Treatment of colorectal cancer is achieved by editing or disrupting genes or related genes expressing BRAF, CD27, CTLA-4, EGFR, mTOR, PD1, PD-L1, Src, VEGF-A, VEGFR2.

[0236] Treatment of renal cancer is achieved by gene or related gene editing or disruption of genes expressing Adenosine A2a receptor, B7-H3, BRAF, CAIX, CD105 / endoglin, CD27, CD70, c-KIT, c-Met / HGFR, CTLA-4, FGF-2, FGFR1, FGFR2, FGFR3, FGFR4, Flt-3, GITR, JAK1, JAK2, JAK3, LAG-3, mTOR, PD1, PDGFR-a, PDGFR-b, PD-L1, PI3K delta, RAF-1, STAT3, Tie-2, TIM-3, VEGF, VEGFR-1, VEGFR-2, VEGFR-3.

[0237] Treatment of leukemia is achieved by gene or related gene editing or disruption of genes expressing BCL-2, BCR-ABL, BTK, CD123, CD19, CD20, CD25, CD30, CD33, CD37, CD47, CD52, c-Kit, CSF2, CTLA-4, NKG2A, P110 delta, PD1, PD-L1, Src.

[0238] Treatment of liver cancer is achieved by gene or related gene editing or disruption of genes expressing Annexin A3, CD105 / endoglin, CDK9, CTLA-4, EGFR, Glypican-3, LAG-3, MET, PD1, PD-L1, RET, VEGFR-2.

[0239] Treatment of lymphoma is achieved by gene or related gene editing or disruption of genes expressing 4-1BB, Bcl-2, BTK, CD16, CD19, CD20, CD22, CD25, CD27, CD30, CD37, CD40, CD52, CD79b, CD80, CSF1R, CTLA-4, CXCR4, HDAC1, HDAC2, HDAC3, HDAC4, HDAC5, HDAC7, HDAC8, HDAC9, IL-13, IL-2R alpha, JAK1, JAK2, mTOR, NF-Kb, PD1, PI3K-alpha, PI3K gamma, PI3K delta, TRAIL, VEGFR-2.

[0240] Treatment of skin cancer is achieved by gene or related gene editing or disruption of genes expressing 4-1BB, Adenosine A2A receptor, B7-H3, BRAF, CD27, CSF1R, CTLA-4, EGFR, GITR, LAG-3, MEK1, MEK2, PD1, PD-L1, TIM-3, VEGF.

[0241] Treatment of pancreatic cancer is achieved by gene or related gene editing or disruption of genes expressing B7-H3, CCR4, CD109, CD40, C-Met, CSF1R, CTLA-4, EGFR, FAK, FLT-3, KIT, KRAS, MEK1, MEK2, MMP-1, MMP-10, MMP-13, MMP-2, MMP-7, MMP-9, mTOR, Myc, NF-KB, PD1, PD-L1, PIM1, PIM3, TGF, Y-secretase.

[0242] Treatment of prostate cancer is achieved by gene or related gene editing or disruption of genes expressing Akt, B7-H3, CD27, CTLA-4, HSP27, MET, mtor, PARP1, PARP2, PD1, PSMA, STEAP-1, Toll-like receptor 3, TROP-2, VEGF-A, VEGFR-1, VEGFR-2, VEGFR-3.

[0243] Treatment of thyroid cancer is achieved by gene or related gene editing or disruption of genes expressing Adenosine-A2A receptor, B7-H3, B-raf, BTK, CD27, C-KIT, C-raf, CSF1R, CTLA-4, CXCR2, EGFR, FGFR1, FGFR2, FGFR3, FGFR4, FLT-3, GITR, HER3, HGF, IDO, JAK1, JAK2, JAK3, KIT, MEK1, MEK2, MET, OX40, PD1, PDGFRA, PDGFRB, RET, STAT3, VEGFR-1, VEGFR-2, VEGFR-3.

[0244] Meanwhile, by the present application, genes (including coding sequences and non-coding sequences) or RNAs on the biological genome related to glycolysis can be taken as targets, and by increasing or inhibiting the expression of related proteins or RNAs, cancer or other metabolic related diseases can be treated. For example, by the present application, genes or related sequences, mRNA or regulatory RNA, etc. of glucose transporter (GLUT), sodium-glucose linked transporter (SGLT) (such as SGLT1, SGLT2), hexokinase (HK) (HK1-4), glucokinase (GK), phosphofructokinase (PFK), phosphofructokinase-2 / fructose-2, 6-bisphosphatase 3 (PFKFB3), pyruvate kinase (PK) (PKL, PKR, PKM1 and PKM2), 3-phosphoglyceraldehyde dehydrogenase (GAPDH), lactate dehydrogenase (LDH), etc. are edited or destroyed, or genes or related sequences, mRNA or regulatory RNA, etc. of antagonistic proteins of the above proteins are edited or destroyed, so as to treat cancer, metabolic related diseases or other diseases.

[0245] Huntington's disease, fragile X syndrome, phenylketonuria, Duchenne's progressive muscular dystrophy, Duchenne's muscular dystrophy, mitochondrial encephalomyopathy, mucopolysaccharidosis type I, mucopolysaccharidosis type II, mucopolysaccharidosis type IIIA, mucopolysaccharidosis type IIIB, mucopolysaccharidosis type IIIA, mucopolysaccharidosis type IIIC, mucopolysaccharidosis type IIID, mucopolysaccharidosis type IVA, mucopolysaccharidosis type IVB, mucopolysaccharidosis type VI, mucopolysaccharidosis type VII, mucopolysaccharidosis type IX, spinal muscular atrophy, Parkinson-plus syndrome, albinism, color blindness, achondroplasia, alkaptonuria, congenital deaf mutism, thalassemia, sickle cell anemia, hemophilia, epilepsy associated with genetic alterations, myoclonus, dystonia, stroke and schizophrenia, antivitamin D rickets, familial colonic polyposis, 21-hydroxylase deficiency, arginase deficiency, Alport syndrome, Angelman syndrome, Tay-Sachs disease, atypical hemolytic uremic syndrome, autoimmune encephalitis, autoimmune hypophysitis, autoimmune insulin receptor disease, beta-ketothiolase deficiency, biotinidase deficiency, cardiac ion channelopathy, primary carnitine deficiency, Castleman disease, Charcot-Marie-Tooth disease, citrullinemia, congenital adrenal hypoplasia, congenital hyperinsulinemic hypoglycemia, congenital myasthenic syndrome, non-dystrophic myotonia syndrome, congenital scoliosis, coronary artery ectasia, congenital pure red cell aplasia, Erdheim-Chester disease, Fabry disease, familial Mediterranean fever, Fanconi anemia, galactosemia, Gaucher disease, generalized myasthenia gravis, Gitelman syndrome, glutaric aciduria type I, glycogen storage disease type I, glycogen storage disease type II, hemophilia, hepatolenticular degeneration, hereditary angioedema, hereditary epidermolysis bullosa, hereditary fructose intolerance, hereditary hypomagnesemia, hereditary multi-infarct dementia, hereditary spastic paraplegia, holocarboxylase synthetase deficiency, homocystinuria, homozygous familial hypercholesterolemia, HHH syndrome, hyperphenylalaninemia, hypophosphatasia, hypophosphatemic rickets, idiopathic cardiomyopathy, idiopathic hypogonadotropic hypogonadism, idiopathic pulmonary arterial hypertension, idiopathic pulmonary fibrosis, IgG4-related disease, inborn error of bile acid synthesis, isovaleric acidemia, Kallmann syndrome, Langerhans cell histiocytosis, Leigh syndrome, Leber hereditary optic neuropathy, long-chain 3-hydroxyacyl-CoA dehydrogenase deficiency, lymphangioleiomyomatosis, lysinuric protein intolerance, lysosomal acid lipase deficiency, maple syrup urine disease, Marfan syndrome, McCune-Albright syndrome, medium-chain acyl-CoA dehydrogenase deficiency, methylmalonic acidemia, multifocal motor neuropathy, multiple acyl-CoA dehydrogenase deficiency, multiple sclerosis, myotonic dystrophy,N-acetylglutamate synthetase deficiency, neonatal diabetes, neuromyelitis optica, Niemann-Pick disease, non-syndromic deafness, Noonan syndrome, ornithine transcarbamylase deficiency, osteogenesis imperfecta, young-onset Parkinson's disease, early-onset Parkinson's disease, paroxysmal nocturnal hemoglobinuria, Peutz-Jeghers syndrome, POEMS syndrome, porphyria, Prader-Willi syndrome, primary combined immunodeficiency, primary hereditary muscular atonia, primary light chain amyloidosis, progressive familial intrahepatic cholestasis, progressive muscular dystrophy, propionic acidemia, pulmonary alveolar proteinosis, pulmonary cystic fibrosis, retinitis pigmentosa, severe congenital neutropenia, severe myoclonic epilepsy in infancy, Dravet syndrome, Silver-Russell syndrome, sitosterolemia, spinal and bulbar muscular atrophy, spinal muscular atrophy, spinocerebellar ataxia, systemic sclerosis, tetrahydrobiopterin deficiency, tuberous sclerosis, primary tyrosinemia, very long chain acyl-CoA dehydrogenase deficiency, Williams syndrome, X-linked agammaglobulinemia, X-linked adrenoleukodystrophy, X-linked lymphoproliferative syndrome, arteriosclerotic cerebral small vessel disease, cerebral amyloid angiopathy, cerebral arteriopathy associated with subcortical infarcts and leukoaraiosis, cerebral arteriopathy associated with subcortical infarcts and leukoaraiosis, cysteine string protein-associated arteriopathy with stroke and leukoaraiosis, pyridoxine-dependent epilepsy, serotonin metabolism AADC enzyme deficiency, AADC deficiency, or hereditary nephritis.

[0246] The neurodegenerative disease is Parkinson's disease, Alzheimer's disease, Huntington's disease, amyotrophic lateral sclerosis, spinocerebellar ataxia, multiple system atrophy, primary lateral sclerosis, Pick's disease, frontotemporal dementia, Lewy body dementia, or progressive supranuclear palsy.

[0247] In the present application, a certain sequence or site, such as a sequence to be inserted and a sequence upstream of a target site, a sequence downstream of a target site, or a sequence flanking a target site to be edited on a nucleic acid, such as a genome, is defined as DNA or RNA in the 5'→3' direction, upstream of the 5' end of the certain sequence or site, and downstream of the 3' end of the certain sequence or site. The upstream sequence is a sequence located before the 5' end of the certain sequence or site, and the downstream sequence is a sequence located after the 3' end of the certain sequence or site.

[0248] The target site of the present application can be one or more; two or more target sites can be located at different sites on two complementary strands of the nucleic acid to be edited (double-stranded DNA, double-stranded RNA, DNA / RNA hybrid chain, genomic double strand) respectively, or can be located at the same site on two complementary strands of the nucleic acid to be edited (double-stranded DNA, double-stranded RNA, DNA / RNA hybrid chain, genomic double strand), as shown in Figure 12, or can be located at different sites on the same strand of the nucleic acid to be edited, such as double-stranded DNA, single-stranded DNA, double-stranded RNA, single-stranded RNA, DNA / RNA hybrid chain, genomic double strand, or a mixture of the above.

[0249] When two target sites are located at the same site on two complementary strands of the nucleic acid to be edited, when one target site is edited by the present application, the other strand will also be inserted with a certain probability of complementary sequence of the inserted sequence on the opposite strand due to the repair mechanism of the genome; that is to say, when two target sites are located at the same site on two complementary strands of the nucleic acid to be edited, editing of any one of the target sites can result in editing of both target sites.

[0250] The basic structure of the framework structure for gene editing provided by the present application is shown in Figure 13, which includes target site upstream sequence, inserted sequence, and target site downstream sequence in the 5'→3' direction.

[0251] Wherein "connection" means direct connection, and "indirect connection" means insertion of other sequences. The inserted sequence in "indirect connection" can be any sequence (deoxyribonucleic acid sequence, ribonucleic acid sequence and / or amino acid sequence), which is related or unrelated to the framework structure provided by the present application. In the present application, "intermediate" means between two sequences, and the two sequences are still complete.

[0252] It should be noted that the insertion site described in the present application is the target site.

[0253] It should be noted that the sequence synthesized according to the inserted sequence on the framework structure of the present application is the complementary sequence of the inserted sequence, but for ease of description, it is also sometimes referred to as the inserted sequence, as appropriate.

[0254] The single-stranded DNA or single-stranded RNA to be edited of the present application can have its target site set on a hypothetical single strand completely complementary thereto, as shown in Figure 14.

[0255] The double-stranded DNA framework, the double-stranded RNA framework or the DNA / RNA hybrid framework of the present application, depending on the nucleic acid target site to be cleaved or edited (the sequence on both sides), the sequence upstream of the target site and the sequence downstream of the target site can be located on the corresponding single strand in the two strands of the double-stranded DNA framework, the double-stranded RNA framework or the DNA / RNA hybrid framework.

[0256] Materials

[0257] 1. pBudORF2-CH plasmid was purchased from Addgene, plasmid number: 51289.

[0258] 2. CD293 medium was purchased from Thermofisher, product number: 11913019.

[0259] 3. PEI transfection reagent was purchased from Serochem, product number: Prime-AQ100-100ML.

[0260] 4. SMS 293-SUPI was purchased from Sino Biological Inc., product number: M293-SUPI-100.

[0261] 5. Potassium acetate was purchased from Sigma-Aldrich, product number: P1190.

[0262] 6. Tris-HCl (pH 7.5) was purchased from Shanghai Shangbao Biological Technology Co., Ltd., product number: T16588.

[0263] 7. Glycerol was purchased from Sigma-Aldrich, product number: G5516.

[0264] 8. Triton X-100 was purchased from Sigma-Aldrich, product number: T8787.

[0265] 9. PMSF protease inhibitor was purchased from Thermofisher, product number: 36978.

[0266] 10. Ni affinity chromatography column (HISTRAP HP) was purchased from Cytiva.

[0267] 11. Imidazole was purchased from Sigma-Aldrich, product number: I5513

[0268] 11. Rabbit anti-his was purchased from Sigma-Aldrich, product number: SAB1306082.

[0269] 12. BSA was purchased from Sigma-Aldrich, product number: A1933.

[0270] 13. Anti-Rabbit IgG (H+L)-Alkaline Phosphatase Goat Anti-Body was purchased from Sigma-Aldrich, product number: A3687.

[0271] 14. PolyFast Transfection Reagent was purchased from MCE, product number: HY-K1014.

[0272] 15. MEGAscript TM T7 Transcription Kit was purchased from Thermofisher, product number: AM1333.

[0273] 16. Opti-MEM TM I Medium was purchased from Thermofisher, product number: A4124802.

[0274] 17. RNase Inhibitor was purchased from Thermofisher, product number: AM2694.

[0275] 18. KOD One TM PCR Master Mix was purchased from Toyobo (Shanghai) Biotech Co., Ltd., product number: KMM-201S.

[0276] 19. Hieff Plus One Step Cloning Kit was purchased from YESEN BIOTECH (SHANGHAI) CO., LTD., product number: 10911ES20.

[0277] 20. Complete Medium was made by 90% DMEM Medium + 10% Fetal Bovine Serum, wherein DMEM Medium was purchased from Thermofisher, product number: 11965092, Fetal Bovine Serum was purchased from Thermofisher, product number: 10100147.

[0278] 21. Blood / Cell / Tissue Genomic DNA Extraction Kit was purchased from TIANGEN BIOTECH (BEIJING) CO., LTD., product catalog number: DP304.

[0279] 22. MEGAscript TMSP6 Transcription Kit was purchased from Thermofisher, product number: AM1330.

[0280] 23. SuperReal PreMix Plus (SYBR Green) was purchased from Tiangen Biochemical Technology (Beijing) Co., Ltd., product catalog number: FP205.

[0281] 24. The chemical synthesis of primers and sequences was completed by Platsen Biotech (Shanghai) Co., Ltd., Aberson (Jiangsu) Biotechnology Co., Ltd., and General Bio (Anhui) Co., Ltd., respectively.

[0282] Example 1 Construction of a vector for expressing ORF2p with a nuclear localization signal for eukaryotes, as shown in Figure 15:

[0283] First, the SV40 poly(A) signal sequence for terminating transcription was obtained, as shown in Seq ID No. 1:

[0284] Add a translation termination codon "TGA" in front of the Seq ID No. 1 sequence to obtain Seq ID No. 2:

[0285] TGAtaacttgtttattgcagcttataatggttacaaataaagcaatagcatcacaaatttcacaaataaagcatttttt tcactgcattctagttgtggtttgtccaaactcatcaatgtatctta, wherein the underlined part is the added translation termination codon, and the sequence is named termination codon + SV40 poly(A) signal sequence, which is obtained by chemical synthesis.

[0286] Add a nuclear localization signal (NLS) in front of the Seq ID No. 2 sequence to obtain Seq ID No. 3:

[0287] wherein the underlined part is the added translation termination codon, and the wavy line part is the nuclear localization signal, and the sequence is named nuclear localization signal + termination codon + SV40 poly(A) signal sequence, which is obtained by chemical synthesis.

[0288] The sequence shown in Seq ID No. 3 is chemically synthesized and constructed into pBudORF2-CH vector by reverse extension and homologous ligation, so that the sequence is directly connected behind the sequence of ORF2p-EN in the vector, and the plasmid is named as pBudORF2-EN-NLS-CH.

[0289] The specific steps are as follows:

[0290] 1. Design primers for amplifying the sequence shown in Seq ID No. 3, wherein the sequence of the forward primer is shown in Seq ID No. 4: 5'-gctggagctgcggatccccaagaaaaagcgca-3', and the sequence of the reverse primer is shown in Seq ID No. 5: 5'-tctgggtcaggttctttaagatacattgatga-3', and the sequence shown in Seq ID No. 3 is subjected to PCR amplification, and the reaction system is shown in Table 1.

[0291] Table 1 Reaction system

[0292] The PCR amplification conditions are as follows: 94℃ 2min; (98℃ 10sec, 60℃ 10sec, 68℃ 2sec) for 40 cycles; 68℃ 5min.

[0293] The amplified product is obtained by gel recovery and purification by a conventional method, and the amplified product is added with sequences homologous to the pBudORF2-CH vector on both sides of the sequence shown in Seq ID No. 3.

[0294] 2. Design PCR primers for amplifying the pBudORF2-CH vector, wherein the sequence of the forward primer is shown in Seq ID No. 6: 5'-tcatcaatgtatcttaaagaacctgacccaga-3', and the sequence of the reverse primer is shown in Seq ID No. 7: 5'-tgcgctttttcttggggatccgcagctccagc-3', and the pBudORF2-CH vector is subjected to PCR amplification, and the reaction system is shown in Table 2:

[0295] Table 2 Reaction system

[0296] The amplification conditions are as follows: 94℃ 2min; (98℃ 10sec, 60℃ 10sec, 68℃ 6sec) for 40 cycles; 68℃ 5min.

[0297] The pBudORF2-CH plasmid vector was obtained by gel recovery and purification using a conventional method, and the plasmid vector has sequences homologous to the synthetic sequence at both ends.

[0298] 3. The amplified product and the amplified pBudORF2-CH vector were connected using a one-step rapid cloning kit, and the specific steps were performed according to the kit instructions, and the reaction system is shown in Table 3:

[0299] Table 3 Connection reaction system

[0300] 4. The recombinant product was transformed into E. coli competent cells (DH5a), and after transformation, the bacteria were plated and sequenced. After sequencing, the plasmid was extracted to obtain the plasmid pBudORF2-EN-NLS-CH.

[0301] Continue to add NLS to the front of ORF2p-EN. Seq ID No. 8 is the translation initiation codon and NLS that needs to be added to the front, and Seq ID No. 8 is as follows: The bold part is the NLS sequence, and the italic part is the translation initiation codon, which is named initiation codon+NLS, and is obtained by chemical synthesis.

[0302] The sequence shown in Seq ID No. 8 is chemically synthesized and constructed into the pBudORF2-EN-NLS-CH vector by reverse expansion and homologous ligation, so that the sequence is directly connected in front of the ORF2p-EN sequence in the vector, and the plasmid is named pBud-NLS-ORF2-EN-NLS-CH.

[0303] The specific steps are as follows:

[0304] 1. Design primers for amplifying the sequence shown in Seq ID No. 8, wherein the forward primer sequence is shown in Seq ID No. 9: The reverse primer sequence is shown in Seq ID No. 10: 5'-tggtgctgccggtcatcactttcctcttcttt-3', and the sequence shown in Seq ID No. 8 is subjected to PCR amplification, and the reaction system is shown in Table 4:

[0305] Table 4 Reaction system

[0306] The amplification conditions are: 94℃ 2min; (98℃ 10sec, 60℃ 10sec, 68℃ 2sec) for 40 cycles; 68℃ 5min.

[0307] The amplified product was obtained by gel recovery and purification using a conventional method, and the amplified product was added with sequences homologous to the pBudORF2-EN-NLS-CH vector on both sides of the synthetic sequence.

[0308] 2. PCR primers for amplifying the pBudORF2-EN-NLS-CH vector were designed, wherein the forward primer is shown in Seq ID No. 11: 5'-aaagaagaggaaagtgatgaccggcagcacca-3', and the reverse primer sequence is shown in Seq ID No. 12: 5'-tggggctagccatggtggtggcggctatgatg-3', and the pBudORF2-EN-NLS-CH vector was subjected to PCR amplification, and the reaction system is shown in Table 5:

[0309] Table 5 Reaction system

[0310] The amplification conditions were: 94℃ 2min; (98℃ 10sec, 60℃ 10sec, 68℃ 6sec) for 40 cycles; 68℃ 5min.

[0311] The pBudORF2-EN-NLS-CH plasmid vector was obtained by gel recovery and purification using a conventional method, and the plasmid vector had sequences homologous to the synthetic sequence on both ends.

[0312] 3. The amplified product and the amplified pBudORF2-EN-NLS-CH vector were connected using a one-step fast cloning kit, and the specific steps were performed according to the kit instructions, and the reaction system is shown in Table 6:

[0313] Table 6 Connection reaction system

[0314] 4. The recombinant product was transformed into E. coli competent cells (DH5a), and after transformation, the bacteria were picked and sequenced, and after the sequence was correct, the plasmid was extracted, and the plasmid pBud-NLS-ORF2-EN-NLS-CH was obtained. The plasmid expresses ORF2p-EN with NLS connected on both ends.

[0315] Example 2. Vector construction of ORF2p with 6×His tag, as shown in Figure 16:

[0316] First, the SV40 poly(A) signal sequence for terminating transcription was obtained, as shown in Seq ID No. 1:

[0317] The translation termination codon "TGA" was added in front of the Seq ID No. 1 sequence to obtain Seq ID No. 2:

[0318] TGA taacttgtttattgcagcttataatggttacaaataaagcaatagcatcacaatttcacaaataaagcattttttt cactgcattctagttgtggtttgtccaaactcatcaatgtatctta, wherein the underlined portion is an added translation termination codon, and the sequence is named as the termination codon + SV40 poly(A) signal sequence.

[0319] A 6xhis tag sequence is added in front of the sequence of Seq ID No. 2 to obtain Seq ID No. 13:

[0320] wherein the underlined portion is an added translation termination codon, and the sequence is named as the termination codon + SV40 poly(A) signal sequence, which is obtained by chemical synthesis.

[0321] The sequence shown in Seq ID No. 13 is chemically synthesized and constructed into the pBudORF2-CH vector by reverse expansion and homologous ligation, so that the sequence is directly connected to the rear of the ORF2p-EN sequence in the vector, and is named as the plasmid pBudORF2-EN-6xhis-CH.

[0322] The specific steps are as follows:

[0323] 2. The primers for amplifying the sequence shown in Seq ID No. 13 are designed, wherein the forward primer sequence is shown in Seq ID No. 14: 5'-gctggagctgcggatccatcaccaccaccatc-3', and the reverse primer sequence is shown in Seq ID No. 15: 5'-tctgggtcaggttctttaagatacattgatga-3', and the sequence shown in Seq ID No. 13 is subjected to PCR amplification, and the reaction system is shown in Table 7:

[0324] Table 7 Reaction system

[0325] The amplification conditions are as follows: 94℃ 2min; (98℃ 10sec, 60℃ 10sec, 68℃ 2sec) for 40 cycles; 68℃ 5min.

[0326] The amplified product was obtained by gel recovery and purification by a conventional method, and the amplified product was added with sequences homologous to the pBudORF2-CH vector on both sides of the synthetic sequence.

[0327] 2. The PCR primers for amplifying the pBudORF2-CH vector were designed, wherein the forward primer is shown in Seq ID No. 16: 5'-tcatcaatgtatcttaaagaacctgacccaga-3', and the reverse primer sequence is shown in Seq ID No. 17: 5'-gatggtggtggtgatggatccgcagctccagc-3', and the pBudORF2-CH vector was subjected to PCR amplification, and the reaction system is shown in Table 8:

[0328] Table 8 Reaction system

[0329] The amplification conditions were: 94℃ 2min; (98℃ 10sec, 60℃ 10sec, 68℃ 6sec) for 40 cycles; 68℃ 5min.

[0330] The pBudORF2-CH plasmid vector was obtained by gel recovery and purification by a conventional method, and the plasmid vector had sequences homologous to the synthetic sequence on both ends.

[0331] 3. The amplified product and the amplified pBudORF2-CH vector were connected using a one-step fast cloning kit, and the specific steps were operated according to the kit instructions, and the reaction system is shown in Table 3:

[0332] Table 3 Connection reaction system

[0333] 4. The recombinant product was transformed into competent cells (DH5α), and after transformation, the bacteria were picked on a plate and sequenced, and after the sequence was correct, the plasmid was extracted, and the plasmid pBudORF2-EN-6×his-CH was obtained.

[0334] Example 3. Preparation of ORF2p-EN

[0335] 1. The pBudORF2-EN-6×his-CH plasmid was transfected and expressed.

[0336] HEK293 cells were cultured and passaged with CD293 medium, then pBudORF2-EN-6xhis-CH plasmid was transfected into HEK293 cells according to the PEI transfection reagent instruction. SMS 293-SUPI feeding solution was added according to the instruction at 1, 3, 5 days after transfection. HEK293 cells were cultured in a flask, and the flask culture conditions were as follows: 5% CO2, temperature 37℃, shaking speed 175 rpm. The reactor culture conditions were as follows: pH 7.2, temperature 37℃, stirring speed 150 rpm, dissolved oxygen 40%. The HEK293 cells were incubated in a shaking incubator, and the cells were harvested by centrifugation at 3000g for 5 min after 7 days of transfection. The added SMS 293-SUPI feeding solution can promote cell survival and increase protein yield.

[0337] 2) Extraction of ORF2p-EN

[0338] The cell lysis solution (100 mM potassium acetate, 50 mM Tris-HCl (pH 7.5), 5% glycerol, 0.3% Triton X-100) was prepared and pre-cooled, and PMSF was added to the cell lysis solution to a concentration of 1 mM before use. The cell lysis solution was added to the cells for lysis, and 40 ml of cell lysis solution was added per liter of cells for lysis (the lysis process can also use RIPA lysis solution or other types of cell lysis solution). After that, the cells were blown and lysed with a gun head, avoiding bubbling and vortexing. Then the cells were further treated with a glass homogenizer, and then ultrasonicated for 3 cycles (work for 15 s, interval for 15 s). Centrifugation at 30000g for 25 min at 4℃, the supernatant was retained, and filtered with a 0.22um filter membrane. The filtered sample was subjected to protein purification by Ni affinity chromatography column (HISTRAP HP).

[0339] The protein purification steps are as follows:

[0340] 1. The Ni affinity chromatography column was washed with 5 times the column volume of deionized water;

[0341] 2. Prepare PBS buffer at pH 7.4, and equilibrate the Ni affinity chromatography column with 5-10 times the column volume of PBS buffer;

[0342] 3. The prepared supernatant was flowed through the Ni affinity chromatography column at a speed of 0.5 ml / min;

[0343] 4. Equilibrate the Ni affinity chromatography column again with the prepared buffer;

[0344] 5. Prepare 0.5M NaCl solutions containing 20mM imidazole, 50mM imidazole, 100mM imidazole, 250mM imidazole and 500mM imidazole, respectively;

[0345] 6. Elute the protein sample on the nickel column with 20 mM imidazole, 50 mM imidazole, 100 mM imidazole, 250 mM imidazole and 500 mM imidazole respectively, and collect the corresponding elution samples.

[0346] The collected elution samples are dialyzed overnight at 4°C using the prepared PBS solution as the buffer. Finally, the obtained samples after dialysis are concentrated by ultrafiltration (using appropriate ultrafiltration tubes), and the obtained target protein is detected by SDS-PAGE (the used primary antibody: rabbit anti-his 1:1500 (5% Milk + 0.1% BSA); the secondary antibody: goat anti-rabbit IgG alkaline phosphatase 1:6000 (5% Milk)), and the protein concentration is confirmed, and then the ORF2p-EN is extracted, purified, freeze-dried, and the purified recombinant ORF2p-EN is obtained.

[0347] Example 4. Cleaving DNA and inserting sequence by ORF2p-EN in vitro

[0348] The 360 bp randomly formed single-stranded DNA to be edited is chemically synthesized, as shown in Seq ID No. 18, wherein 300 bp is the single-stranded DNA to be edited, 60 bp sequence of non-editing sequence not present in the single-stranded DNA to be edited is added at the 5' end for detection, represented by capital letters, and * is the target site complementary strand opposite site: the Seq ID No. 18 sequence is:

[0349] The single-stranded DNA framework is chemically synthesized. The sequence upstream of the target site of the single-stranded DNA framework is complementary to the 3' direction sequence of the target site complementary strand opposite site of the single-stranded DNA to be edited, and is 150 bp long, as shown in Seq ID No. 19:

[0350] ctggcctagaagatatacctccggtgctagaggagggtctcgctaagtatcagccgaaagtggtcttaattgagtataccacgcg tgattagacgttacacatgccgatcttaattaactgcccgttgttcctgcaattcaaattggact; the sequence downstream of the target site of the single-stranded DNA framework is complementary to the 5' direction sequence of the target site complementary strand opposite site on the single-stranded DNA to be edited, and is 150 bp long, as shown in Seq ID No. 20:

[0351] aaatttacgccggactagtatcaaggcgatctggctggtagtgtctggacgacaccgcgttagaaattcttcactaaccacattaactgttgggcaactccagttatactaacaatagagcatcactaggcagccactcgcagacctgca; 60 bp randomly generated sequence to be inserted between the sequence upstream of the target site and the sequence downstream of the target site, represented in capital letters, the complete sequence is shown in Seq ID No. 21:

[0352] The sequence in the 5' direction of the target site complementary strand pair flanking site is directly connected with the sequence in the 3' direction of the target site complementary strand pair flanking site at the target site.

[0353] 50-500 ng of recombinant ORF2p-EN was incubated with 100-200 nM of single-stranded DNA to be edited and 100-200 nM of single-stranded DNA framework in ORF2p-EN buffer (20 mM HEPES at pH 7.0, 50 mM NaCl, 1 mM MgCl2, 1 mM DTT, 0.1 mg / ml BSA, 10% glycerol solution) as the experimental group, and the control group was not added with recombinant ORF2p-EN but was the same as the experimental group, and each group had 3 repeats. The nucleic acid after the reaction was recovered after 3 h.

[0354] The nucleic acid after the extraction reaction was further purified, the insertion efficiency of the sequence to be inserted was detected by ddPCR, and sequencing verification was performed. The experimental results showed that the sequence to be inserted was inserted into the target site of the single-stranded DNA to be edited, and there was a significant difference and statistical significance (P<0.05), as shown in Figure 17.

[0355] As can be seen from Figure 17, the sequence to be inserted in the single-stranded DNA framework was inserted into the nucleic acid in vitro, indicating that the nucleic acid editing method provided by the application in vitro can insert the sequence to be inserted into the nucleic acid sequence in vitro.

[0356] Example 5. Editing of eukaryotic genome using double-stranded DNA framework

[0357] 1. Selecting target site and designing and synthesizing double-stranded DNA framework: a site on the genome was randomly selected for editing. A site in the GAPDH gene, a key gene in the glycolytic pathway, was selected for sequence insertion, and the target site is represented by *, as shown in Seq ID No. 22:

[0358] A double-stranded DNA framework was designed. The sequence upstream of the target site of the double-stranded DNA framework was complementary to the sequence upstream of the genomic target site to be edited, 300 bp in length, as shown in Seq ID No. 23: GTACAAGCGTTTTCTCCCTAAAGGGTGCAGCTGAGCTAGGCAGCAGCAAGCATTCCTGGGGTGGCATAGTGGGGTGGTGAATACCATGTACAAAGCTTGTGCCCAGACTGTGGGTGGCAGTGCCCCACATGGCCGCTTCTCCTGGAAGGGCTTCGTATGACTGGGGGTGTTGGGCAGCCCTGGAGCCTTCAGTTGCAGCCATGCCTTAAGCCAGGCCAGCCTGGCAGGGAAGCTCAAGGGAGATAAAATTCAACCTCTTGGGCCCTCCTGGGGGTAAGGAGATGCTGCATTCGCCCTCTT; the sequence downstream of the target site of the double-stranded DNA framework was complementary to the sequence downstream of the genomic target site to be edited, 300 bp in length, as shown in Seq ID No. 24:

[0359] AATGGGGAGGTGGCCTAGGGCTGCTCACATATTCTGGAGGAGCCTCCCCTCCTCATGCCTTCTTGCCTCTTGTCTCTTAGATTTGGTCGTATTGGGCGCCTGGTCACCAGGGCTGCTTTTAACTCTGGTAAAGTGGATATTGTTGCCATCAATGACCCCTTCATTGACCTCAACTACATGGTGAGTGCTACATGGTGAGCCCCAAAGCTGGTGTGGGAGGAGCCACCTGGCTGATGGGCAGCCCCTTCATACCCTCACGTATTCCCCCAGGTTTACATGTTCCAATATGATTCCACCCAT. The complete sequence is shown in Seq ID No. 25:

[0360] The sequence upstream of the target site and the sequence downstream of the target site on the genome were directly connected at the target site. The 100 bp sequence to be inserted between the sequence upstream of the target site and the sequence downstream of the target site of the double-stranded DNA framework was randomly formed, indicated in bold. The double-stranded DNA framework was chemically synthesized.

[0361] 2. First, prepare the cells to be transfected: the day before, 293T cells were passaged into 6-well plates, and the next day, the 293T cells were grown to a cell density of 40-60%, then the original complete culture medium was removed, the cells were washed with 1 x PBS for 2-3 times, and then 2 mL of fresh Opti-MEM containing 5% fetal bovine serum was added to each culture well TM I medium.

[0362] 3. Prepare the transfection reagent / DNA complex: dilute the PolyFast transfection reagent with 50 μL of serum-free culture medium DMEM, then gently mix with a pipette, react at room temperature for 5 minutes, to obtain one portion of diluted PolyFast transfection reagent, and prepare two portions of the above diluted PolyFast transfection reagent for each culture well.

[0363] For each well of cells in the 6-well plate to be transfected, mix 1 μg of pBud-NLS-ORF2-EN-NLS-CH with 50 μL of serum-free culture medium DMEM, and mix 1 μg of double-stranded DNA framework with 50 μL of serum-free culture medium DMEM, then gently mix with a pipette, react at room temperature for 5 minutes.

[0364] Add one portion of the diluted PolyFast transfection reagent to the serum-free culture medium DMEM containing 1 μg of pBud-NLS-ORF2-EN-NLS-CH and the serum-free culture medium DMEM containing 1 μg of double-stranded DNA framework, respectively, mix gently, and incubate at room temperature for 15 minutes, to obtain one portion of PolyFast transfection reagent-double-stranded DNA framework complex and one portion of PolyFast transfection reagent-pBud-NLS-ORF2-EN-NLS-CH complex.

[0365] 4. Add the above one portion of PolyFast transfection reagent-double-stranded DNA framework complex and one portion of PolyFast transfection reagent-pBud-NLS-ORF2-EN-NLS-CH complex to the same culture well as the experimental group, and the group without adding one portion of PolyFast transfection reagent-pBud-NLS-ORF2-EN-NLS-CH complex but with the same other conditions as the control group, with 3 replicates in each group. Mix gently and incubate the cells at 37°C for 24 hours.

[0366] 5. After 48 hours of culture of the transfected cells, extract the genomic DNA of the cells, detect the insertion efficiency of the to-be-inserted sequence by ddPCR, and verify by sequencing. The efficiency of insertion of the to-be-inserted sequence can be more than 5%, with significant difference and statistical significance (P<0.05), as shown in Figure 18.

[0367] As can be seen from FIG. 18, the to-be-inserted sequence in the double-stranded DNA is inserted into the target site on the genome, indicating that the nucleic acid editing method provided by the application can insert the to-be-inserted sequence into the target site of the eukaryotic genome through the double-stranded DNA framework in the cell.

[0368] Example 6. Editing of eukaryotic genome using single-stranded DNA framework

[0369] 1. Selecting target site and designing and synthesizing single-stranded DNA framework: a site on the genome is randomly selected for editing. The FLT1 gene plays an important role in angiogenesis. A site in the FLT1 gene is selected for sequence insertion, and the target site complementary strand opposite site is indicated by an asterisk, and the target site complementary strand is shown as Seq ID No. 26:

[0370] 2. Designing a single-stranded DNA framework, the sequence upstream of the target site of the single-stranded DNA framework is complementary to the complementary sequence of the sequence upstream of the target site of the genome to be edited, and is 300 bp long, as shown in Seq ID No. 27: TCAGCCTCCCAAAGTGTTGAGACTACTGCATGAGCCACCACACCCAGCCCCTCATGGTAGTCTTGATTTGTATTTCCCTAATGACTGATGATGTTGAGCATCTTTTCGTGTACTCACTGGCCATGCCCACCTGTTTCTGTAAATACAGTGTGTGGGAACAGAGCCATGCCTACTTGATGATGTATTGTCTACGGTTGCTTTTTATACTATAATGCAAGCTTGCCCAACTTGTGGCCCACGGGCTGCATGAATCCCAGGATGGCTTTGAGTGCAGCCCAACACAAACTCGTAAACTTTCTT; the sequence downstream of the target site of the single-stranded DNA framework is complementary to the complementary sequence of the sequence downstream of the target site of the genome to be edited, and is 300 bp long, as shown in Seq ID No. 28:

[0371] The complete sequence is shown in Seq ID No. 29: The sequence upstream of the target site and the sequence downstream of the target site on the genome are directly connected at the target site. The 80 bp to-be-inserted sequence randomly formed between the sequence upstream of the target site and the sequence downstream of the target site of the single-stranded DNA framework is indicated in bold. The single-stranded DNA framework is chemically synthesized.

[0372] 2. First, prepare the cells to be transfected: 293T cells were passaged into 6-well plates one day in advance, and the next day, the 293T cells were grown to a cell density of 40-60% for cell transfection. Then the original complete culture medium was removed, the cells were washed with 1 x PBS for 2-3 times, and then 1 mL of fresh Opti-MEM containing 5% fetal bovine serum was added to each culture well TM I culture medium.

[0373] 3. Prepare the transfection reagent / DNA complex: dilute the PolyFast transfection reagent in 3 μL of PolyFast transfection reagent with 50 μL of serum-free culture medium, then gently mix with a pipette gun, and react at room temperature for 5 minutes to obtain one portion of diluted PolyFast transfection reagent. Prepare 2 portions of the above diluted PolyFast transfection reagent for each culture well.

[0374] For each well of cells in the 6-well plate to be transfected, mix 50 μL of serum-free culture medium with 1 μg of pBud-NLS-ORF2-EN-NLS-CH and 1 μg of single-stranded DNA framework, respectively, and gently mix with a pipette gun, and react at room temperature for 5 minutes.

[0375] Add 1 portion of diluted PolyFast transfection reagent to the serum-free culture medium DMEM containing 1 μg of pBud-NLS-ORF2-EN-NLS-CH and the serum-free culture medium DMEM containing 1 μg of single-stranded DNA framework, respectively, and gently mix, and incubate at room temperature for 15 minutes to obtain 1 portion of PolyFast transfection reagent-single-stranded DNA framework complex and 1 portion of PolyFast transfection reagent-pBud-NLS-ORF2-EN-NLS-CH complex.

[0376] 4. Add the above 1 portion of PolyFast transfection reagent-single-stranded DNA framework complex and 1 portion of PolyFast transfection reagent-pBud-NLS-ORF2-EN-NLS-CH complex to the same culture well as the experimental group, and the group without adding 1 portion of PolyFast transfection reagent-pBud-NLS-ORF2-EN-NLS-CH complex but with the same other conditions as the control group, with 3 replicates in each group. Gently mix and incubate at 37°C for 24 hours.

[0377] 5. Analyze the transfected cells: after the transfected cells are cultured for 48 hours, extract the genomic DNA of the cells, detect the insertion efficiency of the to-be-inserted sequence by ddPCR, and verify by sequencing. The efficiency of insertion of the to-be-inserted sequence can be more than 5%, with significant difference and statistical significance (P<0.05), as shown in Figure 19.

[0378] As can be seen from FIG. 19, the sequence to be inserted of the single-stranded DNA is inserted into the target site, which indicates that the nucleic acid editing method provided by the application can insert the sequence to be inserted into the target site in the genome through the single-stranded DNA framework in the cell.

[0379] Example 7. Editing of eukaryotic genome using DNA / RNA hybridization framework

[0380] 1. Selection of target site and design and synthesis of DNA / RNA hybridization framework: a site on the genome is randomly selected for editing. The HTT gene is the pathogenic gene of the neurodegenerative disease Huntington's disease. A site in the HTT gene is selected for sequence insertion, and the complementary strand of the target site is indicated by * on the opposite side of the site, and the complementary strand of the target site is shown as Seq ID No. 30:

[0381] 2. Design of DNA / RNA hybridization framework, the sequence upstream of the target site of the DNA / RNA hybridization framework is complementary to the sequence upstream of the target site of the genome to be edited or to the complementary sequence thereof, and is 300 bp long, as shown in Seq ID No. 31:

[0382] AAGAGAAGTTCCGACATTTTGTTCCGACACCCCAAACGACCACGTCCCTATCTCCTTGAGTGAGCATATACCACCGTGCCTAGACCAGTGTGCTAGGCCACAGCCTCCACGGGGCAGAAGCTGCAGCTGAAATACTCACCATTGTATACCCAGCCCCTATCCCAGAGCCTAGCACACAGTGGGTGCTCAATCAGTATTTACTGGGTGGATGAACACATATCGTGACCTCTGGTTTTAGAAACAGGGTAGCAGTGGTCTCTGAGAGGCTTAGGTAACACGGCACCTGAGGAGGCATCACCC; the sequence downstream of the target site of the DNA / RNA hybridization framework is complementary to the sequence downstream of the target site of the genome to be edited or to the complementary sequence thereof, and is 300 bp long, as shown in Seq ID No. 32:

[0383] GCATTTGACTCACGGTGACAGGGACATACAGAGGGAAGACAGCTAGAGGTCAGCACCTAGGGTTGGAAACGAGCAGCGCAGCTCCAGAGTCCACACCTCTAACTCTCGGTTCTGCCGCTTCTCTTCCAAATTAATCGGTTCTTTGAATGT CAGTCATTCACATGCAGGGCCTAGAGTAAACCCTGTGCTCCTACCTGGACAAGCTGCTCTGCAGAAGGAGAGACTTCCATTCTTTCCTTGTCACTCCGAAGCTGCCTTTCAGGCTTGTGTCCTTGACCTGCTGCTGCAGCAAGGGCACC. The complete sequence is shown in Seq ID No.33: The upstream and downstream sequences of the target site on the genome are directly linked at the target site. Between the upstream and downstream sequences of the target site in the DNA / RNA hybridization framework is a randomly generated 104 bp insertion sequence, indicated in bold. A T7 promoter is added upstream of the upstream sequence of the target site. The DNA sequence of this T7 promoter + DNA / RNA hybridization framework is chemically synthesized, as shown in Seq ID No. 34. The underlined portion represents the T7 promoter sequence, the italicized portion represents the sequence added to promote T7 transcription, the portion between the italicized and bold portions represents the upstream sequence of the target site, and the portion downstream of the bold portions represents the downstream sequence of the target site. This DNA / RNA hybridization framework was obtained through in vitro transcription.

[0384] 2. First, prepare the cells for transfection: Passage 293T cells into 6-well plates one day in advance. The next day, allow the 293T cells to grow to a cell density of 40-60% to prepare for transfection. Afterward, aspirate the original complete culture medium, wash the cells 2-3 times with 1×PBS, and add 1 mL of fresh Opti-MEM containing 5% fetal bovine serum to each well. TM I. Culture medium.

[0385] 3. Prepare transfection reagent / DNA complex: Add 50 μL of serum-free culture medium to 3 μL of PolyFast transfection reagent to dilute the PolyFast transfection reagent. Then, gently mix by pipetting and react at room temperature for 5 minutes to obtain one diluted PolyFast transfection reagent. Prepare two copies of the diluted PolyFast transfection reagent for each culture well.

[0386] For each well of cells in the 6-well plate to be transfected, 50 μL of serum-free medium was mixed with 1 μg of pBud-NLS-ORF2-EN-NLS-CH and 1 μg of single-stranded DNA framework, respectively, and mixed well by gun blowing, and reacted at room temperature for 5 minutes.

[0387] 1 part of diluted PolyFast transfection reagent was added to 1 μg of pBud-NLS-ORF2-EN-NLS-CH in serum-free medium DMEM and 1 μg of DNA / RNA hybridization framework in serum-free medium DMEM, respectively, mixed gently, incubated at room temperature for 15 minutes, to obtain 1 part of PolyFast transfection reagent-DNA / RNA hybridization framework complex and 1 part of PolyFast transfection reagent-pBud-NLS-ORF2-EN-NLS-CH complex.

[0388] 4. The above 1 part of PolyFast transfection reagent-DNA / RNA hybridization framework complex and 1 part of PolyFast transfection reagent-pBud-NLS-ORF2-EN-NLS-CH complex were added to the same culture well as the experimental group; the group without adding 1 part of PolyFast transfection reagent-pBud-NLS-ORF2-EN-NLS-CH complex but with the same other conditions was the control group, with 3 repeats in each group. After gentle mixing, incubate at 37°C for 24 hours.

[0389] 5. Analysis of transfected cells: After the transfected cells were cultured for 48 hours, the genomic DNA of the cells was extracted, the insertion efficiency of the to-be-inserted sequence was detected by ddPCR, and the verification was performed by sequencing. The insertion efficiency of the to-be-inserted sequence was more than 5%, and there was a significant difference and statistical significance (P<0.05), as shown in FIG. 20.

[0390] As can be seen from FIG. 20, the to-be-inserted sequence of the DNA / RNA hybridization framework is inserted into the target site, indicating that the nucleic acid editing method provided by the application can insert the to-be-inserted sequence into the genomic target site in the cell through the DNA / RNA hybridization framework.

[0391] Example 8

[0392] The junction of the 5' end of the 8th exon and the intron of the GAPDH gene was selected as the target site (the intron sequence is the upstream sequence of the target site, and the 8th exon sequence is the downstream sequence of the target site), and the target site is represented by *, and the sequence is shown in Seq ID No. 35:

[0393] The method of the double-stranded DNA framework of embodiment 5 of the present application was used for gene editing, and the junction of the 5' end of the 8th exon and the intron of the GAPDH gene was selected as the target site (the intron sequence is the upstream sequence of the target site, and the 8th exon sequence is the downstream sequence of the target site), and the sequence to be inserted is the AcGFP1 fluorescent protein sequence, represented by lowercase letters, with the upstream being the upstream sequence of the target site and the downstream being the downstream sequence of the target site, and the sequence is shown in Seq ID No. 36:

[0394] The same experiment group and control group each had 3 replicates, and fluorescence was observed after transfection, and the results are shown in Figure 21.

[0395] As can be seen from Figure 21, some cells in the experimental group emit green fluorescence, indicating that the GFP sequence has been inserted into the genome, while no obvious fluorescence is observed in the control group.

[0396] As can be seen from the above embodiments, the nucleic acid editing system, nucleic acid editing method and application provided by the present application are more secure compared with the prior art of cutting open double strands and then homologous recombination, and can use the endonuclease ORF2p-EN of the human ORF2p which is more human-friendly and exists in normal human bodies, which has a clinical application and other application prospects, and can be applied to the treatment of various diseases related to genes, such as genetic diseases or cancer, etc. in the future. At the same time, the present application omits the retrotransposon sequence and the reverse transcription process of the retrotransposon, so that the entire gene editing process is more controllable.

[0397] The technical features of the above-described embodiments can be combined in any manner. To make the description concise, not all possible combinations of the technical features in the above embodiments are described, but as long as the combinations of the technical features do not contradict, they should be considered as within the scope of the present disclosure.

Claims

1. A nucleic acid editing system, characterized by, The frame structure and the shearing structure are included; The frame structure is one or more of a single-stranded DNA frame, a double-stranded DNA frame, a single-stranded DNA derivative frame, a double-stranded DNA derivative frame, a DNA-RNA hybrid frame, a DNA / RNA hybrid derivative frame, a single-stranded RNA frame, a double-stranded RNA frame, a single-stranded RNA derivative frame, and a double-stranded RNA derivative frame; and the shearing structure is a protein or a polypeptide. The single-stranded DNA frame, the double-stranded DNA frame, the single-stranded DNA derivative frame, and the double-stranded DNA derivative frame are composed of a target site upstream sequence, an insertion sequence, and a target site downstream sequence in the 5'→3' direction; the target site upstream sequence of the single-stranded DNA frame, the double-stranded DNA frame, the single-stranded DNA derivative frame, and the double-stranded DNA derivative frame or a complementary sequence of the target site upstream sequence is used for hybridization with a target site upstream sequence or a complementary sequence of the target site upstream sequence in a target site of a nucleic acid sequence to be sheared or edited after shearing; a target site downstream sequence or a complementary sequence of the target site downstream sequence on the single-stranded DNA frame, the double-stranded DNA frame, the single-stranded DNA derivative frame, and the double-stranded DNA derivative frame is used for hybridization with a target site downstream sequence or a complementary sequence of the target site downstream sequence in the target site of the nucleic acid sequence to be sheared or edited after shearing; the target site upstream sequence and the target site downstream sequence on the single-stranded DNA frame, the double-stranded DNA frame, the single-stranded DNA derivative frame, and the double-stranded DNA derivative frame are directly connected to corresponding sequences in the nucleic acid sequence to be sheared or edited after shearing; and a target site is between the target site upstream sequence and the target site downstream sequence of the nucleic acid sequence to be sheared or edited after shearing or a complementary sequence thereof. The DNA / RNA hybrid frame or the DNA / RNA hybrid derivative frame includes a target site upstream sequence, an insertion sequence, and a target site downstream sequence in the 5'→3' direction of a single-stranded RNA or a single-stranded DNA in the DNA / RNA hybrid frame or the DNA / RNA hybrid derivative frame. The target site upstream sequence of the DNA / RNA hybrid frame or the DNA / RNA hybrid derivative frame or a complementary sequence of the target site upstream sequence is used for hybridization with a target site upstream sequence or a complementary sequence of the target site upstream sequence in a target site of a nucleic acid sequence to be sheared or edited after shearing; a target site downstream sequence or a complementary sequence of the target site downstream sequence on the DNA / RNA hybrid frame or the DNA / RNA hybrid derivative frame is used for hybridization with a target site downstream sequence or a complementary sequence of the target site downstream sequence in the target site of the nucleic acid sequence to be sheared or edited after shearing; the target site upstream sequence and the target site downstream sequence on the DNA / RNA hybrid frame or the DNA / RNA hybrid derivative frame are directly connected to corresponding sequences in the nucleic acid sequence to be sheared or edited after shearing; and a target site is between the target site upstream sequence and the target site downstream sequence of the nucleic acid sequence to be sheared or edited after shearing or a complementary sequence thereof. The single-stranded RNA framework, double-stranded RNA framework, single-stranded RNA derived framework, and double-stranded RNA derived framework comprise a sequence upstream of the target site, a sequence to be inserted, and a sequence downstream of the target site in the 5'→3' direction; The sequence upstream of the target site on the single-stranded RNA framework, double-stranded RNA framework, single-stranded RNA derived framework, and double-stranded RNA derived framework or the complement of the sequence upstream of the target site is used for hybridizing with the sequence upstream of the target site or the complement of the sequence upstream of the target site in the nucleic acid sequence to be cleaved or cleaved for editing, and the sequence downstream of the target site on the single-stranded RNA framework, double-stranded RNA framework, single-stranded RNA derived framework, and double-stranded RNA derived framework or the complement of the sequence downstream of the target site is used for hybridizing with the sequence downstream of the target site or the complement of the sequence downstream of the target site in the nucleic acid sequence to be cleaved or cleaved for editing; the sequence upstream of the target site and the sequence downstream of the target site on the single-stranded RNA framework, double-stranded RNA framework, single-stranded RNA derived framework, and double-stranded RNA derived framework are directly connected in the corresponding sequences in the nucleic acid sequence to be cleaved or cleaved for editing; the sequence upstream of the target site and the sequence downstream of the target site of the nucleic acid sequence to be cleaved or cleaved for editing or the complement thereof are the target site; The cleavage structure is a protein or a polypeptide or DNA or RNA encoding the protein or the polypeptide, wherein the protein is ORF2p-EN, ORF2p, ORF2p-EN homologous protein, ORF2p derived protein, ORF2p-EN derived protein, or ORF2p-EN homologous protein derived protein; the polypeptide is a length part of 50%-99% of the protein; The nucleic acid sequence to be cleaved or cleaved for editing is single-stranded DNA, double-stranded DNA, single-stranded RNA, double-stranded RNA, or DNA / RNA hybrid chain.

2. The nucleic acid editing system of claim 1, wherein, The framework structure and the nucleic acid to be cleaved or edited are completely matched at the 4bp immediately upstream of the target site and the 4bp immediately downstream of the target site; the remaining sequences are complementary to 0%-100%.

3. The nucleic acid editing system of claim 1 or 2, wherein The cleaved or edited nucleic acid is single-stranded DNA, double-stranded DNA, single-stranded RNA, double-stranded RNA, or DNA / RNA hybrid chain in vivo or in vitro; The double-stranded DNA or single-stranded DNA comprises the genome of eukaryotes, prokaryotes, or viruses; the double-stranded RNA or single-stranded RNA comprises the genome of viruses, mRNA, rRNA, miRNA, or siRNA of eukaryotes, prokaryotes, or viruses.

4. The nucleic acid editing system of claim 1 or 2, wherein The framework structure is linear or circular.

5. A vector system characterized by comprising The nucleic acid editing system comprises the framework structure and the cleavage structure.

6. Use of the nucleic acid editing system according to any one of claims 1 to 4 or the vector system according to claim 5 in sequence insertion, sequence deletion, sequence replacement, site insertion, site deletion, site replacement, sequence inversion, and / or sequence inversion correction in any region of the genome.

7. A method of nucleic acid editing, comprising, The method comprises the following steps: 1) selecting the insertion site of the nucleic acid to be cleaved or edited, determining the upstream sequence of the target site and the downstream sequence of the target site on both sides of the insertion site; 2) preparing the vector system according to claim 5; 3a) adding or transferring the vector system into a solution, cell, tissue, organ or organism containing the nucleic acid to be cleaved or edited; or 3b) adding or transferring the vector system into a solution, cell, tissue, organ or organism containing the nucleic acid to be cleaved or edited after combining with the homologous recombination protein and / or reverse transcriptase; or 3c) adding or transferring the vector system into a solution, cell, tissue, organ or organism containing the nucleic acid to be cleaved or edited at the same time with the homologous recombination protein and / or reverse transcriptase; or 3d) adding or transferring the vector system into a solution, cell, tissue, organ or organism containing the nucleic acid to be cleaved or edited and then adding the homologous recombination protein and / or reverse transcriptase; or 3e) adding or transferring the homologous recombination protein and / or reverse transcriptase into a solution, cell, tissue, organ or organism containing the nucleic acid to be cleaved or edited and then adding the vector system.

8. The nucleic acid editing method of claim 7, wherein, The primer is added at the same time as the vector is added in step 3a), 3b), 3c), 3d) or 3e).

9. The nucleic acid editing method of claim 7, wherein, The DNA polymerase or RNA polymerase is added at the same time as the vector is added in step 3a), 3b), 3c), 3d) or 3e).

10. The nucleic acid editing method of claim 7, wherein, The compound RS-1, SCR7, NU7441, Vanillin, Chlorpromazine, Resveratrol, Benomyl, Inositol that improves the efficiency of homologous recombination is added at the same time as the vector is added in step 3a), 3b), 3c), 3d) or 3e).

11. The nucleic acid editing method of any one of claims 7 to 10, wherein, The homologous recombination protein is Rec / Rad51 family, RAD51, RecA, Uvs X, RAD52, RAD51AP1, RAD51AP2, RAD54L, Swi5-Sfr1 complex, BRCA1, BRCA2, PALB2, RAD53, RAD54, RAD55, RAD56, RAD57, BARD1, RecA, RPA, Mre11, NBS, DMC1, SCR7, SRS2, PSF, PARP1, RBM6, RAD58, RAD59, RAD50, TID1 / RDH54, MSH3, Mlh1, XRS2, NBS1, SSBP, PHF1, SMC1, DHX9, RECQ helicase such as RECQ1, P53, Mus81, SWI2 / SNF2 family protein, Hop2 or Mndl.

12. Use of the nucleic acid editing system according to any one of claims 1 to 4 or the vector system according to claim 5 in the preparation of a drug for preventing and / or treating cancer, a genetic-related disease or a neurodegenerative disease.

13. Use according to claim 12, wherein the compound is ###00010### or a pharmaceutically acceptable salt thereof. The cancer is glioma, breast cancer, cervical cancer, lung cancer, gastric cancer, colorectal cancer, duodenal cancer, leukemia, prostate cancer, endometrial cancer, thyroid cancer, lymphoma, pancreatic cancer, liver cancer, melanoma, skin cancer, pituitary tumor, germ cell tumor, meningioma, meningeal cancer, glioblastoma, various astrocytoma, various oligodendroglioma, astrocytic oligodendroglioma, various ependymoma, choroid plexus papilloma, choroid plexus carcinoma, chordoma, various ganglioneuroma, olfactory neuroblastoma, sympathetic nervous system neuroblastoma, pineal cell tumor, pinealoblastoma, medulloblastoma, retinoblastoma, trigeminal schwannoma, facial acoustic neuroma, jugular bulb tumor, angiomatous reticulocyte, craniopharyngioma or granular cell tumor.

14. The use according to claim 12, wherein the compound is ###00009### or a pharmaceutically acceptable salt thereof. Huntington's disease, fragile X syndrome, phenylketonuria, Duchenne's progressive muscular dystrophy, Duchenne's muscular dystrophy, mitochondrial encephalomyopathy, mucopolysaccharidosis type I, mucopolysaccharidosis type II, mucopolysaccharidosis type IIIA, mucopolysaccharidosis type IIIB, mucopolysaccharidosis type IIIA, mucopolysaccharidosis type IIIC, mucopolysaccharidosis type IIID, mucopolysaccharidosis type IVA, mucopolysaccharidosis type IVB, mucopolysaccharidosis type VI, mucopolysaccharidosis type VII, mucopolysaccharidosis type IX, spinal muscular atrophy, Parkinson-plus syndrome, albinism, color blindness, achondroplasia, alkaptonuria, congenital deaf mutism, thalassemia, sickle cell anemia, hemophilia, epilepsy associated with genetic alterations, myoclonus, dystonia, stroke and schizophrenia, antivitamin D rickets, familial colonic polyposis, 21-hydroxylase deficiency, arginase deficiency, Alport syndrome, Angelman syndrome, Tay-Sachs disease, atypical hemolytic uremic syndrome, autoimmune encephalitis, autoimmune hypophysitis, autoimmune insulin receptor disease, beta-ketothiolase deficiency, biotinidase deficiency, cardiac ion channelopathy, primary carnitine deficiency, Castleman disease, Charcot-Marie-Tooth disease, citrullinemia, congenital adrenal hypoplasia, congenital hyperinsulinemic hypoglycemia, congenital myasthenic syndrome, non-dystrophic myotonia syndrome, congenital scoliosis, coronary artery ectasia, congenital pure red cell aplasia, Erdheim-Chester disease, Fabry disease, familial Mediterranean fever, Fanconi anemia, galactosemia, Gaucher disease, generalized myasthenia gravis, Gitelman syndrome, glutaric aciduria type I, glycogen storage disease type I, glycogen storage disease type II, hemophilia, hepatolenticular degeneration, hereditary angioedema, hereditary epidermolysis bullosa, hereditary fructose intolerance, hereditary hypomagnesemia, hereditary multi-infarct dementia, hereditary spastic paraplegia, holocarboxylase synthetase deficiency, homocystinuria, homozygous familial hypercholesterolemia, HHH syndrome, hyperphenylalaninemia, hypophosphatasia, hypophosphatemic rickets, idiopathic cardiomyopathy, idiopathic hypogonadotropic hypogonadism, idiopathic pulmonary arterial hypertension, idiopathic pulmonary fibrosis, IgG4-related disease, inborn error of bile synthesis, isovaleric acidemia, Kallmann syndrome, Langerhans cell histiocytosis, Leigh syndrome, Leber hereditary optic neuropathy, long-chain 3-hydroxyacyl-CoA dehydrogenase deficiency, lymphangioleiomyomatosis, lysinuric protein intolerance, lysosomal acid lipase deficiency, maple syrup urine disease, Marfan syndrome, McCune-Albright syndrome, medium-chain acyl-CoA dehydrogenase deficiency, methylmalonic acidemia, multifocal motor neuropathy, multiple acyl-CoA dehydrogenase deficiency, multiple sclerosis, myotonic dystrophy,N-acetylglutamate synthetase deficiency, neonatal diabetes, neuromyelitis optica, Niemann-Pick disease, non-syndromic deafness, Noonan syndrome, ornithine transcarbamylase deficiency, osteogenesis imperfecta, young-onset Parkinson's disease, early-onset Parkinson's disease, paroxysmal nocturnal hemoglobinuria, Peutz-Jeghers syndrome, POEMS syndrome, porphyria, Prader-Willi syndrome, primary combined immunodeficiency, primary hereditary myotonia, primary light chain amyloidosis, progressive familial intrahepatic cholestasis, progressive muscular dystrophy, propionic acidemia, pulmonary alveolar proteinosis, pulmonary cystic fibrosis, retinitis pigmentosa, severe congenital neutropenia, infantile severe myoclonic epilepsy, Dravet syndrome, Silver-Russell syndrome, sitosterolemia, spinal and bulbar muscular atrophy, spinal muscular atrophy, spinocerebellar ataxia, systemic sclerosis, tetrahydrobiopterin deficiency, tuberous sclerosis, primary tyrosinemia, very long chain acyl-CoA dehydrogenase deficiency, Williams syndrome, Wiskott-Aldrich syndrome, X-linked agammaglobulinemia, X-linked adrenoleukodystrophy, X-linked lymphoproliferative syndrome, arteriosclerotic cerebral small vessel disease, cerebral amyloid angiopathy, cerebral arteriopathy associated with subcortical infarcts and leukoaraiosis, cerebral arteriopathy associated with subcortical infarcts and leukoencephalopathy, cysteine string protein-associated arteriopathy with stroke and leukoencephalopathy, pyridoxine-dependent epilepsy, serotonin metabolism AADC enzyme deficiency, AADC deficiency, or hereditary nephritis.

15. The use according to claim 12, wherein the compound is ###0006### The neurodegenerative disease is Parkinson's disease, Alzheimer's disease, Huntington's disease, amyotrophic lateral sclerosis, spinocerebellar ataxia, multiple system atrophy, primary lateral sclerosis, Pick's disease, frontotemporal dementia, Lewy body dementia or progressive supranuclear palsy.

16. Use of an ORF2p-EN, ORF2p, ORF2p-EN homologous protein, ORF2p-derived protein, ORF2p-EN-derived protein or ORF2p-EN homologous protein-derived protein in the preparation of a substance of a nucleic acid editing system.

17. The use according to claim 16, wherein A nuclear localization signal is further added to the N-terminus, N-part, C-terminus, C-part or sequence middle of the ORF2p-EN, ORF2p, ORF2p-EN homologous protein, ORF2p-derived protein, ORF2p-EN-derived protein or ORF2p-EN homologous protein-derived protein.

18. Use according to claim 16 or 17, wherein A tag is further added to the N-terminus, N-part, C-terminus, C-part or sequence middle of the ORF2p-EN, ORF2p, ORF2p-EN homologous protein, ORF2p-derived protein, ORF2p-EN-derived protein or ORF2p-EN homologous protein-derived protein; the tag is selected from HIS, Flag, GST, Myc, eGFP, eCFP, eYFP, mCherry eGFP, HA, SUMO, MBP, Strep-II Tag, AviTag or SNAP-Tag.

19. A method of nucleic acid editing, comprising, The method comprises the following steps: 1) selecting the insertion site of the nucleic acid to be edited, determining the upstream sequence of the target site and the downstream sequence of the target site on both sides of the insertion site; 2) preparing the nucleic acid editing system according to any one of claims 1 to 4 or the vector system according to claim 5; 3a) when either or both of the nucleic acid to be edited or the frame structure in the nucleic acid editing system is double-stranded, mixing the nucleic acid to be edited with the frame structure to denature and then anneal, so that the single strand of the frame structure is combined with the single strand of the nucleic acid to be edited; or 3b) when one of the nucleic acid to be edited or the frame structure in the nucleic acid editing system is single-stranded and the other is double-stranded, the nucleic acid to be edited, the frame structure and the homologous recombination protein can be mixed to perform strand displacement, so that the single strand of the frame structure is combined with the single strand of the nucleic acid to be edited; or 3c) when either or both of the framework structure in the nucleic acid to be edited or the nucleic acid editing system is double-stranded, mixing the nucleic acid to be edited with the framework structure and adding exonuclease, converting the double-stranded into single-stranded, and then combining the single-stranded of the framework structure with the single-stranded of the nucleic acid to be edited; or 3d) when both of the framework structure in the nucleic acid to be edited or the nucleic acid editing system are single-stranded, directly mixing the two, and combining the single-stranded of the framework structure with the single-stranded of the nucleic acid to be edited; 4) adding a cleavage structure to the hybridization product of the single-stranded of the framework structure and the single-stranded of the nucleic acid to be edited, thereby cleaving the single-stranded of the nucleic acid to be edited; 5) continuing to add DNA polymerase or RNA polymerase to complete the editing of the nucleic acid to be edited.

Citation Information

Patent Citations

  • Gene transcription framework, vector system, genome sequence editing method and application

    CN112708636A

  • Methods and compositions for genome integration

    CN114981409A

  • RNA framework for gene editing and gene editing method

    CN115044583A

  • Targeted insertion via transposition

    WO2024098063A2