CRISPR / Cas9 gene editing system based on SeqCas9 protein and related application of CRISPR / Cas9 gene editing system

By combining the SeqCas9 protein and its mutants with single-stranded guide RNA, a highly efficient fusion protein and base editor were developed, solving the off-target effects and complex PAM sequences of the CRISPR/Cas9 system. This achieved high specificity and high editing activity, expanding the applications of gene editing.

CN121896199APending Publication Date: 2026-04-21SHANGHAI PUDONG HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI PUDONG HOSPITAL
Filing Date
2026-01-07
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing CRISPR/Cas9 gene editing systems suffer from off-target effects and complex PAM sequences, resulting in low editing activity and making them difficult to widely apply.

Method used

By using SeqCas9 protein and its mutants in conjunction with single-stranded guide RNA, highly efficient fusion proteins and base editors have been developed, including SeqCas9 mutant proteins, detectable markers, fusion proteins of cytosine deaminase and adenine deaminase, which enable efficient gene editing by specifically recognizing target sequences.

Benefits of technology

It achieves highly specific, highly editable, and simple PAM sequences, expanding the application scope of gene editing and enabling efficient single-base conversion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121896199A_ABST
    Figure CN121896199A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of gene editing, and particularly relates to a CRISPR (clustered regularly interspaced short palindromic repeats) / Cas9 gene editing system based on SeqCas9 protein and fusion protein of SeqCas9 and application of the CRISPR / Cas9 gene editing system. The method has a wide application prospect in the field of gene editing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of gene editing technology, and more specifically to the SeqCas9 protein, SeqCas9-based fusion proteins, CRISPR / Cas9 gene editing systems based on said protein, and related applications for gene editing. Background Technology

[0002] The CRISPR / Cas9 system is an acquired immune system evolved by bacteria and archaea to defend against invasion by exogenous viruses or plasmids. The CRISPR / Cas9 system contains tracrRNA (trans-activating RNA) and crRNA (CRISPR-derived RNA), which, together with Cas9, form a complex to function. tracrRNA and crRNA can fuse into single-stranded guide RNA (sgRNA) through a linker sequence. Following DNA breaks in eukaryotic cells mediated by this system, two main DNA damage repair mechanisms are responsible for repair: non-homologous end-joining (NHEJ) and homologous recombination (HR). NHEJ repair results in base deletions or insertions, enabling gene knockout; HR repair, when a homologous template is provided, allows for site-specific gene insertion and precise base substitution.

[0003] Beyond basic scientific research, CRISPR / Cas9 gene editing systems hold broad clinical application potential. However, almost all CRISPR / Cas9 systems suffer from off-target effects, leading to the accidental editing of non-target sequences; or the PAM sequences are complex, or the editing activity is low, hindering widespread application. Therefore, a crucial research direction in this field is currently developing more specific editing tools to reduce off-target effects.

[0004] Therefore, finding a CRISPR / Cas system with high editing activity, high specificity, simple PAM sequences, and a narrow base editing window is the hope for solving the above problems. Summary of the Invention

[0005] The inventors discovered in their research that the SeqCas9 protein and its corresponding single-stranded guide RNA can play a gene editing role when used together. Furthermore, they screened for fusion proteins based on the SeqCas9 protein, which also exhibited highly efficient base editing activity, thus realizing the present invention.

[0006] In summary, in a first aspect of the present invention, a fusion protein is provided, comprising: a) A SeqCas9 mutant protein, based on the SeqCas9 protein shown in SEQ ID NO: 1, including mutations that cause the SeqCas9 protein to lose its endonuclease activity or mutations that cause the SeqCas9 protein to have only single-stranded DNA cleavage activity; b) Other proteins or polypeptides; and c) Optional first linker for connecting the SeqCas9 mutant protein to the other protein or polypeptide.

[0007] In a second aspect, the present invention provides a conjugate comprising: a) The fusion protein described in the first aspect; b) Detectable markers; and c) An optional seventh connector for connecting the fusion protein to the detectable marker.

[0008] In a third aspect, the present invention provides a single-stranded guide RNA, wherein the single-stranded guide RNA comprises a guide sequence and a scaffold sequence from the 5' end to the 3' end, and the scaffold sequence is as follows: a) The stent sequence shown in SEQ ID NO: 11; b) A sequence obtained by modifying the sequence shown in SEQ ID NO: 11 while retaining its biological activity.

[0009] In a fourth aspect, the present invention provides an isolated nucleic acid molecule comprising a nucleic acid sequence encoding the fusion protein described in the first aspect or the conjugate described in the second aspect.

[0010] In a fifth aspect, the present invention provides an isolated nucleic acid molecule comprising a nucleic acid sequence encoding the single-stranded guide RNA described in the third aspect.

[0011] In a sixth aspect, the present invention provides a vector comprising a nucleic acid sequence encoding a fusion protein as described in the first aspect or a conjugate as described in the second aspect.

[0012] In a seventh aspect, the present invention provides a vector comprising a nucleic acid sequence encoding the single-stranded guide RNA described in the third aspect.

[0013] In an eighth aspect, the present invention provides a CRISPR / Cas9 gene editing system comprising: 1) A protein component comprising: the SeqCas9 protein with the amino acid sequence as shown in SEQ ID NO: 1, or a conjugate or fusion protein thereof, the fusion protein described in the first aspect, or the conjugate described in the second aspect; and 2) The single-stranded guide RNA described in the third aspect; Furthermore, the protein component and the single-stranded guide RNA bind to each other to form a complex; The conjugate of the SeqCas9 protein comprises: the SeqCas9 protein, a detectable label or a combination of a detectable label and another protein or polypeptide, and optionally an eighth linker for connecting the fusion protein to the detectable label or the combination. The fusion protein of the SeqCas9 protein comprises: the SeqCas9 protein, another protein or polypeptide, and optionally a ninth linker for connecting the fusion protein to the other protein or polypeptide.

[0014] In a ninth aspect, the present invention provides a method for gene editing of a target sequence in an intracellular or in vitro environment, the method comprising: contacting any one of the following (1) to (4) with the target sequence in the intracellular or in vitro environment: (1) The SeqCas9 protein or its conjugate or fusion protein with the amino acid sequence shown in SEQ ID NO: 1, the fusion protein or conjugate described in the first aspect or the second aspect; and the single-stranded guide RNA described in the third aspect; (2) The isolated nucleic acid molecules described in the fourth aspect and, as appropriate, the isolated nucleic acid molecules described in the fifth aspect; (3) The carrier described in the sixth aspect and, as appropriate, the carrier described in the seventh aspect; (4) The CRISPR / Cas9 gene editing system described in aspect eight; The conjugate of the SeqCas9 protein comprises: the SeqCas9 protein, a detectable label or a combination of a detectable label and another protein or polypeptide, and optionally an eighth linker for connecting the SeqCas9 protein to the detectable label or the combination. The fusion protein of the SeqCas9 protein comprises: the SeqCas9 protein, another protein or polypeptide, and optionally a ninth linker for connecting the SeqCas9 protein to the other protein or polypeptide. The SeqCas9 protein or its conjugates or fusion proteins, the fusion protein described in the first aspect or the conjugates described in the second aspect, recognize a protospacer adjacent motif (PAM) located at the 3' end of the target sequence and having the sequence 5'-NNG.

[0015] In a tenth aspect, the present invention provides a cell comprising: the isolated nucleic acid molecule described in the fourth or fifth aspect, or the carrier described in the sixth or seventh aspect.

[0016] In an eleventh aspect, the present invention provides a kit for gene editing of target sequences in intracellular or in vitro environments, comprising: a) Choose any one of (1) to (4) below: (1) The SeqCas9 protein or its conjugate or fusion protein with the amino acid sequence shown in SEQ ID NO: 1, the fusion protein or conjugate described in the first aspect or the second aspect; and the single-stranded guide RNA described in the third aspect; (2) The isolated nucleic acid molecules described in the fourth aspect, and the isolated nucleic acid molecules described in the fifth aspect as appropriate; (3) The carrier described in the sixth aspect, and the carrier described in the seventh aspect as appropriate; (4) The CRISPR / Cas9 gene editing system described in aspect eight; as well as b) Instructions on how to perform gene editing on target sequences in the intracellular or in vitro environment; The conjugate of the SeqCas9 protein comprises: the SeqCas9 protein, a detectable label or a combination of a detectable label and another protein or polypeptide, and optionally an eighth linker for connecting the SeqCas9 protein to the detectable label or the combination. The fusion protein of the SeqCas9 protein comprises: the SeqCas9 protein, another protein or polypeptide, and optionally a ninth linker for connecting the SeqCas9 protein to the other protein or polypeptide.

[0017] The inventors have developed a CRISPR / Cas9 gene editing tool suitable for efficient gene editing in various eukaryotic cell environments. This CRISPR / Cas9 gene editing tool uses the SeqCas9 protein and its corresponding sgRNA, exhibiting high specificity, high editing activity, and simple PAM (Programming Analysis Method). Furthermore, the inventors' experiments have demonstrated that the SeqCas9 protein's editing efficiency at random sites is comparable to that of the SpCas9-HF1 protein, and its editing efficiency at multiple genomic sites is superior to SpCas9-NG, making it more suitable for gene editing development and application research. Further, the inventors have discovered that when developing a base editor based on the SeqCas9 protein, the prepared base editor can effectively achieve single-base conversion. This invention expands the scope of gene editing and has broad application prospects in the field of gene editing. Attached Figure Description

[0018] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the accompanying drawings used in the description of the specific embodiments or the prior art will be briefly introduced below.

[0019] Figure 1 A PAM logo diagram is shown for recognizing PAM sequences by a CRISPR / SeqCas9 gene editing system according to an embodiment of the present invention.

[0020] Figure 2 The results of gene editing efficiency of the CRISPR / SeqCas9 gene editing system according to an embodiment of the present invention after gene editing at six target sites are shown.

[0021] Figure 3 The results show the specificity detection of the CRISPR / SeqCas9 gene editing system according to an embodiment of the present invention in the GFP reporter system HEK293T cell line.

[0022] Figure 4 The results of base editing efficiency of the CRISPR / SeqABE gene editing system according to an embodiment of the present invention after base editing at six target sites are shown.

[0023] Figure 5 The editing efficiency results of the CRISPR / SeqCBE gene editing system according to an embodiment of the present invention after base editing at six target sites are shown. Detailed Implementation

[0024] The present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the following description is merely illustrative and is not intended to limit the scope of the invention; the scope of protection of the invention is defined by the appended claims. Furthermore, those skilled in the art will understand that modifications can be made to the technical solutions of the present invention without departing from its spirit and intent. Unless otherwise specified, the technical means used in the embodiments are conventional means well known to those skilled in the art.

[0025] definition Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the subject matter pertains. Before a detailed description of the invention, the following definitions are provided to better understand the invention.

[0026] In cases where numerical ranges are provided, such as concentration ranges, percentage ranges, or ratio ranges, it should be understood that, unless the context explicitly specifies otherwise, all intermediate values ​​between the upper and lower limits of the range, up to one-tenth of the lower limit unit, and any other values ​​or intermediate values ​​within the range are included in the subject matter. The upper and lower limits of these smaller ranges may be independently included in the smaller ranges, and such embodiments are also included in the subject matter, limited by any specific excluded limit values ​​within the range. Where the range includes one or two limit values, the range excluding any one or both of those included limit values ​​is also included in the subject matter.

[0027] In the context of this invention, many embodiments use the expressions "comprising," "including," or "basically / mainly composed of." The expressions "comprising," "including," or "basically / mainly composed of" are generally understood as open-ended expressions, indicating that they include not only the elements, components, parts, or method steps specifically listed after the expression, but also other elements, components, parts, or method steps. However, in this document, the expressions "comprising," "including," or "basically / mainly composed of" can also be understood as closed-ended expressions in certain cases, indicating that they only include the elements, components, parts, or method steps specifically listed after the expression, and do not include any other elements, components, parts, or method steps. In this case, the expression is equivalent to the expression "composed of."

[0028] To better understand this teaching and without limiting its scope, all figures and other numerical values ​​used in the specification and claims to express quantities, percentages, or proportions should, in all cases, be understood to be modified by the term "about." Therefore, unless otherwise stated, the numerical parameters set forth in the following specification and appended claims are approximate values ​​that may vary depending on the desired properties sought. At a minimum, each numerical parameter should be interpreted based at least on the reported significant figures and by applying common rounding techniques.

[0029] As used herein, the terms "Cas9 protein," "Cas9," and "Cas" are used interchangeably to refer to RNA-guided nucleases, including the Cas9 protein or its functionally active fragments. The Cas9 protein is a protein component of the CRISPR / Cas9 genome editing system that, guided by single-stranded guide RNA (sgRNA), targets and cleaves DNA target sequences, forming DNA double-strand breaks (DSBs). DNA double-strand breaks can activate the cell's inherent repair mechanisms of nonhomologous end-joining (NHEJ) and homologous recombination (HR), thereby repairing DNA damage in the cell. During the repair process, the specific DNA sequence is edited at specific sites.

[0030] The terms “single-stranded guide RNA,” “sgRNA (single guided RNA),” or “mature crRNA” as used herein are used interchangeably and have the meanings commonly understood by those skilled in the art. Generally, a single-stranded guide RNA may comprise a scaffold sequence and a guide sequence, which is also referred to herein as guide RNA (or gRNA). In the context of an endogenous CRISPR system, the guide sequence is also referred to as a spacer sequence. In some cases, the guide sequence is any polynucleotide sequence that is sufficiently complementary to a target sequence to hybridize with said target sequence and guide the specific binding of the CRISPR / Cas9 complex to said target sequence. In some embodiments, the complementarity between the guide sequence and its corresponding target sequence is at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99% when optimal alignment is achieved. Determining optimal alignment is within the capabilities of those skilled in the art. For example, there are publicly available and commercially available comparison algorithms and programs, such as, but not limited to, ClustalW, the Smith-Waterman algorithm in MATLAB, Bowtie, Geneious, Biopython, and SeqMan.

[0031] As used herein, the term "CRISPR / Cas9 complex" refers to a complex formed by the binding of a single-stranded guide RNA or mature crRNA:tracrRNA hybrid to the Cas9 protein or its mutant form, comprising a guide sequence that hybridizes to a target sequence and thereby enables the Cas9 protein or its mutant form to bind to said target sequence. This complex is capable of recognizing and cleaving polynucleotides that hybridize with the single-stranded guide RNA or mature crRNA.

[0032] Therefore, in the formation of the CRISPR / Cas9 complex, the "target sequence" refers to a polynucleotide targeted by a guide sequence designed to be targeted, such as a sequence complementary to the guide sequence, where hybridization between the target sequence and the guide sequence will promote the formation of the CRISPR / Cas9 complex. Perfect complementarity is not required, as long as sufficient complementarity exists to induce hybridization and promote the formation of the CRISPR / Cas9 complex. The target sequence can include any polynucleotide, such as DNA. In some cases, the target sequence is located in the cell nucleus or cytoplasm. In other cases, the target sequence may be located in an organelle of a eukaryotic cell, such as a mitochondrion or chloroplast.

[0033] As used herein, the term "target sequence" or "target polynucleotide" can refer to any endogenous or exogenous polynucleotide for a cell (e.g., a eukaryotic cell). For example, the target polynucleotide can be a polynucleotide present in the nucleus of a eukaryotic cell. The target polynucleotide can be a sequence encoding a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide or useless DNA). In some cases, the target sequence should be associated with a protospacer adjacent motif (PAM). The precise sequence and length requirements for the PAM vary depending on the Cas protein used, but the PAM is typically a 2-5 base sequence adjacent to the protospacer sequence (target sequence). Those skilled in the art can identify the PAM sequence used with a given Cas protein.

[0034] The terms “polynucleotide,” “nucleic acid sequence,” “nucleotide sequence,” or “nucleic acid fragment” used herein are used interchangeably and are single-stranded or double-stranded RNA or DNA polymers, optionally containing synthetic, non-natural, or modified nucleotide bases. Nucleotides are designated by their single-letter names as follows: “A” for adenosine or deoxyadenosine (corresponding to RNA or DNA, respectively), “C” for cytidine or deoxycytidine, “G” for guanosine or deoxyguanosine, “U” for uridine, “T” for deoxythymidine, “R” for purine (A or G), “Y” for pyrimidine (C or T), “K” for G or T, “H” for A, C, or T, “I” for inosine, and “N” for any nucleotide.

[0035] In the sequences used in this invention, degenerate bases are sometimes used to represent bases at one or more positions. Degenerate bases can be represented by the letters R, Y, M, K, S, W, H, B, V, D, and N, where R represents A / G, Y represents C / T, M represents A / C, K represents G / T, S represents C / G, W represents A / T, H represents A / T / C, B represents G / T / C, V represents G / A / C, D represents G / A / T, and N represents A / T / C / G.

[0036] The terms “polypeptide,” “peptide,” and “protein” as used herein are used interchangeably to refer to polymers of amino acid residues. The term applies to amino acid polymers in which one or more amino acid residues are artificial chemical analogs of the corresponding naturally occurring amino acids, and also to naturally occurring amino acid polymers. The terms “polypeptide,” “peptide,” “amino acid sequence,” and “protein” may also include modified forms, including but not limited to glycosylation, lipid linkage, sulfation, γ-carboxylation, hydroxylation, and ADP-ribosylation of glutamate residues.

[0037] As used herein, the term "vector" refers to a nucleic acid delivery vehicle into which polynucleotides can be inserted. A vector is called an expression vector when it enables the expression of a protein encoded by the inserted polynucleotide, or when it enables transcription of the inserted polynucleotide (e.g., to generate mRNA or functional RNA). Vectors can be introduced into host cells through transformation, transduction, or transfection, allowing the genetic material they carry to be expressed in the host cells. Vectors are well-known to those skilled in the art and include, but are not limited to, plasmid vectors and viral vectors. Vectors may also contain various regulatory sequences that regulate expression. The terms "regulatory sequence" and "regulatory element" are used interchangeably herein, referring to a nucleotide sequence located upstream (5' non-coding sequence), midway, or downstream (3' non-coding sequence) of a coding sequence that affects transcription, RNA processing, or stability or translation of the relevant coding sequence. Regulatory sequences may include, but are not limited to, promoter sequences, transcription initiation sequences, enhancer sequences, selection elements, and reporter genes. These regulatory sequences may originate from different sources or from the same source but arranged in a manner different from what is typically found naturally. Additionally, vectors may contain a replication initiation site.

[0038] "Introducing" nucleic acid molecules (such as plasmids, linear nucleic acid fragments, RNA, etc.) or proteins into an organism refers to transforming the cells of an organism with the nucleic acid or protein, enabling the nucleic acid or protein to function within the cell. The term "transformation" as used in this invention includes both stable transformation and transient transformation.

[0039] The terms “identity,” “consistency,” or “homology” used in this article have the generally accepted meanings in the art, and the percentage of sequence identity between two nucleic acid or polypeptide molecules or regions can be calculated using publicly available techniques. Sequence identity can be measured along the full length of a polynucleotide or polypeptide or along a region of the molecule (see, for example, Computational Molecular Biology, Lesk, AM, ed., Oxford University Press, New York, 1988; Biocomputing: Informatics and Genome Projects, Smith, DW, ed., Academic Press, New York, 1993; Computer Analysis of Sequence Data, Part I, Griffin, AM, and Griffin, HG, eds., Humana Press, New Jersey, 1994; Sequence Analysis in Molecular Biology, von Heinje, G., Academic Press, 1987; and Sequence Analysis Primer, Gribskov, M. and Devereux, J., eds., M Stockton Press, New York, 1991). While many methods exist for measuring the identity between two polynucleotides or peptides, the term "identity" is known to those skilled in the art to refer to conserved amino acid substitutions in peptides or proteins that can generally be performed without altering the biological activity of the resulting molecule. Typically, those skilled in the art recognize that a single amino acid substitution in a non-essential region of a peptide does not substantially alter its biological activity (see, for example, Watson et al., Molecular Biology of the Gene, 4th Edition, 1987, The Benjamin / Cummings Pub. co., p. 224).

[0040] As used herein, the term "complementarity" refers to the ability of one nucleic acid sequence to form one or more hydrogen bonds with another nucleic acid sequence via conventional Watson-Crick or other non-conventional types. The complementarity percentage indicates the percentage of residues in one nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with another nucleic acid sequence (e.g., 50%, 60%, 70%, 80%, 90%, and 100% complementarity out of 10). "Complete complementarity" means that all consecutive residues in one nucleic acid sequence form hydrogen bonds with the same number of consecutive residues in another nucleic acid sequence. As used herein, “substantially complementary” refers to a complementarity of at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% in a region having 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50 or more nucleotides, or to two nucleic acids hybridizing under stringent conditions.

[0041] As used in this paper, the hybridization-related term "strict condition" refers to conditions under which a nucleic acid complementary to a target sequence hybridizes primarily with that target sequence and substantially does not hybridize to non-target sequences. Strict conditions are typically sequence-dependent and depend on many factors. Generally, the longer the sequence, the higher the temperature at which it specifically hybridizes to its target sequence. A non-limiting example of a strict condition is described in Tijssen, 1993, *Laboratory Techniques in Biochemistry and Molecular Biology Hybridization With Nucleic Acid Probes*, Section I, Chapter II, "Overview of principles of hybridization and the strategy of nucleic acid probe assay", Elsevier, New York.

[0042] As used herein, the term "hybridization" refers to a reaction in which one or more polynucleotides react to form a complex that is stabilized by hydrogen bonds between the bases of these nucleotide residues. Hydrogen bonds can occur via Watson-Crick base pairing, Hoogstein binding, or any other sequence-specific mechanism. The complex can consist of two strands forming a duplex, three or more strands forming a multi-stranded complex, a single self-hybridizing strand, or any combination thereof. Hybridization can be a step in a broader process, such as the initiation of PCR or the cleavage of a polynucleotide by an enzyme. A sequence capable of hybridizing with a given sequence is called the "complement" of that given sequence.

[0043] Cas9 protein and its fusion protein In this invention, the "biological activity" of the Cas9 protein or its fusion protein refers to the protein's activity in binding to single-stranded guide RNA, its endonuclease activity (including single-stranded cleavage activity and double-stranded cleavage activity), and / or its activity in binding to and cleaving a specific site of a target sequence under the guidance of guide RNA (gRNA), but is not limited thereto.

[0044] As previously stated, the inventors have discovered that the SeqCas9 protein, with an amino acid sequence as shown in SEQ ID NO: 1, can exert a gene editing effect when used in conjunction with single-stranded guide RNA.

[0045] Furthermore, the inventors have further developed an adenine base editor (ABE) and a cytosine base editor (CBE) based on the SeqCas9 protein, which can be used to efficiently perform single base conversions.

[0046] The core components of ABE are nCas9 (D10A) and adenine deaminase (TadA), which together form a fusion protein. The specific working principle of ABE is as follows: When the fusion protein targets genomic DNA under the guidance of sgRNA, adenine deaminase binds to ssDNA, deaminating adenine (A) within a certain range into inosine (I). I is then read and replicated at the DNA level as G, ultimately achieving a direct substitution of A•T base pairs for G•C base pairs.

[0047] The core components of CBE are nCas9 (Cas9 nickase) or dCas9 (dead Cas9) and cytosine deaminase (APOBEC). The Cas9 protein and cytosine deaminase form a fusion protein. The specific working principle of CBE is as follows: When the fusion protein targets genomic DNA under the guidance of sgRNA, cytosine deaminase can bind to the ssDNA of the R-loop region formed by the Cas9 protein, sgRNA, and genomic DNA, deaminate cytosine (C) in a certain range on the ssDNA to uracil (U), and then convert U to thymine (T) through DNA replication or repair, ultimately achieving a direct substitution of C•G base pairs to T•A base pairs.

[0048] Therefore, in a first aspect of the invention, a fusion protein is provided, comprising: a) SeqCas9 mutant protein, based on the amino acid sequence of the SeqCas9 protein shown in SEQ ID NO: 1 and including mutations that cause the SeqCas9 protein to lose its endonuclease activity or mutations that cause the SeqCas9 protein to have only single-stranded DNA cleavage activity; b) Other proteins or peptides; and c) An optional first linker for connecting the SeqCas9 mutant protein to the other protein or polypeptide.

[0049] In a preferred embodiment, the mutation that renders the SeqCas9 protein inactive by nucleases includes mutations D10A and H847A based on the amino acid sequence shown in SEQ ID NO: 1 (i.e., the amino acid sequence of the SeqCas9 protein). In this document, such a SeqCas9 mutant protein may also be referred to as the dSeqCas9 (dead SeqCas9) protein.

[0050] In a preferred embodiment, the mutation that renders the SeqCas9 protein capable of single-stranded DNA cleavage only comprises the mutation D10A based on the amino acid sequence shown in SEQ ID NO: 1. In this document, such a SeqCas9 mutant protein may also be referred to as nSeqCas9 (SeqCas9 nickase) protein.

[0051] In one specific embodiment, the additional protein or polypeptide is selected from one or more of the following: epitope tags, reporter proteins or nuclear localization signal (NLS) sequences, cytosine deaminases, adenine deaminases, cytosine methyltransferases DNMT3A and MQ1, cytosine demethylase Tet1, transcription activators VP64, p65 and RTA, transcription repressors KRAB, histone acetyltransferase p300, histone deacetyltransferase LSD1, Gam protein, and endonuclease FokI.

[0052] In a preferred embodiment, the additional protein or polypeptide is a cytosine deaminase, adenine deaminase, or a dimer of adenine deaminase, wherein the amino acid sequence of the cytosine deaminase is as shown in SEQ ID NO: 4, the amino acid sequence of the adenine deaminase is as shown in SEQ ID NO: 2, and the dimer of the adenine deaminase comprises an adenine deaminase with an amino acid sequence as shown in SEQ ID NO: 2 and an artificially directed evolved adenine deaminase (evolved TadA) with an amino acid sequence as shown in SEQ ID NO: 3, linked by an optional third linker.

[0053] When the fusion protein contains cytosine deaminase, the complex formed by the fusion protein and sgRNA may also be referred to herein as a cytosine base editor (CBE). When the fusion protein contains adenine deaminase, the complex formed by the fusion protein and sgRNA may also be referred to herein as an adenine base editor (ABE).

[0054] In a preferred embodiment, the fusion protein may further comprise a nuclear localization signal (NLS) sequence. The purpose of introducing an NLS sequence into the fusion protein is to increase the expression level of the cytosine base editor within the cell, thereby improving its editing efficiency or increasing the substitution efficiency of the adenine base editor for base A.

[0055] In a preferred embodiment, the fusion protein comprises, from its N-terminus to its C-terminus, an optional nuclear localization signal (NLS) sequence, an optional second adapter, the adenine deaminase or its dimer or cytosine deaminase, the optional first adapter, the SeqCas9 mutant protein, an optional fourth adapter, and an optional nuclear localization signal (NLS) sequence. In a specific embodiment, the amino acid sequence of the nuclear localization signal (NLS) is as shown in SEQ ID NO: 7.

[0056] In one embodiment, the fusion protein further comprises a Gam protein derived from phage Mu, which binds to the end of the DSB to prevent its degradation, thereby inhibiting the NHEJ repair pathway.

[0057] In a further preferred embodiment, the amino acid sequence of the fusion protein is shown in SEQ ID NO: 5, SEQ ID NO: 6 or SEQ ID NO: 9.

[0058] Those skilled in the art will know that uracil DNA glycosylase (UDG) exists in cells. UDG can recognize U•G mismatches and cleave the glycosidic bond between uracil and the phosphate backbone, reversing U to C through the intracellular base excision repair (BER) pathway.

[0059] Therefore, in yet another specific embodiment, the fusion protein comprises cytosine deaminase and one or two uracil DNA glycosylase inhibitors (UGIs), such as uracil DNA glycosylase inhibitors (UGIs) with amino acid sequences as shown in SEQ ID NO: 8.

[0060] In one specific embodiment, the fusion protein comprises, from its N-terminus to its C-terminus, the following in sequence: an optional nuclear localization signal (NLS) sequence, an optional second adapter, the cytosine deaminase (APOBEC), the optional first adapter, the SeqCas9 mutant protein, an optional fifth adapter, the uracil DNA glycosylase inhibitor (UGI), an optional fourth adapter, and an optional nuclear localization signal (NLS) sequence.

[0061] By introducing a uracil DNA glycosylation inhibitor (UGI) into the fusion protein, the CBE of the present invention can utilize UGI to inhibit the repair function of UDG in cells and improve base editing efficiency.

[0062] In addition to its repair function of reversing U to C as described above, UDG can also excise U to form a pyrimidine-free site (AP). Through the action of trans-damage synthesis (TLS) polymerase and DNA replication, there is also a certain probability that C will be converted to other bases. Furthermore, since the formed AP site will produce a gap under AP lyase or spontaneous cleavage, it will form a DSB with the gap generated by nCas9 in the non-edited strand. After passing through the NHEJ repair pathway, it will produce Indels products.

[0063] Therefore, in a preferred embodiment, the fusion protein comprises two uracil DNA glycosylase inhibitors (UGIs), which are optionally linked by a sixth connector. In this case, the fusion protein, from its N-terminus to its C-terminus, comprises, in sequence: an optional nuclear localization signal (NLS) sequence, an optional second connector, the cytosine deaminase, the optional first connector, the SeqCas9 mutant protein, an optional fifth connector, the uracil DNA glycosylase inhibitor (UGI), an optional sixth connector, the uracil DNA glycosylase inhibitor (UGI), an optional fourth connector, and an optional nuclear localization signal (NLS) sequence.

[0064] By introducing two UGIs into the fusion protein, the CBE of the present invention can reduce unnecessary editing products, improve base editing efficiency, and reduce the frequency of C to A or G conversion.

[0065] In a further preferred embodiment, the amino acid sequence of the fusion protein is shown in SEQ ID NO: 10.

[0066] It is understood that multiple protein moieties in the SeqCas9 fusion protein of the present invention may or may not be linked by a linker. Linkers are well known in the art, and examples include, but are not limited to, linkers containing 1-50, for example 10-32 amino acids (such as Glu or Ser) or amino acid derivatives (such as Ahx, β-Ala, GABA or Ava), or PEG linkers.

[0067] Therefore, in one embodiment, the first to sixth connectors can each independently be a connector of 1-50, for example 10-32 amino acids (such as Glu or Ser) or amino acid derivatives (such as Ahx, β-Ala, GABA or Ava), or a PEG connector.

[0068] Derivatized proteins The fusion protein of the first aspect can be further derivatized, for example, by linking it to another molecule (e.g., a detectable marker). Generally, protein derivatization does not adversely affect the protein's desired activity (e.g., activity binding to single-stranded guide RNA, endonuclease activity, activity of binding to and cleaving a target sequence at a specific site guided by the guide RNA). Therefore, the fusion protein of the present invention is also intended to include such derivatized forms. For example, the fusion protein of the present invention can be functionally linked (by chemical coupling, gene fusion, non-covalent linkage, or other means) to one or more other molecular moieties, such as a detectable marker, pharmaceutical reagent, etc.

[0069] Specifically, the fusion protein of the first aspect can be linked to other functional units. For example, it can be linked to a targeting moiety to make the fusion protein targeted. For example, it can be linked to a detectable tag to facilitate the detection of the fusion protein. For example, it can be linked to an epitope tag to facilitate the expression, detection, tracing, and / or purification of the fusion protein.

[0070] Therefore, in a second aspect, the present invention provides a conjugate comprising: a) The fusion protein described in the first aspect; b) Detectable markers; and c) An optional seventh linker for connecting the fusion protein to the detectable tag, for example, a seventh linker with a length of 1-50 amino acids.

[0071] Detectable markers are well known to those skilled in the art, and examples include fluorescent dyes such as fluorescein isothiocyanate (FITC) or DAPI.

[0072] The fusion protein of the present invention can be linked to the detectable marker via a linker, or it can be linked directly to the detectable marker without a linker. Linkers are well known in the art, and examples include, but are not limited to, linkers containing 1-50 amino acids (such as Glu or Ser) or amino acid derivatives (such as Ahx, β-Ala, GABA or Ava), or PEG linkers, etc.

[0073] Single-stranded guide RNA As described in the background section, in the type II CRISPR-Cas9 system, the guide RNA used to guide the Cas9 protein to function consists of two parts: crRNA and tracrRNA. In engineered applications, crRNA and tracrRNA are often linked to form a single-stranded guide RNA (sgRNA). The crRNA originates from the CRISPR array and contains a target-specific spacer sequence and a sequence derived from the repeat region. The tracrRNA contains an anti-repeat region complementary to the repeat-derived sequence in the crRNA, along with its subsequent structurally constant sequence. This part primarily mediates the binding of the Cas9 protein and maintains the stability of the RNA-protein complex. Therefore, the complete sequence of engineered sgRNA is composed of crRNA-related sequences (including sequences derived from the repeat region) and tracrRNA-related sequences (including the anti-repeat region).

[0074] In this paper, the term "guide sequence," also known as the spacer sequence, specifically refers to the target-related sequence in crRNA responsible for complementary pairing with the target DNA sequence for precise targeting. The term "scaffold sequence" refers to the region in sgRNA other than the guide sequence, including sequences derived from the repeat region in crRNA and tracrRNA sequences. The repeat and anti-repeat regions pair bases within the sgRNA to form a stem-loop structure, thereby constructing the RNA secondary structure framework required for Cas9 recognition.

[0075] Therefore, in a third aspect, the present invention provides a single-stranded guide RNA, wherein the single-stranded guide RNA comprises a guide sequence and a scaffold sequence from the 5' end to the 3' end, the scaffold sequence being: a) A stent sequence as shown in SEQ ID NO: 11; or b) A sequence obtained by modifying the sequence shown in SEQ ID NO: 11 while retaining its biological activity.

[0076] In one specific implementation, the modification is one or more of the following: base phosphorylation, base sulfidation, base methylation, base hydroxylation, sequence shortening, and sequence lengthening.

[0077] In a further specific embodiment, the shortening of the sequence and the lengthening of the sequence include the deletion or addition of one, two, three, four, five, six, seven, eight, nine, or ten bases relative to the base sequence.

[0078] In yet another specific implementation, the guide sequence is a sequence of 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides in length that is complementary to the target sequence.

[0079] In a preferred embodiment, the guide sequence is a 20-nucleotide sequence that is complementary to the target sequence.

[0080] In a further embodiment, the single-stranded guide RNA further includes a terminator at the 3' end of the scaffold sequence. As an example, the terminator may be a plurality of terminators, such as at least six (e.g., seven or eight) U.

[0081] The single-stranded guide RNA can bind to the SeqCas9 protein or its conjugate or fusion protein, the first fusion protein or the second conjugate, as shown in SEQ ID NO: 1, to form a complex. This complex can recognize the corresponding PAM (e.g., 5'-NNG) and thereby bind to the target sequence, thereby achieving the cleavage of the target sequence or gene editing.

[0082] Nucleic acid encoding and vectors In a fourth aspect, the present invention provides an isolated nucleic acid molecule comprising: a nucleic acid sequence encoding a fusion protein of the first aspect or a conjugate of the second aspect.

[0083] In one specific implementation, the isolated nucleic acid molecule comprises the nucleic acid sequence shown in SEQ ID NO: 14 or 15 or a degenerate sequence thereof.

[0084] In yet another specific implementation, the isolated nucleic acid molecule further comprises a nucleic acid sequence encoding a single-stranded guide RNA for the third aspect.

[0085] In one specific implementation, the isolated nucleic acid molecule further comprises the nucleic acid sequence shown in SEQ ID NO: 12.

[0086] In the fifth aspect, the isolated nucleic acid molecule contains a nucleic acid sequence encoding a single-stranded guide RNA of the third aspect.

[0087] In one specific implementation, the isolated nucleic acid molecule comprises the nucleic acid sequence shown in SEQ ID NO: 12.

[0088] After the isolated nucleic acid molecules of the present invention are transfected into the corresponding cells using certain tools known in the art, such as expression vectors, the isolated nucleic acid molecules of the present invention can express the fusion protein of the first aspect or the conjugate of the second aspect, and / or the single-stranded guide RNA described above, and perform the corresponding functions, such as gene editing.

[0089] In addition, the isolated nucleic acid molecules of the present invention can express the fusion protein of the first aspect or the conjugate of the second aspect, as well as the single-stranded guide RNA, individually or separately, or can express the expression product in one go, depending on the specific circumstances.

[0090] Furthermore, the expressed product has the corresponding effects and / or functions described above, which will not be repeated here for the sake of brevity.

[0091] In a sixth aspect, the present invention provides a vector comprising a nucleic acid sequence encoding a fusion protein of the first aspect or a conjugate of the second aspect.

[0092] In one specific implementation, the vector comprises the nucleic acid sequence shown in SEQ ID NO: 14 or 15 or a degenerate sequence thereof.

[0093] In yet another specific implementation, the vector further comprises a nucleic acid sequence encoding a single-stranded guide RNA for a third aspect.

[0094] In yet another specific implementation, the vector contains the nucleic acid sequence shown in SEQ ID NO: 12.

[0095] In a further specific embodiment, the vector may be an expression vector, such as a plasmid vector, such as the pUC19 vector, an applicator vector, the pAAV2_ITR vector, a retroviral vector, a lentiviral vector, an adenovirus vector, or an adeno-associated virus vector.

[0096] In a seventh aspect, the present invention provides a vector comprising a nucleic acid sequence encoding a single-stranded guide RNA of the third aspect.

[0097] In one specific implementation, the vector contains the nucleic acid sequence shown in SEQ ID NO: 12.

[0098] As described above, after the vector of the present invention is transfected into cells, the nucleic acid sequence cloned in the vector can be expressed as a fusion protein of the first aspect or a conjugate of the second aspect, and / or the single-stranded guide RNA described above, and perform corresponding functions, such as gene editing.

[0099] Alternatively, multiple vectors, such as two vectors, can be transfected into cells. One vector expresses the fusion protein of the first aspect or the conjugate of the second aspect, while the other vector expresses single-stranded guide RNA. Subsequently, the expressed fusion protein of the first aspect or the conjugate of the second aspect complexes with the expressed single-stranded guide RNA to form a complex, which then performs its corresponding function, such as gene editing.

[0100] Alternatively, the nucleic acid sequence encoding the fusion protein of the first aspect or the conjugate of the second aspect and the nucleic acid sequence encoding the single-stranded guide RNA can be cloned into a vector, so that after the vector is transfected into cells, it expresses both the fusion protein of the first aspect or the conjugate of the second aspect and the single-stranded guide RNA, and performs the corresponding functions, such as gene editing.

[0101] CRISPR / Cas9 gene editing system In an eighth aspect, the present invention provides a CRISPR / Cas9 gene editing system comprising: 1) A protein component comprising: the SeqCas9 protein or its conjugate or fusion protein with the amino acid sequence as shown in SEQ ID NO: 1, the fusion protein of the first aspect, or the conjugate of the second aspect; and 2) The third aspect: single-stranded guide RNA; Furthermore, the protein component and the single-stranded guide RNA bind to each other to form a complex; The conjugate of the SeqCas9 protein comprises: the SeqCas9 protein, a detectable label or a combination of a detectable label and another protein or polypeptide, and optionally an eighth linker for connecting the fusion protein to the detectable label or the combination. The fusion protein of the SeqCas9 protein comprises: the SeqCas9 protein, another protein or polypeptide, and optionally a ninth linker for connecting the SeqCas9 protein to the other protein or polypeptide.

[0102] It is understandable that, in addition to the SeqCas9 protein itself, the SeqCas9 protein can also be bound to other proteins, detectable tags, or combinations thereof, thereby giving the protein other functions.

[0103] Detectable markers are well known to those skilled in the art, and examples include fluorescent dyes such as fluorescein isothiocyanate (FITC) or DAPI.

[0104] In one specific embodiment, the additional protein or polypeptide may be selected from one or more of the following: epitope tags, reporter proteins or nuclear localization signal sequences, cytosine deaminases, adenine deaminases, cytosine methyltransferases DNMT3A and MQ1, cytosine demethylase Tet1, transcription activators VP64, p65 and RTA, transcription repressors KRAB, histone acetyltransferase p300, histone deacetyltransferase LSD1, Gam protein, and endonuclease FokI.

[0105] Epitope tags are well known to those skilled in the art, and examples include, but are not limited to, His, V5, FLAG, HA, Myc, VSV-G, Trx, etc., and those skilled in the art know how to select an appropriate epitope tag according to the desired purpose (e.g., purification, detection, or tracing).

[0106] Reporter proteins are well known to those skilled in the art, and examples include, but are not limited to, GST, HRP, CAT, GFP, HcRed, DsRed, CFP, YFP, and BFP.

[0107] The SeqCas9 protein of the present invention can be linked to the detectable marker, the other protein or polypeptide, or a combination of the other protein or polypeptide and the detectable marker via a linker, or it can be linked directly to the detectable marker, the other protein or polypeptide, or a combination of the other protein or polypeptide and the detectable marker without a linker. Linkers are well known in the art, and examples include, but are not limited to, linkers containing 1-50 amino acids (such as Glu or Ser) or amino acid derivatives (such as Ahx, β-Ala, GABA, or Ava), or PEG linkers. In one embodiment, the eighth and ninth linkers are each independently a linker having 1-50, for example, 10-32 amino acids.

[0108] The CRISPR / Cas9 gene editing system of the present invention can be directly constructed from the SeqCas9 protein or its conjugates or fusion proteins, the fusion protein of the first aspect or the conjugate of the second aspect, and single-stranded guide RNA, as shown in SEQ ID NO:1. Alternatively, it can be constructed from the isolated nucleic acid molecules or vectors described herein, which can subsequently express the fusion protein, conjugates, and single-stranded guide RNA, thereby exerting gene editing effects. The CRISPR / Cas9 gene editing system of the present invention achieves the recognition, localization, cleavage, and gene editing of target sequences through the combined action of the SeqCas9 protein or its conjugates or fusion proteins, the fusion protein of the first aspect or the conjugate of the second aspect, and single-stranded guide RNA.

[0109] The CRISPR / Cas9 gene editing system of this invention can precisely locate target sequences. "Precisely located" has two meanings: first, the CRISPR / Cas9 gene editing system of this invention can recognize and bind to the target sequence; second, the CRISPR / Cas9 gene editing system of this invention can bring other proteins fused with the SeqCas9 protein or its conjugates or fusion proteins, the fusion protein of the first aspect or the conjugate of the second aspect, or proteins that specifically recognize the sgRNA, to the location of the target sequence.

[0110] The CRISPR / Cas9 gene editing system of the present invention has low tolerance for non-target sequences. In this document, "low tolerance" means that the CRISPR / Cas9 gene editing system of the present invention is substantially or completely unable to recognize and bind to non-target sequences, or substantially or completely unable to bring other proteins fused with the SeqCas9 protein or its conjugates or fusion proteins, the fusion protein of the first aspect or the conjugate of the second aspect, or proteins that specifically recognize the sgRNA to the location of the non-target sequence.

[0111] Gene editing methods In a ninth aspect, the present invention provides a method for gene editing of a target sequence in an intracellular or in vitro environment, the method comprising: contacting any one of the following (1) to (4) with the target sequence in the intracellular or in vitro environment: (1) The SeqCas9 protein or its conjugate or fusion protein, the fusion protein of the first aspect or the conjugate of the second aspect, as shown in SEQ ID NO: 1, and the single-stranded guide RNA of the third aspect; (2) The fourth aspect of isolated nucleic acid molecules and, as appropriate, the fifth aspect of isolated nucleic acid molecules; (3) The carrier of the sixth aspect, and the carrier of the seventh aspect, depending on the circumstances; (4) The eighth aspect: the CRISPR / Cas9 gene editing system; The conjugate of the SeqCas9 protein comprises: the SeqCas9 protein, a detectable label or a combination of a detectable label and another protein or polypeptide, and optionally an eighth linker for connecting the SeqCas9 protein to the detectable label or the combination. The fusion protein of the SeqCas9 protein comprises: the SeqCas9 protein, another protein or polypeptide, and optionally a ninth linker for connecting the SeqCas9 protein to the other protein or polypeptide. The SeqCas9 protein or its conjugates or fusion proteins, the first aspect of the fusion protein or the second aspect of the conjugate recognizes a protospacer adjacent motif (PAM) located at the 3' end of the target sequence and having a 5'-NNG sequence.

[0112] The descriptions of the SeqCas9 protein or its conjugates or fusion proteins, detectable markers, other proteins or peptides, epitope tags, reporter proteins, adapters, etc., given above also apply to this aspect of the present invention, and will not be repeated here.

[0113] In one specific implementation, the cell is a prokaryotic cell or a eukaryotic cell, such as an animal cell, and the animal cell is, for example, a mammalian cell, such as a human cell.

[0114] In yet another specific implementation, the gene editing includes one or more of the following: gene knockout of a target sequence, site-specific base alteration, site-specific insertion, regulation of gene transcription, regulation of DNA methylation, DNA acetylation modification, histone acetylation modification, single base conversion, and chromatin imaging tracking.

[0115] In one exemplary embodiment, the single base conversion includes one or more of the following: conversion from guanine G to adenine A, conversion from adenine A to guanine G, conversion from cytosine C to thymine T, conversion from cytosine C to uracil U, and conversion from thymine T to cytosine C.

[0116] Regarding items (2) and (3) above, as can be seen from the above description, the isolated nucleic acid molecule of the fourth aspect and the vector of the sixth aspect of the present invention, in some cases, only contain nucleic acid sequences encoding the SeqCas9 protein or its conjugates or fusion proteins, the fusion protein of the first aspect or the conjugate of the second aspect; in other cases, they contain nucleic acid sequences encoding the SeqCas9 protein or its conjugates or fusion proteins, the fusion protein of the first aspect or the conjugate of the second aspect, and a nucleic acid sequence encoding a single-stranded guide RNA. Therefore, depending on the circumstances, the isolated nucleic acid molecule or the vector may further contain a nucleic acid sequence encoding a single-stranded guide RNA.

[0117] In yet another specific implementation, the guide sequence of the single-stranded guide RNA forms a completely complementary base pairing structure with the target sequence, while forming an incomplete complementary base pairing structure with the non-target sequence.

[0118] In this document, the incomplete base pairing structure refers to a structure that includes a portion of base pairing and a portion of non-base pairing, wherein the non-base pairing includes, for example, base mismatch and / or base bulge.

[0119] In a further specific embodiment, the incomplete base complementary pairing structure includes one or more, for example, two or more base mismatches.

[0120] In a more specific embodiment, the eighth and ninth connectors are each independently a connector having 1-50, for example 10-32 amino acids.

[0121] Therefore, the SeqCas9 protein or its conjugates or fusion proteins of the present invention, the fusion protein of the first aspect or the conjugate of the second aspect, can cleave the target site on the target sequence, and under the cleavage action of the protein, a double-strand break (DSB) occurs in the target sequence. Furthermore, when the method is performed intracellularly, the cleaved target sequence can be repaired through intracellular non-homologous end joining repair or homologous recombination repair pathways, thereby achieving gene editing of the target sequence. In addition, the fusion protein of the first aspect or the conjugate of the second aspect of the present invention may also include nCas9 (Cas9 nickase, D10A) with only single-stranded DNA cleavage activity or dCas9 (catalytically deadCas9) without endonuclease activity. nCas9 with single-stranded DNA cleavage activity only cleaves the single-stranded DNA complementary to the sgRNA, without cleaving the other strand. Under the action of adenine deaminase, adenine (A) on the uncut DNA single strand is converted to hypoxanthine (H) within a certain range. Hypoxanthine (H) is then recognized as guanine (G) during replication, resulting in a base substitution from adenine (A) to guanine (G). Subsequently, the cut DNA single strand uses the uncut DNA single strand as a template for homology-directed repair (HDR), achieving a base substitution from thymine (T) to cytosine (C) on the complementary strand. Similarly, under the action of cytosine deaminase, cytosine (C) on the uncut DNA single strand is converted to uracil (U) within a certain range. Uracil (U) is then recognized as thymine (T) during replication, resulting in a base substitution from cytosine (C) to thymine (T). Next, the cut DNA single strand will use the uncut single strand as a template to repair itself through homologous recombination, achieving the base substitution of guanine (G) to adenine (A) on the complementary single strand.

[0122] The CRISPR / Cas9 gene editing system and gene editing method using this invention have been experimentally shown to have an editing efficiency of 25%-84%. Furthermore, the CRISPR / Cas9 gene editing system exhibits a very low mismatch rate for guide RNAs at positions 4-20. Therefore, the CRISPR / Cas9 gene editing system can edit target genes with high specificity, characterized by high editing efficiency and low off-target rate, and can be widely applied to gene editing in cells or in vitro environments.

[0123] cell In an eleventh aspect, the present invention provides a cell comprising: isolated nucleic acid molecules as described in the fourth or fifth aspect, or a carrier as described in the sixth or seventh aspect.

[0124] As an example, the cell can be a prokaryotic cell or a eukaryotic cell, such as an animal cell. For the animal cell, as an example, it can be a mammalian cell, such as a human cell.

[0125] Reagent test kit In a twelfth aspect, the present invention provides a kit for gene editing of target sequences in intracellular or in vitro environments, comprising: a) Choose any one of (1) to (4) below: (1) The SeqCas9 protein or its conjugate or fusion protein, the fusion protein of the first aspect or the conjugate of the second aspect, as shown in SEQ ID NO: 1, and the single-stranded guide RNA of the third aspect; (2) The fourth aspect of isolated nucleic acid molecules, and, depending on the circumstances, the fifth aspect of isolated nucleic acid molecules; (3) The carrier of the sixth aspect, and the carrier of the seventh aspect, depending on the circumstances; (4) The eighth aspect: the CRISPR / Cas9 gene editing system; as well as b) Instructions on how to perform gene editing on target sequences in the intracellular or in vitro environment; The conjugate of the SeqCas9 protein comprises: the SeqCas9 protein, a detectable label or a combination of a detectable label and another protein or polypeptide, and optionally an eighth linker for connecting the SeqCas9 protein to the detectable label or the combination. The fusion protein of the SeqCas9 protein comprises: the SeqCas9 protein, another protein or polypeptide, and optionally a ninth linker for connecting the SeqCas9 protein to the other protein or polypeptide.

[0126] The descriptions of the SeqCas9 protein or its conjugates or fusion proteins, detectable markers, other proteins or peptides, epitope tags, reporter proteins, adapters, etc., given above also apply to this aspect of the present invention, and will not be repeated here.

[0127] Of course, those skilled in the art will understand that the kit of the present invention may also contain other reagents that facilitate gene editing.

[0128] A brief description of the sequence involved in this invention. SEQ ID NO: 1: SeqCas9 protein sequence SEQ ID NO: 2: Amino acid sequence of WT TadA SEQ ID NO: 3: Amino acid sequence of Evolved TadA SEQ ID NO: 4: Amino acid sequence of APOBEC SEQ ID NO: 5: Amino acid sequence of APOBEC-Linker-nSeqCas9 SEQ ID NO: 6: Amino acid sequence of TadA-Linker-Evolved TadA-Linker-nSeqCas9 SEQ ID NO: 7: Amino acid sequence of NLS SEQ ID NO: 8: Amino acid sequence of UGI SEQ ID NO: 9: Amino acid sequence of NLS-TadA-Linker-Evolved TadA-Linker-nSeqCas9-Linker-NLS SEQ ID NO: 10: Amino acid sequence of NLS-APOBEC-Linker-nSeqCas9-Linker-UGI-Linker-UGI-Linker-NLS SEQ ID NO: 11: Sequence of a single-stranded guide RNA scaffold used in conjunction with the SeqCas9 protein SEQ ID NO: 12: DNA sequence of a single-stranded guide RNA scaffold sequence used in conjunction with the SeqCas9 protein. SEQ ID NO: 13: The coding sequence of the SeqCas9 protein SEQ ID NO: 14: Encoded sequence of NLS-TadA-Linker-Evolved TadA-Linker-nSeqCas9-Linker-NLS SEQ ID NO: 15: Encoded sequence of NLS-APOBEC-Linker-nSeqCas9-Linker-UGI-Linker-UGI-Linker-NLS Example The invention will now be described with reference to the following embodiments, which are intended to be illustrative and not limiting. Those skilled in the art will appreciate that the embodiments provided herein are for the purpose of describing the invention in detail only and are not intended to limit the scope of protection claimed by the invention.

[0129] Unless otherwise specified, the experiments and methods described in the examples were generally performed according to conventional methods well known in the art and described in the various references. Furthermore, for conditions not specifically specified in the examples, conventional conditions or conditions recommended by the manufacturer were followed. Reagents or instruments whose manufacturers are not specified are all commercially available conventional products.

[0130] Example 1 (1) Constructing plasmid pAAV2_Cas9_ITR Download the amino acid sequence of the SeqCas9 protein from NCBI (SEQ ID NO: 1, its gene search number is NZ_UHFK01000003).

[0131] The nucleic acid sequence encoding the SeqCas9 protein was codon-optimized to obtain the gene sequence of the Cas9 protein that is highly expressed in human cells, as shown in SEQ ID NO: 13.

[0132] The gene sequence shown in SEQ ID NO: 13 obtained above was used for gene synthesis and constructed into the pSpCas9(BB)-2A-Puro (PX459) backbone plasmid V2.0 (Addgene platform, catalog#62988) to obtain plasmid pAAV2_Cas9_ITR.

[0133] (2) Constructing the plasmid SeqCas9-PSK-mU6-sgRNA-scaffold The pSKB plasmid (available commercially from the Addgene platform, catalog #62540) was digested with ClaI and XhoI restriction endonucleases. The digestion system consisted of 1 μg pSKB plasmid, 5 μL 10×rCutSmart buffer (from NEB), 1 μL ClaI and 1 μL XhoI restriction endonucleases (from NEB), and water to a final volume of 50 μL. The digestion was incubated at 37°C for 1 hour.

[0134] Then, the enzyme digestion products were electrophoresed on a 1% agarose gel at 120V for 30 minutes.

[0135] The target size DNA fragment was cut from the agarose gel and recovered using a gel recovery kit (Tiangen Biotech (Beijing) Co., Ltd., DP209) according to the manufacturer's instructions. Finally, it was eluted with ultrapure water.

[0136] The DNA sequences of the mU6 sequence and the single-stranded guide RNA scaffold sequence (SEQ ID NO: 12) were synthesized and constructed on a linearized pSKB backbone to obtain the plasmid SeqCas9-PSK-mU6-sgRNA-scaffold.

[0137] (3) Construction of plasmid pAAV2_Cas9-mU6-sgRNA-scaffold_ITR vector The pAAV2_Cas9_ITR plasmid expressing Cas9 protein in (1) and the SeqCas9-PSK-mU6-sgRNA-scaffold plasmid expressing sgRNA in (2) were linearized using PCR.

[0138] For the pAAV2_Cas9_ITR plasmid, the primer sequences are as follows: ctagtccgtttttagcgcg; and acatgtgagcaaaaggcca; For the SeqCas9-PSK-mU6-sgRNA-scaffold plasmid, the primer sequences are: ctggccttttgctcacatgtGATCCGACGCGCCATCT; and acgcgctaaaaacggactagAAAATGCACCCGAATCGGGTGCC.

[0139] The reaction system is as follows:

[0140] The PCR procedure is as follows:

[0141] The PCR products were electrophoresed on a 1% agarose gel at 120V for 30 min. The target DNA fragment was purified using a gel extraction kit according to the manufacturer's instructions. The DNA concentration was measured using a NanoDrop™ Lite spectrophotometer (Thermo Scientific) and stored for later use or at -20 ℃ for long-term storage.

[0142] The linearized pAAV2_Cas9_ITR fragment and the linearized SeqCas9-PSK-mU6-sgRNA-scaffold fragment were homologously recombinated according to the ratio specified in the instructions. The homologous recombinase used was NEBuilder® High Fidelity DNA Assembly Premix (NEB). The reaction system is as follows:

[0143] The reaction conditions are as follows:

[0144] The ligation product was added to Escherichia coli DH5α competent cells (purchased from Shanghai Weidi Biotechnology Co., Ltd.), incubated on ice for 30 min, heat-shocked at 42℃ for 1 min, incubated on ice for 2 min, and then 900 μL of LB medium was added and cultured at 37℃ for 1 hour to activate and reactivate Escherichia coli DH5α competent cells.

[0145] The revived Escherichia coli DH5α competent cells were plated on LB solid plates containing ampicillin resistance and incubated upside down in a 37°C incubator. The resulting Escherichia coli DH5α monoclonal cells were verified by Sanger sequencing.

[0146] Sequencing verified the correct ligation of the E. coli DH5α clone was used for culture, and the plasmid was extracted to obtain the plasmid pAAV2_Cas9-mU6-sgRNA-scaffold_ITR, which was then used for later use.

[0147] (4) Preparation of linearized plasmid pAAV2_Cas9-mU6-sgRNA-scaffold_ITR The plasmid pAAV2_Cas9-mU6-sgRNA-scaffold_ITR prepared in (3) was digested with BbsI restriction endonuclease. The digestion system consisted of 1 μg of plasmid pAAV2_Cas9-mU6-sgRNA-scaffold_ITR, 5 μL of 10×CutSmart buffer (purchased from NEB), 1 μL of BbsI restriction endonuclease (purchased from NEB), and water to a final volume of 50 μL. The digestion system was incubated at 37°C for 1 hour.

[0148] Then, the enzyme digestion products were electrophoresed on a 1% agarose gel at 120V for 30 minutes.

[0149] DNA fragments were excised from agarose gels and recovered using a gel extraction kit (Tiangen Biotech (Beijing) Co., Ltd., DP209) according to the manufacturer's instructions. The fragments were then eluted with ultrapure water. The DNA fragment was the linearized plasmid pAAV2_Cas9-mU6-sgRNA-scaffold_ITR, containing the encoding genes for the SeqCas9 protein and sgRNA scaffold, with a size of 9214 bp.

[0150] The DNA concentration of the recovered linearized plasmid pAAV2_Cas9-mU6-sgRNA-scaffold_ITR was determined using a NanoDrop™ Lite spectrophotometer (Thermo Scientific) for later use or for long-term storage at -20 °C.

[0151] (5) Preparation of plasmid pAAV2_Cas9-mU6-sgRNA_ITR The gRNA sequence was designed, and sticky end sequences (indicated by uppercase letters) corresponding to the flanking sides of the linearized plasmid pAAV2_Cas9-mU6-sgRNA-scaffold_ITR were added to its sense and antisense strands, respectively. The two oligonucleotide single-stranded DNA sequences were then synthesized, and their specific sequences are shown below: gRNA: gctcggagatcatcattgcg Oligo-F:TTTGgctcggagatcatcattgcgGT Oligo-R: TAAAACcgcaatgatgatctccgagc.

[0152] Oligonucleotide single-stranded DNA was annealed to obtain double-stranded DNA. The annealing reaction system consisted of 3 μL 10 μM oligo-F, 3 μL 10 μM oligo-R, and 4 μL water. After vortexing and mixing the annealing system, it was placed in a PCR instrument and the annealing program was run as follows: 95 ℃ for 5 min, 85 ℃ for 1 min, 75 ℃ for 1 min, 65 ℃ for 1 min, 55 ℃ for 1 min, 45 ℃ for 1 min, 35 ℃ for 1 min, 25 ℃ for 1 min, and stored at 4 ℃ with a cooling rate of 0.3 ℃ / s. After annealing, the obtained product was ligated into the linearized pAAV2_Cas9-mU6-sgRNA-scaffold_ITR plasmid obtained in step (4) using DNA ligase (purchased from NEB).

[0153] 1 μL of the obtained ligation product was added to Escherichia coli DH5α competent cells (purchased from Shanghai Weidi Biotechnology Co., Ltd.), incubated on ice for 30 min, heat-shocked at 42℃ for 1 min, incubated on ice for 2 min, and then 900 μL of LB medium was added and cultured at 37℃ for 1 hour to activate and reactivate Escherichia coli DH5α competent cells.

[0154] The revived Escherichia coli DH5α competent cells were plated on LB solid plates containing the corresponding antibiotics and incubated upside down in a 37°C incubator. The resulting Escherichia coli DH5α monoclonal cells were verified by Sanger sequencing.

[0155] Sequencing verified the correct ligation of the E. coli DH5α clone was used for culture, and the plasmid was extracted to obtain the plasmid pAAV2_Cas9-mU6-sgRNA_ITR containing the target sgRNA sequence, which was then used for later use.

[0156] (6) Transfection of HEK293T cell line library containing target sequences by plasmid pAAV2_Cas9-mU6-sgRNA_ITR expressing Cas protein and sgRNA The HEK293T cell line library containing the target sequence of the GFP reporter system was obtained as follows: a 25 bp protospacer (as the target sequence) and a 7 bp random sequence (as the PAM sequence) were inserted between the start codon ATG and the GFP coding sequence, resulting in a frameshift mutation that prevents GFP expression. This GFP gene containing the inserted fragment was started using a CMV promoter and constructed into a lentiviral expression vector. This sequence was then randomly inserted into the genome of HEK293T cells via lentivirus-mediated insertion, creating a stable GFP reporter cell line library. When the target sequence was cut using a gene editing system, the cells' self-repair system caused some cells to recover the GFP reading frame, producing green fluorescence. Flow cytometry analysis was used to statistically analyze the percentage of GFP-positive cells, which allowed for the assessment of the editing capability and specificity of the gene editing system.

[0157] The transfection process includes the following steps: On day 0, the HEK293T cell line library containing the target sequence of the GFP reporter system was plated in a 10cm dish as required for transfection, with the cell density controlled at 30%.

[0158] The HEK293T cell line library containing the target sequence of the GFP reporter system contains the nucleotide sequence CMV-ATG-target site-PAM-GFP, where the PAM sequence is a 7bp random sequence and the target site sequence is GACGGCTCGGAGATCATCATTGCG.

[0159] Day 1, transfection was performed. The transfection process is as follows: Take 2 μg of the plasmid pAAV2_Cas9-mU6-sgRNA_ITR to be transfected and add it to 100 μL of Opti-MEM medium (purchased from Gibco). Gently pipette to mix.

[0160] Gently mix Lipofectamine® 2000 (purchased from Invitrogen) or PEI (purchased from Polysciences), then add 5 μL of Lipofectamine® 2000 or PEI to 100 μL of Opti-MEM medium, mix gently, and let stand at room temperature for 5 min.

[0161] The diluted plasmid and diluted transfection reagent were mixed and gently pipetted to mix. The resulting mixture was allowed to stand at room temperature for 20 min, and then added to the culture medium of the HEK293T cell line library containing the target sequence of the GFP reporter system. The mixture was then placed in a 37°C, 5% CO2 incubator for further culture.

[0162] Five days after culture, the editing of target genes in the HEK293T cell line library by the CRISPR / SeqCas9 system was observed under a fluorescence microscope. Under the microscope, cells transfected with the CRISPR / SeqCas9 system produced green fluorescence, indicating successful editing of the target genes. Subsequently, cells expressing GFP fluorescence were sorted using flow cytometry for further enrichment and growth.

[0163] (7) Preparation of next-generation sequencing libraries The enriched HEK293T cell library was collected, and genomic DNA was extracted using a DNA kit (Tiangen Biotech (Beijing) Co., Ltd., DP304) according to the instructions provided with the DNA kit.

[0164] The first round of PCR for library preparation was performed using a 2×Q5 Mastermix. The PCR primers are shown below: F1 primer: ACACTCTTTCCCTACACGACGCTCTTCCGATCTNNNNgcgagaaaagccttgttt; R1 primer: ACTGGAGTTCAGACGTGTGCTCTTCCGATCTNNNNctgaacttgtggccgtttac.

[0165] The reaction system is as follows:

[0166] The PCR procedure is as follows:

[0167] For the second round of PCR for sequencing library preparation, a 2xQ5 Mastermix was used. The PCR primers are shown below: F2 primer: AATGATACGGCGACCACCGAGATCTACACNNNNNNNNACACTCTTTCCCTACACGAC; R2 primer: CAAGCAGAAGACGGCATACGAGATNNNNNNNNGTGACTGGAGTTCAGACGTGTG.

[0168] The reaction system is as follows:

[0169] The PCR procedure is as follows:

[0170] The second-round PCR products were purified into DNA fragments of the target size using a gel extraction kit following the manufacturer's instructions, and the next-generation sequencing library was prepared.

[0171] (8) Analysis of second-generation sequencing results The prepared next-generation sequencing library was subjected to paired-end sequencing on the high-throughput sequencer HiseqXTen (illumina).

[0172] Based on the results obtained from next-generation sequencing, the PAM sequences of the HEK293T cell library were analyzed, and a PAM logo diagram was drawn, as shown below. Figure 1 As shown, the PAM sequence recognized by the SeqCas9 protein is NNG, which is very simple, indicating that the SeqCas9 protein is advantageous for use in cellular gene editing.

[0173] Example 2 (1)-(4): Same as steps (1)-(4) of Example 1, to prepare linearized plasmid pAAV2_Cas9-mU6-sgRNA-scaffold_ITR.

[0174] (5) Preparation of plasmid pAAV2_Cas9-mU6-sgRNA_ITR Each gRNA was designed, and its sequence is shown in Table 1. Sticky end sequences (indicated by uppercase letters) corresponding to the linearized plasmid pAAV2Cas9-mU6-sgRNA_ITR were added to the sense and antisense strands, respectively. These oligonucleotide single-stranded DNAs were then synthesized, and their specific sequences are also shown in Table 1 below.

[0175] Table 1. Sequences of gRNA and oligonucleotide single-stranded DNA

[0176] Oligonucleotide single-stranded DNA was annealed to obtain double-stranded DNA. The annealing reaction system consisted of 3 μL 10 μM oligo-F, 3 μL 10 μM oligo-R, and 4 μL water. After vortexing and mixing the annealing system, it was placed in a PCR instrument and the annealing program was run as follows: 95℃_5 min, 85℃_1 min, 75℃_1 min, 65℃_1 min, 55℃_1 min, 45℃_1 min, 35℃_1 min, 25℃_1 min, and stored at 4℃ with a cooling rate of 0.3℃ / s. After annealing, the obtained product was ligated into the linearized pAAV2_Cas9-mU6-sgRNA_scaffold_ITR plasmid obtained in step (2) using DNA ligase (purchased from NEB).

[0177] 1 μL of the obtained ligation product was added to Escherichia coli DH5α competent cells (purchased from Shanghai Weidi Biotechnology Co., Ltd.), incubated on ice for 30 min, heat-shocked at 42℃ for 1 min, incubated on ice for 2 min, and then 900 μL of LB medium was added and cultured at 37℃ for 1 hour to activate and reactivate Escherichia coli DH5α competent cells.

[0178] The revived Escherichia coli DH5α competent cells were plated on LB solid plates containing the corresponding antibiotics and incubated upside down in a 37°C incubator. The resulting Escherichia coli DH5α monoclonal cells were verified by Sanger sequencing.

[0179] Sequencing verified the correct ligation of the E. coli DH5α clone was used for culture, and the plasmid was extracted to obtain the plasmid pAAV2_Cas9-mU6-sgRNA_ITR containing the target sgRNA sequence, which was then used for later use.

[0180] (6) Transfection of HEK293T cell line with plasmid pAAV2_Cas9-mU6-sgRNA_ITR expressing Cas protein and sgRNA On day 0, HEK293T cells containing the target sequence were seeded in 6-well plates as needed for transfection, with a cell density of approximately 30%.

[0181] Day 1, transfection was performed. The transfection process is as follows: Take 2 μg of the plasmid pAAV2_Cas9-mU6-sgRNA_ITR to be transfected and add it to 100 μL of Opti-MEM medium (purchased from Gibco). Gently pipette to mix.

[0182] Gently mix the transfection reagents Lipofectamine® 2000 (purchased from Invitrogen) or polyethyleneimine (hereinafter referred to as PEI, purchased from Polysciences), then add 5 μL of Lipofectamine® 2000 or PEI to 100 μL of Opti-MEM medium (purchased from Gibco), gently mix, and let stand at room temperature for 5 min.

[0183] Mix the diluted transfection reagent and diluted plasmid, gently pipette to mix, let stand at room temperature for 20 min, then add to the culture medium containing HEK293T cells to be transfected, and then place the cells in a 37℃, 5% CO2 incubator for 3 days.

[0184] (7) Preparation of next-generation sequencing libraries HEK293T cells were collected three days after editing, and genomic DNA was extracted using a DNA kit (Tiangen Biotech (Beijing) Co., Ltd., DP304) according to the instructions provided with the DNA kit.

[0185] The first round of PCR for library preparation was performed using a 2×Q5 Mastermix. The PCR primers are shown in Table 2. Table 2. List of primers for the first round of PCR in next-generation sequencing

[0186] The reaction system is as follows:

[0187] The PCR procedure is as follows:

[0188] For the second round of PCR for sequencing library preparation, a 2×Q5 Mastermix was used. The PCR primers are shown below: F2 primer: AATGATACGGCGACCACCGAGATCTACACNNNNNNNNACACTCTTTCCCTACACGAC; R2 primer: CAAGCAGAAGACGGCATACGAGATNNNNNNNNGTGACTGGAGTTCAGACGTGTG.

[0189] The reaction system is as follows:

[0190] The PCR procedure is as follows:

[0191] The second-round PCR products were purified into DNA fragments of the target size using a gel extraction kit following the manufacturer's instructions, and the next-generation sequencing library was prepared.

[0192] (8) Analysis of second-generation sequencing results The prepared next-generation sequencing library was subjected to paired-end sequencing on the high-throughput sequencer HiseqXTen (illumina).

[0193] The editing efficiency of six target sites obtained by next-generation sequencing calculations is as follows: Figure 2 As shown in the figure, the X-axis represents the target site, and the Y-axis represents the editing efficiency (Indels%). The figure shows that the gene editing system containing SeqCas9 has an editing efficiency of at least 20%, and the highest can reach at least 80%, indicating that the gene editing system of this invention can be effectively used for cell gene editing.

[0194] Example 3 (1)-(4): Same as steps (1)-(4) of Example 1, to prepare linearized plasmid pAAV2_Cas9-mU6-sgRNA-scaffold_ITR.

[0195] (5) Preparation of plasmid pAAV2_Cas9-mU6-on target sgRNA or pAAV2_Cas9-mU6-mismatch-sgRNA_ITR The sequences of each on-target gRNA and mismatch gRNA were designed, and their corresponding oligonucleotide single-stranded DNA are shown in Table 3. The mismatch bases are shown as bold bases with underlined lines in the sequence listing.

[0196] Table 3. Oligonucleotide single-stranded DNA corresponding to on target gRNA and mismatch gRNA

[0197] The oligonucleotide single-stranded DNA corresponding to the obtained on-target gRNA and the oligonucleotide single-stranded DNA corresponding to different mismatch gRNAs were annealed separately. The annealing reaction system consisted of 3 μL 10 μM oligo-F, 3 μL 10 μM oligo-R, and 4 μL water. After vortexing and mixing the annealing system, it was placed in a PCR instrument and the annealing program was run as follows: 95℃_5 min, 85℃_1 min, 75℃_1 min, 65℃_1 min, 55℃_1 min, 45℃_1 min, 35℃_1 min, 25℃_1 min, and stored at 4℃ with a cooling rate of 0.3℃ / s. After annealing, the obtained products were ligated into the obtained linearized pAAV2_Cas9-mU6-sgRNA_scaffold_ITR plasmid using DNA ligase (purchased from NEB).

[0198] 1 μL of the obtained ligation product was added to Escherichia coli DH5α competent cells (purchased from Shanghai Weidi Biotechnology Co., Ltd.), incubated on ice for 30 min, heat-shocked at 42℃ for 1 min, incubated on ice for 2 min, and then 900 μL of LB medium was added. The cells were then cultured at 37℃ for 1 h to activate and reactivate Escherichia coli DH5α competent cells.

[0199] The revived Escherichia coli DH5α competent cells were plated on LB solid plates containing the corresponding antibiotics and incubated upside down in a 37°C incubator. The resulting Escherichia coli DH5α monoclonal cells were verified by Sanger sequencing.

[0200] Sequencing verified the correct ligation of the E. coli DH5α clone was used for culture, and plasmids were extracted to obtain plasmids pAAV2_Cas9-mU6-on target-sgRNA_ITR expressing the above on target gRNA sequence and plasmids pAAV2_Cas9-mU6-mismatch-sgRNA_ITR expressing the above different mismatch gRNA sequences, which were then used for later use.

[0201] (6) Transfection of HEK293T cell line with plasmids pAAV2_Cas9-mU6-on target-sgRNA_ITR expressing Cas protein and on-target-sgRNA and plasmid pAAV2_Cas9-mU6-mismatch-sgRNA_ITR expressing Cas protein and mismatch-sgRNA. The plasmids pAAV2_Cas9-mU6-ontarget-sgRNA_ITR, expressing the Cas protein and on-target gRNA sequence, and pAAV2_Cas9-mU6-mismatch-sgRNA_ITR, expressing the Cas protein and mismatch gRNA sequence, were transfected separately into the HEK293T cell line containing the target sequence using liposomes. The transfection process included the following steps: On day 0, the HEK293T cell line containing the target sequence of the GFP reporter system was seeded in 6-well plates as required for transfection, with the cell density controlled at 30%.

[0202] The HEK293T cell line, containing the target sequence of the GFP reporter system, contains the nucleotide sequence CMV-ATG-target site-PAM-GFP, where the PAM sequence is GAG, and the target site sequence is: TCATCATTGCGCTGGATCGT (see [link to relevant documentation]). Figure 3 ).

[0203] Day 1, transfection was performed. The transfection process is as follows: Take 2 μg of the plasmid pAAV2_Cas9-mU6-on target-sgRNA_ITR or 2 μg of the plasmid pAAV2_Cas9-mU6-mismatch-sgRNA_ITR and add it to 100 μL of Opti-MEM medium (purchased from Gibco). Gently pipette to mix.

[0204] Gently mix Lipofectamine® 2000 (purchased from Invitrogen) or PEI (purchased from Polysciences), then add 5 μL of Lipofectamine® 2000 or PEI to 100 μL of Opti-MEM medium, mix gently, and let stand at room temperature for 5 min.

[0205] The diluted plasmid and diluted transfection reagent were mixed and gently pipetted to mix. The resulting mixture was allowed to stand at room temperature for 20 min, and then added to the culture medium of the HEK293T cell line containing the target sequence of the GFP reporter system. The cell line was then placed in a 37°C, 5% CO2 incubator for further culture.

[0206] Flow cytometry was used to analyze the editing efficiency and off-target rate of the CRISPR gene editing system of this invention on the target sequence.

[0207] Specifically, HEK293T cell lines cultured in a CO2 incubator for 3 days were collected, and their specificity was detected using a flow cytometer (BDBiosciences FACSCalibur). The GFP positivity rate was analyzed and plotted using FlowJo analysis software.

[0208] The specificity detection results of the CRISPR / SeqCas9 gene editing system of the present invention in the HEK293T cell line containing the target sequence of the GFP reporter system are shown in the figure. Figure 3 The top bar shows a schematic diagram of the GFP reporter system. A specific target sequence and PAM sequence are inserted between the start codon ATG and the GFP coding sequence, causing a GFP frameshift mutation. Therefore, when the gene editing system cuts the target sequence, the cell's own repair system allows some cells to recover the GFP reading frame, producing green fluorescence. Figure 3 In the bar chart below, the Y-axis represents the percentage of GFP-positive cells (%), and the X-axis represents the oligonucleotide single-stranded DNA sequences corresponding to on-target gRNA and mismatch gRNA. Figure 3 As can be seen, the CRISPR / SeqCas9 gene editing system of this invention edited all target sites in the HEK293T cell line of the GFP reporter system. Furthermore, the proportion of gene editing mediated by mismatch gRNA was significantly lower than that mediated by on-target gRNA, indicating that the CRISPR / SeqCas9 gene editing system of this invention has high editing activity, low off-target rate, and high specificity. Moreover, in the research results on the CRISPR / SeqCas9 gene editing system, no obvious mismatches were found in the 4-20 bp dibase mismatches, indicating that the CRISPR / SeqCas9 gene editing system has extremely high requirements for perfect pairing between gRNA and target sequence, exhibiting low error tolerance and high safety in practical applications.

[0209] Example 4 (1) Preparation of linearized plasmid ABEmax PCR was performed using ABEmax plasmid (Addgene platform, catalog # 112101) as a template. The primer sequences were as follows: Primer 1: TGTCTAAGTTGGGCGAAGAAtctggcggctcaaaaagaac Primer 2: CCGATACTGTATGTCTTTTCtgaccccccgctgctgc The reaction system is as follows:

[0210] The PCR procedure is as follows:

[0211] The PCR product was electrophoresed on a 1% agarose gel at 120V for 30 min. The DNA fragment of 5526bp was purified using a gel extraction kit according to the manufacturer's instructions. The DNA concentration was measured using a NanoDrop™ Lite spectrophotometer (Thermo Scientific) and stored for later use or at -20℃ for long-term storage.

[0212] (2) Preparation of plasmid pAAV2_tadA(ABEmax)-SeqCas9_ITR The linearized ABEmax backbone fragment and the synthetically produced humanized SeqCas9 fragment (SEQ ID NO: 13) were homologously recombinated according to the instructions. The homologous recombinase used was NEBuilder® high-fidelity DNA assembly premix (NEB). The reaction system is as follows:

[0213] The reaction conditions are as follows:

[0214] The ligation product was added to Escherichia coli DH5α competent cells (purchased from Shanghai Weidi Biotechnology Co., Ltd.), incubated on ice for 30 min, heat-shocked at 42℃ for 1 min, incubated on ice for 2 min, and then 900 μL of LB medium was added and cultured at 37℃ for 1 hour to activate and reactivate Escherichia coli DH5α competent cells.

[0215] The revived Escherichia coli DH5α competent cells were plated on LB solid plates containing ampicillin resistance and incubated upside down in a 37°C incubator. The resulting Escherichia coli DH5α monoclonal cells were verified by Sanger sequencing.

[0216] Sequencing was used to verify the correct ligation of the E. coli DH5α clone, and the plasmid was extracted to obtain the plasmid pAAV2_tadA(ABEmax)-SeqCas9_ITR, which was then used for later use.

[0217] (3) Preparation of plasmid pAAV2_tadA(ABEmax)-nSeqCas9_ITR A circular PCR reaction was performed using pAAV2_tadA(ABEmax)-SeqCas9_ITR as a template. The primer sequences were as follows: Primer 3: GCCATCGGCACCAATTCTGT Primer 4: CAGGCCGATACTGTATGTCT The reaction system is as follows:

[0218] The PCR procedure is as follows:

[0219] PCR products were electrophoresed on a 1% agarose gel at 120V for 30 min. Following the manufacturer's instructions, the DNA fragments were purified using a gel extraction kit to obtain the target size. DNA concentration was measured using a NanoDrop™ Lite spectrophotometer (ThermoScientific). The fragments were then treated with T4 polynucleotide kinase (T4 PNK) and T4 DNA ligase to introduce the D10A mutation in nSeqCas9. The reaction system is as follows:

[0220] The reaction conditions are as follows:

[0221] Add 1 μL of T4 DNA ligase (purchased from NEB) to the reaction system, vortex to mix, and incubate at room temperature for 2 h.

[0222] The ligation product was added to Escherichia coli DH5α competent cells (purchased from Shanghai Weidi Biotechnology Co., Ltd.), incubated on ice for 30 min, heat-shocked at 42℃ for 1 min, incubated on ice for 2 min, and then 900 μL of LB medium was added and cultured at 37℃ for 1 hour to activate and reactivate Escherichia coli DH5α competent cells.

[0223] The revived Escherichia coli DH5α competent cells were plated on LB solid plates containing ampicillin resistance and incubated upside down in a 37°C incubator. The resulting Escherichia coli DH5α monoclonal cells were verified by Sanger sequencing.

[0224] Sequencing verified the correctly ligated E. coli DH5α clone was shaken to extract the plasmid, resulting in plasmid pAAV2_tadA(ABEmax)-nSeqCas9_ITR, which was then used for later use.

[0225] (4) Linearization preparation of pAAV2_tadA(ABEmax)-nSeqCas9_ITR The pAAV2_tadA(ABEmax)-nSeqCas9_ITR plasmid was digested using MfeI restriction endonuclease (purchased from NEB). The reaction mixture consisted of 2 μg of plasmid pAAV2_tadA(ABEmax)-nSeqCas9_ITR, 5 μL of 10×CutSmart buffer (purchased from NEB), 1 μL of MfeI restriction endonuclease (purchased from NEB), and water to a final volume of 50 μL. The digestion was carried out at 37°C for 2 hours.

[0226] Then, the enzyme digestion products were electrophoresed on a 1% agarose gel at 120V for 30 minutes.

[0227] DNA fragments were excised from the agarose gel and recovered using a gel recovery kit (Tiangen Biotech (Beijing) Co., Ltd., DP209) according to the manufacturer's instructions. Finally, the fragments were eluted with ultrapure water.

[0228] The DNA concentration of the recovered linearized fragment pAAV2_tadA(ABEmax)-nSeqCas9_ITR was determined using a NanoDrop™ Lite spectrophotometer (Thermo Scientific) for later use or for long-term storage at -20°C.

[0229] (5) Preparation of pAAV2_tadA(ABEmax)-nSeqCas9-sgRNA-scaffold_ITR plasmid PCR was performed using SeqCas9-PSK-mU6-sgRNA-scaffold (prepared according to Example 1(2)) as a template. The primer sequences were as follows: Primer 5: caaggcaaggcttgaccgacGATCCGACGCGCCATCTC Primer 6: agcagattcttcatgcaattAAAATGCACCCGAATCGGG The reaction system is as follows:

[0230] The PCR procedure is as follows:

[0231] The PCR products were electrophoresed on a 1.5% agarose gel at 120V for 30 min. The 455bp SeqCas9 sgRNA DNA fragment was purified using a gel extraction kit according to the manufacturer's instructions. The DNA concentration was measured using a NanoDrop™ Lite spectrophotometer (Thermo Scientific) and stored for later use or at -20℃ for long-term storage.

[0232] The linearized pAAV2_tadA(ABEmax)-nSeqCas9_ITR fragment and the SeqCas9 sgRNA DNA fragment were homologously recombinated according to the ratio specified in the manufacturer's instructions. The homologous recombination enzyme used was NEBuilder® high-fidelity DNA assembly premix (purchased from NEB). The reaction system is as follows:

[0233] The reaction conditions are as follows:

[0234] The ligation product was added to Escherichia coli DH5α competent cells (purchased from Shanghai Weidi Biotechnology Co., Ltd.), incubated on ice for 30 min, heat-shocked at 42℃ for 1 min, incubated on ice for 2 min, and then 900 μL of LB medium was added and cultured at 37℃ for 1 hour to activate and reactivate Escherichia coli DH5α competent cells.

[0235] The revived Escherichia coli DH5α competent cells were plated on LB solid plates containing ampicillin resistance and incubated upside down in a 37°C incubator. The resulting Escherichia coli DH5α monoclonal cells were verified by Sanger sequencing.

[0236] Sequencing verified the correct ligation of the E. coli DH5α clone was used for culture, and the plasmid was extracted to obtain the plasmid pAAV2_tadA(ABEmax)-nSeqCas9-sgRNA-scaffold_ITR, which was then used for later use.

[0237] (6) Preparation of plasmid pAAV2_SeqABE-sgRNA_ITR The pAAV2_tadA(ABEmax)-nSeqCas9-sgRNA-scaffold_ITR plasmid was digested with BBSI restriction endonuclease. The digestion system consisted of 2 μg of plasmid pAAV2_tadA(ABEmax)-nSeqCas9-sgRNA-scaffold_ITR, 5 μL of 10×CutSmart buffer (NEB), 1 μL of BBSI restriction endonuclease (NEB), and water to a final volume of 50 μL. The digestion was incubated at 37°C for 2 hours.

[0238] Then, the enzyme digestion products were electrophoresed on a 1% agarose gel at 120V for 30 minutes.

[0239] DNA fragments were excised from the agarose gel and recovered using a gel recovery kit (Tiangen Biotech (Beijing) Co., Ltd., DP209) according to the manufacturer's instructions. Finally, the fragments were eluted with ultrapure water.

[0240] The DNA concentration of the recovered linearized plasmid pAAV2_tadA(ABEmax)-nSeqCas9-sgRNA-scaffold_ITR was determined using a NanoDrop™ Lite spectrophotometer (Thermo Scientific) for later use or for long-term storage at -20°C.

[0241] Endogenous target sequences that meet the PAM requirements of the SeqCas9 protein were randomly selected from the human genome, and the corresponding oligonucleotide single-stranded DNAs are shown in Table 4 (the uppercase letters in the table represent sticky end sequences).

[0242] Table 4. gRNA sequences and corresponding oligonucleotide single-stranded DNA sequences

[0243] Oligonucleotide single-stranded DNA was annealed to obtain double-stranded DNA. The annealing reaction system consisted of 3 μL 10 μM oligo-F, 3 μL 10 μM oligo-R, and 4 μL water. After vortexing and mixing the annealing system, it was placed in a PCR instrument and the annealing program was run as follows: 95℃_5 min, 85℃_1 min, 75℃_1 min, 65℃_1 min, 55℃_1 min, 45℃_1 min, 35℃_1 min, 25℃_1 min, and stored at 4℃ with a cooling rate of 0.3℃ / s. After annealing, the resulting product was ligated into the linearized pAAV2_tadA(ABEmax)-nSeqCas9-sgRNA-scaffold_ITR vector using DNA ligase (purchased from NEB).

[0244] 1 μL of the obtained ligation product was added to Escherichia coli DH5α competent cells (purchased from Shanghai Weidi Biotechnology Co., Ltd.), incubated on ice for 30 min, heat-shocked at 42℃ for 1 min, incubated on ice for 2 min, and then 900 μL of LB medium was added and cultured at 37℃ for 1 hour to activate and reactivate Escherichia coli DH5α competent cells.

[0245] The revived Escherichia coli DH5α competent cells were plated on LB solid plates containing the corresponding antibiotics and incubated upside down in a 37°C incubator. The resulting Escherichia coli DH5α monoclonal cells were verified by Sanger sequencing.

[0246] Sequencing verified the correct ligation of the E. coli DH5α clone was used for culture, and the plasmid was extracted to obtain the plasmid pAAV2_SeqABE-sgRNA_ITR containing the target sgRNA sequence, which was then used for later use.

[0247] (7) Transfection of wild-type HEK293T cell line with pAAV2_SeqABE-sgRNA_ITR plasmid The obtained pAAV2_SeqABE-sgRNA_ITR plasmid was transfected into wild-type HEK293T cell lines using liposomes. The transfection process included the following steps: On day 0, HEK293T cells were seeded in 6-well plates as needed for transfection, with the cell density controlled at 30%.

[0248] Day 1, transfection was performed. The transfection process is as follows: Take 2 μg of the plasmid pAAV2_SeqABE-sgRNA_ITR to be transfected and add it to 100 μL of Opti-MEM medium (purchased from Gibco). Gently pipette to mix.

[0249] Gently mix Lipofectamine® 2000 (purchased from Invitrogen) or PEI (purchased from Polysciences), then add 5 μL of Lipofectamine® 2000 or PEI to 100 μL of Opti-MEM medium, mix gently, and let stand at room temperature for 5 min.

[0250] The diluted plasmid and diluted transfection reagent were mixed and gently blown to mix. The resulting mixture was allowed to stand at room temperature for 20 min, and then added to the culture medium for HEK293T cells. The cells were then placed in a 37°C, 5% CO2 incubator and cultured for 7 days.

[0251] (8) Preparation of next-generation sequencing libraries HEK293T cells were collected seven days after editing, and genomic DNA was extracted using a DNA kit (Tiangen Biotech (Beijing) Co., Ltd., DP304) according to the instructions provided with the DNA kit.

[0252] The first round of PCR for library preparation was performed using a 2×Q5 Mastermix. The PCR primers are shown in Table 5. Table 5. List of PCR primers for each endogenous site

[0253] The reaction system is as follows:

[0254] The PCR procedure is as follows:

[0255] For the second round of PCR library construction, a PCR reaction was performed using a 2×Q5 Mastermix. The PCR primers were the same as the F2 and R2 primers given in Example 1 above.

[0256] The reaction system is as follows:

[0257] The PCR procedure is as follows:

[0258] The DNA fragments from the second round of PCR products were purified using a gel extraction kit according to the manufacturer's instructions, thus completing the preparation of the next-generation sequencing library.

[0259] (9) Analysis of second-generation sequencing results The prepared next-generation sequencing library was subjected to paired-end sequencing on the high-throughput sequencer HiseqXTen (illumina).

[0260] The next-generation sequencing results were processed to obtain the editing ratio of adenine A at each endogenous target site that met the editing requirements. The results are shown in... Figure 4 As can be seen from the figure, the SeqABE base editor successfully performed single-base editing on these endogenous target sites, demonstrating its high base editing efficiency and showcasing its great potential as a base editing tool.

[0261] Example 5 (1) Preparation of linearized plasmid AncBE4max PCR was performed using the AncBE4max plasmid (Addgene platform, catalog # 112100) as a template. The primer sequences were as follows: Primer 1: TGTCTAAGTTGGGCGAAGAAagcggcgggagcggcggg Primer 2: CCGATACTGTATGTCTTTTCgctgccgccgctgctgc The reaction system is as follows:

[0262] The PCR procedure is as follows:

[0263] The PCR products were electrophoresed on a 1% agarose gel at 120V for 30 min. The DNA fragments of the target size were purified using a gel extraction kit according to the manufacturer's instructions. The DNA concentration was measured using a NanoDrop™ Lite spectrophotometer (ThermoScientific) and stored for later use or at -20℃ for long-term storage.

[0264] (2) Preparation of plasmid pAAV2_Anc-APOBEC1-SeqCas9_ITR The linearized AncBE4max backbone fragment and the synthetically produced humanized SeqCas9 fragment (SEQ ID NO:13) were homologously recombinated according to the ratio specified in the instructions. The homologous recombinase used was NEBuilder® High Fidelity DNA Assembly Premix (NEB). The reaction system is as follows:

[0265] The reaction conditions are as follows:

[0266] The ligation product was added to Escherichia coli DH5α competent cells (purchased from Shanghai Weidi Biotechnology Co., Ltd.), incubated on ice for 30 min, heat-shocked at 42℃ for 1 min, incubated on ice for 2 min, and then 900 μL of LB medium was added and cultured at 37℃ for 1 hour to activate and reactivate Escherichia coli DH5α competent cells.

[0267] The revived Escherichia coli DH5α competent cells were plated on LB solid plates containing ampicillin resistance and incubated upside down in a 37°C incubator. The resulting Escherichia coli DH5α monoclonal cells were verified by Sanger sequencing.

[0268] Sequencing was used to verify the correct ligation of the E. coli DH5α clone, and the plasmid was extracted to obtain plasmid pAAV2_Anc-APOBEC1-SeqCas9_ITR, which was then used for later use.

[0269] (3) Preparation of plasmid pAAV2_Anc-APOBEC1-nSeqCas9_ITR A circular PCR reaction was performed using pAAV2_Anc-APOBEC1-SeqCas9_ITR as a template. The primer sequences were as follows: Primer 3: GCCATCGGCACCAATTCTGT Primer 4: CAGGCCGATACTGTATGTCT The reaction system is as follows:

[0270] The PCR procedure is as follows:

[0271] PCR products were electrophoresed on a 1% agarose gel at 120V for 30 min. The DNA fragments of the target size were purified using a gel extraction kit following the manufacturer's instructions. DNA concentration was measured using a NanoDrop™ Lite spectrophotometer (ThermoScientific). T4 PNK treatment and T4 DNA ligase treatment were then performed. The reaction system is as follows:

[0272] The reaction conditions are as follows:

[0273] Add 1 μL of T4 DNA ligase (NEB) to the reaction system, vortex to mix, and incubate at room temperature for 2 h.

[0274] The ligation product was added to Escherichia coli DH5α competent cells (purchased from Shanghai Weidi Biotechnology Co., Ltd.), incubated on ice for 30 min, heat-shocked at 42℃ for 1 min, incubated on ice for 2 min, and then 900 μL of LB medium was added and cultured at 37℃ for 1 hour to activate and reactivate Escherichia coli DH5α competent cells.

[0275] The revived Escherichia coli DH5α competent cells were plated on LB solid plates containing ampicillin resistance and incubated upside down in a 37°C incubator. The resulting Escherichia coli DH5α monoclonal cells were verified by Sanger sequencing.

[0276] Sequencing was used to verify the correct ligation of the E. coli DH5α clone, and the plasmid was extracted to obtain plasmid pAAV2_Anc-APOBEC1-nSeqCas9_ITR, which was then used for later use.

[0277] (4) Linearization preparation of pAAV2_Anc-APOBEC1-nSeqCas9_ITR The pAAV2_Anc-APOBEC1-nSeqCas9_ITR plasmid was digested using MfeI restriction endonuclease (purchased from NEB). The reaction mixture consisted of 2 μg of plasmid pAAV2_Anc-APOBEC1-nSeqCas9_ITR, 5 μL of 10×CutSmart buffer (purchased from NEB), 1 μL of MfeI restriction endonuclease (purchased from NEB), and water to a final volume of 50 μL. The digestion was carried out at 37°C for 2 hours.

[0278] Then, the enzyme digestion products were electrophoresed on a 1% agarose gel at 120V for 30 minutes.

[0279] DNA fragments were excised from the agarose gel and recovered using a gel recovery kit (Tiangen Biotech (Beijing) Co., Ltd., DP209) according to the manufacturer's instructions. Finally, the fragments were eluted with ultrapure water.

[0280] The DNA concentration of the recovered linearized fragment pAAV2_Anc-APOBEC1-nSeqCas9_ITR was determined using a NanoDrop™ Lite spectrophotometer (Thermo Scientific) and stored for later use or at -20°C for long-term preservation.

[0281] (5) Preparation of pAAV2_Anc-APOBEC1-nSeqCas9-sgRNA_ITR plasmid PCR was performed using SeqCas9-PSK-mU6-sgRNA-scaffold (prepared according to Example 1(2)) as a template. The primer sequences were as follows: Primer 5: caaggcaaggcttgaccgacGATCCGACGCGCCATCTC Primer 6: agcagattcttcatgcaattAAAATGCACCCGAATCGGG The reaction system is as follows:

[0282] The PCR procedure is as follows:

[0283] The PCR products were electrophoresed on a 1.5% agarose gel at 120V for 30 min. The 455bp SeqCas9 sgRNA DNA fragment was purified using a gel extraction kit according to the manufacturer's instructions. The DNA concentration was measured using a NanoDrop™ Lite spectrophotometer (Thermo Scientific) and stored for later use or at -20℃ for long-term storage.

[0284] The linearized pAAV2_Anc-APOBEC1-nSeqCas9_ITR fragment and the SeqCas9 sgRNA DNA fragment were homologously recombinated according to the ratio specified in the manufacturer's instructions. The homologous recombination enzyme used was NEBuilder® High Fidelity DNA Assembly Premix (NEB). The reaction system is as follows:

[0285] The reaction conditions are as follows:

[0286] The ligation product was added to Escherichia coli DH5α competent cells (purchased from Shanghai Weidi Biotechnology Co., Ltd.), incubated on ice for 30 min, heat-shocked at 42℃ for 1 min, incubated on ice for 2 min, and then 900 μL of LB medium was added and cultured at 37℃ for 1 hour to activate and reactivate Escherichia coli DH5α competent cells.

[0287] The revived Escherichia coli DH5α competent cells were plated on LB solid plates containing ampicillin resistance and incubated upside down in a 37°C incubator. The resulting Escherichia coli DH5α monoclonal cells were verified by Sanger sequencing.

[0288] Sequencing verified the correctly ligated E. coli DH5α clone was shaken to extract the plasmid, resulting in plasmid pAAV2_Anc-APOBEC1-nSeqCas9-sgRNA-scaffold_ITR, which was then used for later use.

[0289] (6) Preparation of plasmid pAAV2_SeqCBE-sgRNA_ITR The pAAV2_Anc-APOBEC1-nSeqCas9-sgRNA-scaffold_ITR plasmid was digested with BBSI restriction endonuclease. The digestion system consisted of 2 μg of pAAV2_Anc-APOBEC1-nSeqCas9-sgRNA-scaffold_ITR plasmid, 5 μL of 10×CutSmart buffer (NEB), 1 μL of BBSI restriction endonuclease (NEB), and water to a final volume of 50 μL. The digestion was incubated at 37°C for 2 hours.

[0290] Then, the enzyme digestion products were electrophoresed on a 1% agarose gel at 120V for 30 minutes.

[0291] DNA fragments were excised from the agarose gel and recovered using a gel recovery kit (Tiangen Biotech (Beijing) Co., Ltd., DP209) according to the manufacturer's instructions. Finally, the fragments were eluted with ultrapure water.

[0292] The DNA concentration of the recovered linearized plasmid pAAV2_Anc-APOBEC1-nSeqCas9-sgRNA-scaffold_ITR was determined using a NanoDrop™ Lite spectrophotometer (Thermo Scientific) for later use or for long-term storage at -20 °C.

[0293] Endogenous target sequences that meet the PAM requirements of the SeqCas9 protein were randomly selected from the human genome, and the corresponding oligonucleotide single-stranded DNAs are shown in Table 6 (in the table, uppercase letters represent the bases of the sticky end sequence).

[0294] Table 6. gRNA sequences and corresponding oligonucleotide single-stranded DNA sequences

[0295] Oligonucleotide single-stranded DNA was annealed to obtain double-stranded DNA. The annealing reaction system consisted of 3 μL 10 μM oligo-F, 3 μL 10 μM oligo-R, and 4 μL water. After vortexing and mixing the annealing system, it was placed in a PCR instrument and the annealing program was run as follows: 95℃_5 min, 85℃_1 min, 75℃_1 min, 65℃_1 min, 55℃_1 min, 45℃_1 min, 35℃_1 min, 25℃_1 min, and stored at 4℃ with a cooling rate of 0.3℃ / s. After annealing, the resulting product was ligated into the linearized pAAV2_Anc-APOBEC1-nSeqCas9-sgRNA-scaffold_ITR vector using DNA ligase (purchased from NEB).

[0296] 1 μL of the obtained ligation product was added to Escherichia coli DH5α competent cells (purchased from Shanghai Weidi Biotechnology Co., Ltd.), incubated on ice for 30 min, heat-shocked at 42℃ for 1 min, incubated on ice for 2 min, and then 900 μL of LB medium was added and cultured at 37℃ for 1 hour to activate and reactivate Escherichia coli DH5α competent cells.

[0297] The revived Escherichia coli DH5α competent cells were plated on LB solid plates containing the corresponding antibiotics and incubated upside down in a 37°C incubator. The resulting Escherichia coli DH5α monoclonal cells were verified by Sanger sequencing.

[0298] Sequencing verified the correct ligation of the E. coli DH5α clone was used for culture, and the plasmid was extracted to obtain the plasmid pAAV2_SeqCBE-sgRNA_ITR containing the target sgRNA sequence, which was then used for later use.

[0299] (7) Transfection of wild-type HEK293T cell line with pAAV2_SeqCBE-sgRNA_ITR plasmid The obtained pAAV2_SeqCBE-sgRNA_ITR plasmid was transfected into the wild-type HEK293T cell line using liposomes. The transfection process included the following steps: On day 0, HEK293T cells were seeded in 6-well plates as needed for transfection, with the cell density controlled at 30%.

[0300] Day 1, transfection was performed. The transfection process is as follows: Take 2 μg of the plasmid pAAV2_SeqCBE-sgRNA_ITR to be transfected and add it to 100 μL of Opti-MEM medium (purchased from Gibco). Gently pipette to mix.

[0301] Gently mix Lipofectamine® 2000 (purchased from Invitrogen) or PEI (purchased from Polysciences), then add 5 μL of Lipofectamine® 2000 or PEI to 100 μL of Opti-MEM medium, mix gently, and let stand at room temperature for 5 min.

[0302] The diluted plasmid and diluted transfection reagent were mixed and gently blown to mix. The resulting mixture was allowed to stand at room temperature for 20 min, and then added to the culture medium for HEK293T cells. The cells were then placed in a 37°C, 5% CO2 incubator and cultured for 7 days.

[0303] (8) Preparation of next-generation sequencing libraries HEK293T cells were collected seven days after editing, and genomic DNA was extracted using a DNA kit (Tiangen Biotech (Beijing) Co., Ltd., DP304) according to the instructions provided with the DNA kit.

[0304] The first round of PCR for library preparation was performed using a 2×Q5 Mastermix. The PCR primers are shown in Table 7. Table 7. List of PCR primers for each endogenous site

[0305] The reaction system is as follows:

[0306] The PCR procedure is as follows:

[0307] For the second round of PCR library construction, a PCR reaction was performed using a 2×Q5 Mastermix. The PCR primers were the same as the F2 and R2 primers given in Example 1 above.

[0308] The reaction system is as follows:

[0309] The PCR procedure is as follows:

[0310] The DNA fragments from the second round of PCR products were purified using a gel extraction kit according to the manufacturer's instructions, thus completing the preparation of the next-generation sequencing library.

[0311] (9) Analysis of second-generation sequencing results The prepared next-generation sequencing library was subjected to paired-end sequencing on the high-throughput sequencer HiseqXTen (illumina).

[0312] The editing ratio of cytosine C at each endogenous target site that meets the editing requirements was obtained after processing the second-generation sequencing results. The results are shown in... Figure 5 As can be seen from the figure, the SeqCBE base editor successfully performed single-base editing on these endogenous target sites, demonstrating its high base editing efficiency and showcasing its great potential as a base editing tool.

[0313] The sequences used in this invention:

Claims

1. A fusion protein comprising: a) A SeqCas9 mutant protein, based on the SeqCas9 protein shown in SEQ ID NO: 1 and including mutations that cause the SeqCas9 protein to lose its endonuclease activity (preferably mutations D10A and H847A) or mutations that cause the SeqCas9 protein to have only single-stranded DNA cleavage activity (preferably mutation D10A). b) Other proteins or polypeptides; For example, the additional protein or polypeptide is selected from one or more of the following: epitope tags, reporter proteins or nuclear localization signal (NLS) sequences, cytosine deaminases, adenine deaminases, cytosine methyltransferases DNMT3A and MQ1, cytosine demethyltransferase Tet1, transcription activators VP64, p65 and RTA, transcription repressors KRAB, histone acetyltransferase p300, histone deacetyltransferase LSD1, Gam protein, and endonuclease FokI; and c) An optional first linker for connecting the SeqCas9 mutant protein to the other protein or polypeptide; Preferably, the fusion protein comprises, from its N-terminus to its C-terminus, the following in sequence: an optional nuclear localization signal (NLS) sequence, an optional second adapter, the adenine deaminase (e.g., the adenine deaminase with the amino acid sequence shown in SEQ ID NO: 2) or a dimer thereof (e.g., a dimer of the adenine deaminase comprising the amino acid sequence shown in SEQ ID NO: 2, an optional third adapter, and the amino acid sequence shown in SEQ ID NO: 3) or a cytosine deaminase (e.g., the cytosine deaminase with the amino acid sequence shown in SEQ ID NO: 4), the optional first adapter, the SeqCas9 mutant protein, an optional fourth adapter, and an optional nuclear localization signal (NLS) sequence; Preferably, the amino acid sequence of the fusion protein is as shown in SEQ ID NO: 5, SEQ ID NO: 6 or SEQ ID NO:

9.

2. The fusion protein according to claim 1, wherein, The fusion protein comprises a cytosine deaminase and one or two uracil DNA glycosylase inhibitors (UGIs), such as the uracil DNA glycosylase inhibitor (UGI) with the amino acid sequence shown in SEQ ID NO:

8. The fusion protein comprises, from its N-terminus to its C-terminus, an optional nuclear localization signal (NLS) sequence, an optional second adapter, the cytosine deaminase, the optional first adapter, the SeqCas9 mutant protein, an optional fifth adapter, the uracil DNA glycosylase inhibitor (UGI), an optional fourth adapter, and an optional nuclear localization signal (NLS) sequence. Preferably, the fusion protein comprises two uracil DNA glycosylation inhibitors (UGIs), which are optionally linked via a sixth linker; Preferably, each of the first to sixth connectors is independently a connector of 1-50, for example, 10-32 amino acids; Preferably, the amino acid sequence of the fusion protein is shown in SEQ ID NO:

10.

3. A conjugate, said conjugate comprising: a) The fusion protein according to claim 1 or 2; b) Detectable markers; and c) An optional seventh linker for connecting the fusion protein to the detectable tag, for example, a seventh linker with a length of 1-50 amino acids.

4. A single-stranded guide RNA, wherein the single-stranded guide RNA comprises a guide sequence and a scaffold sequence from the 5' end to the 3' end, wherein the scaffold sequence is: a) The stent sequence shown in SEQ ID NO: 11; or b) A sequence obtained by modifying the sequence shown in SEQ ID NO: 11 while retaining its biological activity; For example, the modification is one or more of the following: base phosphorylation, base sulfidation, base methylation, base hydroxylation, sequence shortening, and sequence lengthening; For example, the shortening of the sequence and the lengthening of the sequence include the deletion or addition of one, two, three, four, five, six, seven, eight, nine or ten bases relative to the base sequence; Preferably, the guide sequence is a sequence of 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides (preferably 20 nucleotides) in length and capable of complementary pairing with the target sequence.

5. An isolated nucleic acid molecule comprising: a nucleic acid sequence encoding the fusion protein of claim 1 or 2 or the conjugate of claim 3; For example, the isolated nucleic acid molecule contains the nucleic acid sequence shown in SEQ ID NO: 14 or 15 or a degenerate sequence thereof.

6. The isolated nucleic acid molecule according to claim 5 further comprises a nucleic acid sequence encoding the single-stranded guide RNA of claim 4, such as the nucleic acid sequence shown in SEQ ID NO:

12.

7. An isolated nucleic acid molecule comprising a nucleic acid sequence encoding the single-stranded guide RNA of claim 4, such as the nucleic acid sequence shown in SEQ ID NO:

12.

8. A vector comprising: a nucleic acid sequence encoding the fusion protein of claim 1 or 2 or the conjugate of claim 3, for example, the nucleic acid sequence shown in SEQ ID NO: 14 or 15 or a degenerate sequence thereof; For example, the vector is a plasmid vector such as pUC19 vector, attachor vector, pAAV2_ITR vector, retroviral vector, lentiviral vector, adenovirus vector, or adeno-associated virus vector.

9. The vector according to claim 8, wherein the vector further comprises a nucleic acid sequence encoding the single-stranded guide RNA of claim 4, such as the nucleic acid sequence shown in SEQ ID NO:

12.

10. A vector comprising a nucleic acid sequence encoding the single-stranded guide RNA of claim 4, such as the nucleic acid sequence shown in SEQ ID NO:

12.

11. A CRISPR / Cas9 gene editing system, comprising: 1) A protein component comprising: the SeqCas9 protein with the amino acid sequence as shown in SEQ ID NO: 1, or a conjugate or fusion protein thereof, the fusion protein of claim 1 or 2, or the conjugate of claim 3; and 2) The single-stranded guide RNA as described in claim 4; Furthermore, the protein component and the single-stranded guide RNA bind to each other to form a complex; in, The conjugate of the SeqCas9 protein comprises: the SeqCas9 protein, a detectable tag or a combination of a detectable tag and another protein or polypeptide, and optionally an eighth linker for connecting the SeqCas9 protein to the detectable tag or the combination. The fusion protein of the SeqCas9 protein comprises: the SeqCas9 protein, another protein or polypeptide, and optionally a ninth linker for connecting the SeqCas9 protein to the other protein or polypeptide. For example, the additional protein or polypeptide is selected from one or more of the following: epitope tags, reporter proteins or nuclear localization signal sequences, cytosine deaminases, adenine deaminases, cytosine methyltransferases DNMT3A and MQ1, cytosine demethylase Tet1, transcription activators VP64, p65 and RTA, transcription repressors KRAB, histone acetyltransferase p300, histone deacetyltransferase LSD1, Gam protein, and endonuclease FokI; Preferably, the eighth and ninth connectors are each independently a connector having 1-50, for example 10-32 amino acids.

12. A method for gene editing of a target sequence in an intracellular or in vitro environment, the method comprising: Contact any of the following (1) through (6) with the target sequence in the intracellular or in vitro environment: (1) The SeqCas9 protein or its conjugate or fusion protein with the amino acid sequence shown in SEQ ID NO: 1, the fusion protein of claim 1 or 2 or the conjugate of claim 3; and the single-stranded guide RNA of claim 4; (2) The isolated nucleic acid molecule of claim 5 and the isolated nucleic acid molecule of claim 7; (3) The isolated nucleic acid molecule as described in claim 6; (4) The carrier according to claim 8 and the carrier according to claim 10; (5) The carrier according to claim 9; (6) The CRISPR / Cas9 gene editing system as described in claim 11; The conjugate of the SeqCas9 protein comprises: the SeqCas9 protein, a detectable label or a combination of a detectable label and another protein or polypeptide, and optionally an eighth linker for connecting the SeqCas9 protein to the detectable label or the combination. The fusion protein of the SeqCas9 protein comprises: the SeqCas9 protein, another protein or polypeptide, and optionally a ninth linker for connecting the SeqCas9 protein to the other protein or polypeptide. The protein SeqCas9 or its conjugate or fusion protein, the fusion protein of claim 1 or 2 or the conjugate of claim 3 recognizes a protospacer adjacent motif (PAM) located at the 3' end of the target sequence and having a 5'-NNG sequence. For example, the cell is a prokaryotic cell or a eukaryotic cell, such as an animal cell, and the animal cell is, for example, a mammalian cell, such as a human cell; For example, the gene editing includes one or more of the following: gene knockout of the target sequence, site-specific base alteration, site-specific insertion, regulation of gene transcription, regulation of DNA methylation, DNA acetylation modification, histone acetylation modification, single base conversion, and chromatin imaging tracking. For example, the single base conversion includes one or more of the following: conversion from guanine G to adenine A, conversion from adenine A to guanine G, conversion from cytosine C to thymine T, conversion from cytosine C to uracil U, and conversion from thymine T to cytosine C. Preferably, the eighth and ninth connectors are each independently a connector having 1-50, for example 10-32 amino acids.

13. The method according to claim 12, wherein, The guide sequence of the single-stranded guide RNA forms a completely complementary base pairing structure with the target sequence, while forming an incomplete complementary base pairing structure with non-target sequences. For example, the incomplete base complementary pairing structure includes one or more, such as two or more, base mismatches.

14. A cell comprising: an isolated nucleic acid molecule according to any one of claims 5-7, or a vector according to any one of claims 8-10; For example, the cell is a prokaryotic cell or a eukaryotic cell, such as an animal cell, and the animal cell is, for example, a mammalian cell, such as a human cell.

15. A kit for gene editing of a target sequence in an intracellular or in vitro environment, comprising: a) Choose any one of (1) to (6) below: (1) The SeqCas9 protein or its conjugate or fusion protein with the amino acid sequence shown in SEQ ID NO: 1, the fusion protein of claim 1 or 2 or the conjugate of claim 3; and the single-stranded guide RNA of claim 4; (2) The isolated nucleic acid molecule of claim 5 and the isolated nucleic acid molecule of claim 7; (3) The isolated nucleic acid molecule as described in claim 6; (4) The carrier according to claim 8 and the carrier according to claim 10; (5) The carrier according to claim 9; (6) The CRISPR / Cas9 gene editing system as described in claim 11; as well as b) Instructions on how to perform gene editing on target sequences in the intracellular or in vitro environment; The conjugate of the SeqCas9 protein comprises: the SeqCas9 protein, a detectable label or a combination of a detectable label and another protein or polypeptide, and optionally an eighth linker for connecting the SeqCas9 protein to the detectable label or the combination. The fusion protein of the SeqCas9 protein comprises: the SeqCas9 protein, another protein or polypeptide, and optionally a ninth linker for connecting the SeqCas9 protein to the other protein or polypeptide. Preferably, the eighth and ninth connectors are each independently a connector having 1-50, for example 10-32 amino acids.