Cas9 protein mutants and their applications
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2026-08-14
AI Technical Summary
[0025]本公开通过在远离催化位点的位置引入突变,诱导Cas9蛋白构象改变,获得新的Cas9蛋白突变体,其只切割Cas9-sgRNA靶标链,不切割单链非靶标链,可以用于构建新型碱基编辑器,能对哺乳动物细胞基因组进行高效率的碱基编辑。
Smart Images

Figure SMS_1 
Figure SMS_2 
Figure SMS_3
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of gene editing technology, and in particular to a Cas9 protein mutant and its applications. Background Technology
[0002] The CRISPR / Cas system is an acquired immune system in bacteria used to defend against invading foreign DNA. Its targeting capabilities, DNA unwinding and cleavage abilities, and other properties have been utilized in the development of CRISPR gene editing technology. Cas9 in the CRISPR / Cas9 system is a nuclease whose targeting is determined by the dual RNA guidance structure of crRNA and tracrRNA and the DNA target site PAM. Integrating crRNA and tracrRNA into sgRNA constitutes the Cas9-sgRNA gene editing system. By designing sgRNAs with matching sequences in the crRNA spacer region and combining them with PAM, Cas9 can target specific genomic sites, unwind the target DNA, and cleave the target DNA double strand. Cells primarily repair DNA double-strand breaks (DSBs) through non-homologous end joining (NHEJ) and homologous recombination (HR). Researchers utilize this process for gene insertion, deletion, or replacement, achieving precise gene editing.
[0003] Cas9 contains homologous domains to the HNH and RuvC endonucleases. After DNA unwinding, the HNH domain cleaves the target DNA strand paired with the sgRNA spacer, while the RuvC domain cleaves the single-stranded non-target strand. Researchers introduced a point mutation at the catalytic site of the HNH domain, obtaining the Cas9 H840A mutant, which cleaves only the non-target strand. The Cas9 H840A mutant greatly expands the application scope of the CRISPR / Cas9 gene editing system. Currently, there is an urgent need for novel Cas9 proteins that cleave only the target strand, as well as more gene editing tools to choose from. Summary of the Invention
[0004] The purpose of this disclosure is to provide new Cas9 protein mutants and, based on these, related gene editing tools.
[0005] To achieve the above objectives, this disclosure provides the following technical solutions:
[0006] On one hand, this disclosure provides a Cas9 protein mutant comprising mutations in amino acid residues at positions 921, 922, and / or 923 of the amino acid sequence shown in SEQ ID NO:1, wherein the mutations are selected from one or more of the following:
[0007] (a) Deletion of amino acid residues;
[0008] (b) Substitution of amino acid residues;
[0009] (c) Insertion of amino acid residues;
[0010] Furthermore, the mutation can significantly reduce or eliminate the cleavage activity of the Cas9 protein on non-target chains.
[0011] On the other hand, this disclosure provides a Cas9 protein mutant comprising the amino acid sequence fragments shown in SEQ ID NO: 35, 37, 39, 41, wherein the mutant significantly reduces or eliminates the cleavage activity of the Cas9 protein on non-target chains.
[0012] On the other hand, this disclosure provides a fusion protein comprising the aforementioned Cas9 protein mutant and a heterologous domain fused with the Cas9 protein mutant.
[0013] On the other hand, this disclosure provides a complex comprising the aforementioned Cas9 protein mutant or the aforementioned fusion protein, and guide RNA.
[0014] On the other hand, this disclosure provides a polynucleotide that encodes the aforementioned Cas9 protein mutant or the aforementioned fusion protein.
[0015] On the other hand, this disclosure provides a vector comprising the aforementioned polynucleotide; preferably, the vector comprises a promoter for driving the expression of the polynucleotide.
[0016] In a sixth aspect, this disclosure provides a kit comprising the aforementioned Cas9 protein mutant, the aforementioned fusion protein, the aforementioned complex, the aforementioned polynucleotide, or the aforementioned vector; preferably, the kit further comprises an expression construct encoding a guide RNA backbone.
[0017] On the other hand, this disclosure provides a cell comprising the aforementioned Cas9 protein mutant, the aforementioned fusion protein, the aforementioned complex, the aforementioned polynucleotide, or the aforementioned vector.
[0018] On the other hand, this disclosure provides a pharmaceutical composition comprising the aforementioned Cas9 protein mutant, the aforementioned fusion protein, the aforementioned complex, the aforementioned polynucleotide, the aforementioned vector, or the aforementioned cell.
[0019] On the other hand, this disclosure provides the use of the aforementioned Cas9 protein mutant, the aforementioned fusion protein, the aforementioned complex, the aforementioned polynucleotide, the aforementioned vector, the aforementioned kit, or the aforementioned cell or the aforementioned pharmaceutical composition in nucleic acid modification.
[0020] On the other hand, this disclosure provides a method for modifying a target nucleic acid, the method comprising contacting the target nucleic acid with the following:
[0021] (a) the aforementioned Cas9 protein mutant or the aforementioned fusion protein, and guide RNA; or
[0022] (b) The aforementioned complex.
[0023] On the other hand, this disclosure provides a method for modifying a target nucleic acid, the method comprising introducing the aforementioned Cas9 protein mutant, the aforementioned fusion protein, the aforementioned complex, the aforementioned polynucleotide, the aforementioned vector, the aforementioned cell and / or the aforementioned pharmaceutical composition into a cell or organism containing the target nucleic acid.
[0024] On the other hand, this disclosure provides a method for treating a subject diagnosed with a disease related to or caused by a gene mutation, comprising administering to the subject a therapeutically effective amount of the aforementioned Cas9 protein mutant, the aforementioned fusion protein, the aforementioned complex, the aforementioned polynucleotide, the aforementioned vector, the aforementioned cell, and / or the aforementioned pharmaceutical composition.
[0025] This disclosure induces conformational changes in the Cas9 protein by introducing mutations at locations far from the catalytic site, resulting in novel Cas9 protein mutants that cleave only the Cas9-sgRNA target strand and not the single-stranded non-target strand. These mutants can be used to construct novel base editors capable of highly efficient base editing of mammalian cell genomes.
[0026] Simultaneous mutation of the Cas9 protein mutant and the H840A site in this disclosure inactivates the catalytic domain of the Cas9 protease, constructing dCas9. Base editors (including ABE and CBE) designed using the Cas9 protein mutant of this disclosure exhibit high base editing efficiency in mammalian cells. Fusing the Cas9 protein mutant of this disclosure with a DNA-binding-deficient mutant TX of the 3'→5' exonuclease TREX2 can significantly improve the gene editing efficiency of the 3' ends generated by paired single-cut enzyme cleavage while maintaining its safety. Attached Figure Description
[0027] To more clearly illustrate the specific embodiments of this disclosure or the technical solutions in the prior art, the accompanying drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained from these drawings without creative effort. The accompanying drawings are incorporated in and constitute a part of this specification, illustrating embodiments consistent with this specification, and together with the specification, are used to explain the principles of this specification.
[0028] Figure 1 The diagram illustrates the detection results of STGC, LTGC, and LTGC biases in mouse embryonic stem cells containing the HR reporter system for four SpCas9 insertion / deletion mutants disclosed in this disclosure and after the introduction of H840A into the insertion / deletion mutants. AC represents the four insertion / deletion mutants and the HR reporter system-induced fragmentation after co-transduction with sgRNAg1-1 following the introduction of H840A into the insertion / deletion mutants. Figure A shows the GFP levels detected by flow cytometry. + Cell percentage, i.e., STGC efficiency; Figure B shows GFP detected by flow cytometry. + RFP + The ratio represents LTGC efficiency; Figure C shows the LTGC bias derived from the ratio of the LTGC efficiency of each enzyme combination of g1-1 to the total HR efficiency. Figures DF represent the four insertion / deletion mutants and the HR reporter system-induced breakage after introducing H840A into the insertion / deletion mutants and co-transferring with sgRNA g1-2. Figure D shows the GFP levels detected by flow cytometry. + Cell percentage, i.e., STGC efficiency; Figure E shows GFP detected by flow cytometry. + RFP + The ratio represents the LTGC efficiency; Figure F shows the LTGC bias derived from the ratio of the LTGC efficiency of each nuclease combination g1-2 to the total HR efficiency. SpCas9 serves as a positive control for cleaving double-stranded nucleases.
[0029] Figure 2 The diagram shows the structure of the SpCas9 overexpression vector used to construct a single-copy SpCas9 gene monoclonal cell line in Embodiment 1 of this disclosure.
[0030] Figure 3 A structural diagram of the 3β-HA-330 vector of this disclosure is shown.
[0031] Figure 4 The results of in vitro cleavage of the target strand of Cas9-sgRNA by four SpCas9 insertion / deletion mutants of this disclosure are shown. Figure A shows the results of in vitro cleavage activity detection of SpCas9, SpCas9 H840A, and the four SpCas9 insertion / deletion mutants of this disclosure, where "-" represents the blank control. Figure B shows the results of in vitro cleavage of single-stranded DNA substrates containing the target strand (top) or the non-target strand (bottom) of Cas9-sgRNA containing the 5'-Alexa680 fluorescently labeled target strand, by SpCas9, SpCas9 H840A, and the four SpCas9 insertion / deletion mutants of this disclosure, detecting their single-strand cleavage activity and pattern, where "-" represents the blank control.
[0032] Figure 5The diagram illustrates four SpCas9 insertion / deletion mutants of this disclosure and the NHEJ induction efficiency in mouse embryonic stem cells containing the NHEJ reporter system after introducing H840A into the insertion / deletion mutants. Figure A is a schematic diagram of the NHEJ reporter system. Figure B shows the flow cytometry detection of GFP after inducing fragmentation in mouse embryonic stem cells containing the NHEJ reporter system and co-transforming the four insertion / deletion mutants with H840A and sgRNA g2-2. + Cell proportion results. Positive control for SpCas9 cleavage of double-stranded nuclease.
[0033] Figure 6 The base editors for four SpCas9 insertion / deletion mutants disclosed herein are shown. Figure A is a schematic diagram of the base editing report system. Figure B shows the base editing efficiency of the adenine base editor constructed based on the four SpCas9 insertion / deletion mutants disclosed herein. Figure C shows the base editing efficiency of the cytosine base editor constructed based on the four SpCas9 insertion / deletion mutants disclosed herein. Data are the average of three independent replicates ± SEM.
[0034] Figure 7 The editing efficiency of the adenine base editor constructed from four SpCas9 insertion / deletion mutants of this disclosure at endogenous sites is shown.
[0035] Figure 8 The editing efficiency of the cytosine base editor constructed from four SpCas9 insertion / deletion mutants of this disclosure at endogenous sites is shown.
[0036] Figure 9 This paper illustrates four SpCas9 insertion / deletion mutants fused with the DNA-binding-deficient TREX2 mutant TX, improving the gene editing efficiency of paired cleavage at the 3' end. Figure A shows a schematic diagram of the structure of the mouse embryonic stem cell mutant NHEJ reporter system. gWL2 is on the left side of "Koz-ATG," and PAM is on the Watson strand; gCR6 is on the right side of "Koz-ATG," and PAM is on the Crick strand. Figure B shows the cleavage of the mutant NHEJ reporter system by SpCas9, the four SpCas9 insertion / deletion mutants of this disclosure, and TX fusion, combined with a single sgRNA (gWL2 or gCR6) or a combination of two sgRNAs (gWL2 and gCR6), using flow cytometry to detect GFP. + Cell percentage. Data are the mean of three independent replicates ± SEM. Detailed Implementation
[0037] (I) Definitions or terms
[0038] To facilitate understanding of this disclosure, certain technical and scientific terms are specifically defined below. In this disclosure, unless otherwise stated, the scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. It should be understood that this disclosure is not limited to specific methods, reagents, compounds, compositions, or biological systems, and variations thereof are certainly possible. It should also be understood that the terminology used in this disclosure is for the purpose of describing specific embodiments only and is not intended to be limiting.
[0039] Unless otherwise expressly stated, the singular forms “a,” “an,” “the,” “the,” and similar designations used in this specification and the appended claims include plural designations.
[0040] As used in this article, the conjunction term “and / or” between multiple elements means to include both “and” and “or” meanings. For example, the phrase “A, B and / or C” is intended to cover each of the following: A, B and C; A, B or C; A or C; A or B; B or C; A and C; A and B; B and C; A (alone); B (alone); and C (alone).
[0041] As used herein, the terms “comprises,” “comprising,” “having,” and “containing,” and any variations thereof, are intended to cover non-exclusive inclusion. The term is intended to be open-ended to specify the presence of any of the stated features, elements, integers, steps, or components, but does not exclude the presence or addition of one or more other features, elements, integers, steps, components, or groups thereof. Therefore, the term “comprising” includes the more restrictive terms “consisting of” and “substantially consisting of”.
[0042] The numerical ranges used in this article should be understood as including all numbers within that range. For example, the range 1 to 20 should be understood to include any number, combination of numbers, or subrange from the following group: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20.
[0043] As used herein, the term "about" indicates a range of ±20% of the following value. In some embodiments, the term "about" indicates a range of ±10% of the following value. In some embodiments, the term "about" indicates a range of ±5% of the following value.
[0044] In the description herein, references to “some implementations,” “some methods,” or “some embodiments” describe a subset of all possible embodiments. However, it is understood that “some implementations” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0045] As used herein, the terms “nucleic acid,” “nucleic acid molecule,” “nucleotide,” “nucleotide sequence,” “oligonucleotide,” or “polynucleotide” mean a polymeric compound comprising covalently linked nucleotides. The term “nucleic acid” includes ribonucleic acid (RNA) and deoxyribonucleic acid (DNA), both of which can be single-stranded or double-stranded. DNA includes, but is not limited to, complementary DNA (cDNA), genomic DNA, plasmid or vector DNA, and synthetic DNA. In some embodiments, this disclosure provides polynucleotides encoding any of the polypeptides disclosed herein; for example, this disclosure relates to polynucleotides encoding the Cas9 protein or variants thereof.
[0046] As used herein, the term "gene" refers to an assembly of nucleotides that encode a polypeptide and includes both cDNA and genomic DNA nucleic acid molecules. "Gene" also refers to a nucleic acid fragment that can act as a regulatory sequence preceding (5' non-coding sequence) and following (3' non-coding sequence).
[0047] As used herein, a nucleic acid molecule is “hybridizable” or “hybridizable” with another nucleic acid molecule (such as cDNA, genomic DNA, or RNA) when annealed to it under suitable temperature and solution ionic strength conditions. Hybridization and washing conditions are known and are illustrated, in particular, in Sambrook et al., Molecular Cloning: A Laboratory Manual, 2nd ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor (1989), specifically in Chapter 11 and Table 11.1. Temperature and ionic strength conditions determine the “pattern” of hybridization. Strict conditions can be adjusted to screen for moderately similar fragments (such as homologous sequences from distantly related organisms) to highly similar fragments (such as genes replicating functional enzymes from closely related organisms). For preliminary screening of homologous nucleic acids, low-tightness hybridization conditions corresponding to a Tm of 55°C can be used, such as 5X SSC, 0.1% SDS, 0.25% milk, and no formamide; or 30% formamide, 5X SSC, and 0.5% SDS. Medium-tightness hybridization conditions correspond to a higher Tm, such as 40% formamide and 5X or 6X SCC. High-tightness hybridization conditions correspond to the highest Tm, such as 50% formamide and 5X or 6X SCC. Hybridization requires both nucleic acids to contain complementary sequences; however, base mismatches may exist depending on the tightness of the hybridization.
[0048] As used herein, the term "complementary" is used to describe the relationship between nucleotide bases that are capable of hybridizing with each other. For example, in DNA, adenosine is complementary to thymine, while cytosine is complementary to guanine. Therefore, this disclosure also includes isolated nucleic acid fragments complementary to the complete sequences disclosed or used herein, as well as those nucleic acid sequences that are substantially similar.
[0049] A DNA “coding sequence” is a double-stranded DNA sequence that, when placed under the control of an appropriate regulatory sequence, is transcribed and translated into a polypeptide in vitro or in vivo. An “appropriate regulatory sequence” is a nucleotide sequence located upstream (5' non-coding), inside, or downstream (3' non-coding) of the coding sequence, and that influences transcription, RNA processing, or stability, or translation of the related coding sequence. Regulatory sequences can include promoters, translation leader sequences, introns, polyadenylation recognition sequences, RNA processing sites, effector binding sites, and stem-loop structures. The boundaries of the coding sequence are defined by a start codon at the 5' (amino) end and a translation stop codon at the 3' (carboxyl) end. Coding sequences can include, but are not limited to, prokaryotic sequences, cDNA derived from mRNA, genomic DNA sequences, and even synthetic DNA sequences. If the coding sequence is intended for expression in eukaryotic cells, polyadenylation signals and transcription termination sequences are typically located at the 3' end of the coding sequence. The abbreviation for "open reading frame" is ORF, which refers to a nucleic acid sequence (DNA, cDNA, or RNA) that includes a translation start signal or start codon (such as ATG or AUG) and a stop codon and may be translated into a polypeptide sequence.
[0050] As used herein, the term "homologous recombination" refers to the insertion of a foreign DNA sequence into another DNA molecule, such as inserting a vector into a chromosome. In some cases, the vector targets a specific chromosomal site to perform homologous recombination. For a given homologous recombination, the vector typically contains a sufficiently long region homologous to the chromosomal sequence to allow complementary binding of the vector to the chromosome and incorporation of the vector into the chromosome. Longer homologous regions and greater sequence similarity can improve the efficiency of homologous recombination.
[0051] As used herein, the term “non-homologous end joining (NHEJ)” refers to the repair of double-strand nicks in DNA by directly joining one end of the nick to the other end of the nick without the need for donor template DNA. In the absence of donor template DNA, NHEJ typically results in a small number of random insertions or deletions of nucleotides (“indels” or “indels”) at the site of the double-strand nick. In some cases, cleavage at the target recognition sequence leads to NHEJ at the target recognition site. Nuclease-induced cleavage at the target site in the gene coding sequence, followed by DNA repair via NHEJ, can introduce mutations into the coding sequence, such as frameshift mutations, thereby disrupting gene function.
[0052] Based on the disclosure herein, polynucleotides can be amplified using methods known in the art. Once a suitable host system and growth conditions are established, recombinant expression vectors can be amplified and prepared in large quantities. As described herein, expression vectors that can be used include, but are not limited to, the following vectors or derivatives thereof: human or animal viruses, such as vaccinia virus or adenovirus; insect viruses, such as baculovirus; yeast vectors; bacteriophage vectors (e.g., λ); and plasmid and copious DNA vectors.
[0053] As used herein, the term "operably linked" means that a target polynucleotide, such as a polynucleotide encoding the Cas9 protein, is linked to a regulatory element in a manner that allows for the expression of the polynucleotide sequence. In some embodiments, the regulatory element is a promoter. In some embodiments, the target polynucleotide is operably linked to a promoter on an expression vector.
[0054] As used herein, the terms "promoter," "promoter sequence," or "promoter region" refer to a DNA regulatory region / sequence capable of binding RNA polymerase and involved in initiating transcription of downstream coding or non-coding sequences. In some instances of this disclosure, the promoter sequence includes a transcription start site and extends upstream to include a minimum number of bases or elements used to initiate transcription at a level above the background detectable level. In some embodiments, the promoter sequence includes a transcription start site and a protein-binding domain responsible for RNA polymerase binding. Eukaryotic promoters typically, but not always, contain multiple "TATA" boxes and "CAT" boxes. Various promoters, including inducible promoters, can be used to drive various vectors of this disclosure.
[0055] As used herein, the term "vector" is any tool used to clone and / or transfer nucleic acids into host cells. A vector can be a replicon that may attach to another DNA segment so as to generate replication of the attached segment. A "replicon" is any genetic element (e.g., plasmid, bacteriophage, granulosome, chromosome, virus) that acts as an automated unit for DNA replication in vivo, i.e., capable of replicating under its own control. In some embodiments of this disclosure, the vector is an appendage vector that is removed / lost from the cell population after many cell generations, for example, through asymmetric distribution. The term "vector" includes both viral and nonviral tools for introducing the nucleic acid into cells in vitro, ex vivo, or in vivo. A large number of vectors known in the art can be used to manipulate nucleic acids, integrate responsive elements and promoters into genes, etc. Possible vectors include, for example, plasmids or modified viruses, including, for example, bacteriophages such as λ derivatives, or plasmids such as PBR322 or pUC plasmid derivatives, or Blue script vectors. For example, inserting a DNA fragment corresponding to a response element and promoter into a suitable vector can be accompanied by ligating the suitable DNA fragment to a selected vector with complementary binding ends. Alternatively, the ends of the DNA molecule can be modified by enzyme catalysis or generated at arbitrary sites by attaching nucleotide sequences (linkers) to the DNA ends. Such vectors can be engineered to contain selective marker genes that provide selection for cells that incorporate the markers into their genome. Such markers allow for the identification and / or selection of host cells that incorporate and express the protein encoded by the marker.
[0056] Viral vectors, particularly retroviral vectors, have been used in numerous gene delivery applications in cells and living animals. Viral vectors that can be used include, but are not limited to, retroviruses, adeno-associated viruses, poxviruses, baculoviruses, vaccinia virus, herpes simplex virus, Ebola virus, adenovirus, geminiviruses, and cauliflower mosaic virus vectors. Non-viral vectors include, but are not limited to, plasmids, liposomes, charged lipids (cytotransfectants), DNA-protein complexes, and biopolymers. In addition to nucleic acids, vectors may also include one or more regulatory regions and / or selective markers for selecting, measuring, and monitoring nucleic acid transfer outcomes (tissue to which it was transferred, duration of expression, etc.).
[0057] Vectors can be introduced into desired host cells using known methods, including but not limited to transfection, transduction, cell fusion, and lipid transfection. Vectors may include various regulatory elements, including promoters. In some embodiments, vector design may be based on multiple constructs designed by Mali et al., “Cas9 as a versatile tool for engineeringbiology,” Nature Methods 10:957-63 (2013). In some embodiments, this disclosure provides expression vectors comprising any of the polynucleotides described herein, for example, expression vectors comprising polynucleotides encoding Cas proteins or variants thereof. In some embodiments, this disclosure provides expression vectors comprising polynucleotides encoding Cas9 proteins or variants thereof.
[0058] As used herein, the term "plasmid" refers to an additional chromosomal element that typically carries a gene that is not involved in the central metabolism of the cell and is usually in the form of a circular double-stranded DNA molecule. Such elements can be linear, circular, or supercoiled autonomously replicating sequences, genome-integrated sequences, bacteriophage sequences, or nucleotide sequences derived from any source of single-stranded or double-stranded DNA or RNA, many of which have been linked or recombined into a unique structure capable of introducing a promoter fragment and DNA sequence targeting a selected gene product, along with an appropriate 3' untranslated sequence, into the cell.
[0059] As used herein, the term "transfection" refers to the introduction of a foreign nucleic acid molecule (including a vector) into a cell. A "transfected" cell contains a foreign nucleic acid molecule within its cell, while a "transformed" cell is one in which the foreign nucleic acid molecule induces a change in the cell's phenotypic structure. The transfected nucleic acid molecule may be integrated into the host cell's genomic DNA and / or may be temporarily or permanently maintained extrachromosomally by the cell. A host cell or organism expressing a foreign nucleic acid molecule or fragment is referred to as a "recombinant," "transformed," or "transgenic" organism. In some embodiments, this disclosure provides host cells comprising any of the expression vectors described herein (e.g., expression vectors comprising polynucleotides encoding Cas proteins or variants thereof). In some embodiments, this disclosure provides host cells comprising an expression vector that includes a polynucleotide encoding a Cas9 protein or a variant thereof.
[0060] As used herein, "expression vector" includes a vector capable of expressing DNA operatively linked to regulatory sequences, such as promoter regions, that influence the expression of such DNA fragments. Such additional fragments may contain promoter and terminator sequences and optionally may contain one or more origins of replication, one or more selection markers, enhancers, polyadenylation signals, etc. Expression vectors are generally derived from plasmid or viral DNA, or may contain elements of both. Therefore, an expression vector refers to a recombinant DNA or RNA construct, such as a plasmid, bacteriophage, recombinant virus, or other vector, which, when introduced into a suitable host cell, results in the expression of clonal DNA. Suitable expression vectors are well known to those skilled in the art and include reproducible expression vectors in eukaryotic and / or prokaryotic cells, as well as expression vectors that remain free or are integrated into the host cell genome.
[0061] As used herein, “expression” refers to the process by which a polypeptide is produced through the transcription and translation of polynucleotides. The expression level of a polypeptide can be evaluated using any method known in the art, including, for example, methods for determining the amount of polypeptide produced from host cells. Such methods may include, but are not limited to, quantifying polypeptides in cell lysates by ELISA, Coomassie blue staining following gel electrophoresis, Lowry protein assays, and Bradford white protein assays.
[0062] As used herein, the term "host cell" refers to a cell into which a recombinant expression vector has been introduced, for receiving, maintaining, replicating, and amplifying the vector. The term "host cell" refers not only to the cell into which the expression vector has been introduced (the "parent" cell) but also to its offspring. Because modifications may occur in offspring, for example due to mutations or environmental influences, offspring may differ from the parent cell but are still included within the scope of the term "host cell." When the host cell divides, the nucleic acids contained in the vector replicate, thereby amplifying the nucleic acids. The host cell can be a eukaryotic or prokaryotic cell. Suitable host cells include, but are not limited to, yeast cells.
[0063] As used herein, the terms “peptide,” “polypeptide,” “protein,” and “proteins” are used interchangeably to refer to polymers of amino acids of any length, including coding and non-coding amino acids, chemically or biochemically modified or derived amino acids, and polypeptides having a modified peptide backbone. The starting point of a protein or polypeptide is called the “N-terminus” (or amino-terminus, NH2-terminus, N-terminus, or amine-terminus), which refers to the free amine (-NH2) group of the first amino acid residue of the protein or polypeptide. The end of a protein or polypeptide is called the “C-terminus” (or carboxyl-terminus, C-terminus, C-terminus, or COOH-terminus), which refers to the free carboxyl group (-COOH) of the last amino acid residue of the protein or peptide.
[0064] As used herein, the term "amino acid" refers to a compound that includes a carboxyl group (-COOH) and an amino group (-NH2). "Amino acid" refers to both natural and non-natural (i.e., synthetic) amino acids. Natural amino acids and their three-letter and one-letter abbreviations include: alanine (Ala; A); arginine (Arg, R); asparagine (Asn; N); aspartic acid (Asp; D); cysteine (Cys; C); glutamine (Gln; Q); glutamic acid (Glu; E); glycine (Gly; G); histidine (His; H); isoleucine (Ile; I); leucine (Leu; L); lysine (Lys; K); methionine (Met; M); phenylalanine (Phe; F); proline (Pro; P); serine (Ser; S); threonine (Thr; T); tryptophan (Trp; W); tyrosine (Tyr; Y); and valine (Val; V).
[0065] As used herein, an amino acid "substitution" refers to a polypeptide or protein in which one or more wild-type or naturally occurring amino acids are replaced at that amino acid residue by an amino acid that is different from the wild-type or naturally occurring amino acid. The substituted amino acid can be synthetic or naturally occurring. In some embodiments, the substituted amino acid is a naturally occurring amino acid selected from the group consisting of: A, R, N, D, C, Q, E, G, H, I, L, K, M, F, P, S, T, W, Y, and V. Substitution mutants can be described using an abbreviation system. For example, a substitution mutation in which the fifth (5th) amino acid residue is substituted can be abbreviated as "X5Y", where "X" is the substituted wild-type or naturally occurring amino acid, "5" is the position of the amino acid residue in the amino acid sequence of the protein or polypeptide, and "Y" is the substituted or non-wild-type or non-naturally occurring amino acid.
[0066] As used herein, “isolated” polypeptides, proteins, peptides, or nucleic acids are molecules that have been removed from their natural environment. It should also be understood that “isolated” polypeptides, proteins, peptides, or nucleic acids may be formulated with excipients (such as diluents) or adjuvants and are still considered isolated.
[0067] As used herein, when referring to nucleic acid molecules, peptides, polypeptides, or proteins, the term "recombinant" means a new combination or production of genetic material that is not known to exist in nature. Recombinant molecules can be produced using any well-known techniques existing in the field of recombinant technology, including but not limited to polymerase chain reaction (PCR), gene splicing (e.g., using restriction endonucleases), and solid-phase synthesis of nucleic acid molecules, peptides, or proteins.
[0068] As used herein, when referring to polypeptides or proteins, the term "domain" means a unique functional and / or structural unit within a protein. Domains are sometimes responsible for specific functions or interactions that contribute to the overall function of the protein. Domains can exist in a variety of biological contexts. Similar domains can be found in proteins with different functions. Alternatively, domains with low sequence homology (i.e., less than about 50%, less than about 40%, less than about 30%, less than about 20%, less than about 10%, less than about 5%, or less than about 1% sequence homology) may have the same function. In some embodiments, the Cas9 domain is a RuvC domain. In some embodiments, the Cas9 domain is an HNH domain. In some embodiments, the Cas9 domain is a Rec domain.
[0069] While this disclosure provides specific amino acid or nucleotide sequences, such as those shown in the sequence listing, it should be understood that a specific amino acid or nucleotide sequence includes variants with conserved sequence modifications, such as sequences having at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% homology with it, provided that the biological function or activity of that specific amino acid or nucleotide sequence is not lost.
[0070] As used herein, the term "conserved substitution" or "conserved sequence modification" refers to a nucleotide and amino acid sequence modification that does not eliminate the activity of the Cas9 protein encoded by the nucleotide sequence or containing the amino acid sequence. These conserved sequence modifications include conserved nucleotide and amino acid substitutions, as well as nucleotide and amino acid additions, insertions, and deletions. For example, modifications can be introduced into the sequence listings described herein using standard techniques known in the art, such as gene synthesis and PCR-mediated mutagenesis. Conserved sequence modifications include conserved amino acid substitutions, wherein an amino acid residue is replaced with an amino acid residue having a similar side chain. Families of amino acid residues having similar side chains are defined in the art. These families include amino acids with basic side chains (e.g., lysine, arginine, histidine), amino acids with acidic side chains (e.g., aspartic acid, glutamic acid), amino acids with nonpolar side chains (e.g., glycine, asparagine, glutamine, serine, threonine, tyrosine, cysteine, tryptophan), amino acids with nonpolar side chains (e.g., alanine, valine, leucine, isoleucine, proline, phenylalanine, methionine), amino acids with β-branched side chains (e.g., threonine, valine, isoleucine), and amino acids with aromatic side chains (e.g., tyrosine, phenylalanine, tryptophan, histidine). Therefore, predicted non-essential amino acid residues in the Cas9 protein are preferably substituted by another amino acid residue from the same side chain family. Methods for identifying conserved substitutions of nucleotides and amino acids that contribute to Cas9 protein activity are well known in the art. As described herein, substitutions are represented as “pre-substitution amino acid abbreviation - substitution site - post-substitution amino acid abbreviation”, for example, “S541A” indicates that serine at position 541 is substituted by alanine. The amino acids that can be conservatively substituted are shown in Table 1 below. As described in this article, the representation of amino acid deletion is "the single-letter abbreviation of the amino acid before deletion - the deletion site - Δ", for example, "V922Δ" indicates the deletion of valine at position 922.
[0071] Table 1. Examples of conserved substitutions of amino acids
[0072]
[0073] It should be understood that any amino acid mutation described herein from the first amino acid residue (e.g., A) to the second amino acid residue (e.g., T) (e.g., E923T) also includes mutations from the first amino acid residue to an amino acid residue similar to (e.g., conserved) the second amino acid residue. For example, a mutation from alanine to threonine (e.g., the E923T mutation) can also be a mutation from alanine to an amino acid similar in size and chemical properties to threonine (e.g., serine). Other similar amino acid pairs include, but are not limited to, the following: phenylalanine and tyrosine; asparagine and glutamine; methionine and cysteine; aspartic acid and glutamic acid; and arginine and lysine. Those skilled in the art will recognize that such conserved amino acid substitutions may have a minor effect on protein structure and may be well tolerated without impairing function. However, it should be understood that those skilled in the art will recognize other conserved amino acid residues, and any amino acid mutations to other conserved amino acid residues are also within the scope of this disclosure.
[0074] As used herein, the term "homology" has a universally accepted meaning in the field and is a central concept in comparative biology. The basic meaning of homology is that two samples being compared (e.g., an amino acid sequence or a nucleotide sequence) share a common ancestor. Generally, two traits (states) in two species are considered a pair of homologous traits if either of the following two conditions is met: 1. They are identical to a trait found in the ancestral groups of these species; 2. They are distinct traits with an ancestor-descendant relationship. Amino acid sequence homology can be determined using known methods. For example, amino acid sequence homology (%) can be determined using procedures commonly used in the field (e.g., BLAST, FASTA, etc.) according to initial settings. On the other hand, homology (%) can be determined using any algorithm known in the field, such as those by Needleman et al. (1970) (J. Mol. Biol. 48:444-453) and Myers and Miller (CABIOS, 1988, 4:11-17). Needleman et al.'s algorithm has been integrated into the GAP procedure of the GCG software package (available at www.gcg.com). Homology (%) can be determined, for example, using the BLOSUM 62 matrix or PAM250 matrix, and any of the following: gap weights (16, 14, 12, 10, 8, 6, or 4) and length weights (1, 2, 3, 4, 5, or 6). Additionally, Myers and Miller's algorithm has been integrated into the ALIGN procedure, which is part of the GCG sequence alignment software package. When using the ALIGN procedure to compare amino acid sequences, for example, a PAM120 weighted residue table, gap length penalty, and gap penalty can be used.
[0075] As used herein, the term "fusion protein" refers to a hybrid polypeptide comprising protein domains from at least two different proteins. For example, a protein domain may be located at the N-terminal (N-terminal) portion or the C-terminal (C-terminal) portion of the fusion protein, thereby forming an "N-terminal fusion protein" or a "C-terminal fusion protein," respectively. In an alternative embodiment, the fusion protein is a single-chain polypeptide that may be entirely encoded by a nucleic acid sequence and comprises at least two protein domains covalently linked by peptide linkages or optionally covalently linked by peptide linkers.
[0076] (II) Detailed Technical Solution
[0077] On one hand, this disclosure provides a Cas9 protein mutant comprising mutations in amino acid residues at positions 921, 922, and / or 923 of the amino acid sequence shown in SEQ ID NO:1, wherein the mutations are selected from one or more of the following:
[0078] (a) Deletion of amino acid residues;
[0079] (b) Substitution of amino acid residues;
[0080] (c) Insertion of amino acid residues;
[0081] Furthermore, the mutation can significantly reduce or eliminate the cleavage activity of the Cas9 protein on non-target chains.
[0082] As used herein, the terms “Cas9,” “Cas9 protein,” or “Cas9 nuclease” refer to an RNA-guided nuclease containing the Cas9 protein or fragments thereof (e.g., a protein containing the active or inactive DNA-cutting domain of Cas9, and / or the gRNA-binding domain of Cas9). Cas9 nucleases are sometimes also referred to as casn1 nucleases or CRISPR (clustered regular spaced short palindromic repeats)-associated nucleases. CRISPR is an adaptive immune system that provides protection against mobile genetic elements (viruses, transposable elements, and conjugating plasmids). CRISPR clusters contain spacer regions, sequences complementary to the preceding mobile element, and target invading nucleic acids. CRISPR clusters are transcribed and processed into CRISPR RNA (crRNA). In type II CRISPR systems, proper processing of pre-crRNA requires a trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and the Cas9 protein. tracrRNA serves as a guide for ribonuclease 3-assisted processing of pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA endo-cleaves linear or circular dsDNA targets complementary to the spacer region. The target strand not complementary to the crRNA is cleaved first, followed by 3'-5' exonuclease cleavage. In nature, DNA binding and cleavage typically require both proteins and two RNAs. However, a single guide RNA (“sgRNA” or simply “gNRA”) can be engineered to incorporate aspects of both crRNA and tracrRNA into a single RNA species. See, for example, Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science. 337:816-821 (2012), the entire contents of which are incorporated herein by reference.
[0083] The Cas9 protein mutants described in this disclosure can significantly reduce or eliminate the cleavage activity of Cas9 protein on non-target strands, while retaining or increasing the cleavage activity on target strands, and have multiple application prospects: (1) Combined with deaminases, to establish a single-base editing system. Single nucleotide variations lead to the occurrence of about 2 / 3 of human genetic diseases and are also the genetic basis for many important trait variations in animals and plants, but traditional DNA repair using DSB is difficult to achieve efficient and stable single-base mutations or single-base corrections. Therefore, base editors that can perform single-base editing without relying on DSB have emerged. Currently, base editors mainly include two components: Cas9 protein mutants that only cleave the target strand and deaminases, of which the combination with adenine deaminase is the adenine base editor (ABE), and the combination with cytosine deaminase is the cytosine base editor (CBE). Base editors have shown great potential in basic research and gene therapy. (2) Paired Cas9 single-cleavage enzyme gene knockout technology. Off-target effects caused by non-specific binding of Cas9-sgRNA have always been one of the problems of CRISPR / Cas9 gene editing systems. Taking advantage of the fact that Cas9 mutants can only perform single-strand cutting, two Cas9 protein mutants are used together to act on the opposite strand of adjacent gene sites to generate a single cut, forming a double-strand break (DSB). However, the non-specific binding of two Cas9 protein mutants is unlikely to occur at the same off-target site, so DSB will not be caused. The generated single cut is seamlessly repaired using the single-strand break repair pathway. Paired Cas9 single-cut enzyme gene knockout technology reduces non-specificity by 50 to 1500 times. The inventors used TREX2 nuclease to expand the paired Cas9 single-cut enzyme gene knockout technology, which further improved gene editing efficiency while reducing off-target effects. (3) Combining with mutations that reduce or eliminate the target strand (e.g., H840A) can obtain dCas9 (dead Cas9) without cutting activity, which can be used to activate or inhibit gene expression. CRISPRa (CRISPRactivation) technology allows dCas9 to bind to different transcriptional activators (such as VP64), thereby upregulating the expression levels of target genes. Compared to traditional overexpression techniques, CRISPRa promotes gene expression by efficiently activating endogenous promoters, eliminating the need for additional exogenous expression elements and overcoming limitations imposed by gene transcript size. dCas9-mediated gene regulation is reversible and does not permanently modify genomic DNA. Furthermore, multiple sgRNAs can be designed to simultaneously regulate multiple target genes. Additionally, dCas9 can be fused with fluorescent proteins for targeted gene expression.
[0084] The sequence and structure of the Cas9 nuclease are well known to those skilled in the art (see, for example, Ferretti JJ et al., Proceedings of the
[0085] National Academy of Sciences of the United States of America, 2001, 98(8):4658–4663; Deltcheva E. et al., Nature, 2011, 471:602-607; Jinek M. et al., Science, 2012, 337:816-821, the entire contents of which are incorporated herein by reference. Cas9 orthologs have been described in a variety of species, including but not limited to Cas9 nucleases of Streptococcus pyogenes and Streptococcus thermophilus. Other suitable Cas9 nucleases and sequences will be apparent to those skilled in the art based on this disclosure, and such Cas9 nucleases and sequences include Cas9 sequences from species and loci disclosed in Chylinski K et al., RNA Biology, 2013, 10(5):726-737, the entire contents of which are incorporated herein by reference.
[0086] In some embodiments, the Cas9 protein mutant comprises an amino acid sequence having at least 80%, at least 82%, at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or at least 99.9% homology to the amino acid sequence shown in SEQ ID NO: 1. In some implementations, Cas9 refers to Cas9 derived from the following: *Streptococcus pyogenes* (NCBI Refs: NC_017053.1), *Streptococcus dysgalactiae*, *Streptococcus mutans*, *Staphylococcus aureus*, *Klebsiella pneumoniae*, *Corynebacterium ulcerans* (NCBI Refs: NC_015683.1, NC_017317.1); *Corynebacterium diphtheriae* (NCBI Refs: NC_016782.1, NC_016786.1); and *Spiroplasma syrphidicola* (NCBI Refs: NC_017053.1). Ref: NC_021284.1); Prevotella intermedia (NCBI Ref: NC_017861.1); Spiroplasma taiwanense (NCBI Ref: NC_021846.1); Streptococcus iniae (NCBI Ref: NC_021314.1); Belliella baltica (NCBI Ref: NC_018010.1); Psychroflexus torquisi (NCBI Ref: NC_018721.1); Streptococcus thermophilus (NCBI Ref: YP_820832.1); Listeria innocua (NCBI Ref: NP_472073.1); Campylobacter jejuni Jejuni (NCBI Ref:YP_002344900.1); or Neisseria meningitidis (NCBI Ref:YP_002342100.1).In some embodiments, the Cas9 protein is derived from *Streptococcus pyogenes* and may be simply referred to as SpCas9. The amino acid sequence of the wild-type SpCas9 protein is shown in SEQ ID NO: 1. In some embodiments, the Cas9 protein is a variant of the natural Cas9 protein, such as some SpCas9 variants obtained through directed protein evolution, such as SpCas9 VQR, SpCas9 VRER, or SpCas9 NG (ngSpCas9), which recognize NGA, NGCG, and NG as PAM sequences, respectively.
[0087] In some implementations, the Cas9 nuclease has an inactive (e.g., inactivated) DNA-cutting domain. The nuclease-inactivated Cas9 protein can be interchangeably referred to as the "dCas9" protein ("dead" Cas9). The DNA-cutting domain of Cas9 is known to comprise two subdomains: the HNH nuclease subdomain and the RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA, while the RuvC1 subdomain cleaves the non-complementary strand. Mutations within these subdomains can silence the nuclease activity of Cas9.
[0088] It should be understood that the Cas9 described in this disclosure includes the aforementioned natural Cas9 derived from different bacterial species and variants obtained through artificial directed evolution. The positions of the amino acid residues described in this disclosure include the positions of their corresponding amino acid residues. For ease of description, this disclosure uses the Cas9 protein shown in SEQ ID NO:1 to determine specific amino acid residue positions. Determining the positions of corresponding amino acid residues in another Cas9 protein is known to those skilled in the art. For example, the amino acid sequence of another Cas9 protein is compared with the sequence disclosed in SEQ ID NO:1, and based on this comparison, the Niederman-Onsch algorithm (Niederman and Onsch, 1970, Journal of Molecular Biology 48:443-453) implemented in the Nieder program of the EMBOSS package (EMBOSS: European Molecular Biology Open Software Suite, Rice et al., 2000, Trends in Genetics 16:276-277) (preferably version 5.0.0 or later) is used to determine the amino acid position number corresponding to any amino acid residue in the polypeptide disclosed in SEQ ID NO:2. The parameters used are a vacancy opening penalty of 10, a vacancy extension penalty of 0.5, and an EBLOSUM62 (the EMBOSS version of BLOSUM62) substitution matrix. In some embodiments, the mutations in the Cas9 protein mutants of this disclosure include:
[0089] L921P;
[0090] V922Δ; and / or
[0091] E923△ or E923T.
[0092] In some embodiments, the Cas9 protein mutant comprises a selection of mutations or combinations thereof:
[0093] (1)V922Δ;
[0094] (2)E923Δ;
[0095] (3) L921P and V922Δ;
[0096] (4) V922Δ and E923T;
[0097] (5)L921P;
[0098] (6)E923T;
[0099] (7) L921P and E923T;
[0100] (8) L921P and E923Δ;
[0101] (9) V922Δ and E923Δ;
[0102] (10) L921P, V922Δ and E923T; and
[0103] (11)L921P, V922Δ and E923Δ.
[0104] In some embodiments, the Cas9 protein mutant is selected from the following mutations or combinations of mutations:
[0105] (1)V922Δ;
[0106] (2)E923Δ;
[0107] (3) L921P and V922Δ; and
[0108] (4) V922Δ and E923T.
[0109] In some embodiments, the Cas9 protein mutant comprises a RuvC and an HNH domain, wherein the HNH domain comprises a mutation that reduces or eliminates non-target strand cleavage activity. In some embodiments, the mutation that reduces or eliminates non-target strand cleavage activity comprises H840A.
[0110] On the other hand, this disclosure provides a Cas9 protein mutant comprising the amino acid sequence fragments shown in SEQ ID NO: 35, 37, 39, 41, wherein the mutant significantly reduces or eliminates the cleavage activity of the Cas9 protein on non-target chains.
[0111] In some embodiments, the Cas9 protein mutant comprises an amino acid sequence fragment having at least 80%, at least 82%, at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or at least 99.9% homology to the amino acid sequence fragments shown in SEQ ID NO: 35, 37, 39, 41.
[0112] In some embodiments, the remaining amino acid sequence of the Cas9 protein mutant has at least 80%, at least 82%, at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or at least 99.9% homology to the corresponding sequence in the amino acid sequence shown in SEQ ID NO: 1. In some embodiments, the remaining amino acid sequence of the Cas9 protein mutant is identical to the corresponding sequence in the amino acid sequence shown in SEQ ID NO: 1.
[0113] In some embodiments, the Cas9 protein mutant comprises a RuvC and an HNH domain, wherein the HNH domain contains a mutation that reduces or eliminates target strand cleavage activity. In some embodiments, the mutation that reduces or eliminates target strand cleavage activity comprises H840A.
[0114] On the other hand, this disclosure provides a fusion protein comprising the aforementioned Cas9 protein mutant and a heterologous domain fused with the Cas9 protein mutant.
[0115] The term "heterologous domain" refers to any protein domain that is not present in the Cas9 protein mutant described in this disclosure, such as a domain that is different from the domain present in the Cas9 protein mutant described in this disclosure, or a domain derived from other proteins such as deaminases, methyltransferases, transcription activators, etc.
[0116] In some embodiments, the Cas9 protein may also be modified to include at least one heterologous domain, i.e., Cas9 is fused to one or more heterologous domains. In cases where two or more heterologous domains are fused to Cas9, the two or more heterologous domains may be identical or they may be different. In some embodiments, the heterologous domain is selected from nuclear localization signal domains, cell penetration domains, marker or reporter domains that facilitate detection, DNA or RNA deaminase domains, uracil-DNA glycosylase domains, reverse transcriptase domains, recombinase domains, RNA aptamer-binding domains, nuclease domains, methyltransferase domains, methyltransferase domains, acetyltransferase domains, transcription activator domains, and transcription repressor domains. In some embodiments, the heterologous domain is selected from exonuclease domains and deaminase domains. In some embodiments, the heterologous domain is selected from cytosine deaminase, adenine deaminase, and DNA-binding-deficient 3' to 5' exonucleases.
[0117] In some embodiments, the Cas9 fusion protein provided herein comprises the full-length amino acid sequence of the Cas9 protein, such as one of the Cas9 protein sequences provided above. However, in other embodiments, the fusion protein provided herein does not comprise the full-length Cas9 sequence, but only fragments thereof. For example, in some embodiments, the Cas9 fusion protein provided herein comprises a Cas9 fragment that binds crRNA and tracrRNA or sgRNA, but does not contain a functional nuclease domain, for example, it contains only a truncated form of the nuclease domain or no nuclease domain at all. Exemplary amino acid sequences of suitable Cas9 domains and Cas9 fragments are provided herein, and other suitable sequences of Cas9 domains and Cas9 fragments will be apparent to those skilled in the art. In some embodiments, the Cas9 fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the amino acid length of the corresponding wild-type Cas9 protein. In some embodiments, the Cas9 fragment contains at least 100 amino acids. In some embodiments, the Cas9 fragment is at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, 1300, 1350, 1400, 1450, 1500, 1550 or at least 1600 amino acids of the corresponding wild-type Cas9 protein. In some embodiments, the Cas9 fragment comprises an amino acid sequence having at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, or at least 1200 of the same consecutive amino acid residues as the corresponding wild-type Cas9 protein.
[0118] In some embodiments, one or more heterologous domains may be fused to the N-terminus, C-terminus, internal position, or a combination thereof of the Cas9 protein mutant. Fusion may be direct via chemical bonds or indirect via one or more linkers. A linker is a chemical group that is covalently linked to one or more other chemical groups. Suitable linkers include amino acids, peptides, nucleotides, nucleic acids, organic linker molecules (e.g., maleimide derivatives, N-ethoxybenzylimidazole, biphenyl-3,4',5-tricarboxylic acid, p-aminobenzyloxycarbonyl, etc.), disulfide linkers, and polymer linkers (e.g., PEG). Linkers may include one or more spacer groups, including but not limited to alkylene, alkenylene, alkyneene, alkyl, alkenyl, alkyne, alkoxy, aryl, heteroaryl, aralkyl, areneyl, arynyl, etc. Linkers may be neutral or carry a positive or negative charge. Additionally, linkers may be cleavable, such that the covalent bond connecting the linker to another chemical group can be broken or cleaved under certain conditions, including pH, temperature, salt concentration, light, catalyst, or enzyme. In some embodiments, the linker may be a peptide linker. Peptide linkers can be flexible amino acid linkers (e.g., comprising small, nonpolar or polar amino acids). Non-limiting examples of flexible linkers include LEGGGS (SEQ ID NO:83), TGSG (SEQ ID NO:84), GGSGGGSG (SEQ ID NO:85), and (GGGGS). 1-4 (SEQ ID NO:86). Alternatively, the peptide linker can be a rigid amino acid linker. Such linkers include (EAAAK). 1-4 (SEQ ID NO:87), A(EAAAK) 2-5 A (SEQ ID NO:88), PAPAP (SEQ ID NO:89), and (AP) 6-8 Other examples of suitable connectors are well known in the art, and procedures for designing connectors are readily available (e.g., Crasto et al., Protein Eng., 2000, 13(5):309-312).
[0119] On the other hand, this disclosure provides a complex comprising the aforementioned Cas9 protein mutant or the aforementioned fusion protein, and guide RNA. The guide RNA comprises a polynucleotide consisting of a nucleotide sequence complementary to a target nucleotide sequence upstream of the PAM (Proto-spacer Adjacent Motif) sequence in the target double-stranded polynucleotide, from one base upstream to more than 20 bases upstream and less than 24 bases upstream. In some embodiments, the guide RNA is a sequence of about 15-100 nucleotides comprising at least 10 consecutive nucleotides complementary to the target sequence. In some embodiments, the guide RNA is 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleotides long. In some embodiments, the guide RNA comprises a sequence of 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 consecutive nucleotides complementary to the target sequence. In some embodiments, the target sequence is a DNA sequence. In some embodiments, the target sequence is a sequence from the mammalian genome. In some implementations, the target sequence is a sequence in the human genome. In some implementations, the 3' end of the target sequence is not directly adjacent to the canonical PAM sequence.
[0120] As used herein, “guide RNA,” also known as “gRNA” or “guide RNA,” generally refers to an RNA sequence or molecule (or a set of RNA molecules) that can bind to Cas proteins and help target Cas proteins to a specific location within a target polynucleotide (e.g., DNA or RNA). Guide RNAs can contain crRNA segments and tracrRNA segments. As used herein, the term “crRNA” or “crRNA segment” refers to an RNA molecule or portion thereof that includes a guide sequence, stem sequence, and optionally a 5’ overhang sequence targeting the polynucleotide. The term “tracrRNA” or “tracrRNA segment” refers to an RNA molecule or portion thereof that includes a protein-binding segment (e.g., a protein-binding segment capable of interacting with a CRISPR-associated protein, such as Cas9). Guide RNAs also encompass single guide RNAs (sgRNAs), where the crRNA segment and the tracrRNA segment are located within the same RNA molecule. Guide RNAs also collectively comprise groups of two or more RNA molecules, where the crRNA segment and the tracrRNA segment are located within separate RNA molecules.
[0121] In some embodiments, the Cas9 protein or fusion protein and the guide RNA are mixed under mild conditions and cultured to form the complex, preferably for a period of 0.5 hours to 1 hour. The resulting complex is stable and remains stable even after standing at room temperature for several hours. In some embodiments, the Cas9 protein or fusion protein and the guide RNA form a complex on the target nucleic acid. The Cas9 protein or fusion protein recognizes the PAM sequence and binds to the target nucleic acid at a binding site upstream of the PAM sequence. If the Cas9 protein or fusion protein has endonuclease activity, it cleaves the polynucleotide at this site. The Cas9 protein or fusion protein recognizes the PAM sequence and, starting from the PAM sequence, anneals the double helix structure of the target double-stranded polynucleotide and the complementary nucleotide sequence in the guide RNA to the target double-stranded polynucleotide, thereby unwinding a portion of the double helix structure of the target double-stranded polynucleotide. At this point, the Cas9 protein or fusion protein cleaves the phosphodiester bond of the target double-stranded polynucleotide at the cleavage site upstream of the PAM sequence and / or at the cleavage site upstream of the sequence complementary to the PAM sequence.
[0122] On the other hand, this disclosure provides a polynucleotide that encodes the aforementioned Cas9 protein mutant or the aforementioned fusion protein.
[0123] The polynucleotides disclosed herein can be in DNA or RNA form. DNA form includes cDNA, genomic DNA, or artificially synthesized DNA. DNA can be single-stranded or double-stranded. DNA can be a coding strand or a non-coding strand. In some embodiments, the polynucleotide sequence encoding recombinant Cas9 is codon-optimized for expression in eukaryotic cells. In some embodiments, the polynucleotide sequence encoding stiCas9 is codon-optimized for expression in animal cells. In some embodiments, the polynucleotide sequence encoding recombinant Cas9 is codon-optimized for expression in human cells. In some embodiments, the polynucleotide sequence encoding recombinant Cas9 is codon-optimized for expression in plant cells. Codon optimization involves adjusting codons to match the abundance of tRNA in the expression host to improve the yield and efficiency of recombinant or heterologous protein expression. Codon optimization methods are standard practices in the field and can be performed using software programs such as the codon optimization tool from Integrated DNA Technologies, the codon usage table analysis tool from Entelechon, the Blue Heron software from Genemaker, the Gene Forge software from Aptagen, the DNA Builder software, general codon usage analysis software, the publicly available OPTIMIZER software, and the Optimum Gene algorithm from Genscript.
[0124] On the other hand, this disclosure provides a vector comprising the aforementioned polynucleotide. In some embodiments, the polynucleotide is also operatively linked to a promoter, enhancer, and / or terminator. In some embodiments, the promoter includes: a constitutive promoter, an inducible promoter, a broad-spectrum expression promoter, or a tissue-specific promoter.
[0125] In some embodiments, the vector is an expression vector, more specifically a recombinant expression vector. Suitable expression vectors include viral expression vectors (e.g., viral vectors based on viruses such as vaccinia virus, poliovirus, adenovirus, adeno-associated virus, SV40, herpes simplex virus, human immunodeficiency virus), retroviral vectors (e.g., murine leukosis virus, spleen necrosis virus, and vectors derived from retroviruses such as Rous sarcoma virus, Harvey sarcoma virus, avian leukosis virus, lentivirus, human immunodeficiency virus, myeloproliferative sarcoma virus, and mammary tumor virus), etc.
[0126] On the other hand, this disclosure provides a kit comprising the aforementioned Cas9 protein mutant, the aforementioned fusion protein, the aforementioned complex, the aforementioned polynucleotide, or the aforementioned vector; preferably, the kit further comprises an expression construct encoding a guide RNA backbone. The expression construct contains a cloning site that allows the cloning of a nucleic acid sequence identical to or complementary to the target sequence into the guide RNA backbone.
[0127] As used in this article, the “guide RNA backbone” refers to the sequence within the guide RNA responsible for Cas9 binding, excluding the spacer sequence used to guide Cas9 to the target DNA.
[0128] On the other hand, this disclosure provides a cell comprising the aforementioned Cas9 protein mutant, the aforementioned fusion protein, the aforementioned complex, the aforementioned polynucleotide, or the aforementioned vector.
[0129] In some implementations, cells can be transfected with appropriate molecules (i.e., proteins, DNA, and / or RNA). Suitable transfection methods include nuclear transfection (or electroporation), calcium phosphate-mediated transfection, cationic polymer transfection (e.g., DEAE-glucan or polyethyleneimine), viral transduction, virion transfection, viral particle transfection, liposome transfection, cationic liposome transfection, immunoliposome transfection, non-liposomal lipid transfection, dendritic molecule transfection, heat shock transfection, magnetic transfection, lipid transfection, gene gun delivery, puncture transfection, sonoporosis, phototransfection, and proprietary reagent-enhanced nucleic acid uptake. Transfection methods are well known in the art (see, for example, “Current Protocols in Molecular Biology”, Ausubel et al., John Wiley & Sons, New York, 2003, or “Molecular Cloning: A Laboratory Manual”, Sambrook & Russell, Cold Spring Harbor Press, Cold Spring Harbor, NY, 3rd ed., 2001). In other embodiments, molecules can be introduced into cells via microinjection. For example, molecules can be injected into the cytoplasm or nucleus of the target cell. The amount of each molecule introduced into the cell can vary, but those skilled in the art are familiar with the means used to determine the appropriate amount. Various molecules can be introduced into the cell simultaneously or sequentially. For example, a modified Cas9 protein (or its encoded nucleic acid) and a donor polynucleotide can be introduced simultaneously. Alternatively, one can be introduced first, followed by another.
[0130] Various cells are suitable for use in the methods disclosed herein, including prokaryotic cells (e.g., bacteria) and eukaryotic cells (e.g., animal cells, insect cells, and plant cells). For example, the cells can be human cells, non-human mammalian cells, non-mammal vertebrate cells, invertebrate cells, insect cells, plant cells, yeast cells, or single-celled eukaryotes. In some embodiments, the cells can be single-celled embryos. For example, non-human mammalian embryos include rat, hamster, rodent, rabbit, cat, dog, sheep, pig, cattle, horse, and primate embryos. In other embodiments, the cells can be stem cells, such as embryonic stem cells, ES-like stem cells, fetal stem cells, adult stem cells, etc. In one embodiment, the stem cells are not human embryonic stem cells. Furthermore, stem cells can include those prepared using the techniques disclosed in WO2003 / 046141, which is incorporated herein in its entirety, or Chung et al. (Cell Stem Cell, 2008, 2:113-117). Cells can be in vitro (i.e., in culture), ex vivo (i.e., within a tissue isolated from an organism), or in vivo (i.e., within an organism). In exemplary embodiments, the cells are mammalian cells or mammalian cell lines. In particular embodiments, the cells are human cells or human cell lines.
[0131] Other aspects of this disclosure include animals modified to encode nucleic acids or vectors as described above, or animals permanently modified by a modified SpCas9 variant of this disclosure. For example, the animal may be a model organism (i.e., *Drosophila melanogaster*, a mouse, a mosquito, a rat), or the animal may be a farm animal, a farmed fish, or a pet. As another example, the animal may be a vector for at least one disease. As yet another example, the organism may be a vector for human diseases (i.e., a mosquito, a tick, a bird).
[0132] Other aspects of this disclosure include plants modified using the nucleic acids or vectors described above, or plants transiently or permanently modified by SpCas9 variants modified by this disclosure. For example, the plants can be crops (i.e., rice, soybeans, wheat, tobacco, cotton, alfalfa, low-erucic acid rapeseed, corn, sugar beets, etc.).
[0133] On the other hand, this disclosure provides a pharmaceutical composition comprising the aforementioned Cas9 protein mutant, the aforementioned fusion protein, the aforementioned complex, the aforementioned polynucleotide, the aforementioned carrier, or the aforementioned cell. In some embodiments, the pharmaceutical composition further comprises at least one pharmaceutically acceptable excipient. Pharmaceutically acceptable excipients typically include inactive ingredients used as a medium (e.g., water, capsule shell, etc.), a diluent, or a component constituting a dosage form or pharmaceutical composition comprising a drug such as a therapeutic agent. Pharmaceutically acceptable excipients also include typically inactive ingredients that impart adhesive (i.e., adhesive), disintegrant (i.e., disintegrant), lubricant (lubricant), and / or other functions (i.e., solvent, surfactant, etc.) to the composition.
[0134] On the other hand, this disclosure provides the use of the aforementioned Cas9 protein, the aforementioned fusion protein, the aforementioned complex, the aforementioned polynucleotide, the aforementioned vector, the aforementioned kit, or the aforementioned cell or the aforementioned pharmaceutical composition in nucleic acid modification.
[0135] The aforementioned Cas9 protein mutants, fusion proteins, complexes, polynucleotides, vectors, kits, cells, or pharmaceutical compositions disclosed herein can be used in a variety of therapeutic, diagnostic, industrial, and research applications. In some embodiments, the contents of this disclosure can be used to modify any target chromosomal sequence in cells, animals, or plants to construct gene function models and / or study gene function, investigate target genetic or epigenetic conditions, or study biochemical pathways involving various diseases or conditions. For example, transgenic organisms can be generated to construct models of diseases or conditions in which the expression of one or more nucleic acid sequences associated with the disease or condition is altered. Disease models can be used to study the effects of mutations on organisms, study the development and / or progression of diseases, study the effects of pharmaceutically active compounds on diseases, and / or evaluate the efficacy of potential gene therapy strategies.
[0136] In some embodiments, this disclosure can be used to perform efficient and cost-effective functional genomic screening, which can be used to study the function of genes involved in specific biological processes and how any alterations in gene expression can affect biological processes, or to perform saturation or depth scan mutagenesis of genomic loci linked to cell phenotypes. For example, saturation or depth scan mutagenesis can be used to determine the critical minimum characteristics and discrete vulnerabilities of functional elements required for gene expression, drug resistance, and disease reversal.
[0137] In some implementations, this disclosure can be used for diagnostic testing to determine the presence of a disease or condition and / or to determine treatment options. Examples of suitable diagnostic tests include detecting specific mutations in cancer cells (e.g., specific mutations in EGFR, HER2, etc.), detecting specific mutations associated with specific diseases (e.g., trinucleotide repeats, mutations in β-globin associated with sickle cell disease, specific SNPs, etc.), detecting hepatitis, detecting viruses (e.g., Zika virus), and so on.
[0138] In some embodiments, this disclosure can be used to correct genetic mutations associated with specific diseases or conditions, such as correcting mutations in globin genes associated with sickle cell disease or thalassemia, correcting mutations in adenosine deaminase genes associated with severe combined immunodeficiency (SCID), reducing the expression of the causative gene HTT in Huntington's disease, or correcting mutations in rhodopsin genes for the treatment of retinitis pigmentosa. Such modifications can be performed in vitro.
[0139] In some embodiments, this disclosure can be used to generate crop plants with improved traits or increased resistance to environmental stresses. This disclosure can also be used to generate farm animals or production animals with improved traits. For example, pigs possess many characteristics that make them attractive as biomedical models, particularly in regenerative medicine or xenotransplantation.
[0140] On the other hand, this disclosure provides a method for modifying a target nucleic acid, the method comprising contacting the target nucleic acid with the following:
[0141] (a) the aforementioned Cas9 protein mutant or the aforementioned fusion protein, and guide RNA; or
[0142] (b) The aforementioned complex.
[0143] On the other hand, this disclosure provides a method for modifying a target nucleic acid, the method comprising introducing the aforementioned Cas9 protein mutant, the aforementioned fusion protein, the aforementioned complex, the aforementioned polynucleotide, or the aforementioned vector into a cell or organism containing the target nucleic acid.
[0144] In some embodiments, the target nucleic acid contains a sequence associated with a disease or condition. In some embodiments, the target nucleic acid contains a point mutation associated with a disease or condition. In some embodiments, activity of the Cas9 protein, Cas9 fusion protein, or complex leads to correction of the point mutation. In some embodiments, the target nucleic acid contains a T→C point mutation associated with a disease or condition, and wherein deamination of the mutant C base results in a sequence not associated with the disease or condition. In some embodiments, the target nucleic acid encodes a protein, and wherein the point mutation is located in a codon and results in a change in the amino acid encoded by the mutant codon compared to the wild-type codon. In some embodiments, deamination of the mutant C results in a change in the amino acid encoded by the mutant codon. In some embodiments, deamination of the mutant C results in a codon encoding a wild-type amino acid. In some embodiments, the contact occurs in a subject. In some embodiments, the subject has or has been diagnosed with a disease or condition. In some embodiments, the disease or condition is cystic fibrosis, phenylketonuria, epidermolytic hyperkeratosis (EHK), Charchar-Marie-Toot disease type 4J, neuroblastoma (NB), von Willebrand disease (vWD), congenital myotonia, hereditary renal amyloidosis, dilated cardiomyopathy (DCM), hereditary lymphedema, familial Alzheimer's disease, HIV, prions, chronic infantile neurocutaneous joint syndrome (CINCA), desmin-associated myopathy (DRM), or neoplastic diseases associated with mutant PI3KCA protein, mutant CTNNB1 protein, mutant HRAS protein, or mutant p53 protein.
[0145] On the other hand, this disclosure provides a method for treating a subject diagnosed with a disease related to or caused by a gene mutation, comprising administering to the subject a therapeutically effective amount of the aforementioned Cas9 protein, the aforementioned fusion protein, the aforementioned complex, the aforementioned polynucleotide, the aforementioned vector, the aforementioned cell, and / or the aforementioned pharmaceutical composition. For example, in some embodiments, a method is provided comprising administering to a subject suffering from such a disease (e.g., cancer associated with the PI3KCA point mutation as described above) an effective amount of a corrected point mutation or a Cas9 deaminase fusion protein that introduces a mutation into a disease-related gene. In some embodiments, the disease is a proliferative disease. In some embodiments, the disease is a genetic disease. In some embodiments, the disease is a neoplastic disease. In some embodiments, the disease is a metabolic disease. In some embodiments, the disease is a lysosomal storage disease. Other diseases that can be treated by corrected point mutations or by introducing inactivating mutations into disease-related genes are known to those skilled in the art, and this disclosure is not limited in this respect.
[0146] As used herein, the term "effective dose" or "therapeutic effective dose" includes an amount sufficient to improve or prevent the symptoms or condition of a medical condition. An effective dose also means an amount sufficient to allow or facilitate diagnosis. The effective dose for a particular patient or veterinary subject can vary depending on factors such as the condition to be treated, the patient's overall health, the route and dosage of administration, and the severity of side effects. An effective dose can be the maximum dose or administration regimen that avoids significant side effects or toxicity.
[0147] As used herein, the terms “subject,” “individual,” or “patient” are used interchangeably and refer to a vertebrate, preferably a mammal, more preferably a human. Mammals include, but are not limited to, mice, apes, humans, farm animals, sports animals, and pets.
[0148] The amount of substance administered can depend on the subject being treated, the subject's age, health status, sex, and weight, the type of concurrent treatment (if any), the severity of the condition, the nature of the desired effect, the manner and frequency of treatment, and the prescribing physician's judgment. The frequency of administration can also depend on the pharmacodynamic effect on arterial oxygen partial pressure. However, the optimal dose can be adjusted for individual subjects, as understood by those skilled in the art and determined without tolerance experiments. This typically involves adjusting the standard dose (e.g., reducing the dose if the patient is underweight).
[0149] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Unless otherwise specified, the technical means used in the embodiments are conventional means well known to those skilled in the art. Obviously, the described embodiments are some embodiments of this disclosure, but not all embodiments. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. The following embodiments are used to further illustrate this disclosure, but should not be construed as limiting this disclosure. Any changes, modifications, substitutions, combinations, or simplifications made without departing from the spirit and principle of this disclosure should be considered equivalent substitutions and are included within the protection scope of this disclosure.
[0150] Example 1. Screening for Cas9 protein mutants with novel functions
[0151] (1) Mouse embryonic stem cell HR reporter system
[0152] This patent utilizes an HR reporter system derived from Chandramouly G et al. (BRCA1 and CtIP suppress long-tract gene conversion between sister chromatids. Nature Communications, 2013, 4(1):2404) and introduces it into mouse embryonic stem cells. The mouse embryonic stem cell HR reporter system used can be used to measure Cas9 and nCas9-induced short-tract gene conversion (STGC) and long-tract gene conversion (LTGC) and their bias. The reporter system contains two mutant inactivated GFP genes: the first GFP (TrGFP) is missing its 5' end, and the second GFP (I-SceI-GFP) is interrupted by the recognition site of the rare 18 bp endonuclease I-SceI. Between the two mutant GFPs, two inverted artificial exon expression frames A and B of the RFP gene are inserted, and only cells with rearranged HR will emit light. When a single-terminal DSB is induced by Cas9 at or near the I-SceI site or by nCas9-induced replication-coupled DSB, the cell can utilize the TrGFP of an adjacent sister chromatid as a template for gene conversion during the G2 / S phase of the cell cycle to generate green fluorescence (GFP). + ) cells or green and red fluorescent (GFP) + RFP + Cells, which emit different fluorescence can be detected by flow cytometry. The gene conversion produced by HR includes two forms: 1) STGC, with GFP as the product. + Cells; 2) LTGC, product is GFP + RFP + Cellular DNA repair. STGC typically converts 50-200 bp, while LTGC can convert DNA sequences exceeding 1000 bp. This reporter system can also be used to measure and analyze the selectivity of STGC and LTGC, the DSB repair products induced by different Cas9 nucleases within cells.
[0153] (2) Screening for Cas9 protein mutants
[0154] In mouse embryonic stem cells, SpCas9 knock-in donors containing homologous sequences at both ends and SpCas9 overexpression vectors (vector structure shown) were simultaneously introduced. Figure 2As shown in the figure, a single copy of the SpCas9 gene was precisely knocked into the Rosa26 site using a homologous recombination gene knock-in method, along with an expression vector targeting the mouse Rosa26 site. The amino acid sequence of SpCas9 is shown in SEQ ID NO: 1, the gene sequence of SpCas9 is shown in SEQ ID NO: 2, and the sgRNA spacer sequence used was 5'-ACTCCAGTCTTTCTAGAAGA-3' (SEQ ID NO: 3). After the single clones grew, targeted PCR was performed on the genome of the single clone cell line with the SpCas9 gene precisely knocked into the site using primer pairs TGGAAGAAGCTGTCGTCCACCTTG (SEQ ID NO: 4, forward primer) and CGAGACACGGATCGACCTGTCT (SEQ ID NO: 5, reverse primer). Combined with quantitative PCR, single clone cell lines with precisely inserted single copies of the SpCas9 gene were screened.
[0155] Using the Lipofectamine 2000 liposome transfection method, the LbCas12a expression plasmid (from Addgene) and the expression plasmid pU6-LbCas12a-sgRNA targeting a specific target of the SpCas9 gene were introduced into the cell line. The LbCas12a-sgRNA targets the SpCas9 sequence, and the sgRNA sequence used is 5'-TGATCTGCCGGGTTTCCACCAGC-3' (SEQ ID NO: 6), with a PAM of 5'-TTTG-3' (SEQ ID NO: 7). SpCas9 insertion / deletion mutations were induced, generating various insertion / deletion mutants and forming a pool of insertion / deletion mutant cells.
[0156] The pU6-SpCas9-crRNA vector expressing SpCas9 sgRNA (g1-1) was transfected into a cell pool using the Lipofectamine 2000 liposome transfection method. The sgRNA sequence used was g1-1:5'-GATAACAGGGTAATCAAGG-3' (SEQ ID NO: 8). The viable SpCas9 mutants remaining in the cell pool will utilize this sgRNA to cleave the HR reporter system. After cleavage of the HR reporter system, the cells will produce GFP and RFP fluorescence.
[0157] Cell pools transfected with sgRNA were sorted by flow cytometry to separate GFP-positive cells and GFP / RFP double-positive cells. Genomic DNA was extracted from unsorted cells using a genomic DNA extraction kit (Novozymes: DC102-01). The cleavage site was amplified by PCR using primer pairs (CP4-F: GCTGAACGCCAAGCTGATTAC (SEQ ID NO: 9); CP4-R: ATCCTTCCGGAAATCGGACAC (SEQ ID NO: 10)) to obtain the DNA sequence near the editing site. Next-generation sequencing was performed at Novogene to obtain DNA mutation enrichment data. The ratios of different mutant sequences in unsorted cell pools, GFP-positive cells, and GFP / RFP double-positive cells were statistically analyzed to determine the enrichment levels of different insertion / deletion variants. Mutations with higher enrichment levels than wild-type SpCas9 were considered potential usable SpCas9 insertion / deletion mutants. The SpCas9V922△, SpCas9E923△, SpCas9V922△ / E923T, and SpCas9L921P / V922△ mutants were identified as potential usable insertion / deletion mutants.
[0158] Using the SpCas9 overexpression vector plasmid as a template, PCR amplification was performed using the following primers (introducing corresponding mutations through the primers) to generate different fragments:
[0159] (1)SpCas9-F:5'-CCATGGCCTACCCCTACGAC-3' (SEQ ID NO: 11)
[0160] (2)V922△-R:5'-GCCGGGTTTCCAGCTGTCTCTTGATGAAGCC-3' (SEQ ID NO: 12)
[0161] (3)E923Δ-R:5'-TCTGCCGGGTCACCAGCTGTCTCTTGATGAAGC-3' (SEQ ID NO: 13)
[0162] (4)V922△ / E923T-R:5'-CCGGGTTTCCGGCTGTCTCTTGATGAAGCCGG-3'(SEQ ID NO: 14)
[0163] (5)L921P / V922△-R:5'-CTGCCGGGTTGTCAGCTGTCTCTTGATGAAGCCG-3' (SEQ ID NO: 15)
[0164] (6)SpCas9-R:5'-GCTGATCAGCGGGTTTAAACG-3' (SEQ ID NO: 16)
[0165] (7)V922△-F:5'-AAGAGACAGCTGGAAACCCGGCAGATCACAAAG-3' (SEQ ID NO: 17)
[0166] (8)E923△-F:5'-ACAGCTGGTGACCCGGCAGATCACAAAGC-3' (SEQ ID NO: 18)
[0167] (9)V922△ / E923T-F:5'-AAGAGACAGCCGGAAACCCGGCAGATCACAAAG-3'(SEQ ID NO: 19)
[0168] (10) L921P / V922Δ-F:5'-GAGACAGCTGACAACCCGGCAGATCACAAAGC-3' (SEQ ID NO: 20).
[0169] Fragments (a), (b), (c), and (d) were amplified using primers (1)+(2), (1)+(3), (1)+(4), and (1)+(5), respectively. Fragments (e), (f), (g), and (h) were amplified using primers (6)+(7), (6)+(8), (6)+(9), and (6)+(10), respectively. Fragments (a)+(e), (b)+(f), (c)+(g), and (d)+(h) were mixed and amplified using primer (1)+(6), respectively, to obtain fragments (i), (j), (k), and (l). Fragments (i), (j), (k), and (l) were then combined with 3β-HA-330 (vector structure as shown) cleaved using BamHI+XbaI. Figure 3 As shown, the 6187bp vector DNA fragment (m) generated after the vector sequence (SEQ ID NO: 21) was mixed and homologous recombination was performed using a homologous recombination kit (Wuhan Aibotek) to obtain expression plasmids for the mutants SpCas9V922△, SpCas9E923△, SpCas9 V922△ / E923T, and SpCas9 L921P / V922△. The sequences of the above DNA fragment am are shown in Table 2.
[0170] Table 2. Nucleotide sequences of DNA fragments
[0171]
[0172]
[0173] These amplified fragments were sequenced using the Sanger sequencing method at Qingke Biotechnology Co., Ltd. (sequencing was performed using primers (1) + (6), and all sequences between the two primers were obtained by the sequencing company). The nucleotide and amino acid sequences of four mutants, SpCas9 V922△, SpCas9 E923△, SpCas9 V922△ / E923T, and SpCas9 L921P / V922△, were obtained. The amino acid sequences of the above four Cas9 protein mutants from position 908 to 940 and the corresponding nucleotide sequences are shown in Table 3. The remaining amino acid sequences are the same as those in SEQ ID NO: 1.
[0174] Table 3. Amino acid sequences and corresponding nucleotide sequences of four Cas9 mutants from position 908 to 940.
[0175]
[0176] Example 2. In vitro enzyme digestion assay to verify the in vitro enzyme digestion activity of the Cas9 protein mutant.
[0177] In vitro purification of SpCas9, SpCas9 H840A, and four SpCas9 insertion / deletion mutants SpCas9 V922△, SpCas9-E923△, SpCas9L921P / V922△, and SpCas9 V922△ / E923T proteins (carrying His tags). The pET28(a) vector was cut into linear fragments using restriction endonucleases MluI and XhoI. Using the pET28(a) vector as a template, an 843 bp sequence was amplified by PCR (F primer: 5'-TGATCAGCCCACTGACGCGT-3' (SEQ ID NO: 44); R primer: 5'-GGTATATCTCCTTCTTAAAG-3' (SEQ ID NO: 45)). PCR (F primer: 5'-CTTTAAGAAGGAGATATACCatggacaagaagtacagcatcgg-3' (SEQ ID NO: 46); R primer: 5'-TGGTGGTGGTGGTGCTCGAGgtcacctcccagctgagaca-3' (SEQ ID NO: 47) was then performed on SpCas9, SpCas9 H840A, and four SpCas9 insertion / deletion mutants: SpCas9 V922△, SpCas9-E923△, and SpCas9 H840A. 15bp homologous sequences (uppercase letters in primers indicate homologous sequences) were added to each end of the L921P / V922△ and SpCas9 V922△ / E923T fragments. Seamless cloning (Aibotek, RK21020) was used to construct corresponding Cas protein expression vectors with a 6×His tag fused to the C-terminus. *E. coli* BL21 cells were transformed with the pET28(a) vector encoding Cas9, and single colonies were picked until the bacterial concentration reached OD0.05. 600 When the pH was 0.6–0.8, IPTG was added to a final concentration of 0.5 mM, and induction was performed at 16°C and 180 rpm for 18 h. Cells were collected by centrifugation and lysed by sonication in lysis buffer (20 mM Tris-HCl pH 8.0, 300 mM NaCl, 10 mM imidazole, 0.5 mM PMSF). After lysis, cells were centrifuged twice at 10000 g for 15 min each time, and the supernatant was filtered through a 0.45 μM filter. The supernatant solution was incubated with Ni NTABeads (Tiandi Renhe, SA004100) in a gravity chromatography column for 1 hour, and then the supernatant was discarded. The cells were washed three times at 4°C for 5 min each time with washing buffer (20 mM Tris-HCl pH 8.0, 300 mM NaCl, 20 mM imidazole, 0.5 mM PMSF). The relevant proteins were eluted with elution buffer (20 mM Tris-HCl pH 8.0, 300 mM NaCl, 250 mM imidazole, 0.5 mM PMSF).
[0178] A 3110 bp linear double-stranded DNA substrate containing the SpCas9 target was cleaved in vitro, and its nucleotide sequence is shown in SEQ ID NO: 48. The 3110 bp DNA double-stranded substrate was amplified and recovered by PCR. The sgRNA gdd2 targeting the above DNA double-stranded substrate was constructed, with the gdd2 sequence being 5'-ATCCATGGTGGCGGCGGTTA-3' (SEQ ID NO: 49), and PAM on the Crick strand being 5'-GGG-3'. The gdd2 SpCas9 sgRNA 5'-ATCCATGGTGGCGGCGGTTAGTTTTAGAGCTAGAAATAGCAAGTTAA AATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGC-3' (SEQ ID NO: 50) was synthesized. 200 ng of double-stranded DNA substrate was incubated with 100 nM of the above-purified SpCas9, SpCas9 H840A, and four SpCas9 insertion / deletion mutants SpCas9 V922△, SpCas9-E923△, SpCas9 L921P / V922△, and SpCas9 V922△ / E923T proteins, and 100 nM sgRNA gdd2 in lysis buffer (20 mM HEPES, pH 7.5, 150 mM KCl, 0.5 mM MTT, 0.1 mM EDTA) at 37°C for 30 minutes. No Cas enzyme was added to the blank control. The reaction was terminated with DNA loading buffer (Novizan, P022-01), and the samples were separated by 1% agarose gel electrophoresis and visualized by Safe Green staining (Aibotek, RM19008). The results showed that the SpCas9 protein can cleave linear double-stranded DNA substrates, breaking the 3110 bp substrate into two fragments of 1900 bp and 1210 bp. SpCas9 H840A and the four SpCas9 insertion / deletion mutants SpCas9 V922△, SpCas9 E923△, SpCas9 L921P / V922△, and SpCas9 V922△ / E923T proteins did not possess double-cleavage activity. Figure 4 A).
[0179] Two short DNA double-stranded substrates (70 bp) were prepared, and the single-stranded DNA fragment 5'-T(Alexa 680)TCAGGGTCAGCTTGCCGTAGGTGGCATCGCCCTCGCCCTCGCCGGACACGCTGAACTTGTGGCCGTTTA-3' (Alexa 680 labeled 5' first base) and its reverse complementary single-stranded DNA fragment 5'-T(Alexa 680)AAACGGCCACAAGTTCAGCGTGTCCGGCGAGGGCGAGGGCGATGCCACCTAC GGCAAGCTGACCCTGAA-3' (Alexa 680 labeled 5' first base) was synthesized as shown in SEQ ID NO: 51. The inverse complementary sequences (unlabeled) of the fluorescently labeled single-stranded DNA were synthesized separately. Annealing the fluorescently labeled single-stranded DNA with its complementary sequence formed 5' Alexa 680 fluorescently labeled DNA double-stranded substrates and 3' Alexa 680 fluorescently labeled DNA double-stranded substrates. A SpCas9 sgRNA gX1-F targeting the DNA double-stranded substrate was constructed. The gX1-F sequence was 5'-CAAGTTCAGCGTGTCCGGCG-3' (SEQ ID NO: 53), with PAM on the Crick strand, also 5'-AGG-3'. The gX1-FSpCas9 sgRNA 5'-CAAGTTCAGCGTGTCCGGCGGTTTTAGAGCTAGAAATAGCAAGTTAAA ATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGC-3' (SEQ ID NO: 54) was synthesized. 20 ng of fluorescently labeled double-stranded DNA substrate was incubated with 100 nM SpCas9, SpCas9 H840A, and four SpCas9 insertion / deletion mutants SpCas9 V922△, SpCas9-E923△, SpCas9 L921P / V922△, and SpCas9 V922△ / E923T protein, and 100 nM sgRNA gX1-F in lysis buffer (20 mM HEPES, pH 7.5, 150 mM KCl, 0.5 mM DTT, 0.1 mM EDTA) at 37°C for 30 min. No Cas enzyme was added to the blank control. The reaction was terminated with DNA loading buffer containing 2x TBE (Tomai Biotechnology, TLB-5). Separation was performed by 8% denaturing urea polyacrylamide gel electrophoresis, and visualization was achieved using an Odyssey infrared fluorescence scanning imaging system with an excitation wavelength of 680 nm.
[0180] The results showed that, unlike SpCas9 H840A, which only cleaves the non-target strand of the DNA Cas9-sgRNA target site, the four SpCas9 insertion / deletion mutants SpCas9-V922△, SpCas9-E923△, SpCas9-L921P / V922△, and SpCas9-V922△ / E923T proteins can only cleave the target strand of the Cas9-sgRNA target site and cannot cleave the non-target strand of the target DNA. Figure 4 B).
[0181] Example 3. Validation of the activity of the Cas9 protein mutant using the NHEJ reporter system.
[0182] The mouse embryonic stem cell NHEJ reporter system used can detect the ability of SpCas9 and nCas9 (nickase Cas9) to induce NHEJ repair products after targeted cleavage. This reporter system initiates transcription via the PGK promoter, downstream of which is an enhanced "Koz-ATG" gene, followed by a GFP gene. Initially, the mRNA transcribed by this reporter system is translated from the "Koz-ATG" gene; at this stage, the subsequent GFP fusion gene is in a frameshift state and cannot be translated into active GFP. If SpCas9 or nCas9 is used to target and cleave the DNA sequence between "Koz-ATG" and GFP, inducing double-strand breaks (DSBs) or single-strand breaks (SSBs), the NHEJ pathway repair introduces base insertions or deletions, resulting in full-frame translation of GFP. Therefore, we can directly detect the proportion of cells expressing GFP by flow cytometry, thereby detecting the proportion of NHEJ products. This reporter system can be used to detect the proportion of NHEJ repair products after targeted cleavage by SpCas9 or Cas9 single-nickase enzymes. Figure 5 A).
[0183] The nucleotide sequence of the above NHEJ reporter system is shown in SEQ ID NO: 55, as follows:
[0184]
[0185]
[0186] The bolded part is the PGK promoter sequence, the italicized part is the Kozak-ATG sequence, and the underlined part is the GFP sequence.
[0187] The sgRNA g2-2 was designed and constructed using the DNA sequence between "Koz-ATG" and GFP on the mouse embryonic stem cell NHEJ reporter system. The g2-2 sequence is 5'-GATAACAGGGTAATCCATGG-3' (SEQ ID NO: 56), and PAM is located on the Crick chain as 5'-TGG-3'. The 5'-GATAACAGGGTAATCCATGG-3' (SEQ ID NO: 56) and its reverse complementary sequence were synthesized, annealed, and ligated into the BBSI-digested px330-U6-Chimeric plasmid vector (modified from pX330-U6-Chimeric_BB-CBh-hSpCas9 (Addgene, Cat#42230), with the SpCas9 expression cassette removed) to form the g2-2 expression plasmid. Using the Lipofectamine 2000 liposome transfection method, wild-type SpCas9, and expression plasmids of SpCas9V922△, SpCas9 E923△, SpCas9 L921P / V922△, and SpCas9 V922△ / E923T constructed in Example 1, were co-transfected with the g2-2 expression plasmid into mouse embryonic stem cells containing the NHEJ reporter system. The transfection system (using a 24-well plate as an example) was: 8 x 10⁸ cells per well. 4 The total amount of transfected DNA was 0.5 μg, including 0.25 μg of Cas9 expression plasmid and 0.25 μg of g2-2 expression plasmid. The medium was changed 6-8 hours after transfection, and GFP was counted using flow cytometry 72 hours after transfection. + The proportion of cells. The results showed that the levels of NHEJ repair products induced by the four SpCas9 mutants SpCas9 V922△, SpCas9 E923△, SpCas9 L921P / V922△, and SpCas9 V922△ / E923T were significantly lower than those induced by wild-type SpCas9. Figure 5 B).
[0188] Example 4. Activity verification of Cas9 protein mutants combined with H840A mutation
[0189] (1) HR reporting system detects the activity of Cas9 protein mutants with H840A mutation.
[0190] The primer sequences used below are as follows:
[0191] (1)5'-AAGCCCGAGAACATCGTGAT-3' (SEQ ID NO: 57);
[0192] (2) 5'-GCCCTCTAGACTCGAGCAGC-3' (SEQ ID NO: 58);
[0193] (3) 5'-ACGAGGACATTCTGGAAGATATCGT-3' (SEQ ID NO: 59);
[0194] (4) 5'-ATCACGATGTTCTCGGGCTT-3' (SEQ ID NO: 60).
[0195] First, a single mutant vector of SpCas9 H840A was constructed. The reverse complementary primers for SpCas9 H840A were constructed using the reverse complementary primers that introduced the point mutation: F: 5'-TACGATGTGGACGCCATCGTGCCTCAGAGC-3' (SEQ ID NO: 61), R: 5'-GCTCTGAGGCACGATGGCGTCCACATCGTA-3' (SEQ ID NO: 62). PCR amplification was performed using 3β-HA-330 as a template. Subsequently, the plasmid template was digested with Dpn1 and transformed into E. coli to obtain the single mutant expression vector of SpCas9 H840A. Fragments (c), (d), (e), and (f) were obtained by PCR amplification using paired primers (1) + (2) with SpCas9 V922△, SpCas9 E923△, SpCas9 L921P / V922△, and SpCas9 V922△ / E923T expression vectors as templates, respectively. Fragment (g) was obtained by PCR amplification using paired primers (3) + (4) with SpCas9 H840A expression plasmid as template. Fragments (c), (d), (e), and (f) were mixed with (g) and the vector DNA fragment generated after cutting 3β-HA-330 with EcoRV+XhoI, respectively. Homologous recombination was performed using a homologous recombination kit (Wuhan Aibotek) to construct H840A double mutant expression vectors combining SpCas9 V922△, SpCas9 E923△, SpCas9 L921P / V922△, and SpCas9 V922△ / E923T. Four SpCas9 insertion / deletion mutants—SpCas9 V922△, SpCas9 E923△, SpCas9 L921P / V922△, and SpCas9V922△ / E923T—combined with the H840A double mutant vector were co-transfected with sgRNAs (g1-1 and g1-2, where the sequence of g1-2 is 5'-TCGAGCTGAAGGGCATCGTA-3' (SEQ ID NO: 43)) targeting the I-SceI site near the HR reporter system to verify the activity of these double mutants. The results showed that the double mutant SpCas9 could not induce STGC and LTGC, completely losing its cleavage activity. Figure 1 AF).
[0196] (2) Detection of the activity of Cas9 protein mutants with H840A mutation using the NHEJ reporter system
[0197] Following the method described in Example 4(1), four SpCas9 insertion / deletion mutants—SpCas9-V922△, SpCas9-E923△, SpCas9-L921P / V922△, and SpCas9-V922△ / E923T—were constructed using the H840A double mutant vector. These were co-transfected with sgRNA (g2-2) targeting the DNA sequence between "Koz-ATG" and GFP into mouse embryonic stem cells containing the NHEJ reporter system to verify the activity of the double mutants. The specific procedures were the same as in Example 3. The results showed that the double mutant SpCas9 could not induce the production of NHEJ repair products and had completely lost its cleavage activity. Figure 5 B).
[0198] Example 5. Construction of a base editor for Cas9 protein mutants and verification of base editing efficiency.
[0199] The efficiency of base editor editing was assessed using the mouse embryonic stem cell NHEJ reporter system. This reporter system initiates transcription via the PGK promoter, downstream of which lies an enhanced "Koz-ATG" gene, followed by a GFP gene. Initially, the mRNA transcribed by this reporter system is translated from the "Koz-ATG," and the subsequent GFP fusion gene is in a frameshift state, unable to be translated into active GFP. If the A in the complementary strand GAC of the "Koz-ATG" is edited to C using an adenine base editor, or if the C in the GAC is edited to T using a cytosine base editor, the "Koz-ATG" is disrupted, activating the GFP's own translation initiation ATG, leading to normal GFP gene translation and the production of GFP. + Cells can be counted using flow cytometry. Figure 6 A).
[0200] A SpCas9 sgRNA vector targeting "Koz-ATG" was designed and constructed. PAM was located on the Crick strand as 5'-GGG-3', with a spacer sequence of 5'-ATCCATGGTGGCGGCGGTTA-3' (SEQ ID NO: 49). This 5'-ATCCATGGTGGCGGCGGTTA-3' (SEQ ID NO: 49) sequence was synthesized and ligated into the px330-U6-Chimeric vector to construct the BE sgRNA expression plasmid. The fifth base "A" at the distal end of PAM is the target editing site for the constructed adenine base editor, within the NG-ABE9e editing window (where the 5th to 8th bases at the distal end of PAM have high editing efficiency). The fourth base "C" at the distal end of PAM is the target editing site for the constructed cytosine base editor, within the BE4 max editing window (where the 4th to 8th bases at the distal end of PAM have high editing efficiency).
[0201] The four SpCas9 insertion / deletion mutants disclosed herein—SpCas9 V922△, SpCas9 E923△, SpCas9L921P / V922△, and SpCas9 V922△ / E923T—were introduced into the NG-ABE9e (ABE) and BE4 Max (CBE) vectors. The vector design is described in [reference needed]. Figure 6 B and 6C. ngSpCas9 (Addgene Plasmid #117919) can recognize NG PAM. Using the ngSpCas9 plasmid as a template, four ngSpCas9 insertion / deletion mutants, ngSpCas9V922△, ngSpCas9E923△, ngSpCas9L921P / V922△, and SpCas9V922△ / E923T, were constructed in the same manner as in Example 1 (only the template sequence was different, and the construction method was the same). The ng-ABE9e was replaced in the four ngSpCas9 insertion / deletion mutants ngSpCas9V922△, ngSpCas9E923△, ngSpCas9L921P / V922△, and ngSpCas9V922△ / E923T (see Tu Tianxiang et al., Molecular...). Therapy, 2022, 30(9): 2933–2941) Constructing a Cas9 single-cleavage enzyme in a vector. Figure 6The corresponding adenine base editor in B. The four SpCas9 insertion / deletion mutants SpCas9 V922△, SpCas9 E923△, SpCas9L921P / V922△, and SpCas9 V922△ / E923T were used to replace the Cas9 single-cleavage enzyme in BE4max (Addgene Plasmid#112093) for construction. Figure 6 C corresponding cytosine base editor.
[0202] The specific method for constructing an adenine base editor is as follows: Using four ngSpCas9 insertion / deletion mutant plasmids as templates, four ngSpCas9 insertion / deletion mutant sequences were obtained by PCR amplification using the forward primer 5'-GACAAGAAGTACAGCATCGG-3' (SEQ ID NO: 63) and the reverse primer 5'-GTCACCTCCCAGCTGAGACA-3' (SEQ ID NO: 64). Using the NG-ABE9e plasmid as a template, the NG-ABE9e vector fragment was amplified by PCR using the forward primer 5'-CAGCTGGGAGGTGACTCTGG-3' (SEQ ID NO: 65) and the reverse primer 5'-GCTGTACTTCTTGTCTGACCCC-3' (SEQ ID NO: 66). The four ngSpCas9 insertion / deletion mutant sequences were mixed with the NG-ABE9e vector fragment, and homologous recombination was performed using a homologous recombination kit (Wuhan Aibotek) to construct... Figure 6 B-corresponding adenine base editor vector.
[0203] The specific method for constructing the cytosine base editor is as follows: Using the SpCas9 insertion / deletion mutant plasmid constructed in Example 1 as a template, four SpCas9 insertion / deletion mutant sequences were obtained by PCR amplification using the forward primer 5'-GACAAGAAGTACAGCATCGG-3' (SEQ ID NO: 63) and the reverse primer 5'-GTCACCTCCCAGCTGAGACA-3' (SEQ ID NO: 64). Using the BE4max plasmid as a template, the BE4max vector fragment was amplified by PCR using the forward primer 5'-CAGCTGGGAGGTGACAGCGG-3' (SEQ ID NO: 67) and the reverse primer 5'-GCTGTACTTCTTGTCGCTGCC-3' (SEQ ID NO: 68). The four SpCas9 insertion / deletion mutant sequences were mixed with the BE4max vector fragment, and homologous recombination was performed using a homologous recombination kit (Wuhan Aibotek) to construct... Figure 6 C-corresponding cytosine base editor vector.
[0204] To verify the editing efficiency of the base editor based on four SpCas9 insertion / deletion mutants, the constructed BE sgRNA plasmid and the expression plasmids of the four SpCas9 insertion / deletion mutants were transfected into mouse embryonic stem cells equipped with the NHEJ reporter system using the Lipofectamine 2000 liposome transfection method. The transfection system was as follows (using a 24-well plate as an example): 2 × 10⁶ cells per well. 5 Cells were transfected with a total of 0.5 μg of DNA, including 0.3 μg of base editor expression plasmid and 0.2 μg of BE sgRNA plasmid. To exclude the influence of NHEJ repair products induced by the nicking enzyme itself, we also co-transfected cells with expression plasmids of four insertion / deletion mutants without deaminase as negative controls.
[0205] GFP was analyzed by flow cytometry 72 hours after transfection. + Cell proportion. The results showed that base editors based on the four SpCas9 insertion / deletion mutants all had high base editing efficiency, while only the SpCas9 L921P / V922△ base editor had low base editing efficiency. Figure 6 B and Figure 6 C).
[0206] To detect the editing efficiency of adenine base editors based on four ngSpCas9 and SpCas9 insertion / deletion mutants at endogenous sites, six endogenous adenine base editor sites (ABE sites 1–6) and six cytosine base editor sites (CBE sites 1–6) were designed on the human genome, and corresponding sgRNA expression vectors were designed and constructed (Table 4). These vectors were co-transfected with the corresponding base editor expression vectors into HEK293T cells. The transfection system was as follows (using a 24-well plate as an example): 2 × 10⁶ cells per well. 5 A total of 0.5 μg of DNA was transfected, including 0.3 μg of base editor expression plasmid and 0.2 μg of sgRNA plasmid. Genomic DNA was extracted 72 h after transfection, and the PCR products were deeply sequenced after PCR target amplification. The base editing efficiency of the adenine and cytosine base editors constructed based on the four insertion / deletion mutants of this invention was analyzed. The results showed that the adenine base editor based on the four ngSpCas9 insertion / deletion mutants and the cytosine base editor based on the four SpCas9 insertion / deletion mutants had high base editing efficiency within the editing window. Only ngSpCas9 L921P / V922△ and SpCas9 L921P / V922△ showed lower base editing efficiency after being introduced into the base editor. Figure 7 , 8 ).
[0207] Table 4. sgRNA design for detecting base editing efficiency at endogenous sites
[0208] ABE site 1 TGG GGAATCCCTTCTGCAGCACC SEQ ID NO: 69 ABE site 2 AGA GAGCAAATACCAGAGATAAG SEQ ID NO: 70 ABE site 3 AGC CCCCCCACAGAATATCAAGG SEQ ID NO: 71 ABE site 4 GGT GCTCTGTCAGAAGATCTCCA SEQ ID NO: 72 ABE site 5 GGC GATCAGGAAATAGAGCCACA SEQ ID NO: 73 ABE site 6 TGA GCCACCACAGGGAAGCTGGG SEQ ID NO: 74 CBE site 1 CGG GTCAAGAAAGCAGAGACTGC SEQ ID NO: 75 CBE site 2 GGG GAGTCCGAGCAGAAGAAGAA SEQ ID NO: 76 CBE site 3 TGG GGAATCCCTTCTGCAGCACC SEQ ID NO: 77 CBE site 4 GGG GAACACAAAGCATAGACTGC SEQ ID NO: 78 CBE site 5 TGG GGCCCAGACTGAGCACGTGA SEQ ID NO: 79 CBE site 6 TGG GGCACTGCGGCTGGAGGTGG SEQ ID NO: 80
[0209] Example 6. Gene editing efficiency of Cas9 protein mutant enzyme fused with TREX2
[0210] The gene editing efficiency after pairing single-cut enzyme fusion with TREX2 was detected using the mutant NHEJ reporter system for mouse embryonic stem cells. The mutant NHEJ reporter system inserted an I-SceI sequence between "Koz-ATG" and GFP in Example 1, providing more target options. Similar to Example 1, under normal circumstances, the translation of "Koz-ATG" produces frameshift GFP, and the cell does not emit light. When 3' sgRNA is designed to be generated before and after "Koz-ATG" to cut it, the cell can express normal GFP in the following two situations: (1) When DSBs are generated by cutting between the two start codons ATG, NHEJ repair will connect the cuts together and introduce non-precise sequence insertions and deletions at the repaired cut, so that the sequence undergoes a certain probability of frameshift mutation, and the GFP gene is expressed in full. (2) The first Kozak-ATG is destroyed, and the cell uses the second ATG start codon to express GFP. Figure 9 A).
[0211] Four SpCas9 insertion / deletion mutations were designed and constructed before and after "Koz-ATG" to produce 3'-terminal paired sgRNAs (gWL2 and gCR6) after cleavage. gWL2 PAM is on the Watson strand, and gCR6 PAM is on the Crick strand. The spacer sequences are shown in Table 5. They were synthesized and constructed into the px330-U6-Chimeric vector.
[0212] A C-terminal fusion TX expression plasmid was constructed using SpCas9 and four SpCas9 insertion / deletion mutants disclosed herein: SpCas9 V922△, SpCas9 E923△, SpCas9 L921P / V922△, and SpCas9 V922△ / E923T. TX is a triple mutant of TREX2 (R163A / R165A / R167A), which disrupts TREX2's ability to bind DNA, thus avoiding off-target effects caused by TREX2's non-specific DNA binding. The vector backbone is pcDNA3.1.
[0213] Mouse embryonic stem cells were co-transfected with SpCas9 and the four SpCas9 insertion / deletion mutants disclosed herein, SpCas9 V922△, SpCas9 E923△, SpCas9 L921P / V922△, and SpCas9 V922△ / E923T, with a C-terminus fusion TX expression plasmid. The transfection system was as follows (using a 24-well plate as an example): 2 × 10⁶ cells per well. 5 The total transfection DNA amount was 0.5 μg. When transfecting a single sgRNA, the ratio of Cas enzyme-TX:gRNA:empty U6 was 2:1:1; when transfecting two sgRNAs, the ratio of Cas enzyme-TX:gRNA1:gRNA2 was 2:1:1. GFP was detected by flow cytometry 72 h after transfection. + Cell ratio.
[0214] The results showed that fusing the four SpCas9 insertion / deletion mutants disclosed herein—SpCas9 V922△, SpCas9 E923△, SpCas9 L921P / V922△, and SpCas9 V922△ / E923T—with the 3'→5' exonuclease TREX2 mutant TX can significantly improve the gene editing efficiency at the 3' end generated by paired single-cut enzyme digestion. Figure 9 B). Since the designed sgRNAs gWL2 and gCR6 only produce 3' ends when double-cleaved with SpCas9 single-cleavage enzymes targeting the hybrid strand of sgRNA, the four SpCas9 insertion / deletion mutants disclosed herein, after fusion with TX, can process the 3' ends, thereby improving gene editing efficiency. This further proves that the four SpCas9 insertion / deletion mutants disclosed herein, SpCas9 V922△, SpCas9 E923△, SpCas9 L921P / V922△, and SpCas9 V922△ / E923T, only cleave the sgRNA target strand in vivo and do not cleave the non-target strand of sgRNA.
[0215] Table 5. sgRNA design for detecting gene editing efficiency of paired Cas9 protein mutant enzyme fusion with TREX2
[0216] gWL2 TGG TTCCCTAACCGCCGCCACCA SEQ ID NO: 81 gCR6 TGG CTTGCTGATCATGTGAAGGA SEQ ID NO: 82
Claims
1. A Cas9 protein mutant, wherein the amino acid sequence from position 908 to 940 is as shown in SEQ ID NO: 35, 37, 39 or 41, and the remaining amino acid sequence is the same as in SEQ ID NO:
1.
2. The Cas9 protein mutant according to claim 1, wherein the mutant significantly reduces or eliminates the cleavage activity of the Cas9 protein on non-target strands.
3. The Cas9 protein mutant according to claim 1, wherein the Cas9 protein mutant comprises a RuvC and an HNH domain, wherein the HNH domain comprises a mutant H840A that reduces or eliminates target chain cleavage activity.
4. A fusion protein comprising the Cas9 protein mutant of claim 1, and a heterologous domain fused to the Cas9 protein mutant; wherein the heterologous domain is selected from cytosine deaminase, adenine deaminase, and DNA-binding 3' to 5' exonucleases.
5. A complex comprising the Cas9 protein mutant of claim 1 or the fusion protein of claim 4, and guide RNA.
6. A polynucleotide encoding the Cas9 protein mutant of claim 1 or the fusion protein of claim 4.
7. A vector comprising the polynucleotide of claim 6.
8. The vector according to claim 7, wherein the polynucleotide is operatively linked to a promoter, enhancer, and / or terminator.
9. A kit comprising the Cas9 protein mutant of claim 1, the fusion protein of claim 4, the complex of claim 5, the polynucleotide of claim 6, or the vector of claim 7.
10. The kit of claim 9, further comprising an expression construct encoding a guide RNA backbone.
11. A cell comprising the Cas9 protein mutant of claim 1, the fusion protein of claim 4, the complex of claim 5, the polynucleotide of claim 6, or the vector of claim 7.
12. The cell according to claim 11, wherein the cell is selected from human cells, non-human mammalian cells, plant cells, non-mammal vertebrate cells, invertebrate cells, single-celled eukaryotic organisms, or prokaryotic cells.
13. A pharmaceutical composition comprising the Cas9 protein mutant of claim 1, the fusion protein of claim 4, the complex of claim 5, the polynucleotide of claim 6, the vector of claim 7, or the cell of claim 11.
14. The use of the Cas9 protein mutant of claim 1, the fusion protein of claim 4, the complex of claim 5, the polynucleotide of claim 6, the vector of claim 7, the kit of claim 9, the cell of claim 11, or the pharmaceutical composition of claim 13 in nucleic acid modification; wherein the nucleic acid modification is performed in vitro.
15. A method for modifying a target nucleic acid, the method comprising contacting the target nucleic acid with: (a) a Cas9 protein mutant of claim 1 or a fusion protein of claim 4, and guide RNA; or (b) a complex of claim 5; the method being for non-therapeutic purposes.
16. A method for modifying a target nucleic acid, the method comprising introducing into a cell or organism containing the target nucleic acid the Cas9 protein mutant of claim 1, the fusion protein of claim 4, the complex of claim 5, the polynucleotide of claim 6, the vector of claim 7, and / or the pharmaceutical composition of claim 13; the method being for non-therapeutic purposes.
Citation Information
Patent Citations
towel ring
CN3315821D
cell phone
CN3329834D
Methods for making and using reprogrammed human somatic cell nuclei and autologous and isogenic human stem cells
WO2003046141A2
Engineered CRISPR-Cas9 nuclease with altered PAM specificity
CN118530967A
Engineered CRISPR-Cas9 Nucleases
US20200140835A1