Cas9 protein mutant and application thereof
By introducing amino acid residue mutations at the catalytic site of the HNH domain of the Cas9 protein and constructing new Cas9 protein mutants, the problem of Cas9 protein cutting non-target chains was solved, efficient base editing and editing efficiency were achieved, and the scope of application was broadened.
Patent Information
- Application Number
- CN202510725001.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-05-30
AI Technical Summary
The existing Cas9 protein has the activity of cutting non-target chains during the gene editing process, which limits its application scope and efficiency. There is an urgent need to develop new Cas9 proteins and gene editing tools that only cut target chains.
By introducing mutations in amino acid residues at the catalytic site of the HNH domain of the Cas9 protein, particularly mutations at positions 921, 922, and/or 923, the cleavage activity on non-target chains is reduced or eliminated, and a heterologous domain and guide RNA are fused to construct a new Cas9 protein mutant.
It has achieved efficient base editing in mammalian cells, improved the cutting accuracy and editing efficiency of the target chain while maintaining safety, and broadened the application scope of the CRISPR/Cas9 gene editing system.
Smart Images

Figure SMS_1 
Figure SMS_2 
Figure SMS_3
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of gene editing technology, and in particular to a Cas9 protein mutant and applications thereof. Background Art
[0002] The CRISPR / Cas system is an adaptive immune system in bacteria that defends against invading foreign DNA. Its targeting properties, including its ability to unwind and cleave DNA, have been leveraged in the development of CRISPR gene editing technology. Cas9 in the CRISPR / Cas9 system is an endonuclease whose targeting is determined by the dual RNA guide structure of crRNA and tracrRNA and the target DNA site, the PAM. The crRNA and tracrRNA are integrated into a sgRNA, forming the Cas9-sgRNA gene editing system. By designing an sgRNA with a matching sequence within the crRNA spacer and combining it with the PAM, Cas9 can be targeted to a specific genomic locus, unwinding the target DNA and cleaving both strands. Cells primarily repair double-strand breaks (DSBs) in DNA through non-homologous end joining (NHEJ) and homologous recombination (HR). Researchers exploit this process to insert, delete, or replace genes, achieving precise gene editing.
[0003] Cas9 contains homology domains to the HNH and RuvC endonucleases. After DNA unwinding, the HNH domain is responsible for cleaving the target DNA strand paired with the sgRNA spacer, while the RuvC domain is responsible for cleaving the single-stranded non-target strand. Researchers introduced a point mutation in the catalytic site of the HNH domain, generating the Cas9 H840A mutant, which cleaves only the non-target strand. The Cas9 H840A mutant has greatly expanded the application range of the CRISPR / Cas9 gene editing system. There is an urgent need for new Cas9 proteins that cleave only the target strand, as well as for more alternative gene editing tools. Summary of the Invention
[0004] The purpose of this disclosure is to provide new Cas9 protein mutants and, based on this, related gene editing tools.
[0005] In order to achieve the above objectives, the present disclosure provides the following technical solutions:
[0006] In one aspect, the present disclosure provides a Cas9 protein mutant comprising a mutation at amino acid residue 921, 922, and / or 923 in the amino acid sequence of SEQ ID NO: 1, wherein the mutation is selected from one or more of the following:
[0007] (a) Deletion of amino acid residues;
[0008] (b) substitution of amino acid residues;
[0009] (c) insertion of amino acid residues;
[0010] Furthermore, the mutation can significantly reduce or eliminate the cleavage activity of the Cas9 protein on non-target chains.
[0011] On the other hand, the present disclosure provides a Cas9 protein mutant comprising an amino acid sequence fragment shown in SEQ ID NO: 35, 37, 39, and 41, wherein the mutant significantly reduces or eliminates the cleavage activity of the Cas9 protein on non-target chains.
[0012] On the other hand, the present disclosure provides a fusion protein comprising the aforementioned Cas9 protein mutant and a heterologous domain fused to the Cas9 protein mutant.
[0013] In another aspect, the present disclosure provides a complex comprising the aforementioned Cas9 protein mutant or the aforementioned fusion protein, and a guide RNA.
[0014] In another aspect, the present disclosure provides a polynucleotide encoding the aforementioned Cas9 protein mutant or the aforementioned fusion protein.
[0015] In another aspect, the present disclosure provides a vector comprising the aforementioned polynucleotide; preferably, the vector comprises a promoter driving the expression of the polynucleotide.
[0016] In a sixth aspect, the present disclosure provides a kit comprising the aforementioned Cas9 protein mutant, the aforementioned fusion protein, the aforementioned complex, the aforementioned polynucleotide or the aforementioned vector; preferably, the kit further comprises an expression construct encoding a guide RNA skeleton.
[0017] In another aspect, the present disclosure provides a cell comprising the aforementioned Cas9 protein mutant, the aforementioned fusion protein, the aforementioned complex, the aforementioned polynucleotide, or the aforementioned vector.
[0018] On the other hand, the present disclosure provides a pharmaceutical composition comprising the aforementioned Cas9 protein mutant, the aforementioned fusion protein, the aforementioned complex, the aforementioned polynucleotide, the aforementioned vector or the aforementioned cell.
[0019] On the other hand, the present disclosure provides use of the aforementioned Cas9 protein mutant, the aforementioned fusion protein, the aforementioned complex, the aforementioned polynucleotide, the aforementioned vector, the aforementioned kit or the aforementioned cell or the aforementioned pharmaceutical composition in nucleic acid modification.
[0020] In another aspect, the present disclosure provides a method of modifying a target nucleic acid, the method comprising contacting the target nucleic acid with:
[0021] (a) the aforementioned Cas9 protein mutant or the aforementioned fusion protein, and a guide RNA; or
[0022] (b) The aforementioned complex.
[0023] On the other hand, the present disclosure provides a method for modifying a target nucleic acid, comprising introducing the aforementioned Cas9 protein mutant, the aforementioned fusion protein, the aforementioned complex, the aforementioned polynucleotide, the aforementioned vector, the aforementioned cell and / or the aforementioned pharmaceutical composition into a cell or organism comprising the target nucleic acid.
[0024] On the other hand, the present disclosure provides a method for treating a subject diagnosed with a disease associated with or caused by a gene mutation, comprising administering to the subject a therapeutically effective amount of the aforementioned Cas9 protein mutant, the aforementioned fusion protein, the aforementioned complex, the aforementioned polynucleotide, the aforementioned vector, the aforementioned cell and / or the aforementioned pharmaceutical composition.
[0025] The present invention introduces mutations at positions far away from the catalytic site to induce conformational changes in the Cas9 protein, thereby obtaining new Cas9 protein mutants that only cut the Cas9-sgRNA target chain and do not cut single-stranded non-target chains. This can be used to construct a new base editor that can perform efficient base editing on the genome of mammalian cells.
[0026] The Cas9 protein mutant disclosed herein and the H840A site are mutated simultaneously, which inactivates the Cas9 protease catalytic domain and constructs dCas9. The Cas9 protein mutant disclosed herein is used to design base editors (including ABE and CBE), which have high base editing efficiency in mammalian cells. The Cas9 protein mutant disclosed herein is fused to the DNA binding deletion mutant TX of the 3'→5' exonuclease TREX2, which can greatly improve the gene editing efficiency of the 3' end produced by paired single nickase cleavage and maintain its safety. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the prior art, the following will briefly introduce the drawings required for the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without inventive work. The drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the specification, and together with the specification, are used to explain the principles of the specification.
[0028] Figure 1 Figure 4 shows the detection results of STGC, LTGC, and LTGC Bias in mouse embryonic stem cells containing an HR reporter system after the four SpCas9 insertion / deletion mutants and the introduction of H840A into the insertion / deletion mutants. AC represents four insertion / deletion mutants and the introduction of H840A into the insertion / deletion mutants and the co-transfection of the HR reporter system with sgRNAg1-1 to induce breakage. Figure A shows GFP detected by flow cytometry. + Cell ratio, i.e. STGC efficiency; Figure B shows GFP detected by flow cytometry + RFP + Figure DF represents four insertion / deletion mutants and the HR reporter system induced by the introduction of H840A into the insertion / deletion mutants and co-transfection with sgRNAg1-2 to induce breakage. Figure D shows GFP detected by flow cytometry. + Cell ratio, i.e. STGC efficiency; Figure E shows GFP detected by flow cytometry + RFP + The ratio is the LTGC efficiency. Figure F shows the LTGC bias (LTGC bias) obtained by dividing the LTGC efficiency of each nuclease combination g1-2 by the total HR efficiency. SpCas9 served as a positive control for double-stranded nuclease cleavage.
[0029] Figure 2 The structure of the SpCas9 overexpression vector used to construct a monoclonal cell line of a single-copy SpCas9 gene in Example 1 of the present disclosure is shown.
[0030] Figure 3 A structural diagram of the 3β-HA-330 vector of the present disclosure is shown.
[0031] Figure 4 The results of the detection of the target chain of the Cas9-sgRNA target site by the four SpCas9 insertion / deletion mutants disclosed in the present invention in vitro are shown. Figure A shows the results of SpCas9, SpCas9 H840A and the four SpCas9 insertion / deletion mutants disclosed in the present invention in vitro cleavage and detection of their double-cutting activity, where "-" represents a blank control. Figure B shows the results of SpCas9, SpCas9 H840A and the four SpCas9 insertion / deletion mutants in vitro cleavage of a DNA double-stranded substrate containing a 5'-Alexa680 fluorescently labeled Cas9-sgRNA target site (top) or a non-target chain of the target site (bottom), detecting their single-stranded cutting activity and pattern, where "-" represents a blank control.
[0032] Figure 5Figure 1 shows the NHEJ induction efficiency of four SpCas9 insertion / deletion mutants and the introduction of H840A into the insertion / deletion mutants in mouse embryonic stem cells containing an NHEJ reporter system. Figure A is a schematic diagram of the NHEJ reporter system. Figure B shows the GFP detection by flow cytometry after the four insertion / deletion mutants and the introduction of H840A into the insertion / deletion mutants and the introduction of sgRNA g2-2 into mouse embryonic stem cells containing an NHEJ reporter system to induce breakage. + Results of cell ratios. Positive control for SpCas9 cleaving double-stranded nuclease.
[0033] Figure 6 Figure 2 shows base editors for four SpCas9 indel mutants disclosed herein. Figure A is a schematic diagram of the base editing reporter system. Figure B shows the base editing efficiency of adenine base editors constructed based on the four SpCas9 indel mutants disclosed herein. Figure C shows the base editing efficiency of cytosine base editors constructed based on the four SpCas9 indel mutants disclosed herein. Data are mean ± SEM of three independent replicates.
[0034] Figure 7 The editing efficiency of adenine base editors constructed with four SpCas9 insertion / deletion mutants disclosed in the present invention at endogenous sites is shown.
[0035] Figure 8 The editing efficiency of cytosine base editors constructed with four SpCas9 insertion / deletion mutants disclosed in the present invention at endogenous sites is shown.
[0036] Figure 9 The four SpCas9 insertion / deletion mutants disclosed in the present invention are fused to the DNA-binding-deficient TREX2 mutant TX to improve the gene editing efficiency of the paired cut 3' end. Figure A is a schematic structural diagram of the mutant NHEJ reporter system in mouse embryonic stem cells. gWL2 is on the left side of "Koz-ATG" and PAM is on the Watson chain; gCR6 is on the right side of "Koz-ATG" and PAM is on the Crick chain. Figure B shows SpCas9 and the four SpCas9 insertion / deletion mutants disclosed in the present invention and fused to TX, combined with a single sgRNA (gWL2 or gCR6) or a combination of two sgRNAs (gWL2 and gCR6) to cut the mutant NHEJ reporter system, and GFP detected by flow cytometry. + Cell ratio. Data are the mean ± SEM of three independent repeated experiments. DETAILED DESCRIPTION
[0037] (I) Definitions or terms
[0038] In order to make the present disclosure more easily understood, certain technical and scientific terms are specifically defined below. In the present disclosure, unless otherwise indicated, the scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. It should be understood that the present disclosure is not limited to specific methods, reagents, compounds, compositions or biological systems, and of course, variations thereof are possible. It should also be understood that the terms used in the present disclosure are only for the purpose of describing specific embodiments and are not intended to be limiting.
[0039] As used in this specification and the appended claims, the singular forms "a," "an," "the," "said," and similar referents include plural referents unless the content clearly dictates otherwise.
[0040] As used herein, the conjunction term "and / or" between various elements is intended to include both the meanings of "and" and "or", for example, the phrase "A, B and / or C" is intended to cover each of the following aspects: A, B and C; A, B or C; A or C; A or B; B or C; A and C; A and B; B and C; A (alone); B (alone); and C (alone).
[0041] As used herein, the terms "comprises," "comprising," "having," and "containing," and any variations thereof, are intended to cover a non-exclusive inclusion. The terms are intended to be open-ended, specifying the presence of any stated features, elements, integers, steps, or components, but not excluding the presence or addition of one or more other features, elements, integers, steps, components, or groups thereof. Thus, the term "comprising" encompasses the more restrictive terms "consisting of" and "consisting essentially of."
[0042] As used herein, numerical ranges are to be understood as including all numbers within the range. For example, a range of 1 to 20 is to be understood as including any number, combination of numbers, or subrange from the following group: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20.
[0043] As used herein, the term "about" refers to a range of ±20% of the value that follows. In some embodiments, the term "about" refers to a range of ±10% of the value that follows. In some embodiments, the term "about" refers to a range of ±5% of the value that follows.
[0044] In the description herein, references to “some embodiments,” “some implementations,” or “some examples” describe a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0045] As used herein, the terms "nucleic acid," "nucleic acid molecule," "nucleotide," "nucleotide sequence," "oligonucleotide," or "polynucleotide" mean a polymeric compound comprising covalently linked nucleotides. The term "nucleic acid" includes ribonucleic acid (RNA) and deoxyribonucleic acid (DNA), both of which can be single-stranded or double-stranded. DNA includes, but is not limited to, complementary DNA (cDNA), genomic DNA, plasmid or vector DNA, and synthetic DNA. In some embodiments, the present disclosure provides polynucleotides encoding any of the polypeptides disclosed herein, for example, the present disclosure relates to polynucleotides encoding Cas9 protein or variants thereof.
[0046] As used herein, the term "gene" refers to an assembly of nucleotides that encodes a polypeptide and includes cDNA and genomic DNA nucleic acid molecules."Gene" also refers to a nucleic acid fragment that can act as regulatory sequences before (5' non-coding sequences) and after (3' non-coding sequences) the coding sequence.
[0047] As used herein, a nucleic acid molecule is "hybridizable" or "hybridizes" to another nucleic acid molecule (such as cDNA, genomic DNA, or RNA) when the single-stranded form of the nucleic acid molecule can anneal to the other nucleic acid molecule under suitable conditions of temperature and solution ionic strength. Hybridization and wash conditions are known and are exemplified in Sambrook et al., Molecular Cloning: A Laboratory Manual, 2nd ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor (1989), particularly Chapter 11 and Table 11.1 thereof. Temperature and ionic strength conditions determine the "stringency" of the hybridization. Stringent conditions can be adjusted to screen for moderately similar fragments (e.g., homologous sequences from distantly related organisms) to highly similar fragments (e.g., genes that duplicate functional enzymes from closely related organisms). For preliminary screening of homologous nucleic acids, low stringency hybridization conditions corresponding to a Tm of 55°C can be used, for example, 5X SSC, 0.1% SDS, 0.25% milk and no formamide; or 30% formamide, 5X SSC, 0.5% SDS. Moderate stringency hybridization conditions correspond to higher Tm, for example, 40% formamide and 5X or 6X SCC. High stringency hybridization conditions correspond to the highest Tm, for example, 50% formamide, 5X or 6X SCC. Hybridization requires that the two nucleic acids contain complementary sequences, although mismatches between bases may exist depending on the stringency of the hybridization.
[0048] As used herein, the term "complementary" is used to describe the relationship between nucleotide bases that are capable of hybridizing to each other. For example, with respect to DNA, adenosine is complementary to thymine, while cytosine is complementary to guanine. Thus, the present disclosure also includes isolated nucleic acid fragments that are complementary to the complete sequences disclosed or used herein, as well as those substantially similar nucleic acid sequences.
[0049] DNA " coding sequence " is a double-stranded DNA sequence, which, when placed under the control of an appropriate regulatory sequence, is transcribed and translated into a polypeptide in vitro or in vivo. A "suitable regulatory sequence" refers to a nucleotide sequence located upstream (5' non-coding sequence), inside or downstream (3' non-coding sequence) of a coding sequence, and the nucleotide sequence affects transcription, RNA processing or stability or the translation of the associated coding sequence. Regulatory sequences may include promoters, translation leader sequences, introns, polyadenylation recognition sequences, RNA processing sites, effector binding sites, and stem-loop structures. The boundaries of the coding sequence are determined by the initiation codon at the 5' (amino) end and the translation termination codon at the 3' (carboxyl) end. The coding sequence may include, but is not limited to, prokaryotic sequences, cDNA from mRNA, genomic DNA sequences, or even synthetic DNA sequences. If the coding sequence is intended to be expressed in a eukaryotic cell, polyadenylation signals and transcription termination sequences are typically present at the 3' end of the coding sequence. The abbreviation of "open reading frame" is ORF, which means a nucleic acid sequence (DNA, cDNA or RNA) that includes a translation initiation signal or start codon (such as ATG or AUG) and a stop codon and can be translated into a polypeptide sequence.
[0050] As used herein, the term "homologous recombination" refers to the insertion of a foreign DNA sequence into another DNA molecule, for example, a vector is inserted into a chromosome. In some cases, the vector targets a specific chromosomal site for homologous recombination. For specific homologous recombination, the vector typically contains a sufficiently long region with homology to the chromosomal sequence to allow complementary binding of the vector to the chromosome and to incorporate the vector into the chromosome. Longer regions of homology and greater sequence similarity can improve the efficiency of homologous recombination.
[0051] As used herein, the term "non-homologous end joining (NHEJ)" refers to repairing a double-stranded gap in DNA by directly connecting one end of the gap to the other end of the gap without the need for a donor template DNA. In the absence of a donor template DNA, NHEJ typically results in a small amount of random insertion or deletion of nucleotides ("indel" or "indels") at the site of the double-stranded gap. In some cases, cutting at the target recognition sequence results in NHEJ at the target recognition site. Nuclease-induced cutting of the target site in the gene coding sequence, followed by DNA repair by NHEJ, can introduce mutations into the coding sequence, such as frameshift mutations, thereby disrupting gene function.
[0052] According to the disclosure herein, polynucleotides can be amplified using methods known in the art. Once a suitable host system and growth conditions are established, recombinant expression vectors can be amplified and prepared in large quantities. As described herein, expression vectors that can be used include, but are not limited to, the following vectors or their derivatives: human or animal viruses, such as vaccinia virus or adenovirus; insect viruses, such as baculovirus; yeast vectors; phage vectors (e.g., lambda), and plasmid and cosmid DNA vectors.
[0053] As used herein, the term "operably linked" means that a polynucleotide of interest, such as a polynucleotide encoding a Cas9 protein, is linked to a regulatory element in a manner that allows expression of the polynucleotide sequence. In some embodiments, the regulatory element is a promoter. In some embodiments, the polynucleotide of interest is operably linked to a promoter on an expression vector.
[0054] As used herein, the terms "promoter," "promoter sequence," or "promoter region" refer to a DNA regulatory region / sequence capable of binding RNA polymerase and involved in initiating transcription of downstream coding or non-coding sequences. In some examples of the present disclosure, the promoter sequence includes a transcription start site and extends upstream to include the minimum number of bases or elements used to initiate transcription at a level detectable above background. In some embodiments, the promoter sequence includes a transcription start site, as well as a protein binding domain responsible for RNA polymerase binding. Eukaryotic promoters typically, but not always, contain multiple "TATA" boxes and "CAT" boxes. Various promoters, including inducible promoters, can be used to drive various vectors of the present disclosure.
[0055] As used herein, the term "vector" is any tool for cloning and / or transferring nucleic acid into a host cell. A vector can be a replicon that may be attached to another DNA segment so as to produce the replication of the attached segment." replicon" is any genetic factor (e.g., plasmid, phage, clay, chromosome, virus) that acts as an automatic unit for DNA replication in vivo, i.e., can replicate under its self-control. In some embodiments of the present disclosure, the vector is an additional vector that, after many cell generations, is removed / lost from a cell population by, for example, asymmetric distribution. The term "vector" includes viral and non-viral tools for introducing the nucleic acid into a cell in vitro, in vitro, or in vivo. A large number of vectors known in the art can be used to manipulate nucleic acid, integrate response elements and promoters into genes, etc. Possible vectors include, for example, plasmids or modified viruses, including, for example, phages such as lambda derivatives, or plasmids such as pBR322 or pUC plasmid derivatives, or Blue script vectors. For example, the DNA fragmentation corresponding to response element and promoter will be inserted into the suitable vector and can be accompanied by suitable DNA fragmentation being connected to the selected vector with complementary binding end. Alternatively, the end of DNA molecule can be modified by enzyme catalysis or any site by being connected to this DNA end and producing by nucleotide sequence (joint). Such carriers can be engineered to comprise the selective marker gene that cell is selected, and these cells incorporate the mark into the cytokine genome. Such mark allows identification and / or selection of host cells, and these host cells incorporate and express the protein encoded by this mark.
[0056] Viral vectors, particularly retroviral vectors, have been used in many gene delivery applications in cells and living animals. Operable viral vectors include but are not limited to retrovirus, adeno-associated virus, poxvirus, baculovirus, vaccinia, herpes simplex, Epstein-Barr virus, adenovirus, geminivirus and cauliflower mosaic virus vectors. Non-viral vectors include but are not limited to plasmids, liposomes, charged lipids (cytofectin), DNA-protein complexes and biopolymers. In addition to nucleic acid, the vector may also include one or more regulatory regions and / or selective markers for selecting, measuring and monitoring nucleic acid transfer results (transferred to which tissue, expression duration, etc.).
[0057] The vector can be introduced into the desired host cell by known methods, including but not limited to transfection, transduction, cell fusion and lipofection. The vector may include various regulatory elements, including promoters. In some embodiments, the vector design can be based on a plurality of constructs designed by Mali et al. "Cas9 as a versatile tool for engineering biology (Cas9 as a versatile tool for engineering biology)", Nature Methods 10: 957-63 (2013). In some embodiments, the present disclosure provides an expression vector comprising any polynucleotide described herein, for example, an expression vector comprising a polynucleotide encoding a Cas protein or a variant thereof. In some embodiments, the present disclosure provides an expression vector comprising a polynucleotide encoding a Cas9 protein or a variant thereof.
[0058] As used herein, the term "plasmid" refers to an extrachromosomal element that usually carries genes that are not involved in the central metabolism of the cell and is usually in the form of a circular double-stranded DNA molecule. Such elements can be linear, circular, or supercoiled autonomously replicating sequences of single-stranded or double-stranded DNA or RNA derived from any source, genome-integrating sequences, bacteriophages, or nucleotide sequences in which many nucleotide sequences have been joined or recombined into a unique structure that is capable of introducing a promoter fragment and DNA sequence for a selected gene product, along with appropriate 3' untranslated sequences, into the cell.
[0059] As used herein, the term "transfection" means introducing exogenous nucleic acid molecules (including vectors) into cells."Transfected" cells include exogenous nucleic acid molecules inside the cell, and "transformed" cells are cells in which the exogenous nucleic acid molecules in the cell induce phenotypic changes. The transfected nucleic acid molecules can be integrated into the genomic DNA of the host cell and / or can be maintained outside the chromosome temporarily or for a long time by the cell. Host cells or organisms expressing exogenous nucleic acid molecules or fragments are referred to as "recombinant", "transformed" or "transgenic" organisms. In some embodiments, the present disclosure provides host cells including any expression vector as described herein (for example, an expression vector including a polynucleotide encoding a Cas protein or a variant thereof). In some embodiments, the present disclosure provides host cells including expression vectors including polynucleotides encoding a Cas9 protein or a variant thereof.
[0060] As used herein, "expression vector" comprises a vector capable of expressing DNA operably linked to regulatory sequences such as promoter regions that can affect the expression of such DNA fragments. Such additional fragments may include promoter and terminator sequences, and optionally may include one or more origins of replication, one or more selection markers, enhancers, polyadenylation signals, etc. Expression vectors are generally derived from plasmid or viral DNA, or may include elements of both. Thus, an expression vector refers to a recombinant DNA or RNA construct, such as a plasmid, phage, recombinant virus or other vector, which, when introduced into an appropriate host cell, results in the expression of the cloned DNA. Suitable expression vectors are well known to those skilled in the art and include expression vectors that are replicable in eukaryotic and / or prokaryotic cells, as well as expression vectors that remain episomal or that are integrated into the host cell genome.
[0061] As used herein, "expression" refers to the process of producing a polypeptide through transcription and translation of a polynucleotide. The expression level of a polypeptide can be assessed using any method known in the art, including, for example, methods for measuring the amount of polypeptide produced by a host cell. Such methods may include, but are not limited to, quantification of polypeptide in cell lysates by ELISA, gel electrophoresis followed by Coomassie blue staining, Lowry protein assay, and Bradford assay.
[0062] As used herein, the term "host cell" refers to a cell into which a recombinant expression vector has been introduced for the purpose of receiving, maintaining, replicating, and amplifying the vector. The term "host cell" refers not only to the cell into which the expression vector is introduced ("parent" cell), but also to the offspring of such a cell. Because modifications may occur in offspring, for example due to mutations or environmental influences, the offspring may be different from the parent cell, but are still included within the scope of the term "host cell." When the host cell divides, the nucleic acid contained in the vector replicates, thereby amplifying the nucleic acid. The host cell may be a eukaryotic cell or a prokaryotic cell. Suitable host cells include, but are not limited to, yeast cells.
[0063] As used herein, the terms "peptide," "polypeptide," "protein," and "proteins" are used interchangeably herein to refer to polymeric forms of amino acids of any length, which may include coded and non-coded amino acids, chemically or biochemically modified or derivatized amino acids, and polypeptides with modified peptide backbones. The beginning of a protein or polypeptide is referred to as the "N-terminus" (or amino-terminus, NH2-terminus, N-terminus, or amine-terminus), which refers to the free amine (-NH2) group of the first amino acid residue of the protein or polypeptide. The end of a protein or polypeptide is referred to as the "C-terminus" (or carboxyl-terminus, carboxyl-terminus, C-terminus, or COOH-terminus), which refers to the free carboxyl group (-COOH) of the last amino acid residue of the protein or peptide.
[0064] As used herein, the term "amino acid" refers to a compound that includes a carboxyl (-COOH) group and an amino (-NH2) group. "Amino acid" refers to both natural and unnatural (i.e., synthetic) amino acids. Natural amino acids and their three-letter and one-letter abbreviations include: alanine (Ala; A); arginine (Arg, R); asparagine (Asn; N); aspartic acid (Asp; D); cysteine (Cys; C); glutamine (Gln; Q); glutamic acid (Glu; E); glycine (Gly; G); histidine (His; H); isoleucine (Ile; I); leucine (Leu; L); lysine (Lys; K); methionine (Met; M); phenylalanine (Phe; F); proline (Pro; P); serine (Ser; S); threonine (Thr; T); tryptophan (Trp; W); tyrosine (Tyr; Y); and valine (Val; V).
[0065] As used herein, amino acid " substitution " refers to a polypeptide or protein in which one or more wild-type or naturally occurring amino acids are replaced by an amino acid different from the wild-type or naturally occurring amino acid at the amino acid residue. The substituted amino acid can be a synthetic or naturally occurring amino acid. In some embodiments, the substituted amino acid is a naturally occurring amino acid selected from the group consisting of A, R, N, D, C, Q, E, G, H, I, L, K, M, F, P, S, T, W, Y, and V. Substitution mutants can be described using an abbreviation system. For example, a substitution mutation in which the fifth (5th) amino acid residue is substituted can be abbreviated as "X5Y," where "X" is the wild-type or naturally occurring amino acid replaced, "5" is the position of the amino acid residue in the amino acid sequence of the protein or polypeptide, and "Y" is a substituted or non-wild-type or non-naturally occurring amino acid.
[0066] As used herein, an "isolated" polypeptide, protein, peptide, or nucleic acid is a molecule that has been removed from its natural environment. It is also understood that an "isolated" polypeptide, protein, peptide, or nucleic acid can be formulated with an excipient (such as a diluent) or adjuvant and still be considered isolated.
[0067] As used herein, the term "recombinant" when used to refer to a nucleic acid molecule, peptide, polypeptide, or protein means a new combination of genetic material not known to exist in nature or to be produced thereby. Recombinant molecules can be produced by any well-known technique currently available in the field of recombinant technology, including but not limited to polymerase chain reaction (PCR), gene splicing (e.g., using restriction endonucleases), and solid phase synthesis of nucleic acid molecules, peptides, or proteins.
[0068] As used herein, when used to refer to a polypeptide or protein, the term "domain" means a unique functional and / or structural unit in a protein. A domain is sometimes responsible for a specific function or interaction, contributing to the overall effect of the protein. Domains can exist in a variety of biological contexts. Similar domains can be found in proteins with different functions. Alternatively, domains with low sequence homology (i.e., less than about 50%, less than about 40%, less than about 30%, less than about 20%, less than about 10%, less than about 5% or less than about 1% sequence homology) may have the same function. In some embodiments, the Cas9 domain is a RuvC domain. In some embodiments, the Cas9 domain is a HNH domain. In some embodiments, the Cas9 domain is a Rec domain.
[0069] Although the present disclosure provides some specific amino acid sequences or nucleotide sequences, such as the amino acid sequences or nucleotide sequences shown in the sequence listing, it should be understood that a specific amino acid sequence or nucleotide sequence includes variants with conservative sequence modifications thereof, such as sequences having at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or at least 99.9% homology thereto, as long as the biological function or activity of the specific amino acid sequence or nucleotide sequence is not lost.
[0070] As used herein, the term "conservative substitutions" or "conservative sequence modifications" of a sequence refers to nucleotide and amino acid sequence modifications that do not eliminate the activity of the Cas9 protein encoded by the nucleotide sequence or containing the amino acid sequence. These conservative sequence modifications include conservative nucleotide and amino acid substitutions and nucleotide and amino acid additions, insertions, and deletions. For example, modifications can be introduced into the sequence table described herein by standard techniques known in the art (e.g., gene synthesis and PCR-mediated mutagenesis). Conservative sequence modifications include conservative amino acid substitutions, in which the amino acid residue is replaced with an amino acid residue with a similar side chain. Families of amino acid residues with similar side chains are already defined in the art. These families include amino acids with basic side chains (e.g., lysine, arginine, histidine), amino acids with acidic side chains (e.g., aspartic acid, glutamic acid), amino acids with uncharged polar side chains (e.g., glycine, asparagine, glutamine, serine, threonine, tyrosine, cysteine, tryptophan), amino acids with non-polar side chains (e.g., alanine, valine, leucine, isoleucine, proline, phenylalanine, methionine), amino acids with β-branched side chains (e.g., threonine, valine, isoleucine), and amino acids with aromatic side chains (e.g., tyrosine, phenylalanine, tryptophan, histidine). Therefore, the predicted non-essential amino acid residues in the Cas9 protein are preferably replaced by another amino acid residue from the same side chain family. Methods for identifying nucleotides and amino acid conservative substitutions with Cas9 protein activity are well known in the art. As described herein, the representation of substitutions is "single-letter abbreviation of amino acid before substitution-substitution site-single-letter abbreviation of amino acid after substitution." For example, "S541A" means that the serine at position 541 is replaced by alanine. The amino acids that can be conservatively substituted are shown in Table 1. As described herein, amino acid deletions are represented by "single-letter abbreviation of the amino acid before the deletion-deletion site-Δ", for example, "V922Δ" indicates that valine at position 922 is deleted.
[0071] Table 1. Examples of conservative amino acid substitutions
[0072]
[0073] It should be understood that any amino acid mutation (e.g., E923T) from a first amino acid residue (e.g., A) to a second amino acid residue (e.g., T) as described herein also includes an amino acid residue from the first amino acid residue to an amino acid residue similar to (e.g., conservative) the second amino acid residue. For example, a mutation from alanine to threonine (e.g., an E923T mutation) can also be from alanine to an amino acid similar in size and chemical properties to threonine (e.g., serine). Other similar amino acid pairs include, but are not limited to, the following: phenylalanine and tyrosine; asparagine and glutamine; methionine and cysteine; aspartic acid and glutamic acid; and arginine and lysine. Those skilled in the art will recognize that such conservative amino acid substitutions may have a smaller effect on protein structure and may be well tolerated without compromising function. However, it should be understood that those skilled in the art will recognize other conservative amino acid residues, and any amino acid mutation to other conservative amino acid residues is also within the scope of the present disclosure.
[0074] As used herein, the term "homology" has a generally accepted meaning in the art and is a central concept in comparative biology. The fundamental meaning of homology is that two samples being compared (e.g., an amino acid sequence or a nucleotide sequence) share a common ancestor. Generally speaking, two traits (states) in two species are considered a pair of homologous traits if they meet either of the following two conditions: 1. They are identical to a trait found in the ancestral group of these species; 2. They are different traits with an ancestor-descendant relationship. Amino acid sequence homology can be determined using known methods. For example, amino acid sequence homology (%) can be determined using programs commonly used in the field (e.g., BLAST, FASTA, etc.) according to initial settings. Alternatively, homology (%) can be determined using any algorithm known in the field, such as the algorithm of Needleman et al. (1970) (J. Mol. Biol. 48:444-453) or Myers and Miller (CABIOS, 1988, 4:11-17). The algorithm of Needleman et al. is incorporated into the GAP program of the GCG software package (available at www.gcg.com), and % homology can be determined, for example, using a BLOSUM 62 matrix or a PAM250 matrix, and any of gap weights: 16, 14, 12, 10, 8, 6, or 4 and length weights: 1, 2, 3, 4, 5, or 6. Additionally, the algorithm of Myers and Miller is incorporated into the ALIGN program, which is part of the GCG sequence alignment software package. When utilizing the ALIGN program for comparisons of amino acid sequences, for example, a PAM120 weight residue table, gap length penalty, and gap penalty can be used.
[0075] As used herein, the term "fusion protein" refers to a hybrid polypeptide comprising protein domains from at least two different proteins. For example, one protein domain can be located at the amino-terminal (N-terminal) portion or the carboxyl-terminal (C-terminal) portion of the fusion protein, thereby forming an "amino-terminal fusion protein" or a "carboxyl-terminal fusion protein," respectively. In alternative embodiments, the fusion protein is a single-chain polypeptide that can be completely encoded by a nucleic acid sequence and comprises at least two protein domains covalently linked by a peptide linkage or optionally covalently linked by a peptide linker.
[0076] (II) Detailed technical plan
[0077] In one aspect, the present disclosure provides a Cas9 protein mutant comprising a mutation at amino acid residue 921, 922, and / or 923 in the amino acid sequence of SEQ ID NO: 1, wherein the mutation is selected from one or more of the following:
[0078] (a) Deletion of amino acid residues;
[0079] (b) substitution of amino acid residues;
[0080] (c) insertion of amino acid residues;
[0081] Furthermore, the mutation can significantly reduce or eliminate the cleavage activity of the Cas9 protein on non-target chains.
[0082] As used herein, the terms "Cas9", "Cas9 protein" or "Cas9 nuclease" refer to an RNA-guided nuclease comprising a Cas9 protein or a fragment thereof (e.g., a protein comprising an active or inactive DNA cleavage domain of Cas9, and / or a gRNA binding domain of Cas9). Cas9 nucleases are sometimes also referred to as casn1 nucleases or CRISPR (clustered regularly interspaced short palindromic repeats)-associated nucleases. CRISPR is an adaptive immune system that provides protection against mobile genetic elements (viruses, transposable elements, and conjugative plasmids). The CRISPR cluster contains a spacer, a sequence complementary to the previous mobile element, and targets invading nucleic acids. The CRISPR cluster is transcribed and processed into CRISPR RNA (crRNA). In type II CRISPR systems, trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and Cas9 protein are required for the correct processing of pre-crRNA. tracrRNA can serve as a guide for ribonuclease 3-assisted processing of pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA endo-cuts a linear or circular dsDNA target complementary to the spacer. First, the target chain that is not complementary to the crRNA is cut, and then 3'-5' exonucleolytically cuts. In nature, DNA binding and cutting generally require protein and two kinds of RNA. However, a single guide RNA ("sgRNA" or "gNRA" for short) can be engineered to incorporate aspects of both crRNA and tracrRNA into a single RNA species. See, for example, Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science. 337: 816-821 (2012), the entire contents of which are incorporated herein by reference.
[0083] The Cas9 protein mutant disclosed in the present invention can significantly reduce or eliminate the cleavage activity of the Cas9 protein on non-target chains, retain or improve the cleavage activity on the target chain, and can have multiple application prospects: (1) Combined with deaminase to establish a single-base editing system. Single nucleotide mutations can cause the occurrence of about 2 / 3 of human genetic diseases and are also the genetic basis of important trait variation in many animals and plants. However, traditional DNA repair using DSB triggers is difficult to achieve efficient and stable single-base mutations or single-base corrections. Therefore, base editors that can perform single-base editing and do not rely on DSBs have emerged. Currently, base editors mainly include two elements: Cas9 protein mutants that only cut target chains and deaminases, among which the combination with adenine deaminase is an adenine base editor (ABE), and the combination with cytosine deaminase is a cytosine base editor (CBE). Base editors have shown great potential in basic research and gene therapy. (2) Paired Cas9 single-nickase gene knockout technology. Off-targeting caused by non-specific binding of Cas9-sgRNA has always been one of the problems of CRISPR / Cas9 gene editing system. Taking advantage of the fact that Cas9 mutants can only perform single-strand cutting, two Cas9 protein mutants act together on the opposite strands of adjacent gene sites to produce a single nick, forming a double-strand break (DSB). However, the non-specific binding of two Cas9 protein mutants is unlikely to occur together at the same off-target site, so it will not cause DSB. The single nick produced is seamlessly repaired using the single-strand break repair pathway. Paired Cas9 single nickase gene knockout technology reduces non-specificity by 50 to 1500 times. The inventors used TREX2 nuclease to expand the paired Cas9 single nickase gene knockout technology, further improving the gene editing efficiency while reducing off-targeting. (3) Combining with mutations that reduce or eliminate the target chain (such as H840A) can obtain dCas9 (dead Cas9) without cutting activity, which can be used to activate or inhibit gene expression. CRISPRa (CRISPR activation) technology allows dCas9 to combine with different transcription activators (such as VP64) to increase the expression level of target genes. Compared with traditional overexpression technology, CRISPRa promotes gene expression by activating endogenous promoters for efficient transcription. It does not require the construction of additional exogenous expression elements and is not limited by the size of gene transcripts. dCas9-mediated gene regulation is reversible and does not permanently modify genomic DNA. Moreover, multiple sgRNAs can be designed to regulate multiple target genes at the same time. In addition, dCas9 is fused with fluorescent protein to locate the target gene site.
[0084] The sequence and structure of Cas9 nuclease are well known to those skilled in the art (see, for example, Ferretti JJ et al., Proceedings of the
[0085] National Academy of Sciences of the United States of America, 2001, 98(8):4658–4663; Deltcheva E. et al., Nature, 2011, 471:602-607; Jinek M. et al., Science, 2012, 337:816-821, the entire contents of each of which are incorporated herein by reference). Cas9 orthologs have been described in a variety of species, including but not limited to Cas9 nucleases from Streptococcus pyogenes and Streptococcus thermophilus. Other suitable Cas9 nucleases and sequences will be apparent to those skilled in the art based on this disclosure, and such Cas9 nucleases and sequences include Cas9 sequences from species and loci disclosed in Chylinski K et al., RNA Biology, 2013, 10(5):726-737; the entire contents of which are incorporated herein by reference.
[0086] In some embodiments, the Cas9 protein mutant comprises an amino acid sequence that is at least 80%, at least 82%, at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5% or at least 99.9 homologous to the amino acid sequence of SEQ ID NO: 1. In some embodiments, Cas9 refers to Cas9 from Streptococcus pyogenes (NCBI Refs: NC_017053.1), Streptococcus dysgalactiae, Streptococcus mutans, Staphylococcus aureus, Klebsiella pneumoniae, Corynebacterium ulcerans (NCBI Refs: NC_015683.1, NC_017317.1); Corynebacterium diphtheriae (NCBI Refs: NC_016782.1, NC_016786.1); Spiroplasma syrphidicola (NCBI Refs: NC_016790.1, NC_016791.1); Ref: NC_021284.1; Prevotella intermedia (NCBI Ref: NC_017861.1); Spiroplasma taiwanense (NCBI Ref: NC_021846.1); Streptococcus iniae (NCBI Ref: NC_021314.1); Belliella baltica (NCBI Ref: NC_018010.1); Psychroflexus torquisi (NCBI Ref: NC_018721.1); Streptococcus thermophilus (NCBI Ref: YP_820832.1); Listeria innocua (NCBI Ref: NP_472073.1); Campylobacter jejuni Jejuni) (NCBI Ref: YP_002344900.1); or Neisseria meningitidis (NCBI Ref: YP_002342100.1).In some embodiments, the Cas9 protein is from Streptococcus pyogenes and can be referred to as SpCas9. The amino acid sequence of the wild-type SpCas9 protein is shown in SEQ ID NO: 1. In some embodiments, the Cas9 protein is a variant of the natural Cas9 protein, for example, some SpCas9 variants obtained by protein directed evolution methods, such as SpCas9 VQR, SpCas9 VRER or SpCas9 NG (ngSpCas9), which recognize NGA, NGCG, and NG as PAM sequences, respectively.
[0087] In some embodiments, the Cas9 nuclease has an inactive (e.g., inactivated) DNA cleavage domain. The nuclease-inactivated Cas9 protein can be interchangeably referred to as a "dCas9" protein ("dead" Cas9). The DNA cleavage domain of known Cas9 includes two subdomains, the HNH nuclease subdomain and the RuvC1 subdomain. The HNH subdomain cuts the chain complementary to the gRNA, while the RuvC subdomain cuts the non-complementary chain. Mutations within these subdomains can silence the nuclease activity of Cas9.
[0088] It should be understood that the Cas9 described in the present disclosure includes the aforementioned natural Cas9 derived from different bacterial species and variants obtained by artificial directed evolution. The positions of amino acid residues described in the present disclosure include the positions of corresponding amino acid residues. For the purpose of convenience of description, the present disclosure uses the Cas9 protein shown in SEQ ID NO: 1 to determine specific amino acid residue positions. It is known to those skilled in the art to determine the positions of corresponding amino acid residues in another Cas9 protein. For example, the amino acid sequence of another Cas9 protein is aligned with the sequence disclosed in SEQ ID NO: 1, and based on the alignment, the Needleman-Wunsch algorithm (Needleman and Wunsch, 1970, J. Mol. Biol. 48: 443-453) implemented in the Needle program of the EMBOSS package (EMBOSS: European Molecular Biology Open Software Suite, Rice et al., 2000, Trends in Genet. 16: 276-277) (preferably version 5.0.0 or later) is used to determine the amino acid position number corresponding to any amino acid residue in the polypeptide disclosed in SEQ ID NO: 2. The parameters used are a gap opening penalty of 10, a gap extension penalty of 0.5, and the EBLOSUM62 (EMBOSS version of BLOSUM62) substitution matrix. In some embodiments, the mutations in the Cas9 protein mutants of the present disclosure comprise:
[0089] L921P;
[0090] V922Δ; and / or
[0091] E923△ or E923T.
[0092] In some embodiments, the Cas9 protein mutant comprises a mutation or combination of mutations selected from the following:
[0093] (1) V922Δ;
[0094] (2) E923Δ;
[0095] (3) L921P and V922Δ;
[0096] (4) V922Δ and E923T;
[0097] (5)L921P;
[0098] (6)E923T;
[0099] (7) L921P and E923T;
[0100] (8) L921P and E923Δ;
[0101] (9) V922Δ and E923Δ;
[0102] (10) L921P, V922Δ, and E923T; and
[0103] (11) L921P, V922Δ and E923Δ.
[0104] In some embodiments, the Cas9 protein mutant is selected from the following mutations or combinations of mutations:
[0105] (1) V922Δ;
[0106] (2) E923Δ;
[0107] (3) L921P and V922Δ; and
[0108] (4) V922Δ and E923T.
[0109] In some embodiments, the Cas9 protein mutant comprises RuvC and HNH domains, wherein the HNH domain comprises a mutation that reduces or eliminates non-target chain cleavage activity. In some embodiments, the mutation that reduces or eliminates non-target chain cleavage activity comprises H840A.
[0110] On the other hand, the present disclosure provides a Cas9 protein mutant comprising an amino acid sequence fragment shown in SEQ ID NO: 35, 37, 39, and 41, wherein the mutant significantly reduces or eliminates the cleavage activity of the Cas9 protein on non-target chains.
[0111] In some embodiments, the Cas9 protein mutant comprises an amino acid sequence fragment having at least 80%, at least 82%, at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5% or at least 99.9 homology to the amino acid sequence fragment shown in SEQ ID NO: 35, 37, 39, or 41.
[0112] In some embodiments, the remaining amino acid sequence of the Cas9 protein mutant is at least 80%, at least 82%, at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5% or at least 99.9 homologous to the corresponding sequence in the amino acid sequence of SEQ ID NO: 1. In some embodiments, the remaining amino acid sequence of the Cas9 protein mutant is identical to the corresponding sequence in the amino acid sequence of SEQ ID NO: 1.
[0113] In some embodiments, the Cas9 protein mutant comprises RuvC and HNH domains, wherein the HNH domain comprises a mutation that reduces or eliminates target strand cleavage activity. In some embodiments, the mutation that reduces or eliminates target strand cleavage activity comprises H840A.
[0114] On the other hand, the present disclosure provides a fusion protein comprising the aforementioned Cas9 protein mutant and a heterologous domain fused to the Cas9 protein mutant.
[0115] The term "heterologous domain" refers to any protein domain that is not present in the Cas9 protein mutants described in the present disclosure, for example, a domain present in the Cas9 protein different from the Cas9 protein mutants described in the present disclosure, or a domain from other proteins such as deaminases, methylases, transcription activators, etc.
[0116] In some embodiments, Cas9 protein can also be modified to include at least one heterologous domain, that is, Cas9 is fused to one or more heterologous domains. In the case where two or more heterologous domains are fused to Cas9, two or more heterologous domains can be identical, or they can be different. In some embodiments, the heterologous domain is selected from a nuclear localization signal domain, a cell penetrating domain, a marker or reporter domain for promoting detection, a DNA or RNA deaminase domain, a uracil-DNA-glycosylase domain, a reverse transcriptase domain, a recombinase domain, an RNA aptamer binding domain, a nuclease domain, a methyltransferase domain, a methylase domain, an acetylase domain, an acetyltransferase domain, a transcription activator domain, and a transcription repressor domain. In some embodiments, the heterologous domain is selected from an exonuclease domain and a deaminase domain. In some embodiments, the heterologous domain is selected from a cytosine deaminase, adenine deaminase, and a DNA binding missing 3' to 5' exonuclease.
[0117] In some embodiments, the Cas9 fusion protein provided herein comprises the full-length amino acid sequence of the Cas9 protein, such as one of the sequences of the Cas9 protein provided above. However, in other embodiments, the fusion protein provided herein does not comprise the full-length Cas9 sequence, but only comprises a fragment thereof. For example, in some embodiments, the Cas9 fusion protein provided herein comprises a Cas9 fragment, wherein the fragment binds crRNA and tracrRNA or sgRNA, but does not comprise a functional nuclease domain, such as a nuclease domain comprising only a truncated form or no nuclease domain at all. Suitable exemplary amino acid sequences of Cas9 domains and Cas9 fragments are provided herein, and it will be apparent to those skilled in the art that additional suitable sequences of Cas9 domains and Cas9 fragments are suitable. In some embodiments, the Cas9 fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the amino acid length of the corresponding wild-type Cas9 protein. In some embodiments, the Cas9 fragment comprises at least 100 amino acids in length. In some embodiments, the Cas9 fragment is at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, 1300, 1350, 1400, 1450, 1500, 1550, or at least 1600 amino acids of the corresponding wild-type Cas9 protein. In some embodiments, the Cas9 fragment comprises an amino acid sequence having at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, or at least 1200 identical contiguous amino acid residues of a corresponding wild-type Cas9 protein.
[0118] In some embodiments, one or more heterologous domains can be fused to the N-terminus, C-terminus, internal position or combination thereof of the Cas9 protein mutant. Fusion can be direct via a chemical bond, or can be indirect via one or more joints. A joint is a chemical group that is connected to one or more other chemical groups via at least one covalent bond. Suitable joints include amino acids, peptides, nucleotides, nucleic acids, organic joint molecules (e.g., maleimide derivatives, N-ethoxybenzylimidazole, biphenyl-3,4',5-tricarboxylic acid, p-aminobenzyloxycarbonyl, etc.), disulfide bond joints and polymer joints (e.g., PEG). The joint can include one or more spacer groups, including but not limited to alkylene, alkenylene, alkynylene, alkyl, alkenyl, alkynyl, alkoxy, aryl, heteroaryl, aralkyl, aralkenyl, aralkynyl, etc. The joint can be neutral, or carry a positive charge or a negative charge. In addition, the joint can be cleavable so that the covalent bond of the joint connecting the joint to another chemical group can be broken or cut under certain conditions, and the conditions include pH, temperature, salt concentration, light, catalyst or enzyme. In some embodiments, the joint can be a peptide joint. The peptide linker can be a flexible amino acid linker (e.g., comprising small, non-polar or polar amino acids). Non-limiting examples of flexible linkers include LEGGGS (SEQ ID NO: 83), TGSG (SEQ ID NO: 84), GGSGGGSG (SEQ ID NO: 85), and (GGGGS) 1-4 (SEQ ID NO: 86). Alternatively, the peptide linker can be a rigid amino acid linker. Such linkers include (EAAAK) 1-4 (SEQ ID NO:87), A(EAAAK) 2-5 A (SEQ ID NO: 88), PAPAP (SEQ ID NO: 89) and (AP) 6-8 Additional examples of suitable linkers are well known in the art, and programs for designing linkers are readily available (eg, Crasto et al., Protein Eng., 2000, 13(5):309-312).
[0119] On the other hand, the present disclosure provides a complex comprising the aforementioned Cas9 protein mutant or the aforementioned fusion protein, and a guide RNA. The guide RNA comprises a polynucleotide consisting of a nucleotide sequence complementary to a target nucleotide sequence from 1 base upstream to 20 bases or more and 24 bases or less upstream of a PAM (protospacer adjacent motif) sequence in a target double-stranded polynucleotide. In some embodiments, the guide RNA is about 15-100 nucleotides long and comprises a sequence of at least 10 consecutive nucleotides complementary to the target sequence. In some embodiments, the guide RNA is 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleotides long. In some embodiments, the guide RNA comprises a sequence of 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 consecutive nucleotides that are complementary to the target sequence. In some embodiments, the target sequence is a DNA sequence. In some embodiments, the target sequence is a sequence in the genome of a mammal. In some embodiments, the target sequence is a sequence in the human genome.In some embodiments, the 3' end of the target sequence is not directly adjacent to a canonical PAM sequence.
[0120] As used herein, "guide RNA" is also referred to as "gRNA" or "guide RNA", and generally refers to an RNA sequence or molecule (or a set of RNA molecules) that can bind to a Cas protein and help target the Cas protein to a specific position within a target polynucleotide (e.g., DNA or RNA). Guide RNA can include a crRNA segment and a tracrRNA segment. As used herein, the term "crRNA" or "crRNA segment" refers to an RNA molecule or a portion thereof, which includes a guide sequence, a stem sequence, and optionally a 5' overhang sequence targeting a polynucleotide. The term "tracrRNA" or "tracrRNA segment" refers to an RNA molecule or a portion thereof that includes a protein binding segment (e.g., a protein binding segment that can interact with a CRISPR-associated protein (e.g., Cas9)). Guide RNA also encompasses single guide RNA (sgRNA), in which the crRNA segment and the tracrRNA segment are located in the same RNA molecule. Guide RNA also collectively includes a group of two or more RNA molecules, in which the crRNA segment and the tracrRNA segment are located in separate RNA molecules.
[0121] In some embodiments, the Cas9 protein or fusion protein and the guide RNA are mixed under mild conditions and cultured to form the complex, and the culture time is preferably 0.5 hours or more and 1 hour or less. The formed complex is stable and can maintain stability even when left at room temperature for several hours. In some embodiments, the Cas9 protein or fusion protein and the guide RNA form a complex on the target nucleic acid. The Cas9 protein or fusion protein recognizes the PAM sequence and binds to the target nucleic acid at a binding site located upstream of the PAM sequence. If the Cas9 protein or fusion protein has endonuclease activity, the polynucleotide is cleaved at this site. The Cas9 protein or fusion recognizes the PAM sequence and, starting from the PAM sequence, strips the double helix structure of the target double-stranded polynucleotide. It anneals with a nucleotide sequence in the guide RNA that is complementary to the target double-stranded polynucleotide, thereby unwinding a portion of the double helix structure of the target double-stranded polynucleotide. At this point, the Cas9 protein or fusion cleaves the phosphodiester bond of the target double-stranded polynucleotide at the cleavage site upstream of the PAM sequence and / or at the cleavage site upstream of the sequence complementary to the PAM sequence.
[0122] In another aspect, the present disclosure provides a polynucleotide encoding the aforementioned Cas9 protein mutant or the aforementioned fusion protein.
[0123] The polynucleotides disclosed herein may be in the form of DNA or RNA. DNA forms include cDNA, genomic DNA, or artificially synthesized DNA. DNA may be single-stranded or double-stranded. DNA may be a coding strand or a non-coding strand. In some embodiments, the polynucleotide sequence encoding the recombinant Cas9 is codon-optimized for expression in eukaryotic cells. In some embodiments, the polynucleotide sequence encoding stiCas9 is codon-optimized for expression in animal cells. In some embodiments, the polynucleotide sequence encoding the recombinant Cas9 is codon-optimized for expression in human cells. In some embodiments, the polynucleotide sequence encoding the recombinant Cas9 is codon-optimized for expression in plant cells. Codon optimization is the adjustment of codons to match the tRNA abundance of the expression host to increase the yield and efficiency of recombinant or heterologous protein expression. Codon optimization methods are routine in the art and can be performed using software programs, such as Integrated DNA Technologies' Codon Optimization Tool, Entelechon's Codon Usage Table Analysis Tool, GENEMAKER's Blue Heron software, Aptagen's Gene Forge software, DNA Builder software, Universal Codon Usage Analysis software, the publicly available OPTIMIZER software, and GenScript's Optimum Gene algorithm.
[0124] In another aspect, the present disclosure provides a vector comprising the aforementioned polynucleotide. In some embodiments, the polynucleotide is further operably linked to a promoter, enhancer, and / or terminator. In some embodiments, the promoter comprises a constitutive promoter, an inducible promoter, a broad-spectrum expression promoter, or a tissue-specific promoter.
[0125] In some embodiments, the vector is an expression vector, more specifically a recombinant expression vector. Suitable expression vectors include viral expression vectors (e.g., viral vectors based on viruses such as vaccinia virus, polio virus, adenovirus, adeno-associated virus, SV40, herpes simplex virus, human immunodeficiency virus, retroviral vectors (e.g., murine leukemia virus, spleen necrosis virus, and vectors derived from retroviruses such as Rous sarcoma virus, Harvey sarcoma virus, avian leukemia virus, lentivirus, human immunodeficiency virus, myeloproliferative sarcoma virus, and mammary tumor virus), etc.
[0126] In another aspect, the present disclosure provides a kit comprising the aforementioned Cas9 protein mutant, the aforementioned fusion protein, the aforementioned complex, the aforementioned polynucleotide, or the aforementioned vector; preferably, the kit further comprises an expression construct encoding a guide RNA backbone, wherein the expression construct comprises a cloning site that allows a nucleic acid sequence identical or complementary to a target sequence to be cloned into the guide RNA backbone.
[0127] As used herein, "guide RNA backbone" refers to the sequence within the guide RNA responsible for Cas9 binding, which does not include the spacer sequence used to guide Cas9 to the target DNA.
[0128] In another aspect, the present disclosure provides a cell comprising the aforementioned Cas9 protein mutant, the aforementioned fusion protein, the aforementioned complex, the aforementioned polynucleotide, or the aforementioned vector.
[0129] In some embodiments, cells can be transfected with appropriate molecules (i.e., proteins, DNA, and / or RNA). Suitable transfection methods include nucleofection (or electroporation), calcium phosphate-mediated transfection, cationic polymer transfection (e.g., DEAE-dextran or polyethyleneimine), viral transduction, virosome transfection, virion transfection, liposome transfection, cationic liposome transfection, immunoliposome transfection, non-liposomal lipofection, dendrimer transfection, heat shock transfection, magnetofection, lipofection, gene gun delivery, puncture transfection, sonoporation, phototransfection, and nucleic acid uptake enhanced by proprietary agents. Transfection methods are well known in the art (see, for example, "Current Protocols in Molecular Biology" Ausubel et al., John Wiley & Sons, New York, 2003, or "Molecular Cloning: A Laboratory Manual" Sambrook & Russell, Cold Spring Harbor Press, Cold Spring Harbor, NY, 3rd edition, 2001). In other embodiments, the molecule can be introduced into the cell by microinjection. For example, the molecule can be injected into the cytoplasm or nucleus of the target cell. The amount of each molecule introduced into the cell can vary, but those skilled in the art are familiar with the means for determining the appropriate amount. Various molecules can be introduced into the cell simultaneously or sequentially. For example, the modified Cas9 protein (or its encoding nucleic acid) and the donor polynucleotide can be introduced simultaneously. Alternatively, one can be introduced first and then the other can be introduced into the cell.
[0130] Various cells are suitable for use in methods disclosed herein, including prokaryotic cells (e.g., bacteria) and eukaryotic cells (e.g., animal cells, insect cells, and plant cells). For example, cells can be human cells, non-human mammalian cells, non-mammalian vertebrate cells, invertebrate cells, insect cells, plant cells, yeast cells, or unicellular eukaryotic organisms. In some embodiments, cells can be unicellular embryos. For example, non-human mammalian embryos include rats, hamsters, rodents, rabbits, cats, dogs, sheep, pigs, cattle, horses, and primate embryos. In other embodiments, cells can be stem cells, such as embryonic stem cells, ES-like stem cells, fetal stem cells, adult stem cells, etc. In one embodiment, stem cells are not human embryonic stem cells. In addition, stem cells can include those stem cells prepared by the technology disclosed in WO2003 / 046141 or Chung et al. (Cell Stem Cell, 2008, 2:113-117), which are incorporated herein by their entirety. The cell can be in vitro (i.e., in culture), ex vivo (i.e., in a tissue isolated from an organism), or in vivo (i.e., within an organism). In exemplary embodiments, the cell is a mammalian cell or a mammalian cell line. In specific embodiments, the cell is a human cell or a human cell line.
[0131] Other aspects of the present disclosure include animals engineered to encode nucleic acids or vectors as described above, or animals permanently modified by an engineered SpCas9 variant of the present disclosure. For example, the animal can be a model organism (i.e., Drosophila melanogaster, mouse, mosquito, rat), or the animal can be a farm animal or farmed fish or pet. As another example, the animal can be a vector of at least one disease. As another example, the organism can be a vector for a human disease (i.e., mosquito, tick, bird).
[0132] Still other aspects of the present disclosure include plants modified using nucleic acids or vectors as described above, or plants transiently or permanently modified by modified SpCas9 variants of the present disclosure. For example, the plant can be a crop plant (i.e., rice, soybean, wheat, tobacco, cotton, alfalfa, canola, corn, sugar beet, etc.).
[0133] On the other hand, the present disclosure provides a kind of pharmaceutical composition, which comprises the aforementioned Cas9 protein mutant, the aforementioned fusion protein, the aforementioned complex, the aforementioned polynucleotide, the aforementioned vector or the aforementioned cell. In some embodiments, the pharmaceutical composition further comprises at least one pharmaceutically acceptable excipient. Pharmaceutically acceptable excipients generally include inactive ingredients used as vehicles (such as water, capsule shells, etc.), diluents, or components constituting dosage forms or pharmaceutical compositions, and the dosage forms or pharmaceutical compositions include drugs such as therapeutic agents. Pharmaceutically acceptable excipients also include typical inactive ingredients that impart adhesive function (i.e., adhesive), disintegrating function (i.e., disintegrant), lubricant function (lubricant) and / or other functions (i.e., solvents, surfactants, etc.) to the composition.
[0134] On the other hand, the present disclosure provides the use of the aforementioned Cas9 protein, the aforementioned fusion protein, the aforementioned complex, the aforementioned polynucleotide, the aforementioned vector, the aforementioned kit or the aforementioned cell or the aforementioned pharmaceutical composition in nucleic acid modification.
[0135] The aforementioned Cas9 protein mutants, the aforementioned fusion proteins, the aforementioned complexes, the aforementioned polynucleotides, the aforementioned vectors, the aforementioned kits or the aforementioned cells or the aforementioned pharmaceutical compositions disclosed herein can be used in various therapeutic, diagnostic, industrial and research applications. In some embodiments, the present disclosure can be used to modify any target chromosomal sequence in cells, animals or plants in order to construct a gene function model and / or study the function of a gene, study the target genetic or epigenetic conditions, or study the biochemical pathways involved in various diseases or conditions. For example, a transgenic organism that constructs a model of a disease or condition can be produced, in which the expression of one or more nucleic acid sequences associated with the disease or condition is altered. Disease models can be used to study the effects of mutations on organisms, study the development and / or progression of a disease, study the effects of pharmaceutically active compounds on a disease, and / or evaluate the efficacy of potential gene therapy strategies.
[0136] In some embodiments, the present disclosure can be used to perform efficient and cost-effective functional genomic screening, which can be used to study the function of genes involved in a specific biological process and how any changes in gene expression can affect the biological process, or to perform saturation or deep scanning mutagenesis of genomic loci that are associated with cellular phenotypes. For example, saturation or deep scanning mutagenesis can be used to determine the critical minimum features and discrete vulnerabilities of functional elements required for gene expression, drug resistance, and disease reversal.
[0137] In some embodiments, the present disclosure can be used in diagnostic tests to determine the presence of a disease or condition and / or to determine treatment options. Examples of suitable diagnostic tests include detecting specific mutations in cancer cells (e.g., specific mutations in EGFR, HER2, etc.), detecting specific mutations associated with specific diseases (e.g., trinucleotide repeats, mutations in β-globin associated with sickle cell disease, specific SNPs, etc.), detecting hepatitis, detecting viruses (e.g., Zika virus), and the like.
[0138] In some embodiments, the present disclosure can be used to correct genetic mutations associated with specific diseases or conditions, for example, correcting globin gene mutations associated with sickle cell disease or thalassemia, correcting mutations in the adenosine deaminase gene associated with severe combined immunodeficiency (SCID), reducing the expression of HTT, the causative gene for Huntington's disease, or correcting mutations in the rhodopsin gene for the treatment of retinitis pigmentosa. Such modifications can be performed in ex vivo cells.
[0139] In some embodiments, the present disclosure can be used to generate crop plants with improved traits or increased resistance to environmental stresses. The present disclosure can also be used to generate farm animals or production animals with improved traits. For example, pigs have many characteristics that make them attractive as biomedical models, especially in regenerative medicine or xenotransplantation.
[0140] In another aspect, the present disclosure provides a method of modifying a target nucleic acid, the method comprising contacting the target nucleic acid with:
[0141] (a) the aforementioned Cas9 protein mutant or the aforementioned fusion protein, and a guide RNA; or
[0142] (b) The aforementioned complex.
[0143] In another aspect, the present disclosure provides a method for modifying a target nucleic acid, comprising introducing the aforementioned Cas9 protein mutant, the aforementioned fusion protein, the aforementioned complex, the aforementioned polynucleotide or the aforementioned vector into a cell or organism comprising the target nucleic acid.
[0144] In some embodiments, the target nucleic acid comprises a sequence associated with a disease or condition. In some embodiments, the target nucleic acid comprises a point mutation associated with a disease or condition. In some embodiments, the activity of the Cas9 protein, Cas9 fusion protein, or complex results in correction of the point mutation. In some embodiments, the target nucleic acid comprises a T→C point mutation associated with a disease or condition, and wherein the deamination of the mutant C base results in a sequence not associated with the disease or condition. In some embodiments, the target nucleic acid encodes a protein, and wherein the point mutation is located in a codon and causes the amino acid encoded by the mutant codon to change compared to the wild-type codon. In some embodiments, the deamination of the mutant C results in a change in the amino acid encoded by the mutant codon. In some embodiments, the deamination of the mutant C results in a codon encoding a wild-type amino acid. In some embodiments, the contact is in the subject. In some embodiments, the subject suffers from or has been diagnosed with a disease or condition. In some embodiments, the disease or disorder is cystic fibrosis, phenylketonuria, epidermolytic hyperkeratosis (EHK), Charchar-Marie-Toot disease type 4J, neuroblastoma (NB), von Willebrand disease (vWD), myotonia congenita, hereditary renal amyloidosis, dilated cardiomyopathy (DCM), hereditary lymphedema, familial Alzheimer's disease, HIV, prion disease, chronic infantile neurocutaneous articular syndrome (CINCA), desmin-related myopathy (DRM), a neoplastic disease associated with mutant PI3KCA protein, mutant CTNNB1 protein, mutant HRAS protein or mutant p53 protein.
[0145] On the other hand, the present disclosure provides a method for treating a subject diagnosed with a disease associated with or caused by a gene mutation, comprising administering to the subject a therapeutically effective amount of the aforementioned Cas9 protein, the aforementioned fusion protein, the aforementioned complex, the aforementioned polynucleotides, the aforementioned vector, the aforementioned cell and / or the aforementioned pharmaceutical composition. For example, in some embodiments, a method is provided, comprising administering to a subject suffering from such a disease (e.g., a cancer associated with the PI3KCA point mutation as described above) an effective amount of correction point mutation or introducing a mutated Cas9 deaminase fusion protein in a disease-related gene. In some embodiments, the disease is a proliferative disease. In some embodiments, the disease is a genetic disease. In some embodiments, the disease is a neoplastic disease. In some embodiments, the disease is a metabolic disease. In some embodiments, the disease is a lysosomal storage disease. Other diseases that can be treated by correcting point mutations or introducing inactivating mutations into disease-related genes are known to those skilled in the art, and the present disclosure is not limited in this regard.
[0146] As used herein, the term "effective amount" or "therapeutically effective amount" encompasses an amount sufficient to ameliorate or prevent the symptoms or conditions of a medical condition. An effective amount also means an amount sufficient to allow or facilitate diagnosis. The effective amount for a particular patient or veterinary subject may vary depending on factors such as the condition to be treated, the patient's overall health, the route and dosage of administration, and the severity of side effects. An effective amount can be the maximum dose or dosage regimen that avoids significant side effects or toxic effects.
[0147] As used herein, the terms "subject," "individual," or "patient" are used interchangeably herein and refer to a vertebrate, preferably a mammal, more preferably a human. Mammals include, but are not limited to, mice, apes, humans, farm animals, sport animals, and pets.
[0148] The amount of the substance administered may depend on the subject being treated, the subject's age, health, sex, and weight, the type of concurrent treatment (if any), the severity of the condition, the nature of the desired effect, the mode and frequency of treatment, and the judgment of the prescribing physician. The frequency of administration may also depend on the pharmacodynamic effect on arterial oxygen tension. However, the most preferred dosage may be adjusted for individual subjects, as will be understood by those skilled in the art and can be determined without undue experimentation. This typically involves adjusting the standard dosage (e.g., reducing the dosage if the patient is of low weight).
[0149] In order to make the purpose, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings. Unless otherwise specified, the technical means used in the embodiments are conventional means well known to those skilled in the art. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. In the absence of conflict, the following embodiments and the features in the embodiments can be combined with each other. The following examples are used to further illustrate the present disclosure, but should not be construed as limiting the present disclosure. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present disclosure should be equivalent replacement methods and are included in the scope of protection of the present disclosure.
[0150] Example 1. Screening for Cas9 protein mutants with new functions
[0151] (1) Mouse embryonic stem cell HR reporter system
[0152] This patent uses an HR reporter system derived from Chandramouly G et al. (BRCA1 and CtIP suppress long-tract gene conversion between sister chromatids. Nature Communications, 2013, 4(1): 2404) and is introduced into mouse embryonic stem cells. The mouse embryonic stem cell HR reporter system used can be used to measure short-tract gene conversion (STGC) and long-tract gene conversion (LTGC) induced by Cas9 and nCas9 and their bias. The reporter system contains two mutant inactivated GFP genes: the first GFP (TrGFP) lacks the 5' end, and the second GFP (I-SceI-GFP) is interrupted by an 18bp recognition site for the rare nuclease I-SceI. Between the two mutant GFPs, two head-to-tail inverted artificial exon expression frames A and B of the RFP gene are inserted, and only cells that undergo rearranged HR will emit light. When Cas9 induces DSBs or nCas9-induced replication-coupled single-end DSBs at and near the I-SceI site, cells can use the adjacent sister chromatid TrGFP as a template for gene conversion during the G2 / S phase of the cell cycle to generate green fluorescence (GFP + ) cells or green and red fluorescence (GFP + RFP + ) cells, and cells that emit different fluorescence can be detected by flow cytometry. There are two forms of gene conversion produced by HR: 1) STGC, the product of which is GFP + Cells; 2) LTGC, the product is GFP + RFP + In cells, STGCs typically convert 50-200 base pairs of DNA, while LTGCs can convert over 1000 base pairs of DNA. This reporter system can also be used to measure and analyze the selection bias of STGCs and LTGCs for DSB repair products induced by different Cas9 nucleases in cells.
[0153] (2) Screening of Cas9 protein mutants
[0154] In mouse embryonic stem cells, the SpCas9 knock-in donor and the SpCas9 overexpression vector (vector structure as shown in Figure 2) containing homologous sequences at both ends were simultaneously introduced. Figure 2The method comprises the following steps: (a) an expression vector containing sgRNA targeting the mouse Rosa26 site, and a single copy of the SpCas9 gene is precisely knocked into the Rosa26 site by homologous recombination gene knock-in. The amino acid sequence of SpCas9 is shown in SEQ ID NO: 1, the gene sequence of SpCas9 is shown in SEQ ID NO: 2, and the sgRNA spacer sequence used is 5'-ACTCCAGTCTTTCTAGAAGA-3' (SEQ ID NO: 3). After the monoclone grows out, the primer pair TGGAAGAAGCTGTCGTCCACCTTG (SEQ ID NO: 4, forward primer) and CGAGACACGGATCGACCTGTCT (SEQ ID NO: 5, reverse primer) are used to perform targeted PCR on the genome of the monoclonal cell line with the SpCas9 gene site-directed knock-in. Combined with quantitative PCR technology, a monoclonal cell line with a precisely site-directed insertion of the single copy of the SpCas9 gene is screened.
[0155] The LbCas12a expression plasmid (from Addgene) and the expression plasmid pU6-LbCas12a-sgRNA corresponding to the sgRNA targeting the specific target of the SpCas9 gene were introduced into the cell line using the Lipofectamine 2000 liposome transfection method. LbCas12a-sgRNA targets the SpCas9 sequence, and the sgRNA sequence used is 5'-TGATCTGCCGGGTTTCCACCAGC-3' (SEQ ID NO: 6), and its PAM is 5'-TTTG-3' (SEQ ID NO: 7). SpCas9 insertion and deletion mutations are induced, various insertion and deletion mutants are generated, and an insertion and deletion mutant cell pool is formed.
[0156] Using Lipofectamine 2000, the cell pool was transfected with a pU6-SpCas9-crRNA vector expressing the SpCas9 sgRNA (g1-1). The sgRNA sequence used was g1-1: 5'-GATAACAGGGTAATCAAGG-3' (SEQ ID NO: 8). Active SpCas9 mutants remaining in the cell pool will use this sgRNA to cleave the HR reporter system. After cleavage of the HR reporter system, the cells will produce GFP and RFP fluorescence.
[0157] The cell pool transfected with sgRNA was subjected to flow cytometry sorting, and GFP-positive cells and GFP- and RFP-double-positive cells were sorted respectively. Genomic DNA was extracted using a genomic DNA extraction kit (NoviZan: DC102-01) from the unsorted cell pool cells. The cleavage site was amplified by PCR using a primer pair (CP4-F: GCTGAACGCCAAGCTGATTAC (SEQ ID NO: 9); CP4-R: ATCCTTCCGGAAATCGGACAC (SEQ ID NO: 10)) to obtain the DNA sequence near the editing site. Next-generation sequencing was performed at Novogene to obtain DNA mutation enrichment data. The ratios of different mutant sequences in unsorted cell pools, GFP-positive cells, and GFP- and RFP-double-positive cells were counted, and the enrichment levels of different insertion / deletion variants were analyzed. Mutations with higher enrichment levels than wild-type SpCas9 were considered as potential available SpCas9 insertion / deletion mutants. SpCas9V922△, SpCas9 E923△, SpCas9 V922△ / E923T, and SpCas9 L921P / V922△ mutants were found to be potential available insertion / deletion mutants.
[0158] Using the SpCas9 overexpression vector plasmid as a template, different fragments were generated by PCR amplification using the following primers (corresponding mutations were introduced by the primers):
[0159] (1)SpCas9-F:5'-CCATGGCCTACCCCTACGAC-3' (SEQ ID NO: 11)
[0160] (2)V922△-R:5'-GCCGGGTTTCCAGCTGTCTCTTGATGAAGCC-3' (SEQ ID NO: 12)
[0161] (3)E923△-R:5'-TCTGCCGGGTCACCAGCTGTCTCTTGATGAAGC-3' (SEQ ID NO: 13)
[0162] (4)V922△ / E923T-R:5'-CCGGGTTTCCGGCTGTCTCTTGATGAAGCCGG-3'(SEQ ID NO: 14)
[0163] (5)L921P / V922△-R:5'-CTGCCGGGTTGTCAGCTGTCTCTTGATGAAGCCG-3' (SEQ ID NO: 15)
[0164] (6)SpCas9-R:5'-GCTGATCAGCGGGTTTAAACG-3' (SEQ ID NO: 16)
[0165] (7)V922△-F:5'-AAGAGACAGCTGGAAACCCGGCAGATCACAAAG-3' (SEQ ID NO: 17)
[0166] (8)E923△-F:5'-ACAGCTGGTGACCCGGCAGATCACAAAGC-3' (SEQ ID NO: 18)
[0167] (9)V922△ / E923T-F:5'-AAGAGACAGCCGGAAACCCGGCAGATCACAAAG-3'(SEQ ID NO: 19)
[0168] (10) L921P / V922Δ-F:5'-GAGACAGCTGACAACCCGGCAGATCACAAAGC-3' (SEQ ID NO: 20).
[0169] Amplification was performed using paired primers (1) + (2), (1) + (3), (1) + (4), (1) + (5) to obtain fragments (a), (b), (c), and (d); amplification was performed using paired primers (6) + (7), (6) + (8), (6) + (9), and (6) + (10) to obtain fragments (e), (f), (g), and (h). Fragments (a) + (e), (b) + (f), (c) + (g), (d) + (h) were mixed and amplified using paired primers (1) + (6) to obtain fragments (i), (j), (k), and (l). Fragments (i), (j), (k), and (l) were mixed with 3β-HA-330 (vector structure as shown in FIG. 1 ) cut with BamHI + XbaI. Figure 3 The 6187 bp vector DNA fragment (m) generated after the addition of the vector sequence as shown in SEQ ID NO: 21 was mixed and homologous recombination was performed using a homologous recombination kit (Wuhan Aibotek) to obtain expression plasmids for the mutants SpCas9V922Δ, SpCas9E923Δ, SpCas9 V922Δ / E923T, and SpCas9 L921P / V922Δ. The sequences of the above DNA fragments am are shown in Table 2.
[0170] Table 2. Nucleotide sequences of relevant DNA fragments
[0171]
[0172]
[0173] These amplified fragments were sequenced by the Sanger method at Qingke Biotechnology Co., Ltd. (sequencing was performed using primers (1) + (6), and all sequences between the two primers were obtained by the sequencing company). The nucleotide sequences and amino acid sequences of four mutants, SpCas9 V922Δ, SpCas9 E923Δ, SpCas9 V922Δ / E923T, and SpCas9 L921P / V922Δ, were obtained. The amino acid sequences at positions 908-940 and the corresponding nucleotide sequences of the four Cas9 protein mutants are shown in Table 3. The remaining amino acid sequences are the same as those in SEQ ID NO: 1.
[0174] Table 3. Amino acid sequences at positions 908-940 and corresponding nucleotide sequences of four Cas9 mutants
[0175]
[0176] Example 2. In vitro enzyme cleavage assay to verify the in vitro enzyme cleavage activity of Cas9 protein mutants
[0177] SpCas9, SpCas9 H840A and four SpCas9 insertion / deletion mutants SpCas9 V922△, SpCas9-E923△, SpCas9L921P / V922△ and SpCas9 V922△ / E923T proteins (carrying His tags) were purified in vitro. The pET28(a) vector was cut into linear fragments using restriction endonucleases MluI and XhoI. The 843 bp sequence on the pET28(a) vector was amplified by PCR using the pET28(a) vector as a template (F primer: 5'-TGATCAGCCCACTGACGCGT-3' (SEQ ID NO: 44); R primer: 5'-GGTATATCTCCTTCTTAAAG-3' SEQ ID NO: 45). The 843 bp sequence on the vector was amplified by PCR (F primer: 5'-CTTTAAGAAGGAGATATACCatggacaagaagtacagcatcgg-3' (SEQ ID NO: 46); R primer: 5'-TGGTGGTGGTGGTGCTCGAGgtcacctcccagctgagaca-3' SEQ ID NO: 47). 15bp homologous sequences were added to each end of the L921P / V922△ and SpCas9 V922△ / E923T fragments (the capital letters in the primers indicate the homologous sequences), and the corresponding C-terminal fused 6×His tag Cas protein expression vectors were constructed using seamless cloning (Ibotek, RK21020). Escherichia coli BL21 cells were transformed with the pET28(a) vector encoding Cas9, and single clones were picked and the bacterial solution concentration reached OD 600 When the pH value was between 0.6 and 0.8, IPTG was added to a final concentration of 0.5 mM and induced at 16°C, 180 rpm for 18 hours. Cells were harvested by centrifugation and disrupted by sonication in lysis buffer (20 mM Tris-HCl, pH 8.0, 300 mM NaCl, 10 mM imidazole, 0.5 mM PMSF). After disruption, cells were centrifuged twice at 10,000 g for 15 minutes each time, and the supernatant was filtered through a 0.45 μM filter. The supernatant was incubated with Ni NTA beads (Tiandi Renhe, SA004100) in a gravity chromatography column for 1 hour, and the supernatant was discarded. Cells were then washed three times with wash buffer (20 mM Tris-HCl, pH 8.0, 300 mM NaCl, 20 mM imidazole, 0.5 mM PMSF) at 4°C for 5 minutes each time. The relevant proteins were eluted with elution buffer (20 mM Tris-HCl pH 8.0, 300 mM NaCl, 250 mM imidazole, 0.5 mM PMSF).
[0178] A 3110 bp linear double-stranded DNA substrate containing the SpCas9 target was cleaved in vitro, the nucleotide sequence of which is shown in SEQ ID NO: 48. The 3110 bp double-stranded DNA substrate was amplified and recovered by PCR. An sgRNA gdd2 targeting the above double-stranded DNA substrate was constructed. The gdd2 sequence was 5'-ATCCATGGTGGCGGCGGTTA-3' (SEQ ID NO: 49), and the PAM on the Crick strand was 5'-GGG-3'. The gdd2 SpCas9 sgRNA 5'-ATCCATGGTGGCGGCGGTTAGTTTTAGAGCTAGAAATAGCAAGTTAA AATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGC-3' (SEQ ID NO: 50) was synthesized. 200 ng of double-stranded DNA substrate was incubated with 100 nM of the in vitro-purified SpCas9, SpCas9 H840A, four SpCas9 insertion / deletion mutants SpCas9 V922Δ, SpCas9-E923Δ, SpCas9 L921P / V922Δ, and SpCas9 V922Δ / E923T, and 100 nM sgRNA gdd2 in lysis buffer (20 mM HEPES, pH 7.5, 150 mM KCl, 0.5 mM DTT, 0.1 mM EDTA) at 37°C for 30 min. No Cas enzyme was added for the blank control. The reaction was terminated with DNA loading buffer (Novozymes, P022-01), separated by 1% agarose gel electrophoresis, and visualized by Safe Green (Aibotek, RM19008). The results showed that SpCas9 protein can cut linear double-stranded DNA substrates, cutting the 3110bp substrate into two fragments of 1900bp and 1210bp. SpCas9 H840A and four SpCas9 insertion / deletion mutants SpCas9 V922△, SpCas9 E923△, SpCas9 L921P / V922△ and SpCas9 V922△ / E923T proteins do not have double-cutting activity ( Figure 4 A).
[0179] Two short double-stranded DNA substrates (70 bp) were prepared to synthesize a 5'-T (Alexa 680) labeled single-stranded DNA fragment of SEQ ID NO: 51 (5'-T (Alexa 680) TCAGGGTCAGCTTGCCGTAGGTGGCATCGCCCTCGCCCTCGCCGGACACGCTGAACTTGTGGCCG TTTA-3' (Alexa 680 labeled the first base of the 5' sequence) and its reverse complement (5'-T (Alexa 680) labeled single-stranded DNA of SEQ ID NO: 52 (5'-T (Alexa 680) AAACGGCCACAAGTTCAGCGTGTCCGGCGAGGGCGAGGGCGATGCCACCTAC GGCAAGCTGACCCTGAA-3' (Alexa 680 labeled the first base of the 5' sequence). The reverse complementary sequences (unlabeled) of the fluorescently labeled single-stranded DNA were synthesized separately. The fluorescently labeled single-stranded DNA was annealed with its complementary sequence to form a 5' Alexa 680 fluorescently labeled double-stranded DNA substrate and a 3' Alexa 680 fluorescently labeled double-stranded DNA substrate. A SpCas9 sgRNA gX1-F targeting the double-stranded DNA substrate was constructed. The gX1-F sequence was 5'-CAAGTTCAGCGTGTCCGGCG-3' (SEQ ID NO: 53), and the PAM was on the Crick strand, 5'-AGG-3'. The gX1-FSpCas9 sgRNA 5'-CAAGTTCAGCGTGTCCGGCGGTTTTAGAGCTAGAAATAGCAAGTTAAA ATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGC-3' (SEQ ID NO: 54) was synthesized. 20 ng of fluorescently labeled double-stranded DNA substrate was incubated with 100 nM SpCas9, SpCas9 H840A, four SpCas9 insertion / deletion mutants (SpCas9 V922Δ, SpCas9-E923Δ, SpCas9 L921P / V922Δ, and SpCas9 V922Δ / E923T), and 100 nM sgRNA gX1-F in lysis buffer (20 mM HEPES, pH 7.5, 150 mM KCl, 0.5 mM DTT, 0.1 mM EDTA) at 37°C for 30 minutes. No Cas enzyme was added in the blank control. The reaction was terminated with 2x TBE urea DNA loading buffer (Tomobio, TLB-5), separated by 8% denaturing urea polyacrylamide gel electrophoresis, and visualized using an Odyssey infrared fluorescence scanning imaging system using 680 nm excitation light.
[0180] The results showed that, unlike SpCas9 H840A, which only cuts the non-target strand of the DNACas9-sgRNA target, the four SpCas9 insertion / deletion mutants SpCas9-V922△, SpCas9-E923△, SpCas9-L921P / V922△, and SpCas9-V922△ / E923T proteins can only cut the target strand of the Cas9-sgRNA target, but cannot cut the non-target strand of the target DNA ( Figure 4 B).
[0181] Example 3. NHEJ reporter system verifies the activity of Cas9 protein mutants
[0182] The mouse embryonic stem cell NHEJ reporter system used can detect the ability of SpCas9 and nCas9 (nickase Cas9) to induce NHEJ repair products after targeted cleavage. The reporter system is transcribed by the PGK promoter, and there is an enhanced "Koz-ATG" downstream of the promoter, followed by a GFP gene. In the initial state, the mRNA transcribed and expressed by the reporter system is translated by "Koz-ATG". At this time, the subsequent GFP fusion gene is in a frameshift state and cannot translate active GFP. If SpCas9 or nCas9 is used to target the DNA sequence between "Koz-ATG" and GFP to induce double-strand breaks (DSBs) or single-strand breaks (SSBs), the insertion or deletion of bases will be introduced after repair through the NHEJ pathway, resulting in full-frame translation of GFP. Therefore, we can directly detect the proportion of cells expressing GFP by flow cytometry, thereby detecting the proportion of NHEJ products. This reporter system can be used to detect the proportion of NHEJ repair products after targeted cleavage by SpCas9 or Cas9 single nickase ( Figure 5 A).
[0183] The nucleotide sequence of the NHEJ reporter system is shown in SEQ ID NO: 55, and is as follows:
[0184]
[0185]
[0186] The bold part is the PGK promoter sequence, the italic part is the Kozak-ATG sequence, and the underlined part is the GFP sequence.
[0187] sgRNA g2-2 was designed and constructed based on the DNA sequence between "Koz-ATG" and GFP in a mouse embryonic stem cell NHEJ reporter system. The g2-2 sequence is 5'-GATAACAGGGTAATCCATGG-3' (SEQ ID NO: 56), and the PAM is 5'-TGG-3' on the Crick strand. 5'-GATAACAGGGTAATCCATGG-3' (SEQ ID NO: 56) and its reverse complement were synthesized, annealed, and ligated into the BbsI-digested px330-U6-Chimeric plasmid vector (modified from pX330-U6-Chimeric_BB-CBh-hSpCas9 (Addgene, Cat#42230), with the SpCas9 expression cassette removed) to form the g2-2 expression plasmid. The wild-type SpCas9, the expression plasmids of SpCas9V922△, SpCas9 E923△, SpCas9 L921P / V922△ and SpCas9 V922△ / E923T constructed in Example 1, and the g2-2 expression plasmid were co-transfected into mouse embryonic stem cells containing the NHEJ reporter system using the Lipofectamine 2000 liposome transfection method. The transfection system (taking a 24-well plate as an example) was as follows: 8 x 10 cells per well in 24 wells. 4 The total amount of transfected DNA was 0.5 μg, including 0.25 μg of Cas9 expression plasmid and 0.25 μg of g2-2 expression plasmid. The medium was changed 6-8 hours after transfection, and GFP was counted by flow cytometry 72 hours after transfection. + The results showed that the levels of NHEJ repair products induced by the four SpCas9 mutants SpCas9 V922△, SpCas9 E923△, SpCas9 L921P / V922△ and SpCas9 V922△ / E923T were significantly lower than those of wild-type SpCas9 ( Figure 5 B).
[0188] Example 4. Activity verification of Cas9 protein mutants combined with H840A mutation
[0189] (1) HR reporter system to detect the activity of Cas9 protein mutants with H840A mutation
[0190] The primer sequences used below are as follows:
[0191] (1)5'-AAGCCCGAGAACATCGTGAT-3' (SEQ ID NO: 57);
[0192] (2) 5'-GCCCTCTAGACTCGAGCAGC-3' (SEQ ID NO: 58);
[0193] (3) 5'-ACGAGGACATTCTGGAAGATATCGT-3' (SEQ ID NO: 59);
[0194] (4) 5'-ATCACGATGTTCTCGGGCTT-3' (SEQ ID NO: 60).
[0195] First, a SpCas9 H840A single mutation vector was constructed. The reverse complementary primers used to construct SpCas9H840A, which introduced point mutations, were F: 5'-TACGATGTGGACGCCATCGTGCCTCAGAGC-3' (SEQ ID NO: 61), R: 5'-GCTCTGAGGCACGATGGCGTCCACATCGTA-3' (SEQ ID NO: 62). PCR amplification was performed using 3β-HA-330 as a template. The plasmid template was then digested with Dpn1, and the SpCas9 H840A single mutation expression vector was obtained by transforming Escherichia coli. Fragments (c), (d), (e) and (f) were obtained by PCR amplification using paired primers (1) + (2) with SpCas9 V922△, SpCas9 E923△, SpCas9 L921P / V922△ and SpCas9 V922△ / E923T expression vectors as templates, and fragment (g) was obtained by PCR amplification using paired primers (3) + (4) with SpCas9 H840A expression plasmid as template. Fragments (c), (d), (e) and (f) were mixed with the vector DNA fragment generated after cutting 3β-HA-330 with EcoRV+XhoI, and homologous recombination was performed using a homologous recombination kit (Wuhan Aiboteike) to construct the SpCas9 V922△, SpCas9 E923△, SpCas9 L921P / V922△ and SpCas9 V922△ / E923T combined with H840A double mutation expression vectors. Four SpCas9 insertion / deletion mutants, SpCas9 V922△, SpCas9 E923△, SpCas9 L921P / V922△, and SpCas9V922△ / E923T, were combined with the H840A double mutant vector and co-transfected with sgRNAs targeting the I-SceI site (g1-1 and g1-2, where the sequence of g1-2 is 5'-TCGAGCTGAAGGGCATCGTA-3' (SEQ ID NO: 43)) into mouse embryonic stem cells containing the HR reporter system to verify the activity of the double mutants. The results showed that the double mutant SpCas9 could not induce STGC and LTGC, and completely lost its cleavage activity ( Figure 1 AF).
[0196] (2) Detection of the activity of Cas9 protein mutants with H840A mutation using NHEJ reporter system
[0197] Using the same method as in Example 4(1), four SpCas9 insertion / deletion mutants, SpCas9-V922△, SpCas9-E923△, SpCas9-L921P / V922△, and SpCas9-V922△ / E923T combined with H840A double mutant vectors were constructed, and mouse embryonic stem cells containing an NHEJ reporter system were co-transfected with sgRNA (g2-2) targeting the DNA sequence between "Koz-ATG" and GFP to verify the activity of the double mutants. The specific operation was the same as in Example 3. The results showed that the double mutant SpCas9 could not induce the production of NHEJ repair products and had completely lost its cleavage activity ( Figure 5 B).
[0198] Example 5. Construction of Cas9 protein mutant base editor and verification of base editing efficiency
[0199] The editing efficiency of the base editor was detected using the mouse embryonic stem cell NHEJ reporter system. The reporter system is transcribed by the PGK promoter, and there is an enhanced "Koz-ATG" downstream of the promoter, followed by a GFP gene. In the initial state, the mRNA transcribed and expressed by the reporter system is translated by "Koz-ATG". At this time, the subsequent GFP fusion gene is in a frameshift state and cannot translate active GFP. If the A in the GAC of the complementary chain of "Koz-ATG" is edited to C using an adenine base editor, or the C in GAC is edited to T using a cytosine base editor, the "Koz-ATG" will be destroyed, the translation start ATG of GFP itself will be enabled, the GFP gene will be translated normally, and GFP will be produced. + Cells can be counted by flow cytometry ( Figure 6 A).
[0200] A SpCas9 sgRNA vector targeting the "Koz-ATG" residue was designed and constructed. The PAM residue is located on the Crick strand and is 5'-GGG-3', with a spacer sequence of 5'-ATCCATGGTGGCGGCGGTTA-3' (SEQ ID NO: 49). The 5'-ATCCATGGTGGCGGCGGTTA-3' (SEQ ID NO: 49) sequence was synthesized and ligated into the px330-U6-Chimeric vector to construct the BE sgRNA expression plasmid. The fifth base "A" distal to the PAM is the target editing site for the constructed adenine base editor and falls within the editing window of NG-ABE9e (bases 5 to 8 distal to the PAM have higher editing efficiency). The fourth base "C" distal to the PAM is the target editing site for the constructed cytosine base editor and falls within the editing window of BE4 max (bases 4 to 8 distal to the PAM have higher editing efficiency).
[0201] The four SpCas9 insertion / deletion mutants disclosed herein, SpCas9 V922Δ, SpCas9 E923Δ, SpCas9L921P / V922Δ, and SpCas9 V922Δ / E923T, were introduced into NG-ABE9e (ABE) and BE4 Max (CBE) vectors. The vector designs are shown in Figure 6 B and 6C. ngSpCas9 (Addgene Plasmid #117919) can recognize NG PAM. Using the ngSpCas9 plasmid as a template, four ngSpCas9 insertion / deletion mutants, ngSpCas9 V922Δ, ngSpCas9 E923Δ, SpCas9 L921P / V922Δ, and SpCas9 V922Δ / E923T, were constructed by referring to the method of constructing four SpCas9 insertion / deletion mutants, SpCas9 V922Δ, ngSpCas9 E923Δ, ngSpCas9 L921P / V922Δ, and ngSpCas9 V922Δ / E923T, in Example 1 (only the template sequence was different, and the construction method was the same). The four ngSpCas9 insertion / deletion mutants, ngSpCas9 V922Δ, ngSpCas9 E923Δ, ngSpCas9 L921P / V922Δ, and ngSpCas9 V922Δ / E923T, were replaced with NG-ABE9e (see Tu Tianxiang et al., Molecular Biology, 2012). Therapy, 2022, 30(9): 2933–2941) Cas9 single nickase in the vector, construction Figure 6The corresponding adenine base editors in B. Four SpCas9 insertion / deletion mutants, SpCas9 V922△, SpCas9 E923△, SpCas9L921P / V922△, and SpCas9 V922△ / E923T, were used to replace the Cas9 single nickase in BE4max (Addgene Plasmid #112093) to construct Figure 6 C corresponding cytosine base editor.
[0202] The specific method for constructing the adenine base editor is as follows: using four ngSpCas9 insertion / deletion mutant plasmids as templates, the four ngSpCas9 insertion / deletion mutant sequences were obtained by PCR amplification with the forward primer 5'-GACAAGAAGTACAGCATCGG-3' (SEQ ID NO: 63) and the reverse primer 5'-GTCACCTCCCAGCTGAGACA-3' (SEQ ID NO: 64), and the NG-ABE9e vector fragment was amplified by PCR with the forward primer 5'-CAGCTGGGAGGTGACTCTGG-3' (SEQ ID NO: 65) and the reverse primer 5'-GCTGTACTTCTTGTCTGACCCC-3' (SEQ ID NO: 66). The four ngSpCas9 insertion / deletion mutant sequences were mixed with the NG-ABE9e vector fragment respectively, and homologous recombination was performed using a homologous recombination kit (Wuhan Aiboteike) to construct Figure 6 B Corresponding adenine base editor vector.
[0203] The specific method for constructing the cytosine base editor is as follows: using the SpCas9 insertion / deletion mutant plasmid constructed in Example 1 as a template, four SpCas9 insertion / deletion mutant sequences were obtained by PCR amplification with the forward primer 5'-GACAAGAAGTACAGCATCGG-3' (SEQ ID NO: 63) and the reverse primer 5'-GTCACCTCCCAGCTGAGACA-3' (SEQ ID NO: 64), and the BE4max vector fragment was amplified by PCR using the BE4max plasmid as a template by the forward primer 5'-CAGCTGGGAGGTGACAGCGG-3' (SEQ ID NO: 67) and the reverse primer 5'-GCTGTACTTCTTGTCGCTGCC-3' (SEQ ID NO: 68), and the four SpCas9 insertion / deletion mutant sequences were mixed with the BE4max vector fragment respectively, and homologous recombination was performed using the homologous recombination kit (Wuhan Aiboteike) to construct Figure 6 C corresponding cytosine base editor vector.
[0204] To verify the editing efficiency of base editors based on four SpCas9 insertion / deletion mutants, the above-constructed BE sgRNA plasmid and four expression plasmids of base editors based on four SpCas9 insertion / deletion mutants were transfected into mouse embryonic stem cells with an NHEJ reporter system using Lipofectamine 2000 liposome transfection method. The transfection system is (taking a 24-well plate as an example): 2×10 cells per well in 24 wells 5 The total amount of transfected DNA was 0.5 μg, including 0.3 μg of base editor expression plasmid and 0.2 μg of BE sgRNA plasmid. To exclude the influence of NHEJ repair products induced by the nickase itself, we also co-transfected cells with expression plasmids of four indel mutants without deaminase and BE sgRNA as negative controls.
[0205] 72 h after transfection, GFP expression was analyzed by flow cytometry. + The results showed that the base editors based on the four SpCas9 insertion / deletion mutants all had high base editing efficiency, and only SpCas9 L921P / V922△ had a lower base editing efficiency after being introduced into the base editor ( Figure 6 B and Figure 6 C).
[0206] In order to test the editing efficiency of adenine base editors based on four ngSpCas9 and SpCas9 insertion / deletion mutants at endogenous sites, six endogenous sites of adenine base editors ABE site 1 to 6 and six cytosine base editors CBE site 1 to 6 were designed in the human genome. The corresponding sgRNA expression vectors were designed and constructed (Table 4), and co-transfected with the corresponding base editor expression vectors into HEK293T cells. The transfection system was (taking 24-well plates as an example): 2 × 10 cells per well in 24 wells. 5 The total amount of transfected DNA was 0.5 μg, of which the base editor expression plasmid was 0.3 μg and the sgRNA plasmid was 0.2 μg. The genome was extracted 72 hours after transfection, and the PCR target was amplified and the PCR products were deeply sequenced to analyze the base editing efficiency of the adenine base editor and cytosine base editor constructed based on the four insertion and deletion mutants of the present invention. The results showed that the adenine base editor based on the four ngSpCas9 insertion and deletion mutants and the cytosine base editor based on the four SpCas9 insertion and deletion mutants had higher base editing efficiency within the editing window. Only ngSpCas9 L921P / V922△ and SpCas9 L921P / V922△ had lower base editing efficiency after being introduced into the base editor ( Figure 7 、 8 ).
[0207] Table 4. sgRNA designs for testing base editing efficiency at endogenous sites
[0208] sgRNA name PAM spacer sequence Sequence number ABE site 1 TGG GGAATCCCTTCTGCAGCACC SEQ ID NO: 69 ABE site 2 AGA GAGCAAATACCAGAGATAAG SEQ ID NO: 70 ABE site 3 AGC CCCCCCACAGAATATCAAGG SEQ ID NO: 71 ABE site 4 GGT GCTCTGTCAGAAGATCTCCA SEQ ID NO: 72 ABE site 5 GGC GATCAGGAAATAGAGCCACA SEQ ID NO: 73 ABE site 6 TGA GCCACCACAGGGAAGCTGGG SEQ ID NO: 74 CBE site 1 CGG GTCAAGAAAGCAGAGACTGC SEQ ID NO: 75 CBE site 2 GGG GAGTCCGAGCAGAAGAAGAA SEQ ID NO: 76 CBE site 3 TGG GGAATCCCTTCTGCAGCACC SEQ ID NO: 77 CBE site 4 GGG GAACACAAAGCATAGACTGC SEQ ID NO: 78 CBE site 5 TGG GGCCCAGACTGAGCACGTGA SEQ ID NO: 79 CBE site 6 TGG GGCACTGCGGCTGGAGGTGG SEQ ID NO: 80
[0209] Example 6. Gene Editing Efficiency after Pairing Cas9 Protein Mutant Enzymes with TREX2
[0210] The mutant NHEJ reporter system of mouse embryonic stem cells was used to detect the gene editing efficiency after the paired single nickase fused to TREX2. The mutant NHEJ reporter system inserts the I-SceI sequence between "Koz-ATG" and GFP in Example 1, providing more target selection. Similar to Example 1, under normal circumstances, the translation of "Koz-ATG" produces frameshift GFP, and the cells do not emit light. When the sgRNA with a 3' end is designed before and after "Koz-ATG" to cut it, there are two situations in which the cells can express normal GFP: (1) When DSBs are generated by cutting between the two start codons ATG, NHEJ repair will connect the breaks together and introduce non-precise sequence insertions and deletions at the repaired cuts, causing a certain probability of frameshift mutations in the sequence, resulting in the expression of the GFP gene in full code. (2) The first Kozak-ATG is destroyed, and the cell uses the second ATG start codon to express GFP ( Figure 9 A).
[0211] Four SpCas9 insertion / deletion mutations were designed and constructed before and after "Koz-ATG" to produce paired sgRNAs (gWL2 and gCR6) at the 3' end after cleavage. The gWL2 PAM was on the Watson strand and the gCR6 PAM was on the Crick strand. The spacer sequences are shown in Table 5. They were synthesized and constructed into the px330-U6-Chimeric vector.
[0212] Construct expression plasmids containing TX C-terminally fused to SpCas9 and the four SpCas9 insertion / deletion mutants disclosed herein: SpCas9 V922Δ, SpCas9 E923Δ, SpCas9 L921P / V922Δ, and SpCas9 V922Δ / E923T. TX is a triple mutant of TREX2 (R163A / R165A / R167A). These three mutations disrupt TREX2's DNA binding ability, preventing off-target effects caused by nonspecific DNA binding. The vector backbone is pcDNA3.1.
[0213] SpCas9 and the four SpCas9 insertion / deletion mutants disclosed herein, SpCas9 V922Δ, SpCas9 E923Δ, SpCas9 L921P / V922Δ, and SpCas9 V922Δ / E923T C-terminal fusion TX expression plasmids, were co-transfected into mouse embryonic stem cells with a single sgRNA (gWL2 or gCR6) or two sgRNAs (gWL2 and gCR6). The transfection system was (taking a 24-well plate as an example): 2×10 cells per well in 24 wells. 5 The total amount of transfected DNA was 0.5 μg. When co-transfecting a single sgRNA, the ratio of Cas enzyme-TX: gRNA: empty U6 was 2:1:1; when co-transfecting two sgRNAs, the ratio of Cas enzyme-TX: gRNA1: gRNA2 was 2:1:1. GFP was detected by flow cytometry 72 hours after transfection. + Cell ratio.
[0214] The results showed that fusing the four SpCas9 insertion / deletion mutants SpCas9 V922△, SpCas9 E923△, SpCas9 L921P / V922△, and SpCas9 V922△ / E923T disclosed herein with the 3'→5' exonuclease TREX2 mutant TX can greatly improve the gene editing efficiency of the 3' end produced by paired single nickase cleavage ( Figure 9 B) Because the designed sgRNAs gWL2 and gCR6 only produce 3' ends when double-nicked in combination with the SpCas9 single nickase targeting the sgRNA hybrid strand, the four SpCas9 insertion / deletion mutants disclosed herein can process the 3' ends after fusion with TX, thereby improving gene editing efficiency. This further demonstrates the characteristic of the four SpCas9 insertion / deletion mutants disclosed herein, SpCas9 V922Δ, SpCas9 E923Δ, SpCas9 L921P / V922Δ, and SpCas9 V922Δ / E923T, that they only cleave the sgRNA targeting strand in vivo and do not cleave the sgRNA non-targeting strand.
[0215] Table 5. sgRNA designs for testing gene editing efficiency of paired Cas9 protein mutant enzymes fused to TREX2
[0216] sgRNA name PAM spacer sequence Sequence number wxya TGG TTCCCTAACCGCCGCCACCA SEQ ID NO: 81 gCR6 TGG CTTGCTGATCATGTGAAGGA SEQ ID NO: 82
Claims
1. A Cas9 protein mutant comprising a mutation at amino acid residue 921, 922, and / or 923 of the amino acid sequence of SEQ ID NO: 1, wherein the mutation is selected from one or more of the following: (a) Deletion of amino acid residues; (b) substitution of amino acid residues; (c) insertion of amino acid residues; Furthermore, the mutation can significantly reduce or eliminate the cleavage activity of the Cas9 protein on non-target chains.
2. The Cas9 protein mutant according to claim 1, wherein The mutations include: L921P; V922Δ; and / or E923Δ or E923T; Preferably, the Cas9 protein comprises a mutation or combination of mutations selected from the following: (1) V922Δ; (2) E923Δ; (3) L921P and V922Δ; (4) V922Δ and E923T; (5)L921P; (6)E923T; (7) L921P and E923T; (8) L921P and E923Δ; (9) V922Δ and E923Δ; (10) L921P, V922Δ, and E923T; and (11) L921P, V922Δ, and E923Δ; More preferably, the Cas9 protein comprises a mutation or combination of mutations selected from the group consisting of: (1) V922Δ; (2) E923Δ; (3) L921P and V922Δ; and (4) V922Δ and E923T; Preferably, the Cas9 protein comprises RuvC and HNH domains, and the HNH domain comprises a mutation that reduces or eliminates target chain cleavage activity; More preferably, the mutation that reduces or eliminates target chain cleavage activity comprises H840A.
3. A Cas9 protein mutant comprising an amino acid sequence fragment as shown in SEQ ID NO: 35, 37, 39, or 41, wherein the mutant significantly reduces or eliminates the cleavage activity of the Cas9 protein on non-target strands; Preferably, the Cas9 protein mutant comprises RuvC and HNH domains, and the HNH domain comprises a mutation that reduces or eliminates the target chain cleavage activity; More preferably, the mutation that reduces or eliminates target chain cleavage activity comprises H840A.
4. A fusion protein comprising the Cas9 protein mutant according to any one of claims 1 to 3, and a heterologous domain fused to the Cas9 protein mutant; Preferably, the heterologous domain is fused to the N-terminus, C-terminus, internal position or a combination thereof of the Cas9 protein mutant; Preferably, the Cas9 protein mutant is covalently linked to the heterologous domain via a linker; Preferably, the heterologous domain is selected from the group consisting of a nuclear localization signal domain, a cell penetrating domain, a marker or reporter domain to facilitate detection, a DNA or RNA deaminase domain, a uracil-DNA-glycosylase domain, a reverse transcriptase domain, a recombinase domain, an RNA aptamer binding domain, a nuclease domain, a methyltransferase domain, a methylase domain, an acetylase domain, an acetyltransferase domain, a transcriptional activator domain, and a transcriptional repressor domain; More preferably, the heterologous domain is selected from the group consisting of an exonuclease domain and a deaminase domain; Further preferably, the heterologous domain is selected from cytosine deaminase, adenine deaminase and DNA binding deficient 3' to 5' exonuclease.
5. A complex comprising the Cas9 protein mutant according to any one of claims 1 to 3 or the fusion protein according to claim 4, and a guide RNA; Preferably, the guide RNA is about 15-100 nucleotides in length and comprises a sequence of at least 10 consecutive nucleotides that is complementary to the target sequence; More preferably, the 3' end of the target sequence is not immediately adjacent to a canonical PAM sequence.
6. A biomaterial comprising any one of the following (a) to (d): (a) a polynucleotide encoding the Cas9 protein mutant according to any one of claims 1 to 3 or the fusion protein according to claim 4; (b) a vector comprising the polynucleotide described in (a); preferably, the polynucleotide is operably linked to a promoter, an enhancer and / or a terminator; (c) a kit comprising the Cas9 protein mutant according to any one of claims 1 to 3, the fusion protein according to claim 4, the complex according to claim 5, the polynucleotide according to (a), or the vector according to (b); Preferably, the kit further comprises an expression construct encoding a guide RNA backbone; (d) a cell comprising the Cas9 protein mutant according to any one of claims 1 to 3, the fusion protein according to claim 4, the complex according to claim 5, the polynucleotide according to (a), or the vector according to (b); Preferably, the cell is selected from a human cell, a non-human mammalian cell, a plant cell, a non-mammalian vertebrate cell, an invertebrate cell, a unicellular eukaryotic cell or a prokaryotic cell.
7. A pharmaceutical composition comprising the Cas9 protein mutant according to any one of claims 1 to 3, the fusion protein according to claim 4, the complex according to claim 5, (a) the polynucleotide in the biomaterial according to claim 6, (b) the vector in the biomaterial according to claim 6, or (d) the cell in the biomaterial according to claim (6).
8. Use of the Cas9 protein mutant according to any one of claims 1 to 3, the fusion protein according to claim 4, the complex according to claim 5, the (a) polynucleotide in the biomaterial according to claim 6, the (b) vector in the biomaterial according to claim 6, the (c) kit in the biomaterial according to claim 6, the (d) cell in the biomaterial according to claim 6, or the pharmaceutical composition according to claim 7 in nucleic acid modification.
9. A method for modifying a target nucleic acid, the method comprising contacting the target nucleic acid with: (a) the Cas9 protein mutant according to any one of claims 1 to 3 or the fusion protein according to claim 4, and a guide RNA; or (b) The complex according to claim 5.
10. A method for modifying a target nucleic acid, the method comprising introducing the Cas9 protein mutant according to any one of claims 1 to 3, the fusion protein according to claim 4, the complex according to claim 5, (a) the polynucleotide in the biomaterial according to claim 6, (b) the vector in the biomaterial according to claim 6, and / or the pharmaceutical composition according to claim 7 into a cell or organism comprising the target nucleic acid.
Citation Information
Patent Citations
S. pyogenes cas9 mutant genes and polypeptides encoded by same
CN110462034A
C-NHEJ fixed-point suppression system based on dCas9 and application
CN116286975A
Method for improving gene editing efficiency of paired Cas9 single nickase
CN116970599A
Improved high throughput combined gene modification system and optimized Cas9 enzyme variants
CN118256471A
Engineered CRISPR-Cas9 nuclease with altered PAM specificity
CN118530967A
Cited By
A bovine cell gene editing reagent and a method for constructing a bco2 gene mutant cell strain and application
CN122445674A