Systems and methods for genome editing
The CRISPR-Cpf1 system, utilizing a Cpf1 protein with high sequence identity and a guide RNA, addresses the need for efficient eukaryotic genome editing by enabling site-specific modifications, thereby enhancing genetic manipulation capabilities.
Patent Information
- Application Number
- JP2022064162
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-04-10
- Filing Date
- 2022-04-07
- Publication Date
- 2025-06-30
- Estimated Expiration
- 2038-04-10
AI Technical Summary
There is a need for an efficient CRISPR-Cpf1 system capable of effectively editing the genome of eukaryotic cells, which has not been fully addressed by existing technologies.
The use of a Cpf1 protein with an amino acid sequence having at least 80% sequence identity to specific sequences, combined with a guide RNA, to form a genome editing system that can target and modify specific sequences in the genome of eukaryotic cells.
The system enables site-specific modification of target sequences in eukaryotic genomes, allowing for substitution, deletion, and addition of nucleotides, thereby facilitating efficient genome editing and potential therapeutic applications.
Smart Images

Figure 0007700079000023 
Figure 0007700079000024 
Figure 0007700079000025
Abstract
Description
Technical Field
[0001] The present invention relates to the field of genetic engineering. In particular, the present invention relates to novel eukaryotic genome editing systems and methods. More specifically, the present invention relates to the CRISPR-Cpf1 system capable of efficiently editing the genome of eukaryotic cells and its use.
Background Art
[0002] The CRISPR (Clustered regular interspaced short palindromic repeats) system is an immune system generated during the evolution of bacteria to defend against the invasion of foreign genes. Among them, the type II CRISPR-Cas9 system is a system for DNA cleavage by the Cas9 protein mediated by two small molecule RNAs (crRNA and tracrRNA) or artificially synthesized small molecule RNA (sgRNA), and is the simplest system among the first three (type I, II, III) CRISPR systems discovered. Because of its easy operation, this system was successfully operated in 2013 and successfully achieved eukaryotic genome editing. The CRISPR / Cas9 system has quickly become the most popular technology in life science.
[0003] In 2015, Zhang et al. discovered the CRISPR-Cpf1 system, a new gene editing system different from the CRISPR-Cas9 system, by sequence alignment and tissue analysis. This system requires only one small molecule RNA (crRNA) to mediate genome editing. Although there are thousands of CRISPR systems in nature, only a few systems can successfully perform eukaryotic genome editing.
Summary of the Invention
Problems to be Solved by the Invention
[0004] There is still a need in the art for a CRISPR-Cpf1 system that enables efficient eukaryotic genome editing. **Means for Solving the Problems**
[0005] **Summary of the Invention** In one aspect, the present invention provides the following i) to v): i) a Cpf1 protein and a guide RNA; ii) an expression construct comprising a nucleotide sequence encoding a Cpf1 protein and a guide RNA; iii) a Cpf1 protein and an expression construct comprising a nucleotide sequence encoding a guide RNA; iv) an expression construct comprising a nucleotide sequence encoding a Cpf1 protein and an expression construct comprising a nucleotide sequence encoding a guide RNA; v) an expression construct comprising a nucleotide sequence encoding a Cpf1 protein and a nucleotide sequence encoding a guide RNA A genome editing system for site-specific modification of a target sequence in the genome of a cell, comprising at least one of the above, wherein the Cpf1 protein comprises an amino acid sequence having at least 80% sequence identity to the amino acid sequence of SEQ ID NOs: 1 to 12 or one of SEQ ID NOs: 1 to 12, and the guide RNA can target the Cpf1 protein to a target sequence in the genome of the cell.
[0006] In a second aspect, the present invention provides a method for modifying a target sequence in the genome of a cell, comprising introducing the genome editing system of the present invention into the cell, whereby the guide RNA targets the Cpf1 protein to a target sequence in the genome of the cell, resulting in substitution, deletion and / or addition of one or more nucleotides in the target sequence.
[0007] In a third aspect, the present invention provides a method for treating a disease in a subject in need thereof, the method comprising delivering to the subject an effective amount of the genome editing system of the present invention to modify a gene associated with the disease.
[0008] In a fourth aspect, the present invention provides the use of the genome editing system of the present invention for the manufacture of a pharmaceutical composition for treating a disease in a subject in need thereof, wherein the genome editing system is for modifying a gene associated with the disease.
[0009] In a fifth aspect, the present invention provides a pharmaceutical composition for treating a disease in a subject in need thereof, the pharmaceutical composition comprising the genome editing system of the present invention and a pharmaceutically acceptable carrier, wherein the genome editing system is for modifying a gene associated with the disease.
[0010] In a sixth aspect, the present invention provides a crRNA comprising a crRNA scaffold sequence corresponding to any one of SEQ ID NOs: 25 to 33, or a coding sequence thereof comprising a sequence shown in any one of SEQ ID NOs: 25 to 33. BRIEF DESCRIPTION OF THE DRAWINGS
[0011]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14A - B
Figure 14C
Figure 15
Mode for Carrying Out the Invention
[0012] Detailed Description of the Invention 1. Definitions In the present invention, scientific and technical terms used herein have the meanings generally understood by those skilled in the art, unless otherwise specified. Also, protein and nucleic acid chemistry, molecular biology, cell and tissue culture, microbiology, immunology-related terms, and laboratory procedures used herein are terms widely used in the corresponding fields and routine procedures. For example, standard recombinant DNA and molecular cloning techniques used in the present invention are well known to those skilled in the art and are fully described in the following document: Sambrook, J., Fritsch, E.F. and Maniatis, T., Molecular Cloning: A Laboratory Manual; Cold Spring Harbor Laboratory Press: Cold Spring Harbor, 1989 (hereinafter referred to as "Sambrook" in this specification). On the other hand, for a better understanding of the present invention, definitions and explanations of related terms are provided below.
[0013] "Cpf1 nuclease", "Cpf1 protein" and "Cpf1" are used interchangeably herein and refer to an RNA-guided nuclease comprising a Cpf1 protein or a fragment thereof. Cpf1 is a component of the CRISPR-Cpf1 genome editing system, which targets and cleaves a DNA target sequence under the guidance of a guide RNA (crRNA) to form a DNA double-strand break (DSB). The DSB can activate non-homologous end joining (NHEJ) and homologous recombination (HR), which are native repair mechanisms in living cells, to repair DNA damage in the cell, during which site-specific editing of a specific DNA sequence is achieved.
[0014] "Guide RNA" and "gRNA" can be used interchangeably herein. The guide RNA of the CRISPR-Cpf1 genome editing system typically consists of only a crRNA molecule, which contains a sequence that is identical to the target sequence sufficiently to hybridize to the complement of the target sequence and specifically bind the complex (Cpf1+crRNA) to the target sequence.
[0015] "Genome", as used herein, includes not only chromosomal DNA present in the nucleus, but also organellar DNA present in intracellular components of the cell (e.g., mitochondria, plastids).
[0016] As used herein, "organism" includes any organism suitable for genome editing, with eukaryotes being preferred. Examples of organisms include, but are not limited to, mammals such as humans, mice, rats, monkeys, dogs, pigs, sheep, cows, cats; poultry such as chickens, ducks, geese; plants including monocotyledonous and dicotyledonous plants such as rice, corn, wheat, sorghum, barley, soybeans, peanuts, Arabidopsis thaliana, etc.
[0017] "Genetically modified organisms" or "genetically modified cells" include organisms or cells that contain exogenous polynucleotides or modified genes or expression regulatory sequences within their genomes. For example, the exogenous polynucleotide is stably integrated into the genome of the organism or cell so that it is inherited by subsequent generations that follow the polynucleotide. The exogenous polynucleotide may be integrated into the genome alone or as part of a recombinant DNA construct. A modified gene or expression regulatory sequence means that in the organism genome or cell genome, the sequence contains substitutions, deletions, or additions of one or more nucleotides.
[0018] With respect to a sequence, "exogenous" means a sequence derived from a different species or, when derived from the same species, a sequence that has undergone a significant change in composition and / or locus by intentional human intervention from its natural form.
[0019] "Polynucleotide", "nucleic acid sequence", "nucleotide sequence", or "nucleic acid fragment" are used interchangeably and are single-stranded or double-stranded RNA or DNA polymers that optionally contain synthetic, non-natural, or altered nucleotide bases. Nucleotides are referred to by their single-letter names as follows: "A" is adenosine or deoxyadenosine (corresponding to RNA or DNA respectively), "C" means cytidine or deoxycytidine, "G" means guanosine or deoxyguanosine, "U" represents uridine, "T" means deoxythymidine, "R" means purine (A or G), "Y" means pyrimidine (C or T), "K" means G or T, "H" means A or C or T, "I" means inosine, and "N" means any nucleotide.
[0020] The terms "polypeptide", "peptide", and "protein" are used interchangeably in the present invention to refer to polymers of amino acid residues. These terms apply to amino acid polymers in which one or more amino acid residues are artificial chemical analogs of the corresponding naturally occurring amino acids, as well as to naturally occurring amino acid polymers. The terms "polypeptide", "peptide", "amino acid sequence", and "protein" may also include modified forms including, but not limited to, glycosylation, lipid linkage, sulfation, γ-carboxylation of glutamic acid residues, and ADP-ribosylation.
[0021] The term "identity" of sequences has the meaning recognized in the art, and the percentage of sequence identity between two nucleic acid or polypeptide molecules or regions can be calculated using the disclosed techniques. Sequence identity can be measured along the full length of a polynucleotide or polypeptide, or along a region of the molecule (see, for example, Computational Molecular Biology, Lesk, A.M., ed., Oxford University Press, New York, 1988; Biocomputing: Informatics and Genome Projects, Smith, D.W., ed., Academic Press, New York, 1993; Computer Analysis of Sequence Data, Part I, Griffin, A.M., and Griffin, H.G., eds., Humana Press, New Jersey, 1994; Sequence Analysis in Molecular Biology, von Heinje, G., Academic Press, 1987; and Sequence Analysis Primer, Gribskov, M. and Devereux, J., eds., M Stockton Press, New York, 1991). There are many methods for measuring identity between two polynucleotides or polypeptides, but the term "identity" is well known to those of skill in the art (Carrillo, H. & Lipman, D., SIAM J Applied Math 48: 1073 (1988)).
[0022] Suitable conservative amino acid substitutions in a peptide or protein are known to those skilled in the art and generally can be made without altering the biological activity of the resulting molecule. In general, those skilled in the art recognize that a single amino acid substitution in a non-essential region of a polypeptide does not substantially alter biological activity (see, for example, Watson et al., Molecular Biology of the Gene, 4th Edition, 1987, The Benjamin / Cummings Pub. co., p.224).
[0023] As used herein, an "expression construct" refers to a vector such as a recombinant vector suitable for the expression of a nucleotide sequence of interest in an organism. "Expression" refers to the production of a functional product. For example, expression of a nucleotide sequence can refer to transcription of the nucleotide sequence (e.g., transcription to produce mRNA or functional RNA), and / or translation of the RNA into a precursor or mature protein.
[0024] The "expression construct" of the present invention can be a linear nucleic acid fragment, a circular plasmid, a viral vector, or in some embodiments, a translatable RNA (e.g., mRNA).
[0025] The "expression construct" of the present invention can include regulatory sequences and nucleotide sequences of interest derived from different origins, or regulatory sequences and nucleotide sequences of interest derived from the same origin but arranged in a manner different from that which normally occurs in nature.
[0026] "Regulatory sequence" and "regulatory element" are used interchangeably to refer to a nucleotide sequence located upstream (5' non-coding sequence), in the middle, or downstream (3' non-coding sequence) of a coding sequence and that affects the transcription, RNA processing or stability, or translation of the associated coding sequence. Regulatory sequences include, but are not limited to, promoters, translation leaders, introns, and polyadenylation recognition sequences.
[0027] "Promoter" refers to a nucleic acid fragment capable of controlling the transcription of another nucleic acid fragment. In some embodiments of the present invention, the promoter is a promoter capable of controlling the transcription of a gene in a cell, regardless of whether it is derived from that cell. The promoter may be a constitutive promoter or a tissue-specific promoter or a developmentally regulated promoter or an inducible promoter.
[0028] "Constitutive promoter" refers to a promoter that can generally express a gene in most cell types in most cases. "Tissue-specific promoter" and "tissue-preferential promoter" are used interchangeably, meaning that they are expressed mainly (but not necessarily only) in one tissue or organ and also in specific cells or cell types. "Developmentally regulated promoter" refers to a promoter whose activity is determined by developmental events. "Inducible promoter" selectively expresses an operably linked DNA sequence in response to endogenous or exogenous stimuli (such as environment, hormone, chemical signal, etc.).
[0029] As used herein, the term "operably linked" refers to the linkage of a regulatory element (such as, but not limited to, a promoter sequence, a transcription termination sequence, etc.) to a nucleic acid sequence (such as a coding sequence or an open reading frame) such that the transcription of the nucleotide sequence is controlled and regulated by the transcription regulatory element. Techniques for operably linking a regulatory element region to a nucleic acid molecule are known in the art.
[0030] "Introduction" of a nucleic acid molecule (such as a plasmid, a linear nucleic acid fragment, RNA, etc.) or a protein into an organism means transforming a biological cell using the nucleic acid or the protein so that the nucleic acid or the protein can function in the cell. As used in the present invention, "transformation" includes both stable transformation and transient transformation.
[0031] "Stable transformation" refers to the introduction of an exogenous nucleotide sequence into the genome that results in stable inheritance of the foreign gene. When stably transformed, the exogenous nucleic acid sequence is stably integrated into the genome of both the organism and all subsequent generations.
[0032] "Transient transformation" refers to the introduction of a nucleic acid molecule or protein into a cell that functions without stable inheritance of the exogenous gene. In transient transformation, the exogenous nucleic acid sequence is not integrated into the genome.
[0033] 2. Efficient Genome Editing System In one aspect, the present invention provides the use of a Cpf1 protein comprising an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity to one of SEQ ID NOs: 1-12 in eukaryotic genome editing.
[0034] In another aspect, the present invention provides the following i) - v): i) A Cpf1 protein and a guide RNA; ii) An expression construct comprising a nucleotide sequence encoding a Cpf1 protein and a guide RNA; iii) A Cpf1 protein and an expression construct comprising a nucleotide sequence encoding a guide RNA; iv) An expression construct comprising a nucleotide sequence encoding a Cpf1 protein and an expression construct comprising a nucleotide sequence encoding a guide RNA; v) An expression construct comprising a nucleotide sequence encoding a Cpf1 protein and a nucleotide sequence encoding a guide RNA of at least one of which is a genome editing system for site-specific modification of a target sequence in the genome of a cell, The Cpf1 protein comprises an amino acid sequence having at least 80% sequence identity to the amino acid sequences of SEQ ID NOs: 1 to 12 or to one of SEQ ID NOs: 1 to 12, and the guide RNA can target the Cpf1 protein to a target sequence in the genome of a cell, providing a system.
[0035] In some embodiments of the method of the present invention, the guide RNA is a crRNA. In some embodiments, the coding sequence of the crRNA comprises a crRNA scaffold sequence shown in any one of SEQ ID NOs: 25 to 33. In some preferred embodiments, the crRNA scaffold sequence is SEQ ID NO: 30. In some embodiments, the coding sequence of the cRNA further comprises a sequence (i.e., a spacer sequence) that specifically hybridizes to the complement of the target sequence 3' to the crRNA scaffold sequence.
[0036] In some embodiments, the crRNA is as follows: i) 5'-ATTTCTACtgttGTAGAT (SEQ ID NO: 25)-N x -3'; ii) 5'-ATTTCTACtattGTAGAT (SEQ ID NO: 26)-N x -3'; iii) 5'-ATTTCTACtactGTAGAT (SEQ ID NO: 27)-N x -3'; iv) 5'-ATTTCTACtttgGTAGAT (SEQ ID NO: 28)-N x -3'; v) 5'-ATTTCTACtagttGTAGAT (SEQ ID NO: 29)-N x -3'; vi) 5'-ATTTCTACTATGGTAGAT (SEQ ID NO: 30)-N x -3'; vii) 5'-ATTTCTACTGTCGTAGAT (SEQ ID NO: 31)-N x -3'; viii) 5'-ATTTCTACTTGTGTAGAT (SEQ ID NO: 32)-N x -3'; and ix) 5'-ATTTCTACTGTGGTAGAT (SEQ ID NO: 33)-N x -3' (N x represents a nucleotide sequence consisting of x consecutive nucleotides, where N is independently selected from A, G, C, and T; x is an integer from 18 ≦ x ≦ 35, preferably x = 23) and is encoded by a nucleotide sequence selected from the group consisting of. In some embodiments, the sequence N x (spacer sequence) can specifically hybridize to the complement of the target sequence.
[0037] In some embodiments of the method of the present invention, the cell is a eukaryotic cell, preferably a mammalian cell.
[0038] In some embodiments of the method of the present invention, the Cpf1 protein comprises an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity to one of SEQ ID NOs: 1-12. The Cpf1 protein can target and / or cleave a target sequence in the genome of the cell by crRNA.
[0039] In some embodiments of the method of the present invention, the Cpf1 protein comprises an amino acid sequence having one or more amino acid residue substitutions, deletions, or additions relative to one of SEQ ID NOs: 1-12. For example, the Cpf1 protein comprises an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 amino acid residue substitutions, deletions, or additions relative to one of SEQ ID NOs: 1-12. In some embodiments, the amino acid substitution is a conservative substitution. The Cpf1 protein can target and / or cleave a target sequence in the cell genome by crRNA.
[0040] The Cpf1 protein of the present invention may be derived from a species selected from Agathobacter rectalis, Lachnospira pectinoschiza, Sneathia amnii, Helcococcus kunzii, Arcobacter butzleri, Bacteroidetes oral, Oribacterium sp., Butyrivibrio sp., Proteocatella sphenisci, Candidatus Dojkabacteria, Pseudobutyrivibrio xylanivorans, Pseudobutyrivibrio ruminis.
[0041] In some preferred embodiments of the present invention, the Cpf1 protein is derived from Agathobacter rectalis (ArCpf1), Butyrivibrio sp. (BsCpf1), Helcococcus kunzii (HkCpf1), Lachnospira pectinoschiza (LpCpf1), Pseudobutyrivibrio ruminis (PrCpf1) or Pseudobutyrivibrio xylanivorans (PxCpf1). In some preferred embodiments of the present invention, the Cpf1 protein comprises an amino acid sequence selected from SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 4, SEQ ID NO: 8, SEQ ID NO: 11, SEQ ID NO: 12.
[0042] In some embodiments of the present invention, the Cpf1 protein of the present invention further comprises a nuclear localization sequence (NLS). Generally, one or more NLSs in the Cpf1 protein should have sufficient strength to accumulate the Cpf1 protein in the cell nucleus to an amount that achieves genome editing. Generally, the strength of the nuclear localization activity is determined by the number, position, one or more specific NLSs used in the Cpf1 protein, or a combination of these factors.
[0043] In some embodiments of the present invention, the NLS of the Cpf1 protein of the present invention may be located at the N-terminus and / or the C-terminus. In some embodiments, the Cpf1 protein comprises about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLSs. In some embodiments, the Cpf1 protein comprises about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLSs at the N-terminus or near it. In some embodiments, the Cpf1 protein comprises about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLSs at the C-terminus or near it. In some embodiments, the Cpf1 protein comprises combinations thereof, for example, one or more NLSs at the N-terminus and one or more NLSs at the C-terminus. If there are two or more NLSs, each can be selected independently of the other NLSs. In some preferred embodiments of the present invention, the Cpf1 protein comprises two NLSs. For example, the two NLSs are located at the N-terminus and the C-terminus, respectively.
[0044] Generally, an NLS consists of one or more short sequences of positively charged lysine or arginine exposed on the surface of the protein, although other types of NLSs are also known. Non-limiting examples of NLSs include KKRKV (nucleotide sequence 5'-AAGAAGAGAAAGGTC-3'), PKKKRKV (nucleotide sequence 5'-CCCAAGAAGAAGAGGAAGGTG-3' or CCAAAGAAGAAGAGGAAGGTT), or SGGSPKKKRKV (nucleotide sequence 5'-TCGGGGGGGAGCCCAAAGAAGAAGCGGAAGGTG-3').
[0045] Furthermore, depending on the position of the DNA to be edited, the Cpf1 protein of the present invention may also include other localization sequences, such as cytoplasmic localization sequences, chloroplast localization sequences, mitochondrial localization sequences, and the like.
[0046] In some embodiments of the present invention, to obtain efficient expression in target cells, the nucleotide sequence encoding the Cpf1 protein is codon-optimized for the organism from which the cells to be genome-edited are derived.
[0047] Codon optimization refers to modifying the nucleic acid sequence while maintaining the native amino acid sequence to replace at least one codon of the native sequence (e.g., about 1, 2, 3, 4, 5, 10, 15, 20, 25, 50 codons, or about 1, 2, 3, 4, 5, 10, 15, 20, 25, 50 codons or more than those, or codons exceeding them) with codons that are more frequently or most frequently used in the genes of the host cell. Different species show specific preferences for specific codons of a particular amino acid. Codon selection (differences in codon usage frequency among organisms) is often related to the translation efficiency of messenger RNA (mRNA), which is thought to depend on the nature of the codons being translated and the availability of specific transfer RNA (tRNA) molecules. The advantage of the selected tRNA in the cell generally reflects the codons most frequently used for peptide synthesis. Thus, genes can be customized to be the genes that are most highly expressed in a given organism based on codon optimization. Codon usage frequency tables can be easily obtained, for example, from the Codon Usage Database available at www.kazusa.orjp / codon / , and these tables can be adjusted in various ways. See Nakamura Y. et. al "Codon usage tabulated from the international DNA sequence databases: status for the year 2000 Nucl.Acids Res, 28: 292 (2000).
[0048] The organism from which the cells that can be edited by the method of the present invention are derived is preferably a eukaryote, such as, but not limited to, mammals, such as humans, mice, rats, monkeys, dogs, pigs, sheep, cows, cats; poultry, such as chickens, ducks, geese; plants including monocotyledonous and dicotyledonous plants, such as rice, corn, wheat, sorghum, barley, soybeans, peanuts, and Arabidopsis thaliana.
[0049] In some embodiments of the present invention, the nucleotide sequence encoding the Cpf1 protein is codon-optimized for humans. In some embodiments, the codon-optimized nucleotide sequence encoding the Cpf1 protein is selected from SEQ ID NOs: 13 to 24.
[0050] In some embodiments of the present invention, the nucleotide sequence encoding the Cpf1 protein and / or the nucleotide sequence encoding the guide RNA is operably linked to an expression regulatory element such as a promoter.
[0051] Examples of promoters that can be used in the present invention include, but are not limited to, polymerase (pol) I, pol II, or pol III promoters. Examples of pol I promoters include the chicken RNA pol I promoter. Examples of pol II promoters include, but are not limited to, the immediate early cytomegalovirus (CMV) promoter, the Rous sarcoma virus long terminal repeat (RSV-LTR) promoter, and the immediate early simian virus 40 (SV40) promoter. Examples of pol III promoters include the U6 and H1 promoters. Inducible promoters such as the metallothionein promoter can be used. Other examples of promoters include the T7 phage promoter, the T3 phage promoter, the β-galactosidase promoter, and the Sp6 phage promoter. Promoters that can be used in plants include, but are not limited to, the cauliflower mosaic virus 35S promoter, the maize Ubi-1 promoter, the wheat U6 promoter, the rice U3 promoter, the maize U3 promoter, and the rice actin promoter.
[0052] In another aspect, the present invention provides a crRNA comprising a crRNA scaffold sequence corresponding to any one of SEQ ID NOs: 25 to 33. In some embodiments, the coding sequence of the crRNA comprises a crRNA scaffold sequence shown in any one of SEQ ID NOs: 25 to 33. In some preferred embodiments, the crRNA scaffold sequence is SEQ ID NO: 30. In some embodiments, the coding sequence of the cRNA further comprises a sequence (i.e., a spacer sequence) that specifically hybridizes to the complement of the target sequence 3' to the crRNA scaffold sequence.
[0053] In some embodiments, the crRNA is as follows: i) 5'-ATTTCTACtgttGTAGAT (SEQ ID NO: 25)-N x -3'; ii) 5'-ATTTCTACtattGTAGAT (SEQ ID NO: 26)-N x -3'; iii) 5'-ATTTCTACtactGTAGAT (SEQ ID NO: 27)-N x -3'; iv) 5'-ATTTCTACtttgGTAGAT (SEQ ID NO: 28)-N x -3'; v) 5'-ATTTCTACtagttGTAGAT (SEQ ID NO: 29)-N x -3'; vi) 5'-ATTTCTACTATGGTAGAT (SEQ ID NO: 30)-N x -3'; vii) 5'-ATTTCTACTGTCGTAGAT (SEQ ID NO: 31)-N x -3'; viii) 5'-ATTTCTACTTGTGTAGAT (SEQ ID NO: 32)-N x -3'; and ix) 5'-ATTTCTACTGTGGTAGAT (SEQ ID NO: 33)-N x -3' (N x represents a nucleotide sequence consisting of x consecutive nucleotides, N is independently selected from A, G, C, and T; x is an integer of 18 ≦ x ≦ 35, preferably x = 23) is encoded by a nucleotide sequence selected from the group consisting of. In some embodiments, the sequence N x (spacer sequence) can specifically hybridize to the complement of the target sequence.
[0054] In some embodiments, the crRNA of the present invention is particularly suitable for use in combination with the Cpf1 protein of the present invention for genome editing, particularly genome editing in eukaryotes such as mammals.
[0055] 3. Method for modifying a target sequence in the genome of a cell In another aspect, the present invention provides a method for modifying a target sequence in the genome of a cell, the method comprising introducing the genome editing system of the present invention into the cell, whereby the guide RNA targets the Cpf1 protein to a target sequence in the genome of the cell, resulting in substitution, deletion and / or addition of one or more nucleotides in the target sequence.
[0056] In another aspect, the present invention provides a method for producing a genetically modified cell, the method comprising introducing the genome editing system of the present invention into the cell, whereby the guide RNA targets the Cpf1 protein to a target sequence in the genome of the cell, resulting in substitution, deletion and / or addition of one or more nucleotides in the target sequence.
[0057] In another aspect, the present invention also provides a genetically modified organism comprising a genetically modified cell produced by the method of the present invention or a progeny cell thereof.
[0058] For the design of target sequences or crRNA coding sequences that can be recognized and targeted by the Cpf1 protein and guide RNA (i.e., crRNA) complex, reference can be made to, for example, Zhang et al., Cell 163, 1-13, October 22, 2015. Generally, the 5' end of the target sequence targeted by the genome editing system of the present invention needs to include a protospacer adjacent motif (PAM) 5'-TTTN, 5'-TTN, 5'-CCN, 5'-TCCN, 5'-TCN or 5'-CTN (N is independently selected from A, G, C and T).
[0059] For example, in some embodiments of the present invention, the target sequence has the following structure: 5'-TYYN-N X -3' or 5'-YYN-N X -3' (N is independently selected from A, G, C and T, Y is selected from C and T; x is an integer of 15≤x≤35; N x represents x consecutive nucleotides).
[0060] In some embodiments, the coding sequence of the crRNA comprises a crRNA scaffold sequence shown in any one of SEQ ID NOs: 25 to 33. In some preferred embodiments, the crRNA scaffold sequence is SEQ ID NO: 30. In some embodiments, the coding sequence of the cRNA further comprises a sequence (i.e., a spacer sequence) that specifically hybridizes to the complement of the target sequence 3' to the crRNA scaffold sequence.
[0061] In some embodiments, the crRNA is as follows: i) 5'-ATTTCTACtgttGTAGAT (SEQ ID NO: 25)-N x -3'; ii) 5'-ATTTCTACtattGTAGAT (SEQ ID NO: 26)-N x -3'; iii) 5'-ATTTCTACtactGTAGAT (SEQ ID NO: 27)-N x -3'; iv) 5'-ATTTCTACtttgGTAGAT (SEQ ID NO: 28)-N x -3'; v) 5'-ATTTCTACtagttGTAGAT (SEQ ID NO: 29)-N x -3'; vi) 5'-ATTTCTACTATGGTAGAT (SEQ ID NO: 30)-N x -3'; vii) 5'-ATTTCTACTGTCGTAGAT (SEQ ID NO: 31)-N x -3'; viii) 5'-ATTTCTACTTGTGTAGAT (SEQ ID NO: 32)-N x -3'; and ix) 5'-ATTTCTACTGTGGTAGAT (SEQ ID NO: 33)-N x -3' (N x represents a nucleotide sequence consisting of x consecutive nucleotides, N is independently selected from A, G, C, and T; x is an integer of 18 ≦ x ≦ 35, preferably x = 23), and is encoded by a nucleotide sequence selected from the group consisting of). In some embodiments, the sequence Nx (The spacer array) can specifically hybridize to the complement of the target array.
[0062] In the present invention, the target array to be modified may be located at any position in the genome, for example, in a functional gene such as a gene encoding a protein, or may be located in a gene expression regulatory region such as a promoter region or an enhancer region, whereby modification of gene function or modification of gene expression can be achieved.
[0063] Substitutions, deletions and / or additions in the target array of cells can be detected by T7EI, PCR / RE or sequencing methods.
[0064] In the method of the present invention, the genome editing system can be introduced into cells by various methods well known to those skilled in the art.
[0065] Methods that can be used to introduce the genome editing system of the present invention into cells include, but are not limited to, calcium phosphate transfection, protoplast fusion, electroporation, lipofection, microinjection, virus infection (e.g., baculovirus, vaccinia virus, adenovirus, adeno-associated virus, lentivirus and other viruses), gene gun method, PEG-mediated protoplast transformation, Agrobacterium-mediated transformation.
[0066] Cells edited by the method of the present invention can be cells of mammals such as humans, mice, rats, monkeys, dogs, pigs, sheep, cows, cats; cells of poultry such as chickens, ducks, geese; plants including monocotyledonous and dicotyledonous plants such as rice, corn, wheat, sorghum, barley, soybean, peanut and Arabidopsis thaliana.
[0067] In some embodiments, the method of the present invention is performed in vitro. For example, the cells are isolated cells. In some embodiments, the cells are CAR-T cells. In some embodiments, the cells are induced pluripotent stem cells.
[0068] In other embodiments, the method of the present invention may also be performed in vivo. For example, the cells are cells within an organism, and the system of the present invention can be introduced into the cells in vivo, for example, by a virus-mediated method. For example, the cells may be tumor cells within a patient.
[0069] 4. Therapeutic applications The present invention also encompasses the use of the genome editing system of the present invention in the treatment of diseases.
[0070] By modifying disease-related genes with the genome editing system of the present invention, upregulation, downregulation, inactivation, activation, or mutation correction of disease-related genes can be achieved, thereby enabling disease prevention and / or treatment. For example, in the present invention, the target sequence may be located in the protein-coding region of a disease-related gene, or may be located in a gene expression regulatory region such as a promoter region or an enhancer region, thereby enabling functional modification of the disease-related gene or modification of the expression of the disease-related gene.
[0071] A "disease-related" gene refers to any gene that produces a transcript or translation product at an abnormal level or in an abnormal form in cells derived from a diseased tissue as compared to non-diseased control tissue or cells. If the change in expression is associated with the appearance and / or progression of a disease, it may be a gene that is expressed at an abnormally high level; or it may be a gene that is expressed at an abnormally low level. A disease-related gene also refers to a gene having one or more mutations or genetic variations that are either the direct cause of a disease or show genetic linkage with one or more genes that are the cause of a disease. The transcript or translation product may be known or unknown, and may be at a normal or abnormal level.
[0072] Accordingly, the present invention also provides a method of treating a disease in a subject in need thereof, the method comprising delivering to the subject an effective amount of the genome editing system of the present invention to modify a gene associated with the disease.
[0073] The present invention also provides the use of the genome editing system of the present invention for the manufacture of a pharmaceutical composition for treating a disease in a subject in need thereof, wherein the genome editing system is for modifying a gene associated with the disease.
[0074] The present invention also provides a pharmaceutical composition for treating a disease in a subject in need thereof, the pharmaceutical composition comprising the genome editing system of the present invention and a pharmaceutically acceptable carrier, wherein the genome editing system is for modifying a gene associated with the disease.
[0075] In some embodiments, the subject is a mammal, such as a human.
[0076] Examples of such diseases include, but are not limited to, tumors, inflammation, Parkinson's disease, cardiovascular diseases, Alzheimer's disease, autism, drug addiction, age-related macular degeneration, schizophrenia, genetic diseases, and the like.
[0077] Further examples of the diseases and corresponding disease-associated genes according to the present invention can be found, for example, in Chinese Patent Application CN201480045703.4.
[0078] 5. Kit The scope of the present invention also includes a kit for use in the method of the present invention, the kit comprising the genome editing system of the present invention and instructions. The kit generally includes a label indicating the intended use and / or the method of using the contents of the kit. The term "label" includes any written or recorded material provided on, with, or otherwise in association with the kit. The present invention also provides the following aspects. [1] The following i) to v): i) A Cpf1 protein and a guide RNA; ii) An expression construct containing a nucleotide sequence encoding a Cpf1 protein and a guide RNA; iii) A Cpf1 protein and an expression construct containing a nucleotide sequence encoding a guide RNA; iv) An expression construct containing a nucleotide sequence encoding a Cpf1 protein and an expression construct containing a nucleotide sequence encoding a guide RNA; v) An expression construct containing a nucleotide sequence encoding a Cpf1 protein and a nucleotide sequence encoding a guide RNA A genome editing system for site - specific modification of a target sequence in the genome of a cell, comprising at least one of: The Cpf1 protein comprises an amino acid sequence having at least 80% sequence identity to the amino acid sequences of SEQ ID NOs: 1 - 12 or to one of SEQ ID NOs: 1 - 12, and the guide RNA can target the Cpf1 protein to a target sequence in the genome of the cell. [2] The system according to [1], wherein the Cpf1 protein comprises an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity to one of SEQ ID NOs: 1 - 12. [3] The system according to [1], wherein the Cpf1 protein comprises an amino acid sequence having one or more amino acid residue substitutions, deletions, or additions relative to one of SEQ ID NOs: 1 - 12. For example, the Cpf1 protein comprises an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid residue substitutions, deletions, or additions relative to one of SEQ ID NOs: 1 - 12. [4] The system according to [1], wherein the Cpf1 protein is derived from a species selected from Agathobacter rectalis, Lachnospira pectinoschiza, Sneathia amnii, Helcococcus kunzii, Arcobacter butzleri, Bacteroidetes oral, Oribacterium sp., Butyrivibrio sp., Proteocatella sphenisci, Candidatus Dojkabacteria, Pseudobutyrivibrio xylanivorans, and Pseudobutyrivibrio ruminis. [5] The system according to [1], wherein the Cpf1 protein comprises an amino acid sequence selected from SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 4, SEQ ID NO: 8, SEQ ID NO: 11, and SEQ ID NO: 12. [6] The system according to [1], wherein the nucleotide sequence encoding the Cpf1 protein is codon-optimized for the organism from which the cell to be genome-edited is derived. [7] The system according to [1], wherein the nucleotide sequence encoding the Cpf1 protein is selected from SEQ ID NOs: 13 to 24. [8] The system according to [1], wherein the guide RNA is a crRNA. [9] The system according to [8], wherein the coding sequence of the crRNA comprises a crRNA scaffold sequence shown in any one of SEQ ID NOs: 25 to 33, preferably comprises the crRNA scaffold sequence shown in SEQ ID NO: 30.
[10] The target sequence has the following structure: 5'-TYYN-N X -3' or 5'-YYN-N X -3' (N is independently selected from A, G, C, and T, Y is selected from C and T; x is an integer of 15 ≦ x ≦ 35; N x represents x consecutive nucleotides).
[11] A method for modifying a target sequence in the genome of a cell, comprising the step of introducing the genome editing system according to any one of [1] to
[10] into the cell, whereby the guide RNA targets the Cpf1 protein to a target sequence in the genome of the cell, resulting in substitution, deletion or addition of one or more nucleotides in the target sequence.
[12] The method according to
[11] , wherein the cell is derived from a mammal such as a human, mouse, rat, monkey, dog, pig, sheep, cow, cat; a poultry such as a chicken, duck, goose; a plant including monocotyledonous and dicotyledonous plants such as rice, corn, wheat, sorghum, barley, soybean, peanut and Arabidopsis thaliana.
[13] The method according to any one of
[11] to
[12] , wherein the system is introduced into the cell by the following methods: calcium phosphate transfection, protoplast fusion, electroporation, lipofection, microinjection, virus infection (such as baculovirus, vaccinia virus, adenovirus, adeno-associated virus, lentivirus and other viruses), gene gun method, PEG-mediated protoplast transformation, Agrobacterium-mediated transformation.
[14] A method for treating a disease in a subject in need thereof, comprising delivering to the subject an effective amount of the genome editing system according to any one of [1] to
[10] to modify a gene associated with the disease in the subject.
[15] Use of the genome editing system according to any one of [1] to
[10] for the manufacture of a pharmaceutical composition for treating a disease in a subject in need thereof, wherein the genome editing system is for modifying a gene associated with the disease in the subject.
[16] A pharmaceutical composition for treating a disease in a subject in need thereof, comprising the genome editing system according to any one of [1] to
[10] and a pharmaceutically acceptable carrier, wherein the genome editing system is for modifying a gene associated with the disease.
[17] The method, use or pharmaceutical composition according to
[14] to
[16] , wherein the subject is a mammal, such as a human.
[18] The method, use or pharmaceutical composition according to
[17] , wherein the disease is selected from tumors, inflammation, Parkinson's disease, cardiovascular diseases, Alzheimer's disease, autism, drug poisoning, age-related macular degeneration, schizophrenia, genetic diseases.
[19] A crRNA comprising a crRNA scaffold sequence corresponding to any one of SEQ ID NOs: 25 to 33, or a coding sequence thereof comprising a sequence shown in any one of SEQ ID NOs: 25 to 33.
[20] The crRNA is i) 5'-ATTTCTACtgttGTAGAT (SEQ ID NO: 25)-N x -3'; ii) 5'-ATTTCTACtattGTAGAT (SEQ ID NO: 26)-N x -3'; iii) 5'-ATTTCTACtactGTAGAT (SEQ ID NO: 27)-N x -3'; iv) 5'-ATTTCTACtttgGTAGAT (SEQ ID NO: 28)-N x -3'; v) 5'-ATTTCTACtagttGTAGAT (SEQ ID NO: 29)-N x -3'; vi) 5'-ATTTCTACTATGGTAGAT (SEQ ID NO: 30)-N x -3'; vii) 5'-ATTTCTACTGTCGTAGAT (SEQ ID NO: 31)-N x -3'; viii) 5'-ATTTCTACTTGTGTAGAT (SEQ ID NO: 32)-N x -3'; and ix) 5'-ATTTCTACTGTGGTAGAT (SEQ ID NO: 33)-N x -3' (N x represents a nucleotide sequence consisting of x consecutive nucleotides, N is independently selected from A, G, C and T; x is an integer of 18 ≦ x ≦ 35, preferably x = 23), the crRNA according to
[19] , encoded by a selected nucleotide sequence.
Example
[0079] Materials and Methods Cell Culture and Transfection HEK293T and HeLa cells were cultured in DMEM medium (Gibco) supplemented with 10% FBS (Gibco), 100 U / ml penicillin, and 100 μg / ml streptomycin, and mouse EpiSC cells were cultured in N2B27 medium supplemented with bFGF. Twelve hours before transfection, cells were seeded at a density of 500,000 cells per well in a 12-well plate. After 12 hours, the cell density was approximately 60% - 70%, and then 1.5 μg (Cpf1:crRNA = 2:1) plasmid was transfected into the cells using Lipofectamine LTX and PLUS reagent (Invitrogen), and the medium was replaced with serum-free opti-MEM medium (Gibco) before transfection. After transfection, it was replaced with serum-supplemented DMED medium 8 - 12 hours after transfection. Forty-eight hours after transfection, GFP-positive cells were sorted by FACS for genotyping.
[0080] Genotype Analysis: T7EI Analysis and DNA Sequencing GFP-positive cells were sorted by FACS, lysed at 55 °C for 30 min using Buffer L + 1 / 100 V Proteinase K (Biotool), inactivated at 95 °C for 5 min, and used directly for PCR detection. Appropriate primers were designed near the Cpf1 targeting site, and the amplification products were purified using the DNA Clean & Concentrator™-5 kit (ZYMO Research). 200 ng of the purified PCR product was added to 1 μL of NEBuffer2 (NEB) and diluted to 10 μL with ddH2O. Then, heterodimers were formed by re-annealing according to a previously reported method (Li, W., et al., Simultaneous generation and germline transmission of multiple gene mutations in rat using CRISPR-Cas systems. Nat Biotechnol, 2013. 31(8): p. 684-6). To the re-annealed product, 0.3 μL of T7EI endonuclease and 1 / 10 V of NEBuffer2 (NEB) were added and digested at 37 °C for 1 hour 30 min. Genotype analysis was performed by 3% TAE-gel electrophoresis. Indels were calculated according to the method reported in a previous paper (Cong, L., et al., Multiplex genome engineering using CRISPR / Cas systems. Science, 2013. 339(6121): p. 819-23).
[0081] The PCR products corresponding to the T7EI-positive samples were ligated into the pEASY-T1 or pEASY-B (Transgen) vector, then transformed, seeded, and incubated overnight at 37 °C. For Sanger sequencing, an appropriate amount of single colonies was selected. The genotype of the mutants was determined by alignment with the wild-type genotype.
[0082] Immunofluorescence staining The Cpf1 eukaryotic expression vector was transfected into HeLa cells for 48 hours, then fixed with 4% PFA for 10 minutes at room temperature; washed 3 times with PBST (PBS + 0.3% Triton X100), blocked with 2% BSA (+0.3% Triton X100) for 15 minutes at room temperature, incubated overnight with the primary antibody rat anti-HA (Roche, 1:1000), washed 3 times with PBST, and incubated with a Cy3-labeled fluorescent secondary antibody (Jackson ImmunoResearch, 1:1000) for 2 hours at room temperature. The nuclei were stained with DAPI (Sigma, 1:1000) for 10 minutes and washed 3 times with PBST. The slides were mounted with an aqueous mounting medium (Abcam) and fixed. A Zeiss LSM780 was used for observation.
[0083] Prokaryotic expression and purification of the Cpf1 protein The gene encoding Cpf1 was cloned into the prokaryotic expression vector BPK2103 / 2104. The protein was purified using a 10×His tag fused to the C-terminus. The expression vector was transformed into BL21(DE3) E. coli competent cells (Transgen). Clones in which IPTG induced high expression were selected, inoculated into 300 mL of CmR+LB medium, and cultured at 37 °C with shaking until the OD600 reached approximately 0.4. IPTG was added at a final concentration of 1 mM, and expression was induced at 16 °C for 16 hours. The culture was centrifuged at 8000 rpm at 4 °C for 10 minutes to recover the bacterial pellet. The bacteria were lysed by sonication (total time 15 minutes, sonication 2 seconds, pause 5 seconds, output 100 W) on an ice bath in 40 mL of NPI-10 (+1×EDTA-free Protease Inhibitor Cocktail, Roche; +5% glycerol). After lysis, centrifugation was performed at 8000 rpm at 4 °C for 10 minutes, and the supernatant was recovered. 2 mL of His60 Ni Superflow Resin (Takara) was added to the supernatant, shaken at 4 °C for 1 hour, and purified using a polypropylene column (Qiagen). Using gravity flow, unwanted proteins without the fused His10 tag were completely washed away with 20 mL of NPI-20, 20 mL of NPI-40, and 10 mL of NPI-100, respectively. Finally, the Cpf1 protein with the fused His10 tag was eluted from the Ni column with 6×0.5 mL of NPI-500. The eluted protein of interest was dialyzed overnight against dialysis buffer (50 mM Tris-HCl, 300 mM KCl, 1 mM DTT, 20% glycerol). The protein solution after dialysis was concentrated using a 100 kDa Amicon Ultra-4, PLHK Ultracel-PL ultrafiltration tube (Millipore). The protein concentration was determined using the Micro BCA™ Protein Assay Kit (Thermo Scientific™).
[0084] RNA in vitro transcription An in vitro transcribed crRNA template with the T7 promoter was synthesized (BGI), and crRNA was transcribed in vitro using the HiScribe™ T7 Quick High Yield RNA Synthesis Kit (NEB) according to the protocol. After transcription was completed, the crRNA was purified by Oligo Clean & Concentrator™ (ZYMO Research), measured for concentration using NanoDrop (Thermo Scientific™), and stored at -80 °C.
[0085] In vitro digestion analysis Target sequences with different 5' PAM sequences were synthesized, cloned into the pUC19 or p11-LacY-wtx1 vector, amplified with primers, and purified. 200 ng of the purified PCR product was reacted with 400 ng of crRNA and 50 nM of Cpf1 protein in the NEBuffer3 (NEB) reaction system at 37 °C for 1 hour. Then it was analyzed by electrophoresis on a 12% urea-TBE-PAGE gel or a 2.5% agarose gel.
[0086] Target sequence The target sequences used in the experiment are shown in Table 1 below:
[0087] [Table 1] JPEG0007700079000002.jpg54153
[0088] Example 1. Identification of a novel Cpf1 protein CRISPR / Cpf1 (Clustered Regularly Interspaced Short Palindromic Repeats from Prevotella and Francisella 1) is an acquired immune mechanism found in Prevotella and Francisella, and has been successfully engineered for DNA editing (Zetsche, B., et al., Cpf1 is a single RNA-guided endonuclease of a class 2 CRISPR-Cas system. Cell, 2015. 163(3): p. 759-71.). Unlike CRISPR / Cas9, Cpf1 requires only one crRNA as a guide and does not require tracrRNA, and the Cpf1 protein uses a T-rich PAM sequence.
[0089] The inventors performed a PSI-Blast search in the NCBI database using the previously reported AsCpf1, LbCpf1, and FnCpf1, and selected 25 unreported Cpf1 proteins, 21 of which have direct repeat (DR) sequences.
[0090] Sequence alignment revealed that the sequences and RNA secondary structures of 21 DRs from different Cpf1 proteins were quite conserved (Figure 1). Based on the hypothesis that in addition to the conserved crRNA, the PAM sequences of all Cpf1 proteins were also conserved (i.e., all use T-rich PAM sequences), 12 out of 25 Cpf1 proteins were selected as candidates. Specific information on the 12 selected candidate Cpf1 proteins is shown in Table 2 below:
[0091]
Table 2
[0092] Example 2. Verification of the ability of Cpf1 protein to edit the human genome Codon optimization was performed for expression in humans, and 12 selected protein coding sequences were synthesized (BGI) and cloned into eukaryotic expression vectors (pCAG-2AeGFP-SV40 and pCAG-2AeGFP-SV40_v4) and prokaryotic expression vectors (BPK2103-ccdB and BPK2014-ccdB).
[0093] Through HeLa cell transfection experiments, the inventors were able to find that the selected Cpf1 protein can be clearly expressed in the nucleus (Figure 2).
[0094] Next, co-transfection of 293T cells with crRNA targeting the human Dnmt1 gene, and identification by PCR and T7EI digestion showed that 8 Cpf1 proteins (Ab, Ar, Bo, Bs, Hk, Lp, Ps, Sa) induced mutations at site 3 of the Dnmt1 gene in 293T cells (Figure 3A). DNA sequencing further confirmed that the gene site generated genetic mutations (Figure 3B).
[0095] To further demonstrate that the identified Cpf1 proteins can induce mutations in the mammalian genome, the above experiments were repeated using sites 1 and 2 of the Dnmt1 gene. The results of T7EI digestion revealed that 4 proteins, ArCpf1, BsCpf1, HkCpf1, and LpCpf1, generated significant indels at both sites (Figures 4A and 5A). The results of DNA sequencing showed that in addition to the above 4 proteins, PsCpf1 and SaCpf1 also caused mutations in the gene (Figures 4B and 5B).
[0096] Example 3. Identification of PAM sequences To determine the PAM sequences of the identified Cpf1 proteins, first, the BsCpf1 protein was selected for detailed study.
[0097] BsCpf1 was expressed in E. coli and purified with a His tag (Figure 6A). The purified BsCpf1 protein, along with the corresponding crRNA and dsDNA fragment (human Dnmt1 target site 2), was incubated at 37 °C for 1 hour and then subjected to TBE denaturing PAGE gel electrophoresis, which revealed that the PAM sequence of BsCpf1 is 5'(T)TTN- (Figures 6B - E).
[0098] Next, the PAM sequences of the remaining Cpf1 proteins were identified in a similar manner, and the results are shown in Figure 7. The experimental results indicate that all the selected Cpf1 proteins have in vitro enzyme activity. The PAM sequences of 9 of these Cpf1 proteins are 5'(T)TTN-. The PAM sequences of C6Cpf1, HkCpf1, and PsCpf1 are 5'(T)YYN-.
[0099] To confirm that the PAM sequence for in vivo editing by HkCpf1 is 5'(T)YYN-, the HkCpf1 vector was transfected into 293T cells together with the targeting crRNA vectors at 8 sites on the human AAVS1 gene. The experimental results (Figure 8) show that HkCpf1 was able to edit 7 out of these 8 sites to generate mutations, and the PAM sequence is 5'(T)YYN- (underlined part in Figure 8B).
[0100] Example 4. Genome Editing in Mice To expand the applicability of these Cpf1 proteins, several transfection experiments were conducted in the mouse EpiSC cell line.
[0101] First, the inventors transfected EpiSC cells with a BsCpf1 vector and a crRNA vector (pUC19-crRNA) that targets eight sites on the mouse Tet1 gene, respectively. By T7EI digestion and DNA sequencing, it was confirmed that BsCpf1 caused gene mutations at sites 2, 3, and 7 (Figure 9A). The PAM sequence at site 7 was 5'TTA-, and this result was consistent with the conclusion of Example 3. This indicates that the BsCpf1 protein uses the 5'TTN-PAM sequence.
[0102] Next, the DNA editing abilities of ArCpf1, BsCpf1, HkCpf1, and PxCpf1 were repeatedly verified. Co-transfection was performed using a crRNA vector (pUC19-crRNA) that targets eight sites on the mouse MeCP2 gene. Using T7EI digestion and DNA sequencing, it was demonstrated that ArCpf1 (Figure 10D, E), BsCpf1 (Figure 9B, C), HkCpf1 (Figure 10F, G), and PxCpf1 (Figure 11H, I) can target and edit the mouse genome.
[0103] Furthermore, the editing abilities of ArCpf1, BsCpf1, HkCpf1, LpCpf1, PrCpf1, and PxCpf1 were further verified using the target site 12 of mouse Apob, the target site 4 of mouse MeCP2, the target site 1 of mouse Nrl, and the target site 7 of mouse Nrl. The results are shown in Figure 12, demonstrating that all six proteins can edit the genome. The PAM sequence at site Nrl-7 is 5'TTG, indicating that the PAM sequences of BsCpf1 and PrCpf1 are 5'TTN-.
[0104] The above experimental results demonstrate that all 12 Cpf1 proteins found in the present invention have DNA editing ability, and ArCpf1 (SEQ ID NO: 1), BsCpf1 (SEQ ID NO: 8), HkCpf1 (SEQ ID NO: 4), PxCpf1 (SEQ ID NO: 11), LpCpf1 (SEQ ID NO: 2), and PrCpf1 (SEQ ID NO: 12) enable efficient mammalian genome editing.
[0105] Example 5. Optimization of the crRNA Scaffold for Improving Editing Efficiency In this example, to improve the genome editing efficiency of each Cpf1 protein (Table 3), the crRNA scaffold of newly identified Cpf1 proteins that can be used for mammalian genome editing is optimized.
[0106] The experimental results are shown in FIG. 13. FIG. 13A shows that cells were transfected with each Cpf1 protein and different crRNA plasmids, and PCR and T7EI analysis demonstrated that different crRNA scaffolds affect the editing efficiency of Cpf1 proteins.
[0107] Next, the inventors established a library of crRNA mutants transfected with BsCpf1 or PrCpf1, respectively. PCR and T7EI analysis were used to screen for crRNA mutants that enable Cpf1 to efficiently edit the mammalian genome.
[0108] The editing efficiency of the crRNA31 mutant was significantly higher than that of the wild-type crRNA scaffold (crRNA2) derived from the genome of the strain. Cells were transfected with Cpf1, crRNA31, and crRNA2 plasmids, and PCR and T7EI analysis confirmed that crRNA31 can double the editing efficiency of five Cpf1 proteins, ArCpf1, BsCpf1, HkCpf1, PrCpf1, and PxCpf1, at site MeCP2-4 (FIG. 13B).
[0109] FIGS. 14A and B show that analysis of different target sites by T7EI confirms that the screened crRNA scaffolds can improve the editing efficiency of BsCpf1 and PrCpf1. In the case of PrCpf1, in addition to the crRNA31 scaffold, the crRNA77, crRNA129, and crRNA159 scaffolds can also significantly improve the editing efficiency at site Nrl-1.
[0110] Figure 14C shows the analysis of intracellular GFP editing efficiency by flow cytometry. The crRNA31 scaffold can improve the editing efficiency of BsCpf1 and PrCpf1, and crRNA77 and crRNA159 can also significantly improve the editing efficiency of PrCpf1.
[0111]
Table 3
[0112] Sequence information JPEG0007700079000005.jpg47161JPEG0007700079000006.jpg254161JPEG0007700079000007.jpg254160JPEG0007700079000008.jpg254161JPEG0007700079000009.jpg254162JPEG0007700079000010.jpg254160JPEG0007700079000011.jpg255162JPEG0007700079000012.jpg255160JPEG0007700079000013.jpg255160JPEG0007700079000014.jpg254160JPEG0007700079000015.jpg254161JPEG0007700079000016.jpg254161JPEG0007700079000017.jpg254161JPEG0007700079000018.jpg254161JPEG0007700079000019.jpg254159JPEG0007700079000020.jpg254160JPEG0007700079000021.jpg253163JPEG0007700079000022.jpg4074
Claims
**Claim 1** The following i) to v): i) A Cpf1 protein and a guide RNA; ii) An expression construct containing a nucleotide sequence encoding a Cpf1 protein and a guide RNA; iii) A Cpf1 protein and an expression construct containing a nucleotide sequence encoding a guide RNA; iv) An expression construct containing a nucleotide sequence encoding a Cpf1 protein and an expression construct containing a nucleotide sequence encoding a guide RNA; v) An expression construct containing a nucleotide sequence encoding a Cpf1 protein and a nucleotide sequence encoding a guide RNA A genome editing system for site-specific modification of a target sequence in the genome of a cell, comprising at least one of: The Cpf1 protein comprises the amino acid sequence of SEQ ID NO: 4 or an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 4, and the guide RNA can target the Cpf1 protein to a target sequence in the genome of the cell. The guide RNA is a crRNA, and the coding sequence of the crRNA contains a crRNA scaffold sequence shown in any one of SEQ ID NO: 30 and SEQ ID NOs: 27-29. System. **Claim 2** The system according to claim 1, wherein the Cpf1 protein comprises an amino acid sequence having at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO:
4. **Claim 3** The system according to claim 1, wherein the Cpf1 protein comprises an amino acid sequence having one or more amino acid residue substitutions, deletions, or additions relative to SEQ ID NO:
4. For example, the Cpf1 protein comprises an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid residue substitutions, deletions, or additions relative to SEQ ID NO:
4. **Claim 4** The system according to claim 1, wherein the Cpf1 protein is derived from a species selected from Agathobacter rectalis, Lachnospira pectinoschiza, Sneathia amnii, Helcococcus kunzii, Arcobacter butzleri, Bacteroidetes oral, Oribacterium sp., Butyrivibrio sp., Proteocatella sphenisci, Candidatus Dojkabacteria, Pseudobutyrivibrio xylanivorans, Pseudobutyrivibrio ruminis.
5. The system according to claim 1, wherein the Cpf1 protein comprises the amino acid sequence of SEQ ID NO:
4.
6. The system according to claim 1, wherein the nucleotide sequence encoding the Cpf1 protein is codon-optimized for the organism from which the cell to be genome-edited is derived.
7. The system according to claim 1, wherein the nucleotide sequence encoding the Cpf1 protein is the nucleotide sequence of SEQ ID NO:
16.
8. The target array has the following structure: 5'-TYYN-N X -3' or 5'-YYN-N X -3' (N is independently selected from A, G, C, and T, and Y is selected from C and T; x is an integer such that 15 ≦ x ≦ 35; N x represents x consecutive nucleotides), the system according to claim 1.
9. A method for modifying a target sequence in a cell genome in vitro, comprising introducing the genome editing system according to any one of claims 1 to 8 into the cell, whereby the guide RNA targets the Cpf1 protein to a target sequence in the cell genome, resulting in substitution, deletion or addition of one or more nucleotides in the target sequence.
10. The method according to claim 9, wherein the cell is derived from a mammal such as a human, mouse, rat, monkey, dog, pig, sheep, cow, cat; a poultry such as a chicken, duck, goose; a plant including monocotyledonous and dicotyledonous plants such as rice, corn, wheat, sorghum, barley, soybean, peanut and Arabidopsis thaliana.
11.
12. The method according to any one of claims 9 to 10, wherein the system is introduced into the cell by the following methods: calcium phosphate transfection, protoplast fusion, electroporation, lipofection, microinjection, virus infection (e.g., baculovirus, vaccinia virus, adenovirus, adeno-associated virus, lentivirus and other viruses), gene gun method, PEG-mediated protoplast transformation, Agrobacterium-mediated transformation.
13. A method of treating a disease in a non-human subject in need of treatment for the disease, the method comprising delivering to the non-human subject an effective amount of the genome editing system according to any one of claims 1 to 8 to modify a gene associated with the disease in the non-human subject.
14.
15. Use of the genome editing system according to any one of claims 1 to 8 for the manufacture of a pharmaceutical composition for treating a disease in a subject in need of treatment for the disease, wherein the genome editing system is for modifying a gene associated with the disease in the subject.
16.
17. A pharmaceutical composition for treating a disease in a subject in need of treatment for the disease, comprising the genome editing system according to any one of claims 1 to 8 and a pharmaceutically acceptable carrier, wherein the genome editing system is for modifying a gene associated with the disease.
18. The method according to claim 12, wherein the subject is a mammal.
19. The use according to claim 13, wherein the subject is a mammal.
20. The pharmaceutical composition according to claim 14, wherein the subject is a mammal.
21. The method according to claim 15, wherein the disease is selected from tumors, inflammation, Parkinson's disease, cardiovascular diseases, Alzheimer's disease, autism, drug addiction, age-related macular degeneration, schizophrenia, genetic diseases. The use according to claim 16, wherein the disease is selected from tumors, inflammation, Parkinson's disease, cardiovascular diseases, Alzheimer's disease, autism, drug poisoning, age-related macular degeneration, schizophrenia, and genetic diseases. The pharmaceutical composition according to claim 17, wherein the disease is selected from tumors, inflammation, Parkinson's disease, cardiovascular diseases, Alzheimer's disease, autism, drug poisoning, age-related macular degeneration, schizophrenia, and genetic diseases.
Citation Information
Patent Citations
JPP7100057B
Novel crispr enzymes and systems
WO2016205711A1
Modified crispr RNA and modified single crispr RNA and uses thereof
WO2017004261A1