Engineered CAS12B effector proteins and methods of use thereof
By optimizing the structure of the Cas12b nuclease and introducing mutations in specific amino acid residues, the problem of low gene editing efficiency in the CRISPR-Cas system was solved, and more efficient gene editing effects were achieved.
Patent Information
- Application Number
- CN202211581644.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-12-09
- Filing Date
- 2022-12-09
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2042-12-09
AI Technical Summary
The existing CRISPR-Cas system has the problem of limited gene editing efficiency in genome editing.
By engineering the Cas12b nuclease and introducing specific amino acid residue mutations, including substitutions of positively charged amino acid residues, aromatic ring amino acid residues, and hydrophobic amino acid residues, the interaction between the structural domain of the Cas12b nuclease and the PAM and DNA substrate is optimized, thereby improving its enzymatic activity.
The gene editing efficiency of Cas12b nuclease was significantly improved, and the cutting ability and editing effect of target nucleic acid were enhanced.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims priority to the following international patent application: PCT / CN2021 / 136761 filed on December 9, 2021, the entire contents of which are incorporated herein by reference in their entirety.
[0003] Electronic Sequence Listing Citation
[0004] The contents of the electronic sequence listing (253112000541SEQLIST.xml; file size 111583 bytes; date generated: November 22, 2022) are incorporated herein by reference in their entirety. Technical Field
[0005] The present application generally relates to the field of biotechnology. More specifically, the present application relates to engineered Cas12b effector proteins and engineered gRNA scaffolds with improved activity (e.g., gene editing activity) or eliminated nuclease activity and methods of using the same. Background Art
[0006] Genome editing is an important and useful technology in genome research and various applications. Various genome editing systems are available, including clustered regularly interspaced short palindromic repeats (CRISPR)-Cas systems, transcription activator-like effector nuclease (TALEN) systems, and zinc finger nuclease (ZFN) systems.
[0007] The CRISPR-Cas system is an efficient and cost-effective genome editing technology that is widely applicable to a range of eukaryotic organisms from yeast and plants to zebrafish and humans (reviewed by VanderOost, 2013《Science》339:768-770, and Charpentier and Doudna, 2013《Nature》495:50-51). The CRISPR-Cas system provides adaptive immunity in archaea and bacteria by adopting a combination of Cas effector protein and CRISPR RNA (crRNA). So far, according to the significant function and evolutionary modularity of the system, two categories (class 1 and class 2) of CRISPR-Cas systems including six types (types I to VI) have been characterized. Among the class 2 CRISPR-Cas systems, type II Cas9 system and type VA / B / E / J Cas12a / Cas12b / Cas12e / Cas12f / Cas12j system have been used for genome editing and provide broad prospects for biomedical research. Summary of the Invention
[0008] Current CRISPR-Cas systems have various limitations, including limited gene editing efficiency. The present application provides improved methods and systems for efficient genome editing across multiple genomic sites. In particular, provided herein are engineered Cas12b nucleases, engineered Cas12b effector proteins, engineered gRNAs (e.g., sgRNAs and tracrRNAs) comprising engineered scaffolds, and methods for using engineered Cas12b effector proteins and / or engineered gRNAs (such as in gene editing) with improved enzymatic activity. In one aspect, the application provides engineered Cas12b nucleases comprising one, two, or three types of mutations relative to a reference Cas12b nuclease, wherein the mutations comprise: (1) replacing one or more amino acid residues that interact with a protospacer sequence adjacent motif (PAM) in a reference Cas12b nuclease with a positively charged amino acid residue (e.g., R, H, K); and / or (2) replacing one or more amino acid residues involved in opening a double-stranded DNA (dsDNA) in a reference Cas12b nuclease with an amino acid residue having an aromatic ring (e.g., F, Y, W); and / or (3) replacing one or more amino acid residues that interact with a single-stranded DNA (ssDNA) substrate in the RuvC domain of a reference Cas12b nuclease with a positively charged amino acid residue (e.g., R, H, K) or a hydrophobic amino acid residue (e.g., F, Y, W). In some embodiments, the reference Cas12b nuclease is a wild-type Cas12b nuclease. In some embodiments, the reference Cas12b nuclease is a Cas12b nuclease from Alicyclobacillusacidiphilus (AaCas12b). In some embodiments, the reference Cas12b nuclease comprises the amino acid sequence of SEQ ID NO: 1.
[0009] In some embodiments according to any one of the above-mentioned engineered Cas12b nucleases, the engineered Cas12b nuclease comprises replacing one or more amino acid residues that interact with PAM in the reference Cas12b nuclease with a positively charged amino acid residue. In some embodiments, the one or more amino acid residues that interact with PAM are within 10 (e.g., 9, 8, 7, 6, 5, 4, 3, 2, 1 or less) angstroms of the distance from PAM in the three-dimensional structure. In some embodiments, the one or more amino acid residues that interact with PAM are located at one or more of the following positions: 116, 123, 130, 132, 144, 145, 153, 173, 222, 395, 400, and 475. In some embodiments, the one or more amino acid residues that interact with the PAM include one or more of the following amino acid residues: D116, K123, D130, D132, N144, K145, E153, D173, Q222, D395, N400, and / or E475. In some embodiments, the one or more amino acid residues that interact with the PAM include one or more of the following amino acid residues: D116 and E475. In some embodiments, the positively charged amino acid residue is R or K. In some embodiments, the engineered Cas12b nuclease comprises one or more of the following substitutions: D116R and E475R. In some embodiments, the amino acid residues are numbered according to SEQ ID NO: 1. In some embodiments, the engineered Cas12b nuclease comprises the amino acid sequence of SEQ ID NO: 2 or 3.
[0010] In some embodiments according to any one of the above-mentioned engineered Cas12b nucleases, the engineered Cas12b nuclease comprises replacing one or more amino acid residues involved in opening the DNA double strand in the reference Cas12b nuclease with an amino acid residue having an aromatic ring. In some embodiments, the one or more amino acid residues involved in opening the DNA double strand interact with the last base pair in the PAM relative to the 3' end of the target chain. In some embodiments, the one or more amino acid residues involved in opening the DNA double strand are located at one or more of the following positions: 118 and 119. In some embodiments, the amino acid residue with an aromatic ring is Y, F, or W. In some embodiments, the one or more amino acid residues involved in opening the DNA double strand in the reference Cas12b nuclease with an amino acid residue having an aromatic ring include Q119Y, Q119F, or Q119W. In some embodiments, the amino acid residues are numbered according to SEQ ID NO: 1. In some embodiments, the engineered Cas12b nuclease comprises an amino acid sequence of SEQ ID NO: 4 to 6.
[0011] In some embodiments according to any one of the above-mentioned engineered Cas12b nucleases, the engineered Cas12b nuclease comprises replacing one or more amino acid residues that are located in the RuvC domain and interact with a single-stranded DNA substrate in the reference Cas12b nuclease with a positively charged amino acid residue or a hydrophobic amino acid residue. In some embodiments, the one or more amino acid residues that are located in the RuvC domain and interact with a single-stranded DNA substrate are within 10 (e.g., 9, 8, 7, 6, 5, 4, 3, 2, 1 or less) angstroms from the single-stranded DNA substrate in the three-dimensional structure. In some embodiments, the one or more amino acid residues located in the RuvC domain and interacting with the single-stranded DNA substrate are located at one or more of the following positions: 300, 301, 304, 329, 636, 639, 647, 682, 757, 758, 761, 764, 768, 852, 854, 856, 857, 858, 860, 862, 863, 865, 866, 867, 869, 938, 956, 957, 958, 994, 1093, and 1097. In some embodiments, the one or more amino acid residues located in the RuvC domain and interacting with the single-stranded DNA substrate include one or more of the following amino acid residues: D300, K301, E304, N329, E636, Q639, T647, Q682, I757, E758, E761, E764, K768, E852, Q854, N856, N857, D858, P860, S862, E863, N865, Q866, L867, Q869, E938, E956, G957, E958, I994, Q1093, and W1097. In some embodiments, the engineered Cas12b nuclease comprises replacing one or more of the following amino acid residues with positively charged amino acid residues: E636, Q639, T647, Q682, I757, E758, E761, K768, Q854, N857, D858, N865, Q866, I994, Q1093, and W1097. In some embodiments, the positively charged amino acid residue is R or K. In some embodiments, the engineered Cas12b nuclease comprises one or more of the following substitutions: E636R, Q639R, T647R, Q682R, I757R, E758R, E761R, Q854R, N857K, D858R, I994R, Q1093R, and W1097R. In some embodiments, the engineered Cas12b nuclease comprises substitution of one or more of the following amino acid residues: E758, E761, E863, N865, Q866, Q869, Q956, and Q1093 with a hydrophobic amino acid residue.In some embodiments, the hydrophobic amino acid residue is W, Y, F, or M, such as W, Y or M. In some embodiments, the engineered Cas12b nuclease comprises one or more of the following substitutions: N865W, N865Y, Q866M, Q869M, Q1093W, and Q1093Y. In some embodiments, the amino acid residues are numbered according to SEQ ID NO: 1. In some embodiments, the engineered Cas12b nuclease comprises an amino acid sequence as described in any one of SEQ ID NO: 7 to 19.
[0012] In some embodiments according to any of the above engineered Cas12b nucleases, the engineered Cas12b nuclease comprises any one or a combination of the following substitutions: (1) D116R; (2) E475R; (3) Q119F and E475R; (4) Q119F, E475R, and E758R; (5) Q119Y; (6) Q119F; (7) Q119W; (8) I757R; (9) E758R; (10) E761R; (11) K768R; (12) I757R and E758R; (13) I757R and E761R; (14) I757R and K768R; (15) E758R and E761R; (16) E758R and K768R; (17) E76 1R and K768R; (18) I757R, E758R, and E761R; (19) I757R, E758R, and K768R; (20) I757R, E761R, and K768R; (21) E758R, E761R, and K768R; (22) I757R, E758R, E761R, and K768R ; (23) Q866M; (24) Q869M; and (25) Q866M and Q869M; (26) E636R; (27) Q854R; (28) N857K; (29) N865W; (30) N865Y; (31) Q1093W; (32) Q1093Y and (33) D858R; and wherein the amino acid residues are numbered according to SEQ ID NO: 1. In some embodiments, the engineered Cas12b nuclease comprises any one or a combination of the following substitutions: (1) Q866M+Q869M; (2) Q119F+E475R; and (3) Q119F+E475R+E758R; and wherein the amino acid residues are numbered according to SEQ ID NO: 1. In some embodiments, the engineered Cas12b nuclease comprises an amino acid sequence as described in any one of SEQ ID NOs: 20 to 22.
[0013] In some embodiments according to any of the above-described engineered Cas12b nucleases, the engineered Cas12b nuclease comprises an amino acid sequence having at least about 85% (88%, 90%, 95%, 96%, 97%, 98%, 99% or more) sequence identity to any one of SEQ ID NO: 2 to SEQ ID NO: 22. In some embodiments, the engineered Cas12b nuclease comprises (or consists of, or consists essentially of) an amino acid sequence of any one of SEQ ID NO: 1 to SEQ ID NO: 22.
[0014] In some embodiments according to any of the above-described engineered Cas12b nucleases, the engineered Cas12b nuclease further comprises one or more mutations that increase the flexibility of the flexible region, the flexible region comprising amino acid residues 855 to 859. In some embodiments, the one or more mutations that increase flexibility comprise N856G. In some embodiments, the amino acid positions are numbered according to SEQ ID NO: 1.
[0015] One aspect of the present application provides an engineered Cas12b nuclease comprising any one or more of the following mutations: (1) D116R; (2) E475R; (3) Q119F+E475R; (4) Q119F+E475R+E758R; (5) Q119Y; (6) Q119F; (7) Q119W; (8) I757R; (9) E758R; (10) E761R; (11) K768R; (12) I757R+E758R; (13) I757R+E761R; (14) I757R+K768R; (15) E758R+E761R; (16) E758R+K768R; (17) E761R+K768R. 68R; (18)I757R+E758R+E761R; (19)I757R+E758R+K768R; (20)I757R+E761R +K768R; (21)E758R+E761R+K768R; (22)I757R+E758R+E761R+K768R; (23)Q8 66M;(24) Q869M;(25) Q866M+Q869M;(26) E636R;(27) Q854R;(28) N857K;(29) N865W;(30) N865Y;(31) Q1093W;(32) Q1093Y;and (33) D858R;and wherein amino acid positions are numbered according to SEQ ID NO: 1. In some embodiments, the engineered Cas12b nuclease comprises any one or a combination of the following substitutions: (1) Q866M+Q869M;(2) Q119F+E475R;and (3) Q119F+E475R+E758R;and wherein the amino acid residues are numbered according to SEQ ID NO: 1. In some embodiments, the engineered Cas12b nuclease comprises the following substitutions: Q119F+E475R+E758R; and wherein the amino acid residues are numbered according to SEQ ID NO: 1.
[0016] One aspect of the present application provides an engineered Cas12b nuclease having at least about 85% (e.g., at least about 88%, 90%, 95%, 96%, 97%, 98%, 99% or more) sequence identity to any one of SEQ ID NOs: 2 to 22, or comprising the amino acid sequence of any one of SEQ ID NOs: 2 to 22.
[0017] One aspect of the present application provides an engineered Cas12b effector protein comprising an engineered Cas12b nuclease according to any one of the above-mentioned engineered Cas12b nucleases, or a variant thereof, or a functional derivative thereof. In some embodiments, the engineered Cas12b nuclease or its functional derivative has enzymatic activity. In some embodiments, the engineered Cas12b effector protein is capable of inducing double-strand breaks in DNA molecules. In some embodiments, the engineered Cas12b effector protein is capable of inducing single-strand breaks in DNA molecules. In some embodiments, the engineered Cas12b effector protein comprises an enzymatically inactive mutant of the engineered Cas12b nuclease. In some embodiments, the enzymatically inactive mutant of the engineered Cas12b nuclease comprises a substitution of one or more amino acid residues selected from the group consisting of D570A, E848A, R785A, E848A, R911A, and D977A, and wherein the amino acid residues are numbered according to SEQ ID NO: 1. In some embodiments, the enzymatically inactive mutant of the engineered Cas12b nuclease comprises (or consists of, or consists essentially of) the amino acid sequence of any one of SEQ ID NOs: 79 to 81, or a variant thereof having at least about 85% (e.g., at least about 88%, 90%, 95%, 96%, 97%, 98%, 99% or more) sequence identity to any one of SEQ ID NOs: 79 to 81.
[0018] In some embodiments according to any one of the above-mentioned engineered Cas12b effector proteins, the engineered Cas12b effector protein further comprises a functional domain fused to an engineered Cas12b nuclease or a functional derivative thereof. In some embodiments, the functional domain is selected from the group consisting of a translation initiator domain, a transcription repressor domain, a transactivation domain, an epigenetic modification domain, a nucleobase editing domain, a reverse transcriptase domain, a reporter domain, and a nuclease domain. In some embodiments, the transcription repression domain is a Krüppel-associated box (KRAB) domain, such as comprising an amino acid sequence of SEQ ID NO: 72.
[0019] In some embodiments according to any of the above-described engineered Cas12b effector proteins, the engineered Cas12b effector protein comprises a first polypeptide and a second polypeptide, the first polypeptide comprising the N-terminal portion of the engineered Cas nuclease or a functional derivative thereof, the second polypeptide comprising the C-terminal portion of the engineered Cas nuclease or a functional derivative thereof, wherein the first polypeptide and the second polypeptide are capable of associating with each other in the presence of a guide RNA comprising a guide sequence to form a CRISPR complex that specifically binds to a target nucleic acid comprising a target sequence complementary to the guide sequence. In some embodiments, the engineered Cas12b effector protein comprises a first polypeptide and a second polypeptide, wherein the first polypeptide comprises the N-terminal amino acid residues 1 to X of the engineered Cas12b nuclease or a functional derivative thereof, wherein the second polypeptide comprises residues X+1 to the C-terminus of the engineered Cas12b nuclease or a functional derivative thereof, wherein the first polypeptide and the second polypeptide are capable of associating with each other in the presence of a guide RNA comprising a guide sequence to form a CRISPR complex that specifically binds to a target nucleic acid, wherein the target nucleic acid comprises a target sequence complementary to the guide sequence. In some embodiments, the first polypeptide and the second polypeptide each comprise a dimerization domain. In some embodiments, the first dimerization domain and the second dimerization domain associate with each other in the presence of an inducer.In some embodiments, the first polypeptide and the second polypeptide do not comprise a dimerization domain.
[0020] Another aspect of the present application provides a single guide RNA (sgRNA) comprising the sequence of any one of SEQ ID NOs: 25 to 53.
[0021] Another aspect of the present application provides an engineered CRISPR-Cas12b system comprising: (a) an engineered Cas12b nuclease according to any one of the above-mentioned engineered Cas12b nucleases or an engineered Cas12b effector protein according to any one of the above-mentioned engineered Cas12b effector proteins, or a nucleic acid encoding the same; and (b) a guide RNA comprising a guide sequence complementary to the target sequence of the target nucleic acid, or a nucleic acid encoding the guide RNA, wherein the engineered Cas12b effector protein and the guide RNA are capable of forming a CRISPR complex that specifically binds to the target nucleic acid comprising the target sequence and induces modification of the target nucleic acid. In some embodiments, the guide RNA comprises crRNA and tracrRNA. In some embodiments, the engineered CRISPR-Cas12b system comprises a precursor guide RNA array encoding multiple crRNAs. In some embodiments, the guide RNA is a single guide RNA (sgRNA). In some embodiments, the sgRNA comprises a sequence as described in any one of SEQ ID NOs: 23 to 53. In some embodiments, the engineered CRISPR-Cas12b system comprises one or more vectors encoding the engineered Cas12b nuclease or the engineered Cas12b effector protein. In some embodiments, the one or more vectors are adeno-associated virus (AAV) vectors. In some embodiments, the AAV vector also encodes a guide RNA.
[0022] Another aspect of the present application provides an engineered CRISPR-Cas12b system, comprising: (a) an engineered Cas12b nuclease according to any one of the above-mentioned engineered Cas12b nucleases, or an engineered Cas12b effector protein according to any one of the above-mentioned engineered Cas12b effector proteins, the Cas12b nuclease or effector protein comprising SEQ ID NO: 1 to 22 and 79 to 81 amino acid sequence or its encoding nucleic acid; and (b) gRNA comprising a guide sequence complementary to the target sequence of the target nucleic acid or the nucleic acid encoding the gRNA, wherein the gRNA comprises an engineered scaffold comprising a sequence of any one of SEQ ID NO: 25 to 53; wherein the Cas12b nuclease (e.g., engineered) or its effector protein and gRNA are capable of forming a CRISPR complex that specifically binds to the target nucleic acid and induces modification of the target nucleic acid. In some embodiments, the gRNA comprises crRNA and tracrRNA, and wherein the tracrRNA comprises an engineered scaffold or a portion thereof. In some embodiments, the engineered CRISPR-Cas12b system comprises a precursor gRNA array encoding a variety of crRNAs. In some embodiments, the gRNA is an sgRNA. In some embodiments, the engineered CRISPR-Cas12b system comprises one or more vectors encoding engineered Cas12b nucleases or their effector proteins, or Cas12b nucleases or their effector proteins. In some embodiments, the one or more vectors are AAV vectors. In some embodiments, the one or more vectors further encode the gRNA.
[0023] One aspect of the present application provides a method for detecting a target nucleic acid in a sample, comprising: (a) contacting the sample with an engineered CRISPR-Cas12b system according to any one of the above-mentioned engineered CRISPR-Cas12b systems and a labeled detection nucleic acid, wherein the gRNA comprises a guide sequence complementary to the target sequence of the target nucleic acid, and wherein the labeled detection nucleic acid is single-stranded and does not hybridize with the guide sequence of the guide RNA; and (b) measuring the detectable signal generated by the engineered Cas12b nuclease or its effector protein cleaving the labeled detection nucleic acid, thereby detecting the target nucleic acid.
[0024] One aspect of the present application provides a method for modifying a target nucleic acid comprising a target sequence, comprising contacting the target nucleic acid with an engineered CRISPR-Cas12b system according to any one of the above-mentioned engineered CRISPR-Cas12b systems. In some embodiments, the method is performed in vitro. In some embodiments, the target nucleic acid is present in a cell. In some embodiments, the cell is a bacterial cell, a yeast cell, a plant cell, or an animal cell (e.g., a mammalian cell). In some embodiments, the method is performed in vitro. In some embodiments, the method is performed in vivo. In some embodiments, the target nucleic acid is cut. In some embodiments, the target sequence in the target nucleic acid is changed by the engineered CRISPR-Cas12b system. In some embodiments, the expression of the target nucleic acid is changed by the engineered CRISPR-Cas12b system. In some embodiments, the target nucleic acid is genomic DNA. In some embodiments, the target sequence is associated with a disease or disorder. In some embodiments, the engineered CRISPR-Cas12b system comprises a precursor guide RNA array encoding multiple crRNAs, wherein each crRNA comprises a different guide sequence.
[0025] Another aspect of the present application provides a method for treating a disease or condition associated with a target nucleic acid in a cell of an individual, comprising modifying a target nucleic acid in a cell of the individual using an engineered CRISPR-Cas12b system according to any one of the above-mentioned engineered CRISPR-Cas12b systems, thereby treating the disease or condition. In some embodiments, the disease or condition is selected from the group consisting of cancer, cardiovascular disease, genetic disease, autoimmune disease, metabolic disease, neurodegenerative disease, eye disease, bacterial infection, and viral infection.
[0026] Also provided are engineered cells comprising a modified target nucleic acid, wherein the target nucleic acid has been modified using a method according to any of the methods for modifying a target nucleic acid described above. Also provided are engineered non-human animals comprising one or more engineered cells thereof.
[0027] Also provided are compositions, kits, and articles of manufacture for use in any of the above methods.
[0028] It should be understood that certain features of the present disclosure that are described in the context of separate embodiments for the sake of clarity may also be provided in combination in a single embodiment. Conversely, various features of the present disclosure that are described in the context of individual embodiments for the sake of brevity may also be provided individually or in any suitable subcombination. All combinations of embodiments relating to specific method steps, reagents, or conditions, or components of compositions are specifically contemplated by the present disclosure and are disclosed herein, just as if each and every combination were individually and specifically disclosed. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 Shown are the gene editing efficiencies (% indels) of exemplary AaCas12b variants in which the amino acid residues interacting with the PAM in wild-type AaCas12b were substituted with R. AaCas12b variants with D116R or E475R substitutions showed improved editing efficiency compared to wild-type (WT) AaCas12b.
[0030] Figure 2 The gene editing efficiency of exemplary AaCas12b variants is shown, in which the amino acid residues involved in opening the DNA double strand in wild-type AaCas12b are substituted with aromatic amino acid residues. Compared with WTAaCas12b, AaCas12b variants with Q119Y, Q119F or Q119W substitutions showed improved gene editing efficiency.
[0031] Figure 3 The gene editing efficiency of exemplary AaCas12b variants is shown, in which the amino acid residue located in the RuvC domain and interacting with the single-stranded DNA substrate in the wild-type AaCas12b is replaced by R.
[0032] Figures 4A to 4B Shown are the gene editing efficiencies of exemplary AaCas12b variants, in which the amino acid residues in the wild-type AaCas12b that are located in the RuvC domain and interact with single-stranded DNA are substituted with lysine (K) or arginine (R) residues. Figure 4A The editing efficiency at the genomic locus CCR5-3 is shown, while Figure 4B The editing efficiency at the genomic site RNF2-5 is shown. Compared with WTAaCas12b, AaCas12b variants with E636R, I757R, E758R, E761R, Q854R, D858R, E758K, I994R or N857K or D858K substitutions showed the most improved gene editing efficiency.
[0033] Figure 5 The gene editing efficiency of exemplary AaCas12b variants is shown, in which the amino acid residues in the wild-type AaCas12b that are located in the RuvC domain and interact with the single-stranded DNA substrate are replaced by hydrophobic amino acid residues W, Y, F, or M. Compared with WTAaCas12b, AaCas12b variants with N865W, N865Y, Q866M, Q869M, Q1093W, or Q1093Y substitutions showed the most improved gene editing efficiency.
[0034] Figure 6Shown are the gene editing efficiencies of exemplary AaCas12b variants with combined mutations compared to WTAaCas12b.
[0035] Figure 7 It was shown that the AaCas12b variant Q119F+E475R+E758R had significantly improved gene editing efficiency compared to WTAaCas12b and the corresponding single mutant.
[0036] Figure 8 The amino acid sequence alignment of Cas12b proteins is shown, including AlicyclobacillusacidiphilusCas12b (AaCas12b) (SEQ ID NO: 1), AlicyclobacillusskakegawensisCas12b (AkCas12b) (SEQ ID NO: 54), AlicyclobacillusmacrosporangiidusCas12b (AmCas12b) (SEQ ID NO: 55), Bacillus sp. V3-13Cas12b (Bs3Cas12b) (SEQ ID NO: 56), Bacillus Cas12b (BsCas12b) (SEQ ID NO: 57), Laceyella sediminis Cas12b (LsCas12b) (SEQ ID NO: 58), Bacillushi sashii Cas12b (BhCas12b) (SEQ ID NO: 59), and Bacillus sp. V3-13Cas12b (Bs3Cas12b) (SEQ ID NO: 60). NO: 59), and Spirochaetes bacterium Cas12b (SbCas12b) (SEQ ID NO: 60). The substitutions based on AaCas12b described herein can be performed at the corresponding amino acid positions of any of the Cas12b orthologs described herein.
[0037] Figure 9 Figure 3: sgRNAs with engineered scaffolds significantly improved the gene editing efficiency of the AaCas12b variant Q119F+E475R+E758R. sgRNAs with AaCas12b-Aa-sg scaffold or AacCas12b-sgRNA scaffold (V0) were used as controls.
[0038] Figure 10A Schematic diagram of an exemplary construct encoding the AaCas12b variant Q119F+E475R+E758R+D570A under the control of the CMV promoter and the sgRNA under the control of the U6 promoter. Figure 10BFigure 2 shows the measurement of nuclease activity of AaCas12b(Q119F+E475R+E758R) and AaCas12b(Q119F+E475R+E758R+D570A) expressed as T7EI assay results. sgRNA1 and sgRNA2 specifically recognize the target site in HBG1 / 2. A control sgRNA that does not target any sequence in HBG1 / 2 serves as a negative control.
[0039] Figure 11A Schematic diagram of exemplary constructs encoding the AaCas12b variants Q119F+E475R+E758R+D570A+E848A or Q119F+E475R+E758R+D570A+D977A under the control of the CMV promoter and sgRNA under the control of the U6 promoter. Figure 11B Figure 2 shows the measurement of nuclease activity of AaCas12b (Q119F + E475R + E758R), AaCas12c (Q119F + E475R + E758R + D570A + E848A), and AaCas12d (Q119F / E475R + E758R + D570A + D977A) mediated by sgRNA1 and sgRNA2 that specifically recognize the target site in HBG1 / 2, as expressed in T7EI assay results. A control sgRNA that does not target any sequence in HBG1 / 2 was used as a negative control.
[0040] Figure 12A Schematic diagram of an exemplary construct encoding AaCas12b (Q119F+E475R+E758R+D570A+D977A) fused to KRAB under the control of the CMV promoter and sgRNA under the control of the U6 promoter. Figure 12B Figure 3. Relative mRNA levels of mouse Nav1.7 in mouse N2a cells transfected with AaCas12b (Q119F + E475R + E758R + D570A + D977A)-KRAB fusion protein targeting different sites of the SCN9A gene by different sgRNAs. No sgRNA transfection was used as a control. DETAILED DESCRIPTION
[0041] The present application provides an engineered Cas12b nuclease with increased enzymatic activity (such as gene editing activity) by introducing one, two, or three types of mutations relative to a reference Cas12b nuclease. Also provided are engineered Cas12b nucleases or their effector proteins (e.g., dCas12b) with reduced or eliminated nuclease activity. Also provided are engineered guide RNAs (gRNAs) with engineered scaffold sequences that can increase Cas12b enzymatic activity (e.g., gene editing activity) when used together with Cas12b nucleases (wild type or engineered). Also provided are engineered Cas12b effector proteins, and methods using engineered Cas12b nucleases or engineered Cas12b effector proteins and / or engineered gRNAs.
[0042] I. Definition
[0043] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.
[0044] As used herein, the term "Cas12b protein" is used in its broadest sense and includes a parent or reference Cas12b protein (e.g., AaCas12b comprising SEQ ID NO: 1), derivatives or variants thereof (e.g., engineered Cas12b, dCas12b, or engineered Cas12b effector protein), and functional fragments, such as oligonucleotide-binding fragments thereof.
[0045] As used herein, "effector protein" refers to a protein having activities such as site-specific binding activity, single-stranded DNA cleavage or editing activity, double-stranded DNA cleavage or editing activity, single-stranded RNA cleavage or editing activity, or transcriptional regulatory activity.
[0046] As used herein, "guide RNA" and "gRNA" are used interchangeably in this article and refer to RNA that can form a complex with Cas12b nuclease or effector protein and target nucleic acid (for example, double-helical DNA).Guide RNA can include a single RNA molecule or two or more RNA molecules that associate with each other via the hybridization of the complementary regions in two or more RNA molecules.When used in combination with the Cas nuclease (such as Cas12b) guided by dual RNA, guide RNA includes crRNA and tracrRNA, or single guide RNA (sgRNA). "crRNA" or "CRISPR RNA" includes a guide sequence with enough complementarity to the target sequence of target nucleic acid (for example, double-helical DNA), which guides the sequence-specific binding of CRISPR complex to target nucleic acid. "tracrRNA" or "trans-activation CRISPR RNA" is partially complementary to crRNA and is base-paired with crRNA, and can play a role in the maturation of crRNA. "Single guide RNA" or "sgRNA" are engineered guide RNAs with crRNA and tracrRNA fused to each other in a single molecule.
[0047] As used herein, the term "CRISPR array" refers to a nucleic acid (e.g., DNA) fragment comprising CRISPR repeats and spacers, which starts from the first nucleotide of the first CRISPR repeat and ends at the last nucleotide in the last (terminal) CRISPR repeat. Typically, each spacer in a CRISPR array is located between two repeat regions. As used herein, the term "CRISPR repeats" or "CRISPR direct repeats" or "direct repeats" refers to a plurality of short direct repeat sequences that exhibit very little or no sequence variation in a CRISPR array. Suitably, VI direct repeats can form a stem-loop structure.
[0048] As used herein, "donor template nucleic acid" or "donor template" are used interchangeably to refer to a nucleic acid molecule that can be used by one or more cellular proteins to alter the structure of a target nucleic acid after the CRISPR enzyme described herein alters the target nucleic acid. In some instances, the donor template nucleic acid is a double-stranded nucleic acid. In some instances, the donor template nucleic acid is a single-stranded nucleic acid. In some examples, the donor template nucleic acid is linear. In some instances, the donor template nucleic acid is circular (e.g., a plasmid). In some instances, the donor template nucleic acid is an exogenous nucleic acid molecule. In some instances, the donor template nucleic acid is an endogenous nucleic acid molecule (e.g., a chromosome).
[0049] The terms "nucleic acid," "polynucleotide," and "nucleotide sequence" are used interchangeably to refer to a polymeric form of nucleotides of any length, including deoxyribonucleotides, ribonucleotides, combinations thereof, and analogs thereof. "Oligonucleotide" and "oligonucleotide" are used interchangeably to refer to short polynucleotides having no more than about 50 nucleotides.
[0050] As used herein, "complementarity" refers to the ability of nucleic acid to form hydrogen bonds with another nucleic acid by traditional Watson-Crick base pairing. The percentage of residues that can form hydrogen bonds (i.e., Watson-Crick base pairing) with the second nucleic acid in complementarity percentage indication nucleic acid molecule (e.g., about 5, 6, 7, 8, 9, 10 / 10, respectively about 50%, 60%, 70%, 80%, 90%, and 100% complementarity). "Complete complementarity" means that all continuous residues of a nucleic acid sequence form hydrogen bonds with the same number of continuous residues in the second nucleic acid sequence. As used herein, "substantially complementary" refers to at least about 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% complementarity in the region of about 40, 50, 60, 70, 80, 100, 150, 200, 250 or more nucleotides, or refers to two nucleic acids that hybridize under stringent conditions.
[0051] As used herein, " stringent conditions " for hybridization refer to conditions under which the nucleic acid with complementarity to the target sequence mainly hybridizes with the target sequence, and substantially does not hybridize with non-target sequences. Stringent conditions are typically sequence-dependent and vary according to many factors. Generally, the longer the sequence, the higher the temperature at which the sequence and its target sequence-specific hybridization occur. Non-limiting examples of stringent conditions are described in detail in Tijssen (1993) " Laboratory Techniques in Biochemistry and Molecular Biology - Hybridization With Nucleic Acid Probes " Part I, Chapter II " Overview of principles of hybridization and the strategy of nucleic acid probe assay " Elsevier, New York.
[0052] "Hybridization" refers to the reaction of one or more polynucleotides to form a complex that is stabilized by hydrogen bonding between the bases of the nucleotide residues. Hydrogen bonding can occur by Watson-Crick base pairing, Hoogstein binding, or in any other sequence-specific manner. A sequence that is capable of hybridizing to a given sequence is referred to as the "complement of" the given sequence.
[0053] "Percent (%) sequence identity" with respect to a nucleic acid sequence is defined as the percentage of nucleotides in a candidate sequence that are identical to the nucleotides in a particular nucleic acid sequence, after aligning the sequences to achieve the maximum percent sequence identity by allowing gaps (if necessary). "Percent (%) sequence identity" with respect to a peptide, polypeptide, or protein sequence is the percentage of amino acid residues in a candidate sequence that are identically substituted with the amino acid residues in a particular peptide or amino acid sequence, after aligning the sequences to achieve the maximum percent sequence homology by allowing gaps (if necessary). Alignment for the purpose of determining percent amino acid sequence identity can be achieved in various ways known to those skilled in the art, for example, using publicly available computer software such as BLAST, BLAST-2, ALIGN, or MEGALIGN. TM (DNASTAR) software. Those skilled in the art can determine appropriate parameters for measuring alignment, including any algorithms needed to achieve maximal alignment over the full length of the sequences being compared.
[0054] The terms "polypeptide" and "peptide" are used interchangeably herein to refer to polymers of amino acids of any length. A polymer can be linear or branched, it can contain modified amino acids, and it can be interrupted by non-amino acids. A protein can have one or more polypeptides. The term also encompasses amino acid polymers that have been modified; for example, by disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or any other manipulation, such as conjugation with a labeling component.
[0055] As used herein, "variant" is interpreted to mean a polynucleotide or polypeptide that is different from a reference polynucleotide or polypeptide but retains basic properties. Typical variants of a polynucleotide are different from another reference polynucleotide in nucleic acid sequence. Changes in the nucleic acid sequence of a variant may or may not change the amino acid sequence of a polypeptide encoded by a reference polynucleotide. As discussed below, nucleotide changes may result in amino acid substitutions, additions, deletions, fusions, and truncations in a polypeptide encoded by a reference sequence. Typical variants of a polypeptide are different from another reference polypeptide in amino acid sequence. Typically, the differences are limited so that the sequences of the reference polypeptide and the variant are very similar overall and are identical in many regions. Variant and reference polypeptide may differ in amino acid sequence by one or more substitutions, additions, deletions (in any combination). The substituted or inserted amino acid residues may or may not be amino acid residues encoded by the genetic code. A variant of a polynucleotide or polypeptide may be naturally occurring, such as an allelic variant, or it may be an unknown naturally occurring variant. Non-naturally occurring variants of polynucleotides and polypeptides may be made by mutagenesis techniques, by direct synthesis, and by other recombinant methods known to those skilled in the art.
[0056] As used herein, the term "wild type" has the meaning commonly understood by those skilled in the art to refer to the typical form of an organism, strain, gene, or characteristic as it occurs in nature, as distinguished from mutants or variants. It can be isolated from a natural source and not intentionally modified.
[0057] As used herein, the terms "non-naturally occurring" or "engineered" are used interchangeably and refer to the involvement of human effort. When these terms are used to describe a nucleic acid molecule or polypeptide, it means that the nucleic acid molecule or polypeptide is at least substantially free of at least one other component of its native or naturally found associate.
[0058] As used herein, the term "ortholog" or "ortholog" has the meaning commonly understood by those of ordinary skill in the art. As a further guide, an "ortholog" of a protein referred to herein refers to a protein belonging to a different class that performs the same or a similar function as the protein to which it is an ortholog.
[0059] As used herein, the term "identity" is used to refer to the sequence matching between two polypeptides or between two nucleic acids. When a position in the two sequences being compared is occupied by the same base or amino acid monomer subunit (for example, a position in each of the two DNA molecules is occupied by adenine, or each position in each of the two polypeptides is occupied by lysine), each molecule is identical at that position. The "percentage of identity" between two sequences is a function of the number of matching positions shared by the two sequences divided by the number of positions to be compared × 100. For example, if 6 of the 10 positions of the two sequences match, the two sequences have 60% identity. For example, the DNA sequences CTGACT and CA GGTT share 50% identity (3 matches in a total of 6 positions). Typically, two sequences are compared when aligned to produce maximum identity. This comparison can be achieved by, for example, the method of Needleman et al. (1970) "Journal of Molecular Biology (J.Mol.Biol.)" 48:443-453, which can be easily performed by a computer program (such as Align program (DNAstar company)). The percent identity between two amino acid sequences can also be determined using the algorithm of E.Meyers and W.Miller (Comput. Appl Biosci. 4:11-17 (1988), which is incorporated into the ALIGN program (version 2.0). The PAM120 weighted residual table is used. A gap length penalty of 12 and a gap penalty of 4 are used to determine the percent identity between two amino acid sequences. Additionally, the Needleman-Wunsch (J. Mol. Biol. 48:444-453 (1970), which is incorporated into the GAP program of the GCG software package (available from www.gcg.com), can be used using the Blossum62 matrix or the PAM250 matrix and a gap weight of 16, 14, 12, 10, 8, 6, or 4 and a length weight of 1, 2, 3, 4, 5, or 6 to determine the percent identity between two amino acid sequences.
[0060] As used herein, "cell" should be understood to refer not only to a particular individual cell, but also to the progeny or potential progeny of a cell. Because certain modifications may occur in subsequent generations due to mutation or environmental influences, such progeny may not actually be identical to the parent cell, but are still included within the scope of the term as used herein.
[0061] As used herein, the terms "transduction" and "transfection" include all methods known in the art that use infectious agents (such as viruses) or other methods to introduce DNA into cells to express a protein or molecule of interest. In addition to viruses or virus-like agents, there are chemical-based transfection methods such as those using calcium phosphate, dendrimers, liposomes, or cationic polymers (e.g., DEAE-dextran or polyethyleneimine); non-chemical methods such as electroporation, cell squeezing, sonoporation, optical transfection, transfection by penetration, protoplast fusion, plasmid delivery, or transposons; particle-based methods such as using a gene gun, magnetofection, or magnet-assisted transfection, particle bombardment; and hybrid methods such as nuclear transfection.
[0062] As used herein, the term "transfected" or "transformed" or "transduced" refers to a process in which exogenous nucleic acid is transferred or introduced into a host cell. A "transfected" or "transformed" or "transduced" cell is a cell that has been transfected, transformed, or transduced with exogenous nucleic acid.
[0063] The term "in vivo" refers to inside the body of the organism from which the cell is obtained. "Ex vivo" or "in vitro" means outside the body of the organism from which the cell is obtained.
[0064] As used herein, " treatment (treatment) " or " control (treating) " is the mode of obtaining useful or desired result (comprising clinical outcome).For the purpose of the application, useful or desired clinical outcome includes but is not limited to one or more of the following: alleviate one or more symptoms caused by disease, reduce the degree of disease, stabilizing disease (for example, preventing or delaying the deterioration of disease), preventing or delaying the diffusion of disease (for example, metastasis), preventing or delaying the recurrence of disease, reducing the recurrence rate of disease, delaying or slowing down the progress of disease, improving morbid state, providing the alleviation (partial or complete) of disease, reducing the dosage of one or more other medicines needed for treating disease, delaying the progress of disease, improving quality of life, and / or prolonging survival." treatment " also encompasses the pathological consequences that reduce cancer. The method of the application contemplates any one or more of these treatment aspects.
[0065] As used herein, the term "effective amount" refers to an amount of a compound or composition sufficient to treat a specified condition, disorder, or disease (such as to improve, alleviate, relieve, and / or delay one or more of its symptoms). As understood in the art, an "effective amount" can be in one or more doses, i.e., a single dose or multiple doses may be required to achieve the desired therapeutic endpoint.
[0066] For the purposes of treatment, "subject," "individual," or "patient" are used interchangeably herein and refer to any animal, such as mammals (including humans, livestock and farm animals, as well as zoo, sport, or pet animals, such as dogs, horses, cats, cows, etc.), birds, reptiles, fish, etc. In some embodiments, the individual is a human individual.
[0067] It should be understood that the embodiments of the present application described herein include "consisting of" and / or "consisting essentially of" embodiments.
[0068] Reference herein to "about" a value or parameter includes (and describes) variations with respect to that value or parameter itself. For example, description of "about X" includes description of "X."
[0069] As used herein, reference to "not" a value or parameter generally means and describes "except" a value or parameter. For example, a method is not used to treat type X cancer, which means that the method is used to treat types of cancer other than X.
[0070] As used herein, the term "about XY" has the same meaning as "about X to about Y."
[0071] As used herein and in the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. It should also be noted that the claims may be drafted to exclude any optional elements. As such, this statement is intended to serve as antecedent basis for using exclusive terminology such as "solely," "only," and the like in connection with the recitation of claim elements, or using a "negative" limitation.
[0072] As used herein, the term "and / or," phrases such as "A and / or B," are intended to include both A and B; A or B; A (alone); and B (alone). Likewise, as used herein, the term "and / or," phrases such as "A, B, and / or C," are intended to encompass each of the following embodiments: A, B, and C; A, B, or C; A or C; A or B; B or C; A and C; A and B; B and C; A (alone); B (alone); and C (alone).
[0073] One of ordinary skill in the art will understand that both uracil and thymine can be represented by "t", rather than "u" for uracil and "t" for thymine; in the context of RNA, "t" is used to represent uracil unless otherwise indicated.
[0074] II. Cas12b nuclease and effector proteins
[0075] The application provides engineered Cas12b nucleases and effector proteins with improved activity (such as target binding activity, double-stranded cleavage activity, nickase activity, and / or gene editing activity). Also provided are engineered Cas12b nucleases (dCas12b) with reduced or eliminated nuclease activity. In some embodiments, there is provided an engineered Cas12b effector protein (e.g., Cas12b nuclease, Cas12b nickase, Cas12b fusion effector protein, or split Cas12b effector protein), the engineered Cas12b effector protein comprising any one of the engineered Cas12b nucleases described herein or its functional derivatives.
[0076] Engineered Cas12b nuclease
[0077] In one aspect, the present application provides engineered Cas12b effector proteins with improved activity (e.g., target binding activity, double-stranded cleavage activity, nickase activity, and / or gene editing activity).
[0078] In some embodiments, an engineered Cas12b nuclease is provided, comprising one, two, or three types of mutations relative to a reference Cas12b nuclease, wherein the mutation comprises: (1) replacing one or more amino acid residues that interact with a protospacer sequence adjacent motif (PAM) in a reference Cas12b nuclease with a positively charged amino acid residue (e.g., R, H, K); and / or (2) replacing one or more amino acid residues that participate in opening a double-stranded DNA (dsDNA) in a reference Cas12b nuclease with an amino acid residue having an aromatic ring (e.g., F, Y, W); and / or (3) replacing one or more amino acid residues that interact with a single-stranded DNA substrate in the RuvC domain of a reference Cas12b nuclease with a positively charged amino acid residue (e.g., R, H, K) or a hydrophobic amino acid residue (e.g., F, Y, W, M). In some embodiments, the reference Cas12b nuclease is a naturally occurring wild-type Cas12b nuclease. In some embodiments, the reference Cas12b nuclease is a natural variant Cas12b nuclease. In some embodiments, the reference Cas12b nuclease is a Cas12b nuclease from Alicyclobacillus acidiphilus (AaCas12b). In some embodiments, the reference Cas12b nuclease comprises the amino acid sequence of SEQ ID NO: 1. In some embodiments, the engineered Cas12b nuclease has an increased (e.g., an increase of at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 1-fold, 1.2-fold, 1.5-fold, 2-fold, 5-fold, 10-fold, 20-fold, 50-fold, 100-fold or more) activity (e.g., target binding, double-stranded cutting activity, nickase activity and / or gene editing activity) compared to the reference Cas12b nuclease.
[0079] The engineered Cas12b nuclease may comprise one or more of the mutations described below in Sections A to C below. In some embodiments, one or more of the mutations in the present application may be combined with any of the known Cas12b mutations (such as the mutations described in Section D below) to generate an engineered Cas12b nuclease with improved activity.
[0080] In some embodiments, there is provided an engineered Cas12b nuclease comprising one or more mutations relative to a reference Cas12b nuclease, wherein the one or more mutations comprise replacing one or more amino acid residues interacting with (PAM) in a reference Cas12b nuclease with a positively charged amino acid residue (e.g., R or K). In some embodiments, there is provided an engineered Cas12b nuclease comprising one or more mutations relative to a reference Cas12b nuclease, wherein the one or more mutations comprise replacing one or more amino acid residues participating in opening a double-stranded DNA in a reference Cas12b nuclease with an amino acid residue having an aromatic ring (e.g., W, Y, or F). In some embodiments, there is provided an engineered Cas12b nuclease comprising one or more mutations relative to a reference Cas12b nuclease, wherein the one or more mutations comprise replacing one or more amino acid residues interacting with a single-stranded DNA substrate in a RuvC domain of a reference Cas12b nuclease with a positively charged amino acid residue (e.g., R or K). In some embodiments, an engineered Cas12b nuclease comprising one or more mutations relative to a reference Cas12b nuclease is provided, wherein the one or more mutations comprise replacing one or more amino acid residues in the RuvC domain of the reference Cas12b nuclease that interact with a single-stranded DNA substrate with a hydrophobic amino acid residue (e.g., W, Y, F, or M). In some embodiments, the reference Cas12b nuclease comprises the amino acid sequence of SEQ ID NO: 1.
[0081] In some embodiments, an engineered Cas12b nuclease comprising one or more mutations relative to a reference Cas12b nuclease is provided, wherein the one or more mutations comprise: 1) replacing one or more amino acid residues that interact with a PAM in a reference Cas12b nuclease with a positively charged amino acid residue (e.g., R, H, K), and 2) replacing one or more amino acid residues involved in opening a double-stranded DNA in a reference Cas12b nuclease with an amino acid residue having an aromatic ring (e.g., F, Y, W). In some embodiments, an engineered Cas12b nuclease comprising one or more mutations relative to a reference Cas12b nuclease is provided, wherein the one or more mutations comprise: 1) replacing one or more amino acid residues that interact with a PAM in a reference Cas12b nuclease with a positively charged amino acid residue (e.g., R, H, K), and 2) replacing one or more amino acid residues that interact with a single-stranded DNA substrate in the RuvC domain of a reference Cas12b nuclease with a positively charged amino acid residue (e.g., R, H, K) or a hydrophobic amino acid residue (e.g., F, Y, W, M). In some embodiments, an engineered Cas12b nuclease comprising one or more mutations relative to a reference Cas12b nuclease is provided, wherein the one or more mutations comprise: 1) replacing one or more amino acid residues involved in opening the double-stranded DNA in the reference Cas12b nuclease with an amino acid residue having an aromatic ring (e.g., F, Y, W), and 2) replacing one or more amino acid residues that interact with a single-stranded DNA substrate in the RuvC domain of the reference Cas12b nuclease with a positively charged amino acid residue (e.g., R, H, K) or a hydrophobic amino acid residue (e.g., F, Y, W, M). In some embodiments, the reference Cas12b nuclease comprises the amino acid sequence of SEQ ID NO: 1.
[0082] In some embodiments, an engineered Cas12b nuclease comprising one or more mutations relative to a reference Cas12b nuclease is provided, wherein the one or more mutations comprise: 1) replacing one or more amino acid residues that interact with a PAM in the reference Cas12b nuclease with a positively charged amino acid residue (e.g., R, H, K), 2) replacing one or more amino acid residues involved in opening a double-stranded DNA in the reference Cas12b nuclease with an amino acid residue having an aromatic ring (e.g., F, Y, W), and 3) replacing one or more amino acid residues that interact with a single-stranded DNA substrate in the RuvC domain of the reference Cas12b nuclease with a positively charged amino acid residue (e.g., R, H, K). In some embodiments, the reference Cas12b nuclease comprises the amino acid sequence of SEQ ID NO: 1.
[0083] In some embodiments, an engineered Cas12b nuclease comprising one or more mutations relative to a reference Cas12b nuclease is provided, wherein the one or more mutations comprise: 1) replacing one or more amino acid residues that interact with a PAM in the reference Cas12b nuclease with a positively charged amino acid residue (e.g., R, H, K), 2) replacing one or more amino acid residues involved in opening a double-stranded DNA in the reference Cas12b nuclease with an amino acid residue having an aromatic ring (e.g., F, Y, W), and 3) replacing one or more amino acid residues that interact with a single-stranded DNA substrate in the RuvC domain of the reference Cas12b nuclease with a hydrophobic amino acid residue (e.g., F, Y, W, M). In some embodiments, the reference Cas12b nuclease comprises the amino acid sequence of SEQ ID NO: 1.
[0084] Mutation as described herein can be designed based on the structure of reference Cas12b nuclease.The crystal structure of the acid soil alicyclobacillus (Alicyclobacillus cidoterrestris) Cas12b that is bound to sgRNA as a binary complex and is bound to target DNA as a ternary complex has been described in Yang H. et al.《Cell (Cell)》167:1814-1828(2016) and Liu L. et al.《Molecular Cell (Mol.Cell)》65:310-322(2017).In short, the crystal structure shows 2 discontinuous REC (recognition, residues 15 to 386, 658 to 783) and NUC (nuclease, residues 1 to 14, 387 to 658 and 784 to 1129) leaves, each of which is composed of several domains.crRNA (or single guide RNA, sgRNA) is incorporated into the central channel between the two leaves. PAM recognition is sequence-specific and occurs primarily through interactions with the REC1 (helix-1) and WED-II (OBD-II) domains. The sgRNA-target DNA heteroduplex binds primarily to the REC lobe in a sequence-independent manner.
[0085] It will be understood that other Cas12b orthologs, such as BhCas12b (SEQ ID NO: 59), Bs3Cas12b (SEQ ID NO: 56), LsCas12b (SEQ ID NO: 58), SbCas12b (SEQ ID NO: 60), AkCas12b (SEQ ID NO: 54), AmCas12b (SEQ ID NO: 55), BsCas12b (SEQ ID NO: 57), and DiCas12b, etc., have domain structures similar to AaCas12b (SEQ ID NO: 1) and other exemplary reference Cas12b proteins described herein, and engineered Cas12b proteins can be designed based on any of the orthologs using the cleavage positions corresponding to the exemplary engineered AaCas12b proteins described herein. When the amino acid sequences of two polypeptides are aligned with each other, the corresponding positions refer to the positions in the two polypeptides that are aligned with each other. See the present application. Figure 8 . Moreover, Figure S2 in Teng F. et al. Cell Discovery (2019) 5:23 provides an alignment of AaCas12b, AkCas12b, AmCas12b, Bs3Cas12b, BsCas12b, LsCas12b, BhCas12b, and SbCas12b, which is incorporated herein by reference in its entirety.
[0086] A. Replace one or more amino acid residues that interact with PAM in the reference Cas12b with positively charged amino acid residues.
[0087] In some embodiments, the engineered Cas12b nuclease comprises replacing one or more amino acid residues that interact with (PAM) in the reference Cas12b nuclease with a positively charged amino acid residue (e.g., R, H, K). In some embodiments, the engineered Cas12b nuclease comprises one, two, three, four, five, or six amino acid substitutions.
[0088] In some embodiments, the one or more amino acid residues that interact with the PAM in the reference Cas12b nuclease are amino acids that are within 15 (e.g., 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, 1, or less) angstroms of the PAM in the three-dimensional structure. In some embodiments, the one or more amino acid residues that interact with the PAM in the reference Cas12b nuclease are amino acids that are within 10 angstroms of the PAM in the three-dimensional structure. In some embodiments, the one or more amino acid residues that interact with the PAM in the reference Cas12b nuclease are amino acids that are within 9 angstroms of the PAM in the three-dimensional structure. In some embodiments, the one or more amino acid residues that interact with the PAM are located at one or more of the following positions: 116, 123, 130, 132, 144, 145, 153, 173, 222, 395, 400, and 475. In some embodiments, the one or more amino acid residues that interact with a PAM comprise one or more of the following amino acid residues: D116, K123, D130, D132, N144, K145, E153, D173, Q222, D395, N400, and E475. In some embodiments, the one or more amino acid residues that interact with a PAM comprise one or more of the following amino acid residues: D116 and E475. In some embodiments, the amino acid residues are numbered according to SEQ ID NO: 1.
[0089] In the context of this application, D116 refers to the 116th amino acid D (aspartic acid) in the referenced amino acid sequence. The three-letter and one-letter abbreviations of commonly used amino acids are listed below:
[0090] Ala(A) Leu(L) Gln(Q) Ser(S) Arg(R) Lys(K) Glu(E) Thr(T) Asn(N) Met(M) Gly(G) Trp(W) Asp(D) Phe(F) His(H) Tyr(Y) Cys(C) Pro(P) Ile(I) Val(V)
[0091] As used herein, "amino acid at position X, wherein the amino acids are numbered according to SEQ ID NO: 1" refers to an amino acid residue located at a position of the reference enzyme Cas12b that corresponds to position X in SEQ ID NO: 1 when the amino acid sequences of the reference enzyme Cas12b and SEQ ID NO: 1 are aligned based on sequence homology. For example, Figure 8 Homology alignments of the amino acid sequences of Cas12b orthologs (SEQ ID NO: 1 and SEQ ID NO: 54 to SEQ ID NO: 60) are shown. One skilled in the art can readily use known software (such as Clustal Omega) to compare and align the amino acid sequence of any reference Cas12b nuclease with SEQ ID NO: 1 to determine the amino acid position corresponding to position X in SEQ ID NO: 1.
[0092] In some embodiments, the positively charged amino acid residue is R, H, or K. In some embodiments, the positively charged amino acid residue is R. In some embodiments, the positively charged amino acid residue is K.
[0093] In some embodiments, the substitution of one or more amino acid residues in the reference Cas12b nuclease that interacts with the PAM with a positively charged amino acid residue is one or more of the following substitutions: D116R, K123R, D130R, D132R, N144R, K145R, E153R, D173R, Q222R, D395R, N400R, and E475R. In some embodiments, the substitution of one or more amino acid residues in the reference Cas12b nuclease that interacts with the PAM with a positively charged amino acid residue is one or more of the following substitutions: D116R and E475R. In some embodiments, the engineered Cas12b nuclease comprises a D116R mutation. In some embodiments, the engineered Cas12b nuclease comprises an E475R mutation. In some embodiments, the amino acid residues are numbered according to SEQ ID NO: 1.
[0094] In some embodiments, the engineered Cas12b nuclease comprises an amino acid sequence having at least about 85% sequence identity (such as at least about 87%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or any one of 99% sequence identity) to the amino acid sequence of SEQ ID NO: 2 or SEQ ID NO: 3. In some embodiments, the engineered Cas12b nuclease comprises an amino acid sequence of SEQ ID NO: 2 or SEQ ID NO: 3.
[0095] B. Replace one or more amino acid residues involved in opening the DNA double strand in the reference Cas12b nuclease with amino acid residues having an aromatic ring
[0096] In some embodiments, the engineered Cas12b nuclease comprises replacing one or more amino acid residues involved in opening the DNA double strand in the reference Cas12b nuclease with an amino acid residue (e.g., F, Y, W) having an aromatic ring. In some embodiments, the engineered Cas12b nuclease comprises the replacement of one, two, three, four, five, or six amino acid residues.
[0097] In some embodiments, the one or more amino acid residues involved in opening the DNA double strand interact with the last base pair in the PAM relative to the 3' end of the target strand. For example, the PAM sequence recognized by AaCas12b is a 5'-TTN-3' base pair. The last base pair in the PAM relative to the 3' end of the target strand is a base pair formed by the N base at the 3' end of the PAM sequence, followed by the sequence of the target site.
[0098] In some embodiments, the one or more amino acid residues involved in opening the DNA double strand are located at one or more of the following positions: 118, and / or 119, such as Q118 and / or Q119. In some embodiments, the amino acid residues are numbered according to SEQ ID NO: 1.
[0099] In some embodiments, the amino acid residue having an aromatic ring is Y, F, or W. In some embodiments, the amino acid residue involved in opening the double-stranded DNA is substituted with F, Y, or W. In some embodiments, the engineered Cas12b nuclease comprises any of the following: i) Q118Y, Q118F, or Q118W; and / or ii) Q119Y, Q119F, or Q119W. In some embodiments, the amino acid residue numbering is according to SEQ ID NO: 1.
[0100] In some embodiments, the substitution of one or more amino acid residues involved in opening the DNA double strand in the reference Cas12b nuclease with an amino acid having an aromatic ring is a Q119Y, Q119F or Q119W substitution. In some embodiments, the amino acid residues are numbered according to SEQ ID NO: 1.
[0101] In some embodiments, the engineered Cas12b nuclease comprises an amino acid sequence having at least about 85% sequence identity (such as at least about 88%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or any one of 99% sequence identity) to the amino acid sequence of SEQ ID NO:4, SEQ ID NO:5, or SEQ ID NO:6. In some embodiments, the engineered Cas12b nuclease comprises an amino acid sequence of SEQ ID NO:4, SEQ ID NO:5, or SEQ ID NO:6.
[0102] C. Substituting one or more amino acid residues in the RuvC domain of the reference Cas12b nuclease that interact with the single-stranded DNA substrate with a positively charged amino acid residue or a hydrophobic amino acid residue.
[0103] In some embodiments, the engineered Cas12b nuclease comprises replacing one or more amino acid residues in the reference Cas12b nuclease that are located in the RuvC domain and interact with a single-stranded DNA substrate with a positively charged amino acid residue (e.g., R, H, K). In some embodiments, the engineered Cas12b nuclease comprises replacing one or more amino acid residues in the reference Cas12b nuclease that are located in the RuvC domain and interact with a single-stranded DNA substrate with a hydrophobic amino acid residue (e.g., F, Y, W, M). In some embodiments, the engineered Cas12b nuclease comprises the replacement of one, two, three, four, five, or six amino acid residues.
[0104] In some embodiments, the distance between the one or more amino acid residues located in the RuvC domain and interacting with the single-stranded DNA substrate and the single-stranded DNA substrate in the three-dimensional structure is within 15 (e.g., 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, 1 or less) angstroms. In some embodiments, the one or more amino acid residues located in the RuvC domain and interacting with the single-stranded DNA substrate are within 10 angstroms of the single-stranded DNA substrate in the three-dimensional structure. In some embodiments, the one or more amino acid residues located in the RuvC domain and interacting with the single-stranded DNA substrate are within 9 angstroms of the single-stranded DNA substrate in the three-dimensional structure. The RuvC domain is the active domain of the Cas12b protein responsible for segmenting single-stranded DNA or double-stranded DNA. In the primary sequence of the protein, the RuvC domain includes the first RuvC domain (RuvC-1), the second RuvC domain (RuvC-II), and the third RuvC domain (RuvC-III).
[0105] In some embodiments, the one or more amino acid residues located in the RuvC domain and interacting with the single-stranded DNA substrate are located at one or more of the following positions: 300, 301, 304, 329, 636, 639, 647, 682, 757, 758, 761, 764, 768, 852, 854, 856, 857, 858, 860, 862, 863, 865, 866, 867, 869, 938, 956, 957, 958, 994, 1093, and 1097. In some embodiments, the one or more amino acid residues located in the RuvC domain and that interact with the single-stranded DNA substrate include one or more of the following amino acid residues: D300, K301, E304, N329, E636, Q639, T647, Q682, 1757, E758, E761, E764, K768, E852, Q854, N856, N857, D858, P860, S862, E863, N865, Q866, L867, Q869, E938, E956, G957, E958, 1994, Q1093, and W1097. In some embodiments, the one or more amino acid residues in the RuvC domain that interact with a single-stranded DNA substrate include one or more of the following amino acid residues: D300, K301, E636, Q639, T647, Q682, I757, E758, E761, K768, Q854, N857, D858, N865, Q866, Q869, I994, Q1093, and W1097. In some embodiments, the one or more amino acid residues in the RuvC domain that interact with a single-stranded DNA substrate include one or more of the following amino acid residues: E636, I757, E758, E761, Q854, N857, D858, N865, Q866, Q869, and Q1093. In some embodiments, the amino acid residues are numbered according to SEQ ID NO: 1.
[0106] In some embodiments, the engineered Cas12b nuclease comprises replacing one or more amino acid residues in the RuvC domain of the reference Cas12b nuclease and interacting with the single-stranded DNA substrate with a positively charged amino acid residue (e.g., R, H, K). In some embodiments, the positively charged amino acid residue is R. In some embodiments, the positively charged amino acid residue is K. In some embodiments, the engineered Cas12b nuclease comprises one or more of the following substitutions: D300R, K301R, E304R, N329R, E636R, Q639R, T647R, Q682R, I757R, E758R, E761R, E764R, K768R, E852R, Q854R, N856R, N857R, D858R, P860R, S862R, E863R, N865R , Q866R, L867R, Q869R, E938R, E956R, G957R, E958R, I994R, Q1093R, W109R7R, E636K, Q639K, T647K, Q682K, I757K, E758K, E761K, Q854K, N857K, D858K, N865K, Q866K, I994K, Q1093K and W1097K, wherein the amino acid residues are numbered according to SEQ ID NO: 1. In some embodiments, the engineered Cas12b nuclease comprises one or more of the following substitutions: D300R, K301R, E636R, Q639R, T647R, Q682R, I757R, E758R, E761R, K768R, Q854R, N857R, D858R, N865R, Q866R, I994R, Q1093R, W1097R, E635K, Q639K, T647K, Q682K, I757K, E758K, E761K, Q854K, N857K, D858K, N865K, I994K, Q1093K, and W1097K, wherein the amino acid residues are numbered according to SEQ ID NO: 1. In some embodiments, the engineered Cas12b nuclease comprises one or more of the following substitutions: E636R, I757R, E758R, E761R, Q854R, D858R, E636K, I757K, E758K, E761K, Q854K, N857K, and D858K, wherein the amino acid residues are numbered according to SEQ ID NO: 1. In some embodiments, the engineered Cas12b nuclease comprises E636R, I757R, E758R, E761R, Q854R, and D857R, wherein the amino acid residues are numbered according to SEQ ID NO: 1.In some embodiments, the engineered Cas12b nuclease comprises one or more of the following substitutions: E636K, I757K, E758K, E761K, Q854K, N857K, and D858K, wherein the amino acid residues are numbered according to SEQ ID NO: 1.
[0107] In some embodiments, the substitution of one or more amino acid residues in the reference Cas12b nuclease that are located in the RuvC domain and interact with the single-stranded DNA substrate is one or more of the following substitutions: E636R, I757R, E758R, E761R, Q854R, N857K, and D858R, wherein the amino acid residues are numbered according to SEQ ID NO: 1. In some embodiments, the engineered Cas12b nuclease comprises an amino acid sequence having at least about 85% sequence identity to the amino acid sequence of any one of SEQ ID NOs: 7 to 13, for example, at least about 88%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity. In some embodiments, the engineered Cas12b nuclease comprises an amino acid sequence of any one of SEQ ID NOs: 7 to 13.
[0108] In some embodiments, the engineered Cas12b nuclease comprises replacing one or more amino acid residues in the RuvC domain of the reference Cas12b nuclease and interacting with the single-stranded DNA substrate with a hydrophobic amino acid residue. In some embodiments, the hydrophobic amino acid residue is A, M, L, I, V, C, Y, F, or W. In some embodiments, the hydrophobic amino acid residue is W, Y, F, or M. In some embodiments, the hydrophobic amino acid residue is W, Y, or M. In some embodiments, the engineered Cas12b nuclease comprises one or more of the following substitutions: i) E758W, E758Y, E758F, or E758M, ii) E761W, E761Y, E761F, or E761M, iii) E863W, E863Y, E863F, or E863M, iv) N865W, N865Y, N865F, or N865M, v) Q866W, Q866F, Q866Y, or Q866M, vi) Q869W, Q869Y, Q869F, or Q869M, vii) E956W, E956Y, E956MF, or E956M, and viii) Q1093W, Q1093F, Q1093Y, or Q1093M; wherein the amino acid residues are numbered according to SEQ ID NO: 1. In some embodiments, the engineered Cas12b nuclease comprises one or more of the following substitutions: i) E758W, E758Y, or E758M, ii) E761Y, iii) N865W, N865F, or N865Y, iv) Q866M, v) Q869M, and vi) Q1093W, Q1093F, Q1093Y, or Q1093M; wherein the amino acid residues are numbered according to SEQ ID NO: 1. In some embodiments, the engineered Cas12b nuclease comprises one or more of the following substitutions: i) N865W or N865Y, ii) Q866M, iii) Q869M, and iv) Q1093W or Q1093Y; wherein the amino acid residues are numbered according to SEQ ID NO: 1.
[0109] In some embodiments, the engineered Cas12b nuclease comprises one or more of the following substitutions: 865W, 865Y, 866M, 869M, 1093W, and 1093Y. In some embodiments, the substitution of one or more amino acid residues in the reference Cas12b nuclease that are located in the RuvC domain and interact with the single-stranded DNA substrate is one or more of the following substitutions: N865W, N865Y, Q866M, Q869M, Q1093W, and / or Q1093Y. In some embodiments, the engineered Cas12b nuclease comprises Q866M and Q869M substitutions. In some embodiments, the amino acid residues are numbered according to SE Q ID NO: 1.
[0110] In some embodiments, the engineered Cas12b nuclease comprises an amino acid sequence having at least about 85% sequence identity (such as at least about 88%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or any one of 99% sequence identity) to the amino acid sequence of any one of SEQ ID NO: 7 to SEQ ID NO: 20. In some embodiments, the engineered Cas12b nuclease comprises an amino acid sequence of any one of SEQ ID NO: 14 to SEQ ID NO: 20.
[0111] D. Other mutations
[0112] Any one or more of the mutations described in sections A to C above can be combined with any one or more of the known mutations that increase Cas12b activity (such as target binding activity, target specificity, double-stranded cleavage activity, nickase activity, and / or gene editing activity). Exemplary mutations can be found in, for example, the following files WO2022120520, WO2022040909, WO2022042557, CN113308451A, and CN112195164A, the contents of which are incorporated herein by reference in their entirety.
[0113] In some embodiments, reference Cas12b protein comprises the following one or more from N end to C end: the first WED domain (WED-I), the first REC domain (REC1), the second WED structure (WED-II), the first RuvC domain (RuvC-I), the BH domain, the second REC domain (REC2), the second RuvC domain (RuvCII), the first Nuc domain (Nuc-I), the 3rd RuvC structural domain (RucIII) and the second Nuc domain (NucII). In some embodiments, other one or more mutations (e.g., insertion, deletion, replacement) can be present in one or more such domains.
[0114] In some embodiments, the engineered Cas12b nuclease further comprises one or more flexible region mutations that increase the flexibility of the flexible region in the reference Cas12b nuclease. The flexible region in the reference Cas12b nuclease can be determined using any method known in the art. In some embodiments, multiple flexible regions are determined based solely on the amino acid sequence of the reference enzyme. In some embodiments, multiple flexible regions are determined based on the structural information of the reference enzyme, including, for example, secondary structure, crystal structure, NMR structure, etc.
[0115] In some embodiments, multiple flexible partitions are determined using a program selected from the following group: PredyFlexy, FoldUnfold, PROFbval, Flexserv, FlexPred, DynaMine, and Disomine. In some embodiments, multiple flexible regions are located at random coils. In some embodiments, multiple flexible regions are in the DNA and / or RNA interaction domains of the reference Cas12b nuclease. In some embodiments, the length of the flexible region is at least about 5 (e.g., 5) amino acids.
[0116] In some embodiments, the engineered Cas12b nuclease comprises one or more mutations that increase the flexibility of the flexible region corresponding to amino acid residues 855 to 859, wherein the amino acid residue numbering is based on SEQ ID NO: 1, wherein compared to the reference Cas12b nuclease, the engineered Cas12b nuclease has an increase (e.g., an increase of at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 1 times, 1.2 times, 1.5 times, 2 times, 5 times, 10 times, 20 times, 50 times, 100 times or more) activity (e.g., target binding, double-stranded cleavage activity, nikase activity and / or gene editing activity). In some embodiments, the reference Cas12b nuclease is AaCas12b. In some embodiments, the reference Cas12b nuclease comprises the amino acid sequence of SEQ ID NO: 1. In some embodiments, the one or more mutations are included in the insertion of one or more (e.g., 2) G residues in the flexible region. In some embodiments, one or more G residues are inserted N-terminally to a flexible amino acid residue in a flexible region, wherein the flexible amino acid residue is selected from the group consisting of G, S, N, D, H, M, T, E, Q, K, R, A, and P. In some embodiments, the flexible amino acid residues are selected according to the following priority order: G>S>N>D>H>M>T>E>Q>K>R>A>P. In some embodiments, the one or more mutations comprise replacing a hydrophobic amino acid residue in a flexible region with a G group, wherein the hydrophobic amino acid residue is selected from the group consisting of L, I, V, C, Y, F, and W. In some embodiments, the one or more mutations that increase flexibility comprise N856G.
[0117] E. Mutation Combinations
[0118] Combinations of engineered enzymes obtained from the methods described in sections A to D of this specification with multiple amino acid substitutions in the exemplary sequence table are within the scope of this application. In some embodiments, the engineered Cas12b nuclease comprises one or more mutations (e.g., substitutions) described in sections AD above.
[0119] In some embodiments, the engineered Cas12b nuclease comprises a substitution or combination of substitutions at any of the following amino acid residue positions: (1) 116; (2) 475; (3) 119 and 475; (4) 119, 475, and 758; (5) 119; (6) 636; (7) 757; (8) 758; (9) 761; (10) 768; (11) 858; (12) 854; (13) 857; (14) 119, 475, and 758; (15) 768; (16) 757 and 758; (17) 75 7 and 761; (18) 757 and 768; (19) 758 and 761; (20) 758 and 768; (21) 761 and 768; (22) 757, 758 and 761; (23) 757, 758 and 768; (24) 757, 761 and 768; (25) 758, 761 and 768; (26) 757, 758, 761 and 768; (27) 865; (28) 866; (29) 869; (30) 1093; and (31) 866 and 869, wherein amino acid positions are numbered according to SEQ ID NO: 1.
[0120] In some embodiments, the engineered Cas12b nuclease comprises a substitution or combination of substitutions at any of the following amino acid residues: (1) D116; (2) E475; (3) Q119 and E475; (4) Q119, E475, and E758; (5) Q119; (6) E636; (7) I757; (8) E758; (9) E761; (10) K768; (11) D858; (12) Q854; (13) N857; (14) Q119, E475, and E758; (15) K768; (16) I757 and E758; (17) I757 and E761; (18) I757 and K768; (19) E758 and E761; (20) E758 and K768; (21) E761 and K768; (22) I757, E758, and E761; (23) I757, E758, and K768; (24) I757, E761, and K768; (25) E758, E761, and K768; (26) I757, E758, E761, and K768; (27) N865; (28) Q866; (29) Q869; (30) Q1093; and (31) Q866 and Q869; wherein amino acid positions are numbered according to SEQ ID NO: 1. In some embodiments, the engineered Cas12b nuclease comprises a substitution or combination of substitutions at any one of the following amino acid residues: (1) Q866+Q869; (2) Q119+E475; (3) Q119+E475+E758; and wherein the amino acid residues are numbered according to SEQ ID NO: 1. In some embodiments, the substitution at amino acid position D116 and / or E475 is substituted with a positively charged amino acid residue (such as R or K). In some embodiments, the substitution at amino acid position Q119 is substituted with an amino acid residue having an aromatic side chain (such as Y, F, or W). In some embodiments, the substitution at amino acid position E636, I757, E758, E761, K768, Q854, D858, and / or N857 is substituted with a positively charged amino acid residue (such as R or K). In some embodiments, the substitution at amino acid position N865, Q866, Q869, and / or Q1093 is with a hydrophobic amino acid residue (such as W, Y, or M).
[0121] In some embodiments, the engineered Cas12b nuclease comprises any of the following amino acid residues or combinations thereof: (1) 116R; (2) 475R; (3) 119F and 475R; (4) 119F, 475R, and 758R; (5) 119Y; (6) 119F; (7) 119W; (8) 636R; (9) 757R; (10) 758R; (11) 761R; (12) 854R; (13) 857K; (14) 768R; (15) 757R and 758R; (16) 757R and 761R; (17) 757R and 768R; (18) 758R and 761 R; (19) 758R and 768R; (20) 761R and 768R; (21) 757R, 758R and 761R; (22) 757R, 758R and 768R; (23) 757R, 761R and 768R; (24) 758R, 761R and 768R; (25) 757R, 758R, 761R and 768R; (26) 865W; (27) 865Y; (28) 866M; (29) 869M; (30) 1093W; (31) 1093Y; (32) 866M and 869M; and (33) 858R; wherein amino acid positions are numbered according to SEQ ID NO: 1.
[0122] In some embodiments, the engineered Cas12b nuclease comprises any of the following substitutions or combinations thereof: (1) D116R; (2) E475R; (3) Q119F+E475R; (4) Q119F+E475R+E758R; (5) Q119Y; (6) Q119F; (7) Q119W; (8) I757R; (9) E758R; (10) E761R; (11) K768R; (12) I757R+E758R; (13) I757R+E761R; (14) I757R+K768R; (15) E758R+E761R; (16) E758R+K768R; (17) E761R+K768R 8R; (18) I757R + E758R + E761R; (19) I757R + E758R + K768R; (20) I757R + E761R + K768R; (21) E758R + E761R + K768R; (22) I757R + E758R + E761R + K768R; (23) Q866M; (24) Q869M; (25) Q866M + Q869M; (26) E636R; (27) Q854R; (28) N857K; (29) N865W; (30) N865Y; (31) Q1093W; (32) Q1093Y; and (33) D858R; wherein the amino acid positions are according to SEQ ID NO: 1. In some embodiments, the engineered Cas12b nuclease comprises any one or a combination of the following substitutions: (1) Q866M+Q869M; (2) Q119F+E475R; (3) Q119F+E475R+E758R; and wherein the amino acid residues are numbered according to SEQ ID NO: 1.
[0123] In some embodiments, the engineered Cas12b nuclease comprises one or more of the following substitutions: D116R, K123R, D130R, D132R, N144R, K145R, E153R, D173R, Q222R, D395R, N400R, and E475R. In some embodiments, the engineered Cas12b nuclease comprises one or more of the following substitutions: Q118Y, Q118F, Q118W, Q119Y, Q119F, and Q119W. In some embodiments, the engineered Cas12b nuclease comprises one or more of the following substitutions: D300R, K301R, E304R, N329R, E636R, Q639R, T647R, Q682R, I757R, E758R, E761R, E764R, K768R, E852R, Q854R, N856R, N857R, D858R, P860R, S862R, E863R, N865R, Q866R, L867R, Q869R, E938R, E956R, G957R, E958R, I944R, Q1093R, and / or W1097R. In some embodiments, the engineered Cas12b nuclease comprises one or more of the following substitutions: E636K, Q639K, T647K, Q682K, I757K, E758K, E761K, Q854K, N857K, D858K, N865K, Q866K, I994K, Q1093K, and W1097K. In some embodiments, the engineered Cas12b nuclease comprises one or more of the following substitutions: E758W, E758Y, E758F, E758M, E761W, E761Y, E761F, E761M, E863W, E863Y, E863F, E863M, N865W, N865Y, N865F, N865M, Q866W, Q866Y, Q866F, Q866M, Q869W, Q869Y, Q869F, Q869M, E956W, E956Y, E956F, E956M, Q1093W, Q1093Y, Q1093F, and Q1093M. In some embodiments, amino acid positions are numbered according to SEQ ID NO: 1.
[0124] In some embodiments, the engineered Cas12b nuclease includes amino acid substitutions at Q866 and Q869. In some embodiments, the engineered Cas12b nuclease includes amino acid substitutions Q866M and Q869M. In some embodiments, the amino acid positions are numbered according to SEQ ID NO: 1. In some embodiments, the engineered Cas12b nuclease includes an amino acid sequence having at least about 85% sequence identity (such as at least about 88%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or any one of 99% sequence identity) to the amino acid sequence of SEQ ID NO: 20. In some embodiments, the engineered Cas12b nuclease includes an amino acid sequence having at least about 85% sequence identity (such as at least about 88%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or any one of 99% sequence identity) to the amino acid sequence of SEQ ID NO: 20.
[0125] In some embodiments, the engineered Cas12b nuclease includes amino acid substitutions at Q119 and E475. In some embodiments, the engineered Cas12b nuclease includes amino acid substitutions Q119F and E475R. In some embodiments, the amino acid positions are numbered according to SEQ ID NO: 1. In some embodiments, the engineered Cas12b nuclease includes an amino acid sequence having at least about 85% sequence identity (such as at least about 88%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or any one of 99% sequence identity) to the amino acid sequence of SEQ ID NO: 21. In some embodiments, the engineered Cas12b nuclease includes an amino acid sequence having at least about 85% sequence identity (such as at least about 88%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or any one of 99% sequence identity) to the amino acid sequence of SEQ ID NO: 21.
[0126] In some embodiments, the engineered Cas12b nuclease is included in amino acid substitutions at Q119, E475, and E758. In some embodiments, the engineered Cas12b nuclease includes amino acid substitutions Q119F, E475R, and E758R. In some embodiments, the amino acid positions are numbered according to SEQ ID NO: 1. In some embodiments, the engineered Cas12b nuclease includes an amino acid sequence having at least about 85% sequence identity (such as at least about 88%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or any one of 99% sequence identity) to the amino acid sequence of SEQ ID NO: 22. In some embodiments, the engineered Cas12b nuclease includes an amino acid sequence having at least about 85% sequence identity (such as at least about 88%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or any one of 99% sequence identity) to the amino acid sequence of SEQ ID NO: 22.
[0127] Reference Cas12b nuclease
[0128] In some embodiments, the reference Cas12 nuclease is AaCas12b or an ortholog thereof. In some embodiments, the reference Cas12b nuclease is a naturally occurring Cas12b nuclease. In some embodiments, the reference Cas12b nuclease is a wild-type Cas12b nuclease. In some embodiments, the reference Cas12b nuclease is an engineered Cas12b nuclease.
[0129] Cas12b nucleases from various organisms can be used as reference Cas12b nucleases to provide the engineered Cas12b nucleases and effector proteins of the present application. In some embodiments, the reference Cas12b nuclease has enzymatic activity. In some embodiments, the reference Cas12b is a nuclease that segments the double strands of a target double-helical nucleic acid (e.g., duplex DNA). In some embodiments, the reference Cas12b is a nickase that segments the single strands of a target double-helical nucleic acid (e.g., duplex DNA). In some embodiments, the reference Cas12b nuclease is an enzyme that is inactive (e.g., dCas12b). Orthologs having a certain sequence identity (e.g., at least about 60%, 70%, 80%, 85%, 90%, 95%, 98%, or higher) with Cas12b or its functional derivatives can be used as the basis for designing the engineered Cas12b nuclease or effector proteins of the present application. In some embodiments, the reference Cas12b nuclease is a mutant Cas12b, but does not include any mutation described in sections AE above.
[0130] In some embodiments, engineered Cas12b nucleases are functional variants based on naturally occurring Cas12b nucleases. In some embodiments, functional variants have one or more mutations, such as amino acid substitutions, insertions, and deletions. For example, compared to wild-type naturally occurring Cas12b nucleases, functional variants can include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more amino acid substitutions. In some embodiments, one or more substitutions are conservative substitutions. In some embodiments, functional variants have all domains of naturally occurring Cas12b nucleases. In some embodiments, functional variants do not have one or more domains of naturally occurring Cas12b nucleases.
[0131] VB type CRISPR-Cas12b (also referred to as C2c1) system has been identified as dual RNA guided (i.e. crRNA and tracrRNA) DNA endonuclease system, with different features from Cas9 and Cas12a (Shmakov, S. et al. "Molecular Cell" 60,385-397 (2015)). First, it is reported that when reconstructed with crRNA / tracrRNA double helix, Cas12b generates staggered ends away from (PAM) sites in vitro. Secondly, although the RuvC domain of Cas12b is similar to the RuvC domain of Cas9 and Cas12a, its putative Nuc domain does not have sequence or structural similarity with the HNH domain of Cas9 and the Nuc domain of Cas12a. In addition, the Cas12b protein is smaller than the most widely used SpCas9 and Cas12a (e.g., AacCas12b: 1,129 amino acids (aa); SpCas9: 1,369aa; AsCas12a: 1,353aa; LbCas12a: 1,228aa), which makes Cas12b suitable for adeno-associated virus (AAV)-mediated in vivo delivery in gene therapy. Compared to small Cas9 proteins (such as SaCas9 and CjCas9), Cas12b recognizes simpler PAM sequences (e.g., AacCas12b: 5′-TTN-3′; while SaCas9: 5′-NNGRRT-3′, CjCas9: 5′-NNNNRYAC-3′), which significantly increases the targeting range of Cas12b in the genome. Additionally, Cas12b has minimal off-target effects and can therefore be used as a safer option for therapeutic and clinical applications.
[0132] Cas12b (C2c1) nucleases from various organisms can be used as reference Cas12b nucleases to provide engineered Cas12b effector proteins of the present application. Exemplary Cas12b nucleases have been described in, for example, Shmakov, S. et al., Molecular Cell 60, 385-397 (2015); Shmakov, S. et al., Nat. Rev. Microbiol. 15, 169-182 (2017); WO2016205764, and WO2020 / 087631, the contents of which are incorporated herein by reference in their entirety.
[0133] In some embodiments, the engineered Cas12b effector protein is based on a reference Cas12b protein (e.g., a Cas12b nuclease) selected from the group consisting of: a Cas12b protein from Alicyclobacillus acidiphilus (AaCas12b), a Cas12b from Alicyclobacillus kakegawensis (AkCas12b), a Cas12b from Alicyclobacillus macrosporangiidus (AmCas12b), a Cas12b from Bacillus thuringiensis (BhCas12b), a BsCas12b from Bacillus sp., a Bs3Cas12b from Bacillus sp., a Cas12b from Desulfovibrio inopinatus (DiCas12b), a Cas12b from Leysella sedimentosa (LsCas12b), a Cas12b from a spirochete bacterium (SbCas12b), a Cas12b from Tuberibacillus thermophilus (Tuberibacillus sp. calidus) Cas12b (TcCas12b), and its functional derivatives. The sequences of naturally occurring Cas12b proteins are known in, for example, UniProtKB IDs: T0D7A2, A0A6I3SPI6, and A0A6I7FUC4, which are incorporated herein by reference in their entirety.
[0134] In some embodiments, the reference Cas12b protein is a Cas12b nuclease from Alicyclobacillus acidiphilus (AaCas12b) or a functional derivative thereof. In some embodiments, the engineered Cas12b effector protein is based on a reference Cas12b protein comprising an amino acid sequence having at least about 85% (e.g., at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or any one of 99%) sequence identity to the amino acid sequence of SEQ ID NO: 1. In some embodiments, the engineered Cas12b effector protein is based on a reference Cas12b nuclease comprising the amino acid sequence of SEQ ID NO: 1.
[0135] It should be noted that orthologs having a certain sequence identity (e.g., at least about 60%, 70%, 80%, 85%, 90%, 95%, 98%, or more) with a reference Cas12b protein or fragment thereof can be used as the basis for designing the engineered Cas12b effector protein of the present application. Those skilled in the art can determine the percentage of sequence identity of orthologs of Cas12b or its fragments suitable for the present application based on purpose and application. Methods for determining sequence identity values can be found in Lesk, A.M., ed., Computational Molecular Biology, Oxford University Press, New York, 1988; Smith, D.W., ed., Biocomputing: Informatics and Genome Projects, Academic Press, New York, 1993; Griffin, A.M. and Griffin, H.G., eds., Computer Analysis of Sequence Data, Part I, Humana Press, New Jersey, 1994; von Heinje, G., Sequence Analysis in Molecular Biology, Academic Press, 1987; and Sequence Analysis Primer, Gribskov, M. and Devereux, J., eds., Stockton Press, New York, 1991). Various Cas12b orthologs have been described in WO2020 / 087631, the contents of which are incorporated herein by reference in their entirety. In some embodiments, the engineered Cas12b effector protein is based on a reference Cas12b protein comprising an amino acid sequence having at least about 85% (e.g., at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity to the amino acid sequence of any one of SEQ ID NO: 54 to SEQ ID NO: 60.
[0136] Activity of engineered Cas12b
[0137] In some embodiments, the engineered Cas12b nuclease has increased activity compared to the reference Cas12b nuclease. In some embodiments, the activity is target DNA binding activity. In some embodiments, the activity is site-specific nuclease activity. In some embodiments, the activity is double-stranded DNA cleavage activity. In some embodiments, the activity is single-stranded DNA cleavage activity, including, for example, site-specific DNA cleavage activity or non-specific DNA cleavage activity. In some embodiments, the activity is single-stranded RNA cleavage activity, such as site-specific RNA cleavage activity or non-specific RNA cleavage activity. In some embodiments, activity is measured in vitro. In some embodiments, activity is measured in cells (such as bacterial cells, plant cells, or eukaryotic cells). In some embodiments, activity is measured in mammalian cells (such as rodent cells or human cells). In some embodiments, activity is measured in human cells (such as 293T cells). In some embodiments, activity is measured in mouse cells (such as Hepa1-6 cells). In some embodiments, compared with the reference Cas12b nuclease, the engineered Cas12b nuclease has an activity relative to the reference Cas12b nuclease height of at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 1 times, 1.2 times, 1.5 times, 2 times, 3 times, 4 times, 5 times, 10 times, 20 times, 50 times or more. The site-specific nuclease activity of the engineered Cas12b nuclease can be measured using methods known in the art, including, for example, PCR, sequencing or gel shift assays, as described in the examples provided herein. In some embodiments, the activity is gene editing activity in a cell. In some embodiments, the cell is a bacterial cell, a plant cell, or a eukaryotic cell. In some embodiments, the cell is a mammalian cell, such as a rodent cell or a human cell. In some embodiments, the cell is a 293T cell. In some embodiments, activity is measured in mouse cells (such as Hepa1-6 cells). In some embodiments, the activity is the indel formation activity at the target genomic site in the cell, such as the site-specific cutting of the target nucleic acid by the engineered Cas12b nuclease and the non-homologous end joining (NHEJ) mechanism for DNA repair. In some embodiments, the activity is the insertion of an exogenous nucleic acid sequence at the target genomic site in the cell, for example, the site-specific cutting of the target nucleic acid by the engineered Cas12b nuclease and the homologous recombination (HR) mechanism for DNA repair. In some embodiments, the homologous recombination after cutting with the engineered Cas12b nuclease further includes introducing a donor template.In some embodiments, compared with the reference Cas12b nuclease, the gene editing (e.g., indel formation) activity of the engineered Cas12b nuclease at the genomic site of the cell (e.g., human cells (such as 293T cells) or mouse Hepa1-6 cells) increases by at least about 20%, an increase of 30%, 40%, 60%, 70%, 80%, 90%, 1.5 times, 2 times, 3 times, 4 times, 5 times, 10 times, or more. In some embodiments, the engineered Cas12b nuclease is capable of editing a greater number (e.g., 2, 3, 4, 5, 10, 20, 50, 100 or more) of genomic sites than the reference Cas12b nuclease. In some embodiments, the consensus PAM sequence of the engineered Cas12b nuclease is identical to that of the reference Cas12b nuclease. In some embodiments, the engineered Cas12b nuclease recognizes more (e.g., 1, 2, 3, 4, 5, 10, 20, 50, 100, or more) PAM sequences compared to a reference Cas12b nuclease.
[0138] Any method known in the art can be used to determine the cutting or gene editing efficiency of engineered Cas12b nuclease in cells, including such as T7 endonuclease 1 (T7E1) determination, PCR, target DNA sequencing (including such as Sanger sequence, and second generation sequencing), deletion tracking insertion and deletion (TIDE) determination, or by amplicon analysis (IDAA) for indel detection.See, for example, Sentmanat MF et al. "Overview of validation strategies for CRISPR-Cas9 editing (Asurvey of validation strategies for CRISPR-Cas9 editing)" Scientific Reports 2018, 8, article number 888, the contents of which are incorporated herein by reference in their entirety. In some embodiments, for example, as described in the examples herein, targeted next generation sequencing (NGS) is used to measure the gene editing efficiency of engineered Cas12b nuclease in cells. Exemplary genomic sites for determining the cutting and gene editing efficiency of engineered Cas12b nuclease include but are not limited to CCR5, AAVS, CD34, RNF2, SCN9A, HBG1 / 2 and EMX1. In some embodiments, the gene editing efficiency of the engineered Cas12b nuclease can cut or edit at least about 1, 2, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 100 or more loci (compared to the average cutting or gene editing efficiency of the reference Cas12b nuclease). In some embodiments, the cutting or gene editing efficiency (e.g., indel rate) of the engineered Cas12b nuclease is at least about 10%, 20%, 30%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 1 times, 2 times, 5 times, 10 times, 20 times, 50 times or more (compared to the reference Cas12b nuclease).
[0139] Engineered Cas12b effector proteins
[0140] The application also provides an engineered Cas12b effector protein based on any of the engineered Cas12b nucleases, variants (e.g., dCas12b) or functional derivatives described herein. In some embodiments, the engineered Cas12b effector protein comprises (or consists of, or is essentially composed of) any of the engineered Cas12b nucleases, variants, or functional derivatives described herein. In some embodiments, the engineered Cas12b effector protein comprises a functional derivative of an engineered Cas12b nuclease, such as any of the functional derivatives described in the "Functional Derivatives" section below.
[0141] In some embodiments, the engineered Cas12b effector protein has enzymatic activity. In some embodiments, the engineered Cas12b effector protein is a nuclease that cuts the double strand of a target double-helical nucleic acid (e.g., duplex DNA). In some embodiments, the engineered Cas12b effector protein is a nickase, i.e., a single strand that cuts a target double-helical nucleic acid (e.g., duplex DNA). In some embodiments, the engineered Cas12b effector protein comprises an enzyme-inactivated mutant (dCas12b) of an engineered Cas12b nuclease. Mutations at one or more amino acid residues in the active site of the Cas12b nuclease can result in an enzyme-dead Cas12b (dCas12b). For example, D570A, E848A, R785A, E848A, R911A, and / or D977A mutants of AaCas12b (SEQ ID NO: 1) are significantly reduced (e.g., reduced by at least about 60%, 70%, 80%, 90%, 95% or more) or have no nuclease activity in human cells. See, for example, Teng F. et al. Cell Discovery, 4, Article No.: 63 (2018), the contents of which are incorporated herein by reference in their entirety. In some embodiments, the engineered Cas12b effector protein comprises an engineered Cas12b having one or more mutations corresponding to D570A, E848A, R785A, E848A, R911A, and D977A of AaCas12b. In some embodiments, one or more mutations selected from D570A, E848A, R785A, E8486A, R911A, and D977A are further introduced into an AaCas12b comprising a Q119F+E475R+E758R mutation. In some embodiments, the enzymatically inactive mutant of the engineered Cas12b nuclease comprises the amino acid sequence of any one of SEQ ID NOs: 79 to 81. In some embodiments, the engineered Cas12b effector protein comprises an engineered Cas12b having a mutation corresponding to the R785A mutation of AaCas12b. In some embodiments, the engineered Cas12b effector protein comprises an engineered Cas12b having a mutation corresponding to the R911A mutation of AaCas12b. In some embodiments, the engineered Cas12b effector protein comprises an engineered Cas12b having a mutation corresponding to the D977A mutation of AaCas12b. In some embodiments, the engineered Cas12b effector protein comprises an engineered Cas12b having a mutation corresponding to the E848A mutation of AaCas12b. In some embodiments, the engineered Cas12b effector protein comprises an engineered Cas12b having a mutation corresponding to the D570A mutation of AaCas12b.In some embodiments, the engineered Cas12b effector protein comprises an engineered Cas12b having mutations corresponding to the D570A+E848A mutations of AaCas12b or the D570A+D977A mutations of AaCas12a.
[0142] In some embodiments, an engineered Cas12b nickase is provided. In some embodiments, an engineered Cas12b fusion effector protein is provided, comprising an engineered Cas12b nuclease or variant or a functional derivative thereof fused to a functional domain (e.g., an enzyme-inactive mutant of an engineered Cas12b nuclease, such as SEQ ID NO: 79 to 81), wherein the functional domain is such as a translation initiator domain, a transcription repressor domain (e.g., a Krüppel-associated box (KRAB) domain), a transactivation domain, an epigenetic modification domain, a core base editing domain (e.g., a cytosine base editor (CBE) or adenine base editor (ABE) domain), a reverse transcriptase domain, a reporter domain (e.g., a fluorescent domain), or a nuclease domain (e.g., a ZFN domain). In some embodiments, an engineered Cas12b base editor is provided, comprising a catalytically inactive variant of any one of the engineered Cas12b nucleases described herein (e.g., SEQ ID NO: 79 to 81) fused to a cytosine deaminase domain or an adenosine deaminase domain. In some embodiments, an engineered Cas12b base editor is provided, comprising a catalytically inactive variant of any one of the engineered Cas12b nucleases described herein (e.g., SEQ ID NO: 79 to 81), which is fused to a KRAB domain or a functional fragment thereof, such as a ZIM3KRAB domain (SEQ ID NO: 72). In some embodiments, an engineered Cas12b lead editor is provided, comprising a catalytically inactive variant of any one of the engineered Cas12b nucleases described herein (e.g., SEQ ID NO: 79 to 81) fused to a reverse transcriptase domain. In some embodiments, a split Cas12b effector protein system is provided.
[0143] Variants / functional derivatives
[0144] The application also provides variants and functional derivatives according to any one of the engineered Cas12b nucleases described herein. In some embodiments, there is provided an engineered Cas12b effector protein, which comprises (or is composed of, or is substantially composed of) a functional variant of an engineered Cas12b nuclease as described herein. In some embodiments, compared with the amino acid sequence of the corresponding engineered Cas12b nuclease (e.g., SEQ ID NO: 2 to 22), the amino acid sequence of the functional variant has at least one amino acid residue difference (e.g., with deletion, insertion, substitution, and / or fusion). In some embodiments, the functional variant has one or more mutations, such as amino acid substitutions, insertions, and / or deletions. For example, compared with the engineered Cas12b nuclease, the functional variant can include any one of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more amino acid substitutions. In some embodiments, one or more substitutions are conservative substitutions. In some embodiments, the functional variant has all domains of the engineered Cas12b nuclease. In some embodiments, the functional variant lacks one or more domains of the engineered Cas12b nuclease.
[0145] For any of the Cas12b variant proteins described herein (e.g., nickase Cas12b protein, inactivated or catalytically inactive Cas12b (dCas12b), fusion Cas12b), the Cas12b variant can include the same parameters (e.g., domains, percent sequence identity, etc.) as any of the Cas12b protein sequences described herein.
[0146] Exemplary mutations in Cas12b functional variants are described in WO2016205764, WO2016205749, and WO2020 / 087631, the contents of which are incorporated herein by reference in their entirety.
[0147] catalytic activity
[0148] In some embodiments, the functional variants of the engineered Cas12b nuclease have different catalytic activities compared to their non-mutated forms. In some embodiments, the mutation (e.g., amino acid substitution, insertion, and / or deletion) is in the catalytic domain (e.g., RuvC domain) of the Cas12b effector protein. In some embodiments, the variant comprises mutations in multiple catalytic domains. The Cas12b effector protein that cuts one chain of a double-stranded target nucleic acid but does not cut the other chain is referred to herein as a "nickase" (e.g., "nickase Cas"). In some embodiments, the engineered Cas12b effector protein comprises (or is composed of, or is essentially composed of) a nickase mutant of the engineered Cas12b nuclease. The Cas12b protein that substantially does not have nuclease activity is referred to herein as a dead Cas12b protein ("dCas12b") (note that in the case of a fusion Cas12b effector protein, nuclease activity can be provided by a heterologous polypeptide-a fusion partner, which is described in more detail below). In some embodiments, a Cas12b effector protein is considered to substantially lack all DNA cleavage activity when the DNA cleavage activity of a mutant Cas12b is less than about any of 25%, 20%, 10%, 5%, 1%, 0.1%, 0.01% or less relative to its non-mutated form.
[0149] In some embodiments, the engineered Cas12b nuclease is dCas12b. In some embodiments, the engineered Cas12b functional variant includes a mutation corresponding to the D570A of AaCas12b (SEQ ID NO: 1). In some embodiments, the engineered Cas12b functional variant includes a mutation corresponding to the E848A of AaCas12b. In some embodiments, the engineered Cas12b functional variant includes a mutation corresponding to the R785A of AaCas12b. In some embodiments, the engineered Cas12b functional variant includes a mutation corresponding to the E848A of AaCas12b. In some embodiments, the engineered Cas12b functional variant includes a mutation corresponding to the R991A of AaCas12b. In some embodiments, the engineered Cas12b functional variant includes a mutation corresponding to the D977A of AaCas12b. In some embodiments, the engineered Cas12b functional variant includes a mutation corresponding to the D573A of BthCas12b. In some embodiments, the catalytically inactive or substantially inactive variant of AaCas12b (Q119F+E475R+E758R) further comprises one or more substitutions selected from D570A, E848A, and D977A, wherein the amino acid positions correspond to SEQ ID NO: 22. In some embodiments, dCas12b comprises the amino acid sequence of any one of SEQ ID NOs: 79 to 81.
[0150] Split Cas12b effector protein
[0151] The CRISPR-Cas12b system described herein can include any polypeptide pair comprising a split Cas12b portion in this section (also referred to herein as a "split Cas12b polypeptide"). Exemplary split Cas12b protein systems have been described in, for example, PCT / CN2020 / 111057 and PCT / CN2021 / 114339, the contents of which are incorporated herein by reference in their entirety.
[0152] In some embodiments, a split Cas12b effector protein is provided, comprising a first polypeptide and a second polypeptide, the first polypeptide comprising the N-terminal portion of any one of the engineered Cas12b nucleases described herein or variants or functional derivatives thereof (also referred to as "parent Cas12b proteins" in this section), the second polypeptide comprising the C-terminal portion of the engineered Cas12b nucleases or variants or functional derivatives thereof, wherein the first polypeptide and the second polypeptide are capable of associating with each other in the presence of a guide RNA comprising a guide sequence to form a CRISPR complex that specifically binds to a target nucleic acid comprising a target sequence complementary to the guide sequence. In some embodiments, the first polypeptide and the second polypeptide each comprise a dimerization domain. In some embodiments, the first dimerization domain and the second dimerization domain associate with each other in the presence of an inducer (e.g., rapamycin). In some embodiments, the first polypeptide and the second polypeptide do not comprise any dimerization domain. In some embodiments, the split Cas12b effector protein is autoinducible.
[0153] The split Cas12b portion is designed based on any of the engineered Cas12b nucleases described herein, or variants or functional variants thereof.
[0154] Cas12b protein has various domains. In some embodiments, the parent Cas12b protein comprises from N-terminus to C-terminus: a first WED domain (WED-I; also referred to as an OBD-I domain), a first REC domain (REC1), a second WED domain (WED-II; also referred to as an OBD-II domain), a first RuvC domain (RuvC-I), a bridge helix (BH) domain, a second RuvC domain (RuvC-II), a first Nuc domain (Nuc-I; also referred to as a UK-I domain), a third RuvC domain (RuvC-III), and a second Nuc domain (Nuc-II; also referred to as a UK-II domain). Domain boundaries can be determined using methods known in the art, such as based on crystal structures of naturally occurring Cas12b proteins (e.g., PDB ID Nos: 5U30, 5U31, 5U33, 5U34, and 5WQE for AaCas12b), and / or sequence homology to known functional domains in parent Cas12b proteins. In some embodiments, AaCas12b has the following domains: a WEB-I domain (amino acid residues 1 to 14), a REC1 domain (amino acid residues 15 to 386), a WED-II domain (amino acid residues 387 to 518), a RuvC-I domain (amino acid residues 519 to 628), a BH domain (amino acid residues 629 to 658), a REC2 domain (amino acid residues 659 to 783), a RuvC-II domain (amino acid residues 784-900), a Nuc-I domain (amino acid residues 901 to 974), a RuvC-III domain (amino acid residues 975 to 993), and a Nuc-II domain (amino acid residues 994 to 1129), wherein the amino acid numbering is based on SEQ ID NO: 1.
[0155] The engineered Cas12b nuclease or its variant or functional derivative is split, i.e., two split Cas12b parts substantially comprise functional Cas12b. Cas12b can serve as a genome editing enzyme (when forming a complex with target DNA and guide RNA), such as a single-stranded or double-stranded nuclease for cutting double-stranded nucleic acids, or it can be a catalytically dead Cas12b (dCas12b), which is substantially a DNA binding protein with very little or no catalytic activity due to the typical mutations in its catalytic domain. Mutations at one or more amino acid residues in the active site of reference Cas12b can result in catalytically dead Cas12b, such as D570A, R785A, E848A, R911A, and / or D977A mutants of AaCas12b.
[0156] The Cas12b portion of the split described herein can be designed by dividing (i.e., splitting) an engineered Cas12b nuclease or a variant or functional derivative thereof (referred to herein as "parent Cas12b protein," such as any one of SEQ ID NOs: 2 to 22 and 79 to 81) (e.g., a full-length Cas12b protein or a functional variant thereof) into two halves at the location of the split, where the N-terminal portion of the parent Cas12b protein is separated from the C-terminal portion. In some embodiments, the N-terminal portion comprises amino acid residues 1 to X of the parent Cas12b protein, and the C-terminal portion comprises amino acid residues X+1 to the C-terminal. In this example, the numbering is consecutive, but this may not always be necessary, as amino acids (or nucleotides encoding them) can be trimmed from the end of either of the cleaved ends, and / or mutations (e.g., insertions, deletions, and substitutions) at internal regions of the polypeptide chain are also contemplated, provided that sufficient DNA binding activity and (if desired) DNA nickase or double-stranded cleavage activity of the reconstructed Cas12b protein is retained, such as at least about 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95% or more of any one of the activities compared to the parent Cas12b protein.
[0157] Also envisioned are Cas12b portions with some N and / or C-terminal truncations or deletions and / or splits with internal mutations relative to the engineered Cas12b nucleases described herein. One skilled in the art can readily use the information of the exemplary split Cas12b polypeptides described herein to design corresponding split Cas12b polypeptides based on other Cas12b proteins and functional variants, for example, by using standard sequence alignment tools.
[0158] The position of the split can be located in a flexible region (such as a loop). Preferably, the position of the split occurs in a place where the interruption of the amino acid sequence does not result in partial or complete destruction of structural features (e.g., α helix or β sheet). Unstructured regions (regions that do not appear in the crystal structure because the degree of structure in these regions is not enough to "freeze" in the crystal) are generally preferred options. It is envisioned that lobes can be produced in unstructured regions on the surface exposed to the parent Cas12b protein.
[0159] In some embodiments, the parent Cas12b protein is not split at or near the amino acid residue participating in the interaction with the guide RNA and / or target RNA (e.g., within about 10, 8, 6, 5, 4, 3, 2, or 1 amino acid residues). For example, amino acid residues 4 to 9, 118 to 122, 143 to 144, 442 to 446, 573 to 574, 742 to 746, 753 to 754, 792 to 796, 800 to 819, 835 to 839, 897 to 900, and 973 to 978 of the AaCas12b protein participate in the interaction with a single guide RNA and / or target DNA, wherein numbering is based on SEQ ID NO: 1.
[0160] In some embodiments, the parent Cas12b protein splits at an amino acid residue within the amino acid residues corresponding to amino acid residues 516 to 793 of the AaCas12b protein, wherein the numbering is based on SEQ ID NO: 1. In some embodiments, the parent Cas12b protein splits at an amino acid residue adjacent to the WED-II domain and the RuvC-I domain. In some embodiments, the parent Cas12b protein splits at an amino acid residue within the amino acid residues corresponding to amino acid residues 516 to 519 of the AaCas12b protein, wherein the numbering is based on SEQ ID NO: 1. In some embodiments, the parent Cas12b protein splits at an amino acid residue adjacent to the BH domain and the REC2 domain. In some embodiments, the parent Cas12b protein splits at an amino acid residue within the amino acid residues corresponding to amino acid residues 621 to 627 of the AaCas12b protein, wherein the numbering is based on SEQ ID NO: 1. In some embodiments, the parent Cas12b protein splits at an amino acid residue adjacent to the REC2 domain and the RuvC-II domain. In some embodiments, the parent Cas12b protein is split at an amino acid residue within the amino acid residues corresponding to amino acid residues 777 to 793 of the AaCas12b protein, wherein the numbering is based on SEQ ID NO: 1. In some embodiments, the parent Cas12b protein is split within the RCE2 domain. In some embodiments, the parent Cas12b protein is split at an amino acid residue within the amino acid residues corresponding to amino acid residues 659 to 664, 676 to 684, or 702 to 706 of the AaCas12b protein, wherein the numbering is based on SEQ ID NO: 1.
[0161] In some embodiments, the parent Cas12b protein is cleaved at an amino acid residue that is no more than about 20 (e.g., no more than about any one of 18, 16, 14, 12, 10, 8, 7, 6, 5, 4, 3, 2, or 1) amino acid residues from the amino acid residue corresponding to amino acid residue 518 of the AaCas12b protein, wherein the numbering is based on SEQ ID NO: 1. In some embodiments, the parent Cas12b protein is cleaved at an amino acid residue that is no more than about 20 (e.g., no more than about any one of 18, 16, 14, 12, 10, 8, 7, 6, 5, 4, 3, 2, or 1) amino acid residues from the amino acid residue corresponding to amino acid residue 658 of the AaCas12b protein, wherein the numbering is based on SEQ ID NO: 1. In some embodiments, the parent Cas12b protein is cleaved at the amino acid residue corresponding to amino acid residue 658 of the AaCas12b protein, wherein the numbering is based on SEQ ID NO: 1. In some embodiments, the parent Cas12b protein is cleaved at an amino acid residue that is within no more than about 20 (e.g., no more than about any one of 18, 16, 14, 12, 10, 8, 7, 6, 5, 4, 3, 2, or 1) amino acid residues of the amino acid residue corresponding to amino acid residue 783 of the AaCas12b protein, wherein the numbering is based on SEQ ID NO: 1. In some embodiments, the parent Cas12b protein is cleaved at the amino acid residue that is corresponding to amino acid residue 783 of the AaCas12b protein, wherein the numbering is based on SEQ ID NO: 1.
[0162] In some embodiments, the N-terminal portion of the parent Cas12b protein comprises WED-I, REC1, WED-II, RuvC-I, and BH domains of the AaCas12b protein, and wherein the C-terminal portion of the parent Cas12b protein comprises REC2, RuvC-II, Nuc-I, RuvC-III, and Nuc-II domains of the AaCas12b protein. In some embodiments, the N-terminal portion of the parent Cas12b protein comprises amino acid residues 1 to 658 of the parent Cas12b protein, and the C-terminal portion of the parent Cas12b protein comprises amino acid residues 659 to 1129 of the parent Cas12b protein, wherein the amino acid residues are numbered according to SEQ ID NO: 1.
[0163] In some embodiments, the N-terminal portion of the parent Cas12b protein comprises WED-I, REC1, WED-II, RuvC-I, BH, and REC2 domains of the parent Cas12b protein, and wherein the C-terminal portion of the parent Cas12b protein comprises RuvC-II, Nuc-I, RuvC-III, and Nuc-II domains of the parent Cas12b protein. In some embodiments, the N-terminal portion of the parent Cas12b protein comprises amino acid residues 1 to 783 of the parent Cas12b protein, and the C-terminal portion of the parent Cas12b protein comprises amino acid residues 784 to 1129 of the parent Cas12b protein, wherein the amino acid residues are numbered according to SEQ ID NO: 1.
[0164] In some embodiments, the N-terminal portion of the parent Cas12b protein comprises WED-I, REC1, WED-II, RuvC-I, and BH domains of the parent Cas12b protein, wherein the C-terminal portion of the parent Cas12b protein comprises RuvC-II, Nuc-I, RuvC-III, and Nuc-II domains of the parent Cas12b protein, and wherein the REC2 domain of the parent Cas12b protein is split between the N-terminal portion of the parent Cas12b protein and the C-terminal portion of the parent Cas12b protein.
[0165] In some embodiments, the N-terminal portion of the parent Cas12b protein comprises WED-I, REC1, and WED-II domains of the parent Cas12b protein, and wherein the C-terminal portion of the parent Cas12b protein comprises RuvC-I, BH, REC2, RuvC-II, Nuc-I, RuvC-III, and Nuc-II domains of the parent Cas12b protein. In some embodiments, the N-terminal portion of the parent Cas12b protein comprises amino acid residues 1 to 518 of the parent Cas12b protein, and the C-terminal portion of the parent Cas12b protein comprises amino acid residues 519 to 1129 of the parent Cas12b protein, wherein the amino acid residues are numbered according to SEQ ID NO: 1.
[0166] Splitting point is usually designed on a computer and cloned into a construct. The two split Cas12b parts (N-terminal and C-terminal parts) together form a functional Cas12b protein, which preferably comprises at least about 70% or more of the amino acid sequence of the parent Cas12b protein, such as at least about 75%, 80%, 85%, 90%, 95%, 98%, 99%, or more of the amino acid sequence of the parent Cas12b protein. Some pruning and mutants are envisioned. The non-functional domain can be completely removed. For all split Cas12b systems, the two split Cas12b parts can be combined together, and the Cas12b function of the prestige can be restored or reconstructed. The activity of reconstructed Cas12b protein or CRISPR complex (Cas12b+ guide RNA complex) can be assessed using methods known in the art. For example, T7 endonuclease I (T7EI) can be used to determine the nuclease activity in the cell. Gene editing activity can also be assessed by DNA sequencing.
[0167] In some embodiments, the parent Cas12b protein is split into more than two parts.
[0168] The Cas12b effector proteins of the split can each include one or more dimerization domains. In some embodiments, the first polypeptide includes a first dimerization domain fused to the first split Cas12b effector portion, and the second polypeptide includes a second dimerization domain fused to the second split Cas12b effector portion. The dimerization domain can be fused to the split Cas12b effector portion via a peptide linker (e.g., a flexible peptide linker, such as a GS linker) or a chemical bond. In some embodiments, the dimerization domain is fused to the N-terminal end of the split Cas12b effector portion. In some embodiments, the dimerization domain is fused to the C-terminal end of the split Cas12b effector portion.
[0169] In some embodiments, the split Cas12b effector protein does not comprise any dimerization domain.
[0170] In some embodiments, the dimerization domain promotes the association of two split Cas12b effector moieties. In some embodiments, the split Cas12b effector moieties are induced to associate or dimerize into a functional Cas12b effector protein by an inducer. In some embodiments, the split Cas12b effector protein comprises an inducible dimerization domain. In some embodiments, the dimerization domain is not an inducible dimerization domain, i.e., the dimerization domain dimerizes in the absence of an inducer.
[0171] Inducer can be an induction energy source or an induction molecule different from guide RNA (e.g., sgRNA). The effect of the inducer is to reconstruct two split Cas12b effector parts into functional Cas12b effector proteins via the induced dimerization of the dimerization domain. In some embodiments, the inducer combines two split Cas12b effector parts together by the induced association of the inducible dimerization domain. In some embodiments, in the absence of an inducer, the two split Cas12b effector parts do not associate with each other to reconstruct into functional Cas12b effector proteins. In some embodiments, in the absence of an inducer, the two split Cas12b effector parts can associate with each other to reconstruct into functional Cas12b effector proteins in the presence of a guide RNA (e.g., sgRNA).
[0172] The inducer of the present application can be heat, ultrasound, electromagnetic energy, or chemical compound.In some embodiments, the inducer is an antibiotic, a small molecule, a hormone, a hormone derivative, a steroid, or a steroid derivative.In some embodiments, the inducer is abscisic acid (ABA), doxycycline (DOX), cumate, rapamycin, 4-hydroxytamoxifen (4OHT), estrogen, or ecdysone.In some embodiments, the Cas12b effector system of splitting is a system controlled by inducer, which is selected from: an inducible system based on antibiotics, an inducible system based on electromagnetic energy, an inducible system based on small molecules, an inducible system based on nuclear receptors, and an inducible system based on hormones. In some embodiments, the Cas12b effector system of splitting is a system controlled by an inducer, which is selected from: tetracycline (Tet) / DOX inducible system, light inducible system, ABA inducible system, cumate repressor / operator system, 4OHT / estrogen inducible system, ecdysone-based inducible system, and FKBP12 / FRAP (FKBP12-rapamycin complex) inducible system. This inducer is also discussed herein and in PCT / US2013 / 051418, the contents of which are incorporated herein by reference in their entirety. FRB / FKBP / rapamycin system has been described in Paulmurugan and Gambhir《Cancer Research (Cancer Res)》August 15, 2005, 65;7413;With Crabtree et al.《Chemistry & Biology (Chemistry & Biology)》13,99-107, January 2006, their contents are incorporated herein by reference in their entirety.
[0173] In some embodiments, the split Cas12b effector protein pair is separate and inactive until induced dimerization of the dimerization domain (e.g., FRB and FKBP), which results in reassembly of a functional Cas12b effector nuclease. In some embodiments, a first split Cas12b effector protein comprising a first half of an inducible dimer (e.g., FRB) is delivered and / or positioned separately from a second split Cas12b effector protein comprising a second half of an inducible dimer (e.g., FKBP).
[0174] Other exemplary FKBP-based inducible systems that can be used for the inducer-controlled split Cas12b effector system described herein include, but are not limited to: FKBP dimerized with calcineurin (CNA) in the presence of FK506; FKBP dimerized with CyP-Fas in the presence of FKCsA; FKBP dimerized with FRB in the presence of rapamycin; GyrB dimerized with GryB in the presence of coumermycin; GAI dimerized with GID1 in the presence of gibberellins; or Snap-tag dimerized with HaloTag in the presence of HaXS.
[0175] Alternatives within the FKBP family itself are also contemplated, such as FKBPs that homodimerize (i.e., one FKBP dimerizes with another FKBP) in the presence of FK1012.
[0176] In some embodiments, the dimerization domain is FKBP and the inducer is FK1012. In some embodiments, the dimerization domain is GryB and the inducer is coumermycin. In some embodiments, the dimerization domain is ABA and the inducer is gibberellin.
[0177] In some embodiments, the split Cas12b effector portion can be self-inducible (i.e., self-activated or self-induced) to associate / dimerize into a functional Cas12b effector protein in the absence of an inducer. Without being bound by any theory or hypothesis, the self-induction of the split Cas12b effector portion can be mediated by binding to a guide RNA (such as sgRNA). In some embodiments, the first polypeptide and the second polypeptide do not comprise a dimerization domain. In some embodiments, the first polypeptide and the second polypeptide comprise a dimerization domain.
[0178] In some embodiments, the reconstituted Cas12b effector protein of the split Cas12b effector system described herein (including inducer-controlled systems and autoinducible systems) has an editing efficiency of at least about 70% (such as at least about 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or more efficiency, or any of 100%) of the editing efficiency of the parent Cas12b effector protein.
[0179] In some embodiments, the reconstituted Cas12b effector protein of the inducer-controlled split-Cas12b effector system described herein has an editing efficiency of no more than about 50% (such as no more than about any of 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, 5%, or less, or 0% efficiency) in the absence of an inducer of the editing efficiency of the parental Cas12b effector protein (i.e., due to autoinduction).
[0180] Fusion of Cas12b effector protein
[0181] The present application also provides engineered Cas12b effector proteins comprising additional protein domains and / or components, such as linkers, nuclear localization / export sequences, functional domains, and / or reporter proteins.
[0182] In some embodiments, the engineered Cas12b effector protein is a protein complex comprising one or more heterologous protein domains (e.g., about or more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more domains) in addition to the nucleic acid targeting domain of the engineered Cas12b nuclease or a variant or functional derivative thereof. In some embodiments, the engineered Cas12b effector protein is a fusion protein comprising one or more heterologous protein domains (e.g., about or more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more domains) fused to the engineered Cas12b nuclease or a variant or functional derivative thereof.
[0183] In some embodiments, the engineered Cas12b effector protein of the present application may include (e.g., via a fusion protein, such as via one or more peptide linkers, e.g., a GS peptide linker, etc.) one or more functional domains or be associated with one or more functional domains (e.g., via co-expression of multiple proteins). In some embodiments, one or more functional domains are enzyme domains. These functional domains can have various activities, such as DNA and / or RNA methylase activity, demethylase activity, transcription activation activity, transcription repression activity, transcription release factor activity, histone modification activity, RNA cleavage activity, DNA cleavage activity, nucleic acid binding activity, and switch activity (e.g., light inducible). In some embodiments, one or more functional domains are transcription activation domains (i.e., transactivation domains) or repressor domains. In some embodiments, transcription activation domains or repressor domains can recruit chromatin modifiers. In some embodiments, one or more functional domains are histone modification domains. In some embodiments, one or more functional domains are transposase domains, HR (homologous recombination) machine domains, recombinase domains, and / or integrase domains. In some embodiments, the functional domain is a Krüppel-associated box (KRAB), VP64, VP16, Fok1, P65, HSF1, MyoD1, biotin-APEX, APOBEC1, AID, PmCD A1, Tad1, and M-MLV reverse transcriptase. In some embodiments, the functional domain is selected from the group consisting of a translation initiator domain, a transcription repressor domain, a transactivation domain, an epigenetic modification domain, a nucleobase editing domain (e.g., a CBE or ABE domain), a reverse transcriptase domain, a reporter domain (e.g., a fluorescent domain), and a nuclease domain. In some embodiments, the functional domain is a KRAB domain, such as the KRAB domain of ZIM3. In some embodiments, the KRAB structure comprises the amino acid sequence of SEQ ID NO: 72.
[0184] In some embodiments, the positioning of one or more functional domains in the engineered Cas12b effector protein allows the correct spatial orientation of the functional domain to affect the target with the functional effect to which it is attributed. For example, if the functional domain is a transcriptional activator (e.g., VP16, VP64, or p65), the transcriptional activator is placed in a spatial orientation that allows it to affect target transcription. Similarly, the transcriptional repressor is positioned to affect target transcription, and the nuclease (e.g., Fok1) is positioned to cut or partially cut the target. In some embodiments, the functional domain (e.g., KRAB domain, such as comprising SEQ ID NO: 72) is positioned at the N-terminus of the engineered Cas12b effector protein (e.g., SEQ ID NO: 79 to 81, such as SEQ ID NO: 81). In some embodiments, the functional domain (e.g., KRAB domain, such as comprising SEQ ID NO: 72) is positioned at the C-terminus of the engineered Cas12b effector protein (e.g., SEQ ID NO: 79 to 81, such as SEQ ID NO: 81). In some embodiments, the engineered Cas12b effector protein comprises a first functional domain at the N-terminus and a second functional domain at the C-terminus. In some embodiments, the engineered Cas12b effector protein comprises a catalytically inactive mutant of any one of the engineered Cas12b nucleases described herein fused to one or more functional domains (e.g., KRAB domains) (e.g., any one of SEQ ID NOs: 79 to 81). In some embodiments, the engineered Cas12b effector protein is a transcriptional activator. In some embodiments, the engineered Cas12b effector protein comprises an enzymatically inactive variant of any one of the engineered Cas12b nucleases described herein fused to a transactivation domain (e.g., any one of SEQ ID NOs: 79 to 81). In some embodiments, the transactivation domain is selected from the group consisting of: VP64, p65, HSF1, VP16, MyoD1, HSF1, RTA, SET7 / 9, and combinations thereof. In some embodiments, the transactivation domain comprises VP64, p65, and HSF1. In some embodiments, the engineered Cas12b effector protein comprises two split Cas12b effector polypeptides, each of which is fused to a transactivation domain. In some embodiments, the engineered Cas12b effector protein further comprises one or more nuclear localization sequences (e.g., any one of SEQ ID NOs: 61, 62, and 82).
[0185] In some embodiments, the engineered Cas12b effector protein is a transcriptional repressor. In some embodiments, the engineered Cas12b effector protein comprises an enzymatically inactive variant of any one of the engineered Cas12b nucleases described herein fused to a transcriptional repressor domain (e.g., KRAB): any one of SEQ ID NO: 79 to 81. In some embodiments, the transcriptional repressor domain is selected from: Krüppel associated box (KRAB), EnR, NuE, NcoR, SID, SID4X, and combinations thereof. In some embodiments, the engineered Cas12b effector protein comprises two split Cas12b effector polypeptides, each of which is fused to a transcriptional repressor domain. In some embodiments, the engineered Cas12b effector protein further comprises one or more nuclear localization sequences (e.g., SEQ ID NO: any one of 61, 62, and 82).
[0186] In some embodiments, the engineered Cas12b effector protein is a base editor, such as a cytosine editor or an adenosine editor. In some embodiments, the engineered Cas12b effector protein comprises an enzymatically inactive variant of any one of the engineered Cas12b nucleases described herein fused to a core base editing domain (such as a cytosine base editor (CBE) domain or an adenosine base editor (AB E) domain) (e.g., SEQ ID NO: 79 to 81). In some embodiments, the core base editing domain is a DNA editing domain. In some embodiments, the core base editing domain has deaminase activity. In some embodiments, the core base editing domain is a cytosine deaminase domain. In some embodiments, the core base editing domain is an adenosine deaminase domain. Exemplary base editing based on Cas nucleases has been described in, for example, WO2018 / 165629A1 and WO2019 / 226953A1, the contents of which are incorporated herein by reference in their entirety. Exemplary CBE domains include, but are not limited to, activation-induced cytidine deaminase or AID (e.g., hAID), apolipoprotein B mRNA editing complex, or APOBEC (e.g., rat APOBEC1, hAPOBEC3A / B / C / D / E / F / G), and PmCDA1. Exemplary ABE domains include, but are not limited to, TadA, ABE8, and variants thereof (see, e.g., Gaudelli et al., 2017 Nature 551:464-471; and Richter et al., 2020 Nature Biotechnology 38:883-891, the contents of each of which are incorporated herein by reference in their entirety). In some embodiments, the functional domain is an APOBEC1 domain, e.g., a rat APOBEC1 domain. In some embodiments, the functional domain is a TadA domain. In some embodiments, the engineered Cas12b effector protein further comprises one or more nuclear localization sequences (e.g., any one of SEQ ID NOs: 61, 62, and 82).
[0187] In some embodiments, the engineered Cas12b effector protein is a lead editor. The lead editor based on Cas9 has been described in, for example, A.Anzalone et al. Nature 2019, 576 (7785): 149-157, the contents of which are incorporated herein by reference in their entirety. In some embodiments, the engineered Cas12b effector protein comprises any one of the engineered Cas12b nucleases described herein fused to a reverse transcriptase domain. In some embodiments, the functional domain is a reverse transcriptase domain. In some embodiments, the reverse transcriptase domain is an M-MLV reverse transcriptase or a variant thereof, such as an M-MLV reverse transcriptase with one or more mutations in D200N, T306K, W313F, T330P, and L603W. In some embodiments, an engineered CRISP R / Cas12b system comprising a lead editor is provided. In some embodiments, the engineered CRISPR / Cas12b system further comprises a second Cas12b nickase, for example, based on the same engineered Cas12b nuclease as the lead editor. In some embodiments, the engineered CRISPR / Cas12b system comprises a lead editor guide RNA (pegRNA) comprising a primer binding site and a reverse transcriptase (RT) template sequence.
[0188] In some embodiments, the application provides a split Cas12b effector system having one or more (e.g., 1, 2, 3, 4, 5, 6, or more) functional domains, which are associated with (i.e., combined with or fused to) one or two split Cas12b effector portions. The functional domain can be provided as a fusion within the construct as part of the first and / or second split Cas12b effector protein. The functional domain is typically fused to other portions (e.g., split Cas12b effector portions) in the split Cas12b effector protein via a peptide linker (such as a GS linker). The functional domain can be used to replace the function of the split Cas12b effector system based on a catalytically dead Cas12b effector.
[0189] In some embodiments, the engineered Cas12b effector protein comprises one or more nuclear localization sequences (NLS) and / or one or more nuclear export sequences (NES). Exemplary NLS sequences include, for example, PKKKRKV (SEQ ID NO: 82), PKKKRKVPG (SEQ ID NO: 61), and ASPKKKRKV (SEQ ID NO: 62). The NLS and / or NES can be operably linked to the N-terminus and / or C-terminus of the engineered Cas12b effector protein or the polypeptide chain in the engineered Cas12b effector protein.
[0190] In some embodiments, the engineered Cas12b effector protein can encode additional components, such as a reporter protein. In some embodiments, the engineered Cas12b effector protein comprises a fluorescent protein, such as GFP. This system can allow imaging of genomic loci (see, for example, "Dynamic Imaging of Genomic Loci in Living Human Cells by an Optimized CRISPR / Cas System" Chen B et al. "Cell" 2013). In some embodiments, the engineered Cas12b effector protein is an inducible splitting Cas effector system that can image genomic loci.
[0191] Engineered CRISPR-Cas12b system
[0192] Also provided is an engineered CRISPR-Cas12b system comprising: (a) any one of the engineered Cas12b nuclease or variants or derivatives thereof (e.g., any one of SEQ ID NOs: 2 to 22 and 79 to 81) or the engineered Cas12b effector proteins described herein (e.g., engineered Cas12b nucleases, nickases, split Cas12b proteins, transcriptional repressors, transcriptional activators, base editors, or lead editors), or nucleic acids encoding the same; and (b) a guide RNA comprising a guide sequence complementary to a target sequence of a target nucleic acid, or one or more nucleic acids encoding the guide RNA, wherein the engineered Cas12b nuclease or engineered Cas12b effector protein and the guide RNA are capable of forming a CRISPR complex that specifically binds to a target nucleic acid comprising a target sequence and induces modification of the target nucleic acid.In some embodiments, an engineered CRISPR-Cas12b system is provided, comprising: (a) an engineered Cas12b nuclease or an effector protein thereof, comprising one, two, or three types of mutations relative to a reference Cas12b nuclease, wherein the mutations comprise: (1) substitution of one or more amino acid residues in the reference Cas12b nuclease that interact with the PAM (e.g., one or more of the following positions: 116, 123, 130, 132, 144, 145, 153, 173, 222, 395, 400, and 475) with a positively charged amino acid residue (e.g., R, H, K); and / or (2) substitution of one or more amino acids in the reference Cas12b nuclease that are involved in opening the DNA double strand with an amino acid residue having an aromatic ring (e.g., F, Y, W). and / or (3) replacing one or more amino acid residues in the RuvC domain of a reference Cas12b nuclease that interacts with an ssDNA substrate (e.g., one or more of the following positions: 300, 301, 304, 329, 636, 639, 647, 682, 757, 758, 761, 764, 768, 852, 854, 856, 857, 858, 860, 862, 863, 865, 866, 867, 869, 938, 956, 957, 958, 994, 1093, and 1097) with a positively charged amino acid (e.g., R, H, K) or a hydrophobic amino acid residue (e.g., F, Y, W, M), wherein the reference Cas12b nuclease comprises SEQ ID NO:1 amino acid sequence, or a nucleic acid encoding an engineered Cas12b nuclease or its effector protein; and (b) gRNA comprising a guide sequence complementary to the target sequence of the target nucleic acid or a nucleic acid encoding the gRNA, wherein the engineered Cas12b nuclease or its effector protein and the gRNA are capable of forming a CRISPR complex that specifically binds to a target nucleic acid comprising the target sequence and induces modification of the target nucleic acid. In some embodiments, the engineered CRISPR-Cas12b system comprises one or more nucleic acids encoding the engineered Cas12b effector protein and / or guide RNA of the engineered nuclease or its variant or derivative. In some embodiments, the gRNA comprises cr RNA and tracrRNA. In some embodiments, the engineered CRISPR-Cas12b system comprises a precursor guide RNA array that can be processed into multiple crRNAs, for example, by the engineered Cas12b nuclease or its variant or derivative or the engineered Cas12b effector protein. In some embodiments, the gRNA is an sgRNA. In some embodiments, the sgRNA comprises the backbone sequence of any one of SEQ ID NOs: 23 to 53.In some embodiments, the engineered CRISPR-Cas12b system comprises one or more vectors encoding the engineered Cas12b nuclease or its variant or derivative or the engineered Cas12b effector protein and / or guide RNA. In some embodiments, the engineered Cas12b nuclease or its variant or derivative or the engineered Cas12b effector protein and / or guide RNA are encoded by one or more vectors (such as adeno-associated virus (AAV) vectors). In some embodiments, the engineered CRISPR-Cas12b system comprises a ribonucleoprotein (RNP) complex comprising an engineered Cas12b nuclease or its variant or derivative or an engineered Cas12b effector protein bound to a guide RNA.
[0193] In some embodiments, an engineered CRISPR-Cas12b system is provided, comprising: (a) a Cas12b nuclease comprising the amino acid sequence of SEQ ID NO: 1, or an effector protein thereof (e.g., a nickase, a split Cas12b protein, a transcriptional repressor, a transcriptional activator, a base editor, or a guide editor), or any one of the engineered Cas12b nucleases or variants or derivatives thereof described herein (e.g., any one of SEQ ID NOs: 2 to 22 and 79 to 81), or an engineered Cas12b effector protein (e.g., a nickase, a split Cas12b protein, a transcriptional repressor, a transcriptional activator, a base editor, or a guide editor), or an encoding nucleic acid thereof; and (b) a gRNA comprising a guide sequence complementary to a target nucleic acid or a target sequence of a nucleic acid encoding the gRNA, wherein the gRNA comprises an engineered scaffold comprising SEQ ID NO: 1. NO: The sequence of any one of 25 to 53; wherein i) Cas12b nuclease or its effector protein or engineered Cas12b nuclease or its variant or derivative or engineered Cas12b effector protein, and ii) gRNA is capable of forming a CRISPR complex that specifically binds to the target nucleic acid and induces modification of the target nucleic acid.In some embodiments, an engineered CRISPR-Cas12b system is provided, comprising: (a) a Cas12b nuclease comprising the amino acid sequence of SEQ ID NO: 1, or an effector protein thereof (e.g., a nickase, a split Cas12b protein, a transcriptional repressor, a transcriptional activator, a base editor, or a guide editor), or three types of mutations relative to a reference Cas12b nuclease, wherein the mutations comprise: (1) substitution of one or more amino acid residues interacting with a PAM in the reference Cas12b nuclease with a positively charged amino acid residue (e.g., R, H, K), such as one or more of the following positions: 116, 123, 130, 132, 144, 145, 153, 173, 222, 395, 400, and 475); and / or (2) substitution of one or more amino acid residues involved in opening a double-stranded DNA strand in the reference Cas12b nuclease with an amino acid residue having an aromatic ring (e.g., F, Y, W). and / or (3) replacing one or more amino acid residues in the RuvC domain of a reference Cas12b nuclease that interacts with an ssDNA substrate (e.g., one or more of the following positions: 300, 301, 304, 329, 636, 639, 647, 682, 757, 758, 761, 764, 768, 852, 854, 856, 857, 858, 860, 862, 863, 865, 866, 867, 869, 938, 956, 957, 958, 994, 1093, and 1097) with a positively charged amino acid (e.g., R, H, K) or a hydrophobic amino acid residue (e.g., F, Y, W, M), wherein the reference Cas12b nuclease comprises SEQ ID NO: 1, or a nucleic acid encoding a Cas12b nuclease or an effector protein thereof; and (b) a gRNA comprising a guide sequence complementary to the target sequence of the target nucleic acid or the nucleic acid encoding the gRNA, wherein the gRNA comprises an engineered scaffold comprising the sequence of any one of SEQ ID NOs: 25 to 53; wherein i) the Cas12b nuclease or its effector protein or its engineered Cas12b nuclease or effector protein, and ii) the gRNA is capable of forming a CRISPR complex that specifically binds to the target nucleic acid and induces modification of the target nucleic acid.In some embodiments, an engineered CRISPR-Cas12b system is provided, comprising: (a) a Cas12b nuclease or a Cas12b effector protein comprising an amino acid sequence of any one of SEQ ID NOs: 1 to 22 and 79 to 81 or a nucleic acid encoding thereof; and (b) a gRNA comprising a guide sequence complementary to a target sequence of a target nucleic acid or a nucleic acid encoding the gRNA, wherein the gRNA comprises an engineered scaffold comprising a sequence of any one of SEQ ID NOs: 25 to 53; wherein the Cas12b nuclease or Cas12b effector protein and the gRNA are capable of forming a CRISPR complex that specifically binds to the target nucleic acid and induces modification of the target nucleic acid. In some embodiments, the gRNA comprises crRNA and tracrRNA, and wherein the tracrRNA comprises an engineered scaffold or a portion thereof. In some embodiments, the engineered CRISPR-Cas12b system comprises an array of precursor gRNAs encoding a variety of crRNAs. In some embodiments, the gRNA is an sgRNA. In some embodiments, the engineered CRISPR-Cas12b system comprises one or more vectors encoding engineered Cas12b nucleases, engineered Cas12b effector proteins, Cas12b nucleases, or Cas12b effector proteins. In some embodiments, one or more vectors are AAV vectors. In some embodiments, one or more vectors further encode gRNA.
[0194] PAM
[0195] In some embodiments, the engineered Cas12b nuclease or variant or derivative thereof, engineered Cas12b effector protein, Cas12b nuclease, or Cas12b effector protein Cas12b recognizes a PAM comprising (or consisting of) a 5′-TTN-3′ sequence, wherein N is A, T, G, or C. In some embodiments, the PAM comprises or consists of 5′-TTC-3′, 5′-TTTA-3′, 5′-TTT-3′, or 5′-TTG-3′.
[0196] guide RNA
[0197] The engineering CRISPR-Cas12b system of the application can include any suitable guide RNA.Guide RNA (gRNA) can include a guide sequence (or spacer) that can hybridize with the target sequence in the target nucleic acid of interest (such as the genomic locus of interest in the cell). In some embodiments, gRNA includes CRISPR RNA (crRNA) sequence, and the crRNA sequence includes a guide sequence. In some embodiments, crRNA as described herein includes a direct repeat sequence (DR) and a spacer sequence. In certain embodiments, crRNA includes a direct repeat sequence connected to a guide sequence or a spacer sequence, is essentially composed of a direct repeat sequence connected to a guide sequence or a spacer sequence, or is composed of a direct repeat sequence connected to a guide sequence or a spacer sequence. In certain embodiments, a direct repeat sequence can be located upstream (i.e., 5') of a guide sequence or a spacer sequence. In other embodiments, a direct repeat sequence can be located downstream (i.e., 3') of a guide sequence or a spacer sequence. In some embodiments, crRNA includes a direct repeat sequence, a spacer sequence, and a direct repeat sequence (DR-spacer-DR), which is a typical configuration of a precursor crRNA (pre-crRNA) configuration. In some embodiments, crRNA includes truncated direct repeat sequences and spacer sequences, which are typical sequences of processed or mature crRNA. In some embodiments, crRNA includes mutant DR sequences and spacer sequences. In some embodiments, gRNA includes transactivating CRISPR RNA (tracrRNA) sequences. In some embodiments, tracrRNA is fused to crRNA at the 5' end of the DR sequence. In some embodiments, the guide RNA is a single guide RNA (sgRNA). In some embodiments, gRNA or sgRNA includes tracrRNA and crR NA. In some embodiments, the sgRNA includes SEQ ID NO: The sequence of any one of 23 to 53. In some embodiments, the tracrRNA includes SEQ ID NO: The sequence of any one of 23 to 53 or a portion thereof.
[0198] In some embodiments, gRNA comprises non-homologous crRNA sequences and / or tracrRNA sequences that are not naturally found in the CRISPR locus with reference to Cas12b proteins. For example, AaCas12b, AkCas12b, AmCas12b, and BhCas12b homologous tracrRNA and crRNA sequences, BsCas12b, BsCas12b, LsCas12b, and SbCas12b and exemplary sgRNA sequences thereof are described in Figures S4 and S8 of Cell Discovery (2019) 5:23 by Teng F. et al., the contents of which are incorporated herein by reference in their entirety.
[0199] In some embodiments, the CRISPR-Cas12b system described herein comprises one or more (e.g., 1, 2, 3, 4, 5, 10, 15 or more) gRNAs (e.g., crRNA, tracrRNA, or sgRNA) or nucleic acids encoding them. In some embodiments, two or more gRNAs target different target sites, such as 2 target sites of the same target DNA or gene, or 2 target sites of 2 different target DNAs or genes.
[0200] The sequence and length of gRNA as described herein can be optimized. In some embodiments, the optimal length of gRNA can be determined by identifying the processed form of crRNA or by the empirical length study of crRNA. In some embodiments, gRNA includes base modifications, such as in gRNA scaffold regions.
[0201] The spacer does not need to be fully complementary, provided that the gRNA (e.g., crRNA or sgRNA) has sufficient complementarity to play a role (i.e., directing the Cas12b nuclease (e.g., engineered) or its effector protein to the target site). The editing or cutting efficiency of the Cas12b nuclease (e.g., engineered) or its effector protein mediated by gRNA can be adjusted by introducing one or more mismatches (e.g., 1 or 2 mismatches between the spacer sequence and the target sequence, including the position of the mismatch along the spacer / target sequence). When the mismatch (e.g., double mismatch) is located at the more central part of the spacer (i.e., not at the 3' or 5' end of the spacer), the cutting efficiency is more affected. Therefore, by selecting the mismatch position along the spacer sequence, the editing or cutting efficiency of the Cas12b nuclease (e.g., engineered) or its effector protein can be adjusted. For example, if the editing or cutting of the target sequence is desired to be less than 100% (e.g., in a cell population), 1 or 2 mismatches between the spacer sequence and the target sequence can be introduced into the spacer sequence.
[0202] In some embodiments, the guide sequence or spacer is designed to have at least one mismatch with the target sequence such that the heteroduplex formed between the guide sequence and the target sequence includes an unpaired C in the guide sequence relative to the target A, or an unpaired A in the guide sequence relative to the target C, for deamination on the target sequence (e.g., for base editing). In some embodiments, except for this AC or CA mismatch, the degree of complementarity is about or greater than about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99% or more when optimally aligned using a suitable alignment algorithm. The guide sequence may have a suitable length. In some embodiments, the length of the guide or spacer sequence is about 10nt to about 50nt. In some embodiments, the length of the guide or spacer sequence is at least about 16 nucleotides, preferably about 16 to about 100 nucleotides, more preferably about 16 to about 50 nucleotides (e.g., any of about 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 nucleotides). In some embodiments, the spacer is about 16 to about 27 nucleotides, such as any of about 17 to about 24 nucleotides, about 18 to about 24 nucleotides, or about 18 to 22 nucleotides. In some embodiments, the guide sequence is between about 18 to about 35 nucleotides, including, for example, any of 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 nucleotides.
[0203] In some embodiments, the guide or spacer sequence is at least about 60% complementary to the target sequence (e.g., at least about any of 70%, 75%, 80%, 85%, 90%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%). In some embodiments, there are at least about 15 (e.g., at least about 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, or more) base pairings between the spacer sequence and the target sequence of the target nucleic acid (e.g., DNA).
[0204] Optimal alignment can be determined using any suitable algorithm for aligning sequences, non-limiting examples of which include the Smith-Waterman algorithm, the Needleman-Wunsch algorithm, algorithms based on the Burrows-Wheeler transform (e.g., Burrows-Wheeler aligner), ClustalW, Clustal X, BLAT, Novolign (Novocraft Technologies; available at www.Novocraft.com), ELAND (Illumina, San Diego, CA), SOAP (available at SOAP.genomis.org.cn), and Maq (available at Maq.sourceforge.net). The ability of a guide sequence (within a nucleic acid-targeting guide RNA) to direct sequence-specific binding of a nucleic acid-targeting complex to a target nucleic acid sequence can be assessed by any suitable assay. For example, components of a nucleic acid-targeting CRISPR system sufficient to form a nucleic acid-targeting complex (including a guide sequence to be tested) can be provided to a host cell having a corresponding target nucleic acid sequence, for example by transfection with a vector encoding the components of the nucleic acid-targeting complex, followed by assessment of preferential targeting (e.g., cleavage) within the target nucleic acid sequence, for example by Surveyor analysis as described herein. Similarly, cleavage of a target nucleic acid sequence can be assessed in a test by providing a target nucleic acid sequence, components of a nucleic acid-targeting complex (including a guide sequence to be tested and a control guide sequence different from the test guide sequence), and comparing the binding or cleavage rate at the target sequence between the test and control guide sequence reactions. Other assays are possible, and will occur to those skilled in the art.
[0205] As used herein, target nucleic acid is used interchangeably with target sequence or target nucleic acid sequence to refer to a specific nucleic acid comprising a nucleic acid sequence that is complementary to all or part of the spacer in crRNA or gRNA. In some embodiments, the target nucleic acid comprises a gene or a sequence within a gene. In some embodiments, the target nucleic acid comprises a non-coding region (e.g., a promoter). In some embodiments, the target nucleic acid is single-stranded. In some embodiments, the target nucleic acid is double-stranded. Target nucleic acid can be selected to target any target nucleic acid sequence, such as a DNA or RNA sequence (e.g., mRNA).
[0206] The target nucleic acid should be associated with a PAM (i.e., a short sequence recognized by the CRISPR complex). According to the properties of the CRISPR-Cas protein, the following target sequence should be selected, and its complementary sequence in the DNA duplex (the complementary sequence of the target sequence) is located upstream or downstream of the PAM. In one embodiment of the present application, the complementary sequence of the target sequence is downstream or 3' of the PAM. The exact sequence and length requirements of the PAM depend on the Cas12b protein used.
[0207] "TracrRNA" sequence or similar terms include any polynucleotide sequence that has sufficient complementarity to hybridize with a crRNA sequence. In some embodiments, when ideally aligned, the degree of complementarity between the tracrRNA sequence and the crRNA sequence along the length of the shorter of the two is about or greater than about 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97.5%, 99% or more. In some embodiments, the length of the tracr sequence is about or greater than about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50 or more nucleotides. In some embodiments, the tracr sequence and the crRNA sequence are contained within a single transcript such that hybridization between the two produces a transcript with a secondary structure, such as one or more hairpins. Typically, the degree of complementarity refers to the ideal alignment of the guide sequence and the tracr sequence along the length of the shorter of the two sequences. Ideal alignment can be determined by any suitable alignment algorithm, and may further take secondary structure into account.
[0208] Any gRNA scaffold or tracrRNA or DR sequence that can mediate the combination of Cas12b protein described herein with corresponding gRNA (such as crRNA) can be used in the present application.In some embodiments, gRNA scaffold or tracrRNA or DR sequence comprises a stem-loop structure near the 5' or 3' end (next to the spacer sequence). "Stem-loop structure" refers to a nucleic acid with the following secondary structure, which includes a nucleotide region known or predicted to form a double-stranded (stem) portion and is connected at one end by a connecting region (loop) of basic single-stranded nucleotides. The term "hairpin" structure is also used herein to refer to a stem-loop structure. This structure is well known in the art, and these terms are used according to their well-known meanings in the art. The stem-loop structure does not require precise base pairing. Therefore, the rod can include one or more base mispairings. Alternatively, the base pairing may be precise, i.e., does not include any mispairing.
[0209] In some embodiments, the gRNA scaffold or tracrRNA or DR is a "functional variant" of the wild-type backbone or tracrRNA or DR, such as a "functional truncated version," "functionally extended version," or "functionally substituted version." A "functional variant" of a gRNA scaffold or tracrRNA or DR is a 5' and / or 3' extended (functional extended version) or truncated (functional truncated version) variant of a reference backbone or tracrRNA or DR (e.g., a parent DR), and / or a substitution (functional replacement version) of one or more nucleotides relative to a reference backbone or tracrRNA or DR (e.g., a parent DR). The function mediates the binding of a Cas12b nuclease (e.g., engineered) or its effector protein to the corresponding sgRNA or crRNA. The gRNA scaffold or tracrRNA or DR functional variant typically retains a stem-loop secondary structure or portion thereof that can be used to bind to a Cas12b nuclease (e.g., engineered) or its effector protein. In some embodiments, the gRNA scaffold or tracrRNA or DR or its functional variant comprises at least two (e.g., 2, 3, 4, 5 or more) stem-loop secondary structures or portions thereof that can be used to bind to a Cas12b nuclease (e.g., engineered) or its effector protein.
[0210] In some embodiments, DR or its functional variant comprises at least about 16 nucleotides (nt), such as 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40 or more nucleotides. In some embodiments, DR comprises about 20 nt to about 40 nt, for example, about 20 nt to about 30 nt, about 22 nt to about 40 nt, about 23 nt to about 38 nt, about 23 nt to about 36 nt, or about 30 nt to about 40 nt. In some embodiments, DR includes 22 nt, 23 nt or 24 nt. In some embodiments, DR includes 35 nt, 36 nt or 37 nt. In some embodiments, the sgRNA scaffold or its functional variant comprises about 50 nt to about 180 nt, for example, about 70 nt to about 140 nt, or any one of about 90 nt to about 120 nt.
[0211] In some embodiments, the sgRNA comprises a backbone sequence comprising a stem-loop structure (e.g., 1, 2, 3, 4, or more stem-loops) near the 5' end of the spacer sequence. In some embodiments, the stem comprises at least about 4 bp comprising complementary X and Y sequences, but stems of more, such as 5, 6, 7, 8, 9, 10, 11, or 12, or fewer, such as 3 or 2 base pairs, are also contemplated. Thus, for example, X2-10 and Y2-10 (wherein X and Y represent any complementary nucleotide groups) are contemplated. In some embodiments, the stem made of X and Y nucleotides, together with the loop, will form a complete hairpin in the entire secondary structure; and, this may be advantageous, and the number of base pairs may be any number that forms a complete hairpin. In some embodiments, any complementary X:Y base pairing sequence (e.g., with respect to length) may be tolerated as long as the secondary structure of the entire guide molecule is preserved. In some embodiments, the loop connecting the stem made of X:Y base pairs may be any sequence of the same length (e.g., 4 or 5 nucleotides) or longer that does not interrupt the overall secondary structure of the guide molecule. In some embodiments, the stem comprises about 5-7bp, which comprises complementary X and Y sequences, but stems of more or fewer base pairs can also be considered. In some embodiments, non-Watson-Crick base pairing is envisioned, wherein such pairing generally maintains the structure of a stem-loop at this position. In some embodiments, the stem contained in the backbone sequence comprises (e.g., consisting of) 5 pairs of complementary bases that hybridize to each other, and the loop length is 6, 7, 8, or 9 nucleotides. In some embodiments, the rod may comprise at least 2, at least 3, at least 4, or at least 5 base pairs. In some embodiments, the stem-loop structure comprises a first stem nucleotide chain of 5 nucleotides in length; a second stem nucleotide chain of 5 nucleotides in length, wherein the first and second stem nucleotide chains can hybridize to each other; and a cyclic nucleotide chain arranged between the first and second stem nucleotide chains, wherein the cyclic nucleotide chain comprises 6, 7, or 8 nucleotides.
[0212] In some embodiments, the natural hairpin or stem-loop structure of the guide molecule is extended or replaced by an extended stem-loop. In some cases, it has been shown that the extension of the stem can enhance the assembly of the guide molecule with the CRISPR-Cas protein (Chen et al., Cell. (2013); 155(7): 1479-1491). In some embodiments, the stem of the stem-loop is extended by at least 1, 2, 3, 4, 5 or more complementary base pairs (i.e., corresponding to the addition of 2, 4, 6, 8, 10 or more nucleotides in the guide molecule). In some embodiments, they are located at the end of the stem, adjacent to the loop of the stem-loop.
[0213] As used herein, the secondary structures of two or more sgRNAs or tracrRNAs are substantially identical or substantially different, which means that the sgRNAs or tracrRNAs comprise stems and / or loops that differ in length by no more than 1, 2, or 3 nucleotides; the nucleotide sequences of the sgRNAs or tracrRNAs differ by no more than 1, 2, 3, 4, 5, 6, 7, or 8 nucleotides in terms of nucleotide type (A, U, G, or C) when compared by sequence alignment. In some embodiments, the secondary structures of two or more sgRNAs or tracrRNAs are substantially identical or non-substantially different, which means that the sgRNAs or tracrRNAs contain stems that differ by at most a pair of complementary bases and / or loops that differ in length by at most one nucleotide, and / or contain stems of the same length but with base mismatches.
[0214] In some embodiments, the gRNA scaffold sequence that can direct any engineered Cas12b effector protein of the present application to the target site includes one or more nucleotide changes selected from nucleotide additions, insertions, deletions and / or deletions, and substitutions that do not result in substantial differences in secondary structure compared to the backbone sequence described in any one of SEQ ID NOs: 23 to 53 or a functionally truncated version thereof. In some embodiments, the gRNA scaffold comprises a sequence of any one of SEQ ID NOs: 25 to 53 or a variant thereof comprising a difference of up to about 10 nt (e.g., 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 nt).
[0215] In some embodiments, the guide RNA comprises crRNA. In some embodiments, the engineered CRISPR-Cas12b system comprises a precursor guide RNA array encoding multiple crRNAs. In some embodiments, the Cas12b effector protein cleaves the precursor guide RNA array to produce multiple crRNAs. In some embodiments, the engineered CRISPR-Cas12b system comprises a precursor guide RNA array encoding multiple crRNAs, wherein each crRNA comprises a different guide sequence. In some embodiments, the crRNA encoded by the precursor guide RNA array is associated with the tracrRNA.
[0216] Constructs and vectors
[0217] Also provided herein are constructs, vectors, and expression systems encoding any of the engineered Cas12b effector proteins described herein (including engineered Cas12b nucleases). In some embodiments, the construct, vector, or expression system further comprises one or more gRNAs (e.g., sgRNAs) or crRNA arrays.
[0218] A "vector" is a composition of matter that contains an isolated nucleic acid and can be used to deliver the isolated nucleic acid to the interior of a cell. Many vectors are known in the art, including but not limited to linear polynucleotides, polynucleotides associated with ions or amphiphilic compounds, plasmids, and viruses. Typically, a suitable vector contains a replication origin that functions properly in at least one organism, a promoter sequence, convenient restriction endonuclease sites, and one or more selectable markers. The term "vector" should also be interpreted to include non-plasmid and non-viral compounds that facilitate the transfer of nucleic acids into cells, such as polylysine compounds, liposomes, and the like.
[0219] In some embodiments, the vector is a viral vector. Examples of viral vectors include, but are not limited to, adenoviral vectors, adeno-associated viral vectors, lentiviral vectors, retroviral vectors, vaccinia viral vectors, herpes simplex viral vectors, and derivatives thereof. In some embodiments, the vector is a phage vector. Viral vector technology is well known in the art and is described in, for example, Sambrook et al. (2001 "Molecular Cloning: A Laboratory Manual" Cold Spring Harbor Laboratory, New York) and other virology and molecular biology manuals.
[0220] Many virus-based systems have been developed to transfer genes into mammalian cells. For example, retroviruses provide a convenient platform for gene delivery systems. Heterologous nucleic acids can be inserted into vectors and packaged into retroviral particles using techniques known in the art. The recombinant virus can then be isolated and delivered to engineered mammalian cells in vitro or ex vivo. Many retroviral systems are known in the art. In some embodiments, adenoviral vectors are used. Many adenoviral vectors are known in the art. In some embodiments, lentiviral vectors are used. In some embodiments, self-inactivating lentiviral vectors are used.
[0221] In certain embodiments, the vector is an adeno-associated virus (AAV) vector, such as AAV2, AAV8, or AAV9, which can be expressed as a vector containing at least 1×10 5 A single dose of adenovirus or adeno-associated virus containing particles (also called particle units, pu) is administered. In some embodiments, the dose is at least about 1×10 6 particles, at least about 1×10 7 particles, at least about 1×10 8 particles, or at least about 1×10 9 Delivery methods and dosages are described, for example, in WO2016205764 and U.S. Patent No. 8,454,972, the contents of each of which are incorporated herein by reference in their entirety.
[0222] In some embodiments, the vector is a recombinant adeno-associated virus (rAAV) vector. For example, in some embodiments, a modified AAV vector can be used for delivery. The modified AAV vector can be based on one or more of several capsid types, including AAV1, AV2, AAV5, AAV6, AAV8, AAV8.2, AAV9, AAVrh10, modified AAV vectors (e.g., modified AAV2, modified AAV3, modified AAV6), and pseudotyped AAV (e.g., AAV2 / 8, AAV2 / 5, and AAV2 / 6). Exemplary AAV vectors and techniques that can be used to produce rAAV particles are known in the art (see, e.g., Aponte-Ubillus et al. (2018) Appl. Microbiol. Biotechnol. 102(3):1045-54; Zhong et al. (2012) J. Genet. Syndr. Gene Ther. S1:008; West et al. (1987) Virology 160:38-47 (1987); Tratschin et al. (1985) Mol. Cell. Biol. l. Cell. Biol.)》5:3251-60; U.S. Patent Nos. 4,797,368 and 5,173,414; and International Publication Nos. WO2015 / 054653 and WO93 / 24641, each of which is incorporated herein by reference).
[0223] Any of the known AAV vectors for delivering Cas9 and other Cas12b proteins can be used to deliver the engineered Cas12b nuclease or effector protein or system of the present application.
[0224] Methods for introducing vectors into mammalian cells are known in the art. Vectors can be transferred into host cells by physical, chemical, or biological methods.
[0225] Physical methods for introducing vectors into host cells include calcium phosphate precipitation, lipofection, particle bombardment, microinjection, electroporation, and the like. Methods for producing cells containing vectors and / or exogenous nucleic acids are well known in the art. See, for example, Sambrook et al. (2001) Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory, New York. In some embodiments, vectors are introduced into cells by electroporation.
[0226] Biological methods for introducing heterologous nucleic acids into host cells include the use of DNA and RNA vectors. Viral vectors have become the most widely used method for inserting genes into mammalian (e.g., human) cells.
[0227] Chemical means for introducing vectors into host cells include colloidal dispersion systems, such as macromolecular complexes, nanocapsules, microspheres, beads, and lipid-based systems, including oil-in-water emulsions, micelles, mixed micelles, and liposomes. An exemplary colloidal system used as an in vitro delivery vehicle is a liposome (e.g., an artificial membrane vesicle). In some embodiments, the engineered CRISPR-Cas12b system is delivered as an RNP in a nanoparticle.
[0228] In some embodiments, a vector or expression system encoding a CRISPR-Cas12b system or components thereof comprises one or more selectable or detectable markers that provide a means for isolating or efficiently selecting cells that contain the CRISPR-Cas12b system and / or have been modified by the CRISPR-Cas12b system (e.g., at an early stage and on a large scale).
[0229] Reporter genes can be used to identify potential functionality of transfected cells and evaluation regulatory sequences. Typically, a reporter gene is a gene that is not present in or is not expressed by a recipient organism or tissue, and encodes a polypeptide whose expression is demonstrated by some easily detectable properties (e.g., enzymatic activity). The expression of the reporter gene is measured at the appropriate time after DNA has been introduced into the recipient cell. Suitable reporter genes can include genes encoding luciferase, beta galactosidase, chloramphenicol acetyltransferase, secretory alkaline phosphatase, or a green fluorescent protein gene (e.g., Ui-Tei et al. FEBS Letters 479:79-82 (2000)).
[0230] Other methods for confirming the presence of heterologous nucleic acid in host cells include, for example, molecular biological assays well known to those skilled in the art, such as Southern and Northern blotting, RT-PCR and PCR; and biochemical assays, such as detecting the presence or absence of specific peptides by immunological methods such as ELISA and Western blotting.
[0231] In some embodiments, the nucleic acid sequence encoding engineered Cas12b nuclease or effector protein and / or instructing RNA is operably connected to a promoter. In some embodiments, the promoter is an endogenous promoter relative to the cell engineered using the engineered CRISPR-Cas12b system. For example, the nucleic acid encoding the engineered Cas12b effector protein can be knocked into the genome of the engineered mammalian cell downstream of the endogenous promoter using any method known in the art. In some embodiments, the endogenous promoter is a promoter of abundant protein (such as beta actin). In some embodiments, the endogenous promoter is an inducible promoter, for example, it can be induced by the endogenous activation signal of the engineered mammalian cell. In some embodiments (wherein the engineered mammalian cell is a T cell), the promoter is a T cell activation dependent promoter (such as an IL-2 promoter, a NFAT promoter, or a NFκB promoter).
[0232] In some embodiments, the promoter is a heterologous promoter relative to the cell engineered using the engineered CRISPR-Cas12b system. Various promoters for gene expression in mammalian cells have been explored, and any of the promoters known in the art can be used in the present application. Promoters can be roughly classified as constitutive promoters or regulatable promoters, such as inducible promoters.
[0233] In some embodiments, the nucleic acid sequence encoding engineered Cas12b effector proteins and / or guide RNA is operably connected to a constitutive promoter. A constitutive promoter allows heterologous genes (also referred to as transgenics) to be constitutively expressed in host cells. Exemplary constitutive promoters contemplated herein include, but are not limited to, cytomegalovirus (CMV) promoters, human elongation factor-1α (hEF1α), ubiquitin C promoter (Ub iC), phosphoglycerokinase promoter (PGK), simian virus 40 early promoter (SV40), and the chicken beta actin promoter (CAG) coupled to the CMV early enhancer. In some embodiments, the promoter is a CAG promoter, which includes cytomegalovirus (CMV) early enhancer elements, promoters, the first exon and the first intron of the chicken beta actin gene, and the splicing acceptor of the rabbit beta globin gene.
[0234] In some embodiments, the nucleic acid sequence encoding the engineered CRISPR-Cas12b protein and / or guide RNA is operably linked to an inducible promoter. Inducible promoters belong to the category of regulated promoters. Inducible promoters can be induced by one or more conditions, such as the physical conditions, microenvironment, or physiological state of the host cell; inducer (i.e., inducer); or a combination thereof. In some embodiments, the induction conditions are selected from: inducer, radiation (such as ionizing radiation, light), temperature (such as heat), redox state, tumor environment, and the activation state of cells to be engineered by the engineered CRISPR-Cas12b system. In some embodiments, the promoter can be induced by small molecule inducers (such as chemical compounds). In some embodiments, small molecules are selected from: doxycycline, tetracycline, alcohol, metal, or steroid. Chemically induced promoters have been most widely explored. This promoter includes promoters whose transcriptional activity is regulated by the presence or absence of small molecule chemicals (such as doxycycline, tetracycline, alcohol, steroid, metal, and other compounds). The doxycycline inducible system with a reverse tetracycline-controlled transactivator (rtTA) and a tetracycline responsive element promoter (TRE) is currently the most mature system. WO9429442 describes the strict control of gene expression in eukaryotic cells by tetracycline-responsive promoters. WO9601313 discloses a tetracycline-regulated transcription modulator. Additionally, tetracycline technology (such as tetracycline-regulated (Tet-on) system) has been described on websites such as TetSystems.com. Any of the known chemically regulated promoters can be used to drive the expression of engineered CRISPR-Cas12b proteins and / or guide RNAs in this application.
[0235] In some embodiments, the nucleic acid sequence encoding engineered Cas12b nuclease or effector protein is codon optimized. In some embodiments, the expression construct encodes a tag (e.g., 10xHis tag) operably connected to the C-terminal end of the engineered Cas12b nuclease or effector protein. In some embodiments, the Cas12b construct encoding each engineered split is a fluorescent protein, such as GFP or RFP. Reporter protein can be used to assess the colocalization and / or dimerization (e.g., using a microscope) of the engineered split Cas12b protein. The nucleic acid sequence encoding the engineered Cas12b effector protein can be fused to the nucleic acid sequence encoding the additional component using a sequence encoding a self-cleavage peptide (e.g., T2A, P2A, E2A, or F2A peptide).
[0236] In some embodiments, an expression construct for a mammalian cell (e.g., a human cell) is provided, the expression construct comprising a nucleic acid sequence encoding an engineered Cas12b nuclease or effector protein. In some embodiments, the expression construct comprises a codon-optimized sequence encoding an engineered Cas12b nuclease or effector protein inserted into a pCAG-2A-eGFP vector such that the Cas12b protein is operably linked to eGFP. In some embodiments, a second vector is provided for expressing a guide RNA (e.g., sgRNA, crRNA, or pre-crRNA array) in a mammalian cell (e.g., a human cell). In some embodiments, the sequence encoding the guide RNA is expressed in a pUC19-U6-Aa-sgRNA vector backbone.
[0237] In some embodiments, the nucleic acid encoding Cas12b protein and the nucleic acid encoding gRNA are located on different vectors. In some embodiments, the nucleic acid encoding Cas12b protein and the nucleic acid encoding gRNA are located on the same vector. In some embodiments, the nucleic acid encoding Cas12b protein and the nucleic acid encoding gRNA are controlled by different promoters (such as CMV promoter and U6 promoter). In some embodiments, the nucleic acid encoding Cas12b protein is located upstream of the nucleic acid encoding gRNA. In some embodiments, the nucleic acid encoding Cas12b protein is located downstream of the nucleic acid encoding gRNA. In some embodiments, the nucleic acid encoding Cas12b protein and the nucleic acid encoding gRNA are simultaneously contacted with the target nucleic acid or introduced into the cell. In some embodiments, the nucleic acid encoding Cas12b protein and the nucleic acid encoding gRNA are sequentially contacted with the target nucleic acid or introduced into the cell, for example, before the nucleic acid encoding gRNA, the nucleic acid encoding Cas12a protein is introduced, or after the nucleic acid encoding gRNA, the nucleic acid encoding Cas12b protein is introduced. In some embodiments, the cell has expressed Cas12b protein. In some embodiments, only the nucleic acid encoding gRNA is introduced into the cell. In some embodiments, the cell has expressed gRNA.In some embodiments, only the nucleic acid encoding Cas12b protein is introduced into the cell.
[0238] III. How to use
[0239] One aspect of the present application provides methods for detecting a target nucleic acid or modifying a nucleic acid in vitro, ex vivo, or in vivo using any of the engineered Cas12b nucleases or effector proteins or CRISPR-Cas12b systems described herein, as well as methods for treating or diagnosing a target nucleic acid or modifying a nucleic acid using an engineered Cas12b nuclease or effector protein or CRISPR-Cas12b system. Also provided are uses of the engineered Cas12b effector protein or CRISPR-Cas12b system described herein for detecting or modifying nucleic acids in cells, and for treating or diagnosing a disease or condition in a subject; and compositions comprising any one of one or more components of the engineered Cas12b nuclease or effector protein or engineered CRISPR-Cas12b system for the manufacture of a medicament for detecting or modifying nucleic acids in cells and for treating or diagnosing a disease or condition in a subject.
[0240] Modification method
[0241] In some embodiments, the application provides a method for modifying a target nucleic acid comprising a target sequence, comprising contacting the target nucleic acid with any one of the engineered CRISPR-Cas12b systems described herein or a component thereof. For example, when the Cas12b protein or the nucleic acid encoding it is already present, only gRNA or the nucleic acid encoding it needs to be further provided; when the gRNA or the nucleic acid encoding it is already present, only Cas12b protein or the nucleic acid encoding it needs to be further provided. In some embodiments, a method for modifying a target nucleic acid comprising a target sequence is provided, comprising contacting the target nucleic acid with a CRISPR-Cas12b system (e.g., engineered, non-naturally occurring) (e.g., in vitro, ex vivo, or in vivo), wherein the CRISPR-Cas12b system comprises: (a) an engineered Cas12b nuclease or its effector protein (e.g., a nickase, a splitting Cas12b nuclease, a cleaving ... s12b protein, transcription repressor, transcription activator, base editor or guide editor), which comprises one, two or three types of mutations relative to a reference Cas12b nuclease, wherein the mutations include: (1) replacing one or more amino acid residues in the reference Cas12b nuclease that interacts with the PAM (e.g., one or more of the following positions: 116, 123, 130, 132, 144, 145, 153, 173, 222, 395, 400 and 475) with an amino acid residue having an aromatic ring (e.g., F, Y, W) or more of: 118 and 119); and / or (3) replacing one or more amino acid residues in the RuvC domain of the reference Cas12b nuclease that interacts with the ssDNA substrate (e.g., one or more of the following positions: 300, 301, 304, 329, 636, 639, 647, 682, 757, 758, 761, 764, 768, 852, 854, 856, 857, 858, 860, 862, 863, 865, 866, 867, 869, 938, 956, 957, 958, 994, 1093, and 1097) with a positively charged amino acid residue (e.g., R, H, K) or a hydrophobic amino acid residue (e.g., F, Y, W, M), wherein the reference Cas12b nuclease comprises SEQ IDNO:1 amino acid sequence, or a nucleic acid encoding an engineered Cas12b nuclease or its effector protein; and (b) gRNA comprising a guide sequence complementary to the target sequence of the target nucleic acid or a nucleic acid encoding the gRNA, resulting in the target nucleic acid being modified by the engineered Cas12b nuclease or its effector protein. In some embodiments, the gRNA comprises a backbone comprising any one of SEQ ID NOs: 23 and 25 to 53.In some embodiments, the engineered Cas12b nuclease or its effector protein comprises the amino acid sequence of any one of SEQ ID NOs: 2 to 22 and 79 to 81. In some embodiments, a method of modifying a target nucleic acid comprising a target sequence is provided, the method comprising contacting the target nucleic acid with a CRISPR-Cas12b system (e.g., engineered, non-naturally occurring) (e.g., in vitro, ex vivo, or in vivo), wherein the CRISPR-Cas12b system comprises one, two, or three types of mutations relative to a reference Cas12b nuclease, wherein the mutations comprise: (1) substitution of one or more amino acid residues in the reference Cas12b nuclease that interacts with the PAM with a positively charged amino acid residue (e.g., R, H, K), such as one or more of the following positions: 116, 123, 130, 132, 144, 145, 153, 173, 222, 395, 400, and 475; and / or (2) substitution of the reference Cas12b nucleic acid with an amino acid residue having an aromatic ring (e.g., F, Y, W). (i) one or more amino acid residues in the RuvC domain of a reference Cas12b nuclease that interacts with an ssDNA substrate (e.g., one or more of the following positions: 300, 301, 304, 329, 636, 639, 647, 682, 757, 758, 761, 764, 768, 852, 854, 856, 857, 858, 860, 862, 863, 865, 866, 867, 869, 938, 956, 957, 958, 994, 1093, and 1097) that are involved in opening a double-stranded DNA in the enzyme; and / or (ii) one or more amino acid residues in the RuvC domain of a reference Cas12b nuclease that interacts with an ssDNA substrate (e.g., one or more of the following positions: 300, 301, 304, 329, 636, 639, 647, 682, 757, 758, 761, 764, 768, 852, 854, 856, 857, 858, 860, 862, 863, 865, 866, 867, 869, 938, 956, 957, 958, 994, 1093, and 1097) that are replaced with a positively charged amino acid (e.g., R, H, K) or a hydrophobic amino acid residue (e.g., F, Y, W, M), wherein the reference Cas12b nuclease comprises SEQ IDNO:1 amino acid sequence, or encoding Cas12b nuclease (e.g., engineered) or its effector protein nucleic acid; and (b) gRNA, which comprises a guide sequence complementary to the target sequence of the target nucleic acid or a nucleic acid encoding the gRNA, wherein the gRNA comprises an engineered scaffold, and the engineered scaffold comprises the sequence of any one of SEQ ID NO:25 to 53; wherein the hybridization of the guide sequence and the target sequence of the target nucleic acid mediates the contact of the Cas12b nuclease (e.g., engineered) or its effector protein with the target sequence of the target nucleic acid, resulting in the target nucleic acid being modified by the Cas12b nuclease (e.g., engineered) or its effector protein. In some embodiments, the engineered Cas12b nuclease or its effector protein comprises SEQ ID NO:2 to 22 and 79 to 81.In some embodiments, the method also includes providing a repair / donor template comprising a repair / donor nucleic acid, wherein the nucleic acid of the repair / donor can be incorporated into the target nucleic acid of the modification at the target sequence (for example, by homologous recombination). In some embodiments, the modification of the target nucleic acid repairs the mutation (for example, loss of function mutation) in the target nucleic acid to a wild-type (or non-detrimental) sequence. In some embodiments, the modification of the target nucleic acid introduces an exogenous sequence. In some embodiments, the method is carried out in vitro. In some embodiments, the target nucleic acid is present in a cell. In some embodiments, the cell is a bacterial cell, a yeast cell, a plant cell, or an animal cell (for example, a mammalian cell, such as a human or mouse cell). In some embodiments, the method is carried out in vitro. In some embodiments, the method is carried out in vivo.
[0242] In some embodiments, the target nucleic acid is cut or the target sequence in the target nucleic acid is changed (e.g., base editing) by an engineered CRISPR-Cas12b system. In some embodiments, the expression of the target nucleic acid is changed by an engineered CRISPR-Cas12b system. In some embodiments, the target nucleic acid is a genomic DNA, such as in a cell. In some embodiments, the target sequence is associated with a disease or illness. In some embodiments, the method for modifying the target sequence treats a disease or illness associated with the target sequence. In some embodiments, the engineered CRISPR-Cas12b system includes a precursor guide RNA array encoding multiple crRNAs, wherein each crRNA includes different guide sequences.
[0243] In some embodiments, the present application provides a method for treating a disease or condition associated with a target nucleic acid in a cell of an individual, comprising modifying a target nucleic acid in a cell of the individual using any of the methods for modifying a target nucleic acid described herein, thereby treating the disease or condition. In some embodiments, the disease or condition is selected from the group consisting of cancer, cardiovascular disease, genetic disease, autoimmune disease, metabolic disease, neurodegenerative disease, eye disease, bacterial infection, and viral infection.
[0244] The engineered CRISPR-Cas12b system described herein can modify the target nucleic acid in the cell in various ways, depending on the type of engineered Cas12b effector protein in the CRISPR-Cas12b system. In some embodiments, the method induces site-specific cutting in the target nucleic acid. In some embodiments, the method cuts the genomic DNA in the cell (such as a bacterial cell, a plant cell, or an animal cell (e.g., a mammalian cell). In some embodiments, the method kills the cell by cutting the genomic DNA in the cell. In some embodiments, the method cuts the viral nucleic acid in the cell. In some embodiments, the method base edits the target nucleic acid, for example, repairing harmful or disease-related mutations to non-disease-related sequences. In some embodiments, the method enhances (e.g., increases at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 1 times, 2 times, 5 times, 10 times, 20 times or more) expression of the target nucleic acid (e.g., fixing harmful mutations that downregulate expression). In some embodiments, the methods reduce (e.g., reduce by at least about any of 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 1-fold, 2-fold, 5-fold, 10-fold, 20-fold, or more) expression of a target nucleic acid (e.g., fix a deleterious mutation that upregulates expression).
[0245] In some embodiments, the method changes (such as increases or decreases) the expression level of the target nucleic acid in the cell. In some embodiments, the method increases the expression level of the target nucleic acid in the cell (e.g., using an engineered Cas12b effector protein based on an inactive Cas12b protein fused to a transactivation domain (e.g., any of SEQ ID NOs: 79 to 81). In some embodiments, the method reduces the expression level of the target nucleic acid in the cell (e.g., using an engineered Cas12b effector protein based on an inactive Cas12b protein fused to a transcriptional repressor domain (such as a KR AB domain) (e.g., any of SEQ ID NOs: 79 to 81). In some embodiments, the method introduces an epigenetic modification into the target nucleic acid in the cell (e.g., using an engineered Cas12b effector protein based on an inactive Cas12b protein fused to an epigenetic modification domain (e.g., any of SEQ ID NOs: 79 to 81). In some embodiments, the method introduces base editing into a target nucleic acid in a cell, for example, using an engineered Cas12b effector protein based on an enzymatically inactive Cas12b protein (e.g., any of SEQ ID NOs: 79 to 81) fused to a cytosine deaminase domain or an adenosine deaminase domain (e.g., TadA) or a functional fragment thereof. The engineered Cas12b system described herein can be used to introduce other modifications into a target nucleic acid, depending on the functional domains comprised by the engineered Cas12b effector protein.
[0246] In some embodiments, the method changes the target sequence in the target nucleic acid in the cell. In some embodiments, the method introduces mutations into the target nucleic acid in the cell. In some embodiments, the method uses one or more endogenous DNA repair pathways in the cell, such as non-homologous end joining (NHEJ) or homology-directed recombination (HDR), to repair the double-strand breaks induced in the target DNA due to sequence-specific cutting by the CRISPR complex. Exemplary mutations include but are not limited to insertions, deletions, substitutions, and frameshifts. In some embodiments, the method inserts donor DNA at the target locus. In some embodiments, the insertion of donor DNA causes a selection marker or reporter protein to be introduced into the cell. In some embodiments, the insertion of donor DNA causes gene knock-in. In some embodiments, the insertion of donor DNA causes knockout mutations. In some embodiments, the insertion of donor DNA causes substitution mutations, such as single nucleotide substitutions. In some embodiments, the method induces phenotypic changes in cells.
[0247] In some embodiments, the engineered CRISPR-Cas12b system is used as part of a genetic circuit, or for inserting a genetic circuit into the genomic DNA of a cell. The engineered split Cas12b effector protein controlled by the inducer described herein can be used as a component of a genetic circuit. The genetic circuit can be used for gene therapy. The methods and techniques for designing and using genetic circuits are known in the art. Reference can be made to, for example, Brophy, Jennifer AN, and Christopher A. Voigt. "Principles of genetic circuit design" Nature methods 11.5 (2014): 508.
[0248] The engineered CRISPR-Cas12b system described herein can be used to modify a wide range of target nucleic acids. In some embodiments, the target nucleic acid is in a cell. In some embodiments, the target nucleic acid is genomic DNA. In some embodiments, the target nucleic acid is extrachromosomal DNA. In some embodiments, the target nucleic acid is exogenous to the cell. In some embodiments, the target nucleic acid is a viral nucleic acid, such as viral DNA. In some embodiments, the target nucleic acid is a plasmid in a cell. In some embodiments, the target nucleic acid is a horizontally transferred plasmid. In some embodiments, the target nucleic acid is RNA, such as mRNA.
[0249] In some embodiments, the target nucleic acid is an isolated nucleic acid, such as an isolated DNA. In some embodiments, the target nucleic acid is present in a cell-free environment. In some embodiments, the target nucleic acid is an isolated vector, such as a plasmid. In some embodiments, the target nucleic acid is an isolated linear DNA fragment.
[0250] The methods described herein are applicable to any suitable cell type. In some embodiments, the cell is a bacterium, a yeast cell, a fungal cell, an algae cell, a plant cell, or an animal cell. (For example, a mammalian cell, such as a human cell). In some embodiments, the cell is a cell isolated from a natural source, such as a tissue biopsy. In some embodiments, the cell is a cell isolated from a cell line cultured in vitro. In some embodiments, the cell is from a primary cell line. In some embodiments, the cell is from an immortalized cell line. In some embodiments, the cell is a genetically engineered cell.
[0251] In some embodiments, the cell is an animal cell from an organism including, but not limited to, cat, dog, mouse, rat, hamster, cow, sheep, goat, horse, pig, deer, chicken, duck, goose, rabbit, and fish.
[0252] In some embodiments, the cell is a plant cell from an organism selected from the group consisting of corn, wheat, barley, oats, rice, soybean, oil palm, safflower, sesame, tobacco, flax, cotton, sunflower, pearl millet, millet, sorghum, rapeseed, hemp, vegetable crops, forage crops, cash crops, tree crops, and biomass crops.
[0253] In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a mouse cell, such as a Neuro 2A (N2a) cell. In some embodiments, the cell is a human cell. In some embodiments, the human cell is a human embryonic kidney 293T (HEK293T or 293T) cell or a HeLa cell. In some embodiments, the mammalian cell is selected from the group consisting of an immune cell, a liver cell, a tumor cell, a stem cell, a neuronal cell, a zygote, a muscle cell, and a skin cell.
[0254] In some embodiments, the cell is an immune cell selected from the group consisting of cytotoxic T cells, helper T cells, natural killer (NK) T cells, iNK-T cells, NK-T-like cells, γδT cells, tumor-infiltrating T cells, and dendritic cells (DC) activated T cells. In some embodiments, the method produces modified immune cells such as CAR-T cells, CAR-NK cells, or TCR-T cells.
[0255] In some embodiments, the cell is an embryonic stem (ES) cell, an induced pluripotent stem (iPS) cell, a gamete progenitor cell, a gamete, a zygote, or a cell in an embryo.
[0256] The methods described herein can be used to modify target cells in vivo, in vitro, or in vitro, and can be performed in such a way that the cells are altered so that, once modified, the descendants or cell lines of the modified cells retain the altered phenotype. The modified cells and descendants can be part of a multicellular organism, such as a plant or animal, that is subjected to in vitro or in vivo applications such as genome editing and gene therapy.
[0257] In some embodiments, the modification method is carried out in vitro. In some embodiments, after the engineered CRISPR-Cas12b system is introduced into the cell, the modified cell (e.g., mammalian cell) is propagated in vitro. In some embodiments, the modified cell is cultured so that it is propagated for at least about 1 day, 2 days, 3 days, 4 days, 5 days, 6 days, 7 days, 10 days, 12 days, or any one of 14 days. In some embodiments, the modified cell is cultivated for no more than about 1 day, 2 days, 3 days, 4 days, 5 days, 6 days, 7 days, 10 days, 12 days, or any one of 14 days. In some embodiments, the modified cell is further evaluated or screened by PCR or sequencing to select cells with one or more ideal phenotypes or properties.
[0258] In some embodiments, the target sequence is a sequence associated with a disease or condition. Exemplary diseases or conditions include, but are not limited to, cancer, blood disorders, cardiovascular diseases, genetic diseases, autoimmune diseases, metabolic diseases, neurological diseases, neurodegenerative diseases, eye diseases, bacterial infections, and viral infections. In some embodiments, the disease or condition is graft versus host disease (GvHD) or host versus graft disease (HvG). In some embodiments, the disease or condition is a genetic disease. In some embodiments, the disease or condition is a monogenic disease or condition. In some embodiments, the disease or condition is a polygenic disease or condition.
[0259] In some embodiments, the target sequence has a mutation compared to the wild-type sequence.In some embodiments, the target sequence has a single nucleotide polymorphism (SNP) associated with a disease or condition.
[0260] In some embodiments, the donor DNA encoding biological product inserted into the target nucleic acid is selected from: reporter protein, antigen-specific receptor, therapeutic protein, antibiotic resistance protein, RNAi molecule, cytokine, kinase, antigen, antigen-specific receptor, chimeric receptor, cytokine receptor, and suicide polypeptide. In some embodiments, the donor DNA encodes a therapeutic protein, such as a cytokine. In some embodiments, the donor DNA encodes a therapeutic protein that can be used for gene therapy. In some embodiments, the donor DNA encodes a therapeutic antibody. In some embodiments, the donor DNA encodes an engineered receptor, such as a chimeric antigen receptor (CAR) or an engineered TCR. In some embodiments, the donor DNA encodes a therapeutic RNA, such as a small RNA (e.g., siRNA, shRNA, or miRNA) or a long non-coding RNA (lincRNA).
[0261] The methods described herein can be used to perform multiple gene editing or regulation at two or more (e.g., 2, 3, 4, 5, 6, 8, 10 or more) different target loci. In some embodiments, the method detects or modifies multiple target nucleic acids or target nucleic acid sequences. In some embodiments, the method comprises contacting the target nucleic acid with a guide RNA comprising multiple (e.g., 2, 3, 4, 5, 6, 8, 10 or more) crRNA sequences, wherein each crRNA comprises different target sequences.
[0262] Also provided are engineered cells comprising a modified target nucleic acid, produced using any of the modification methods described herein. The engineered cells can be used for cell therapy. Autologous or allogeneic cells can be used to prepare the engineered cells using the methods described herein for cell therapy.
[0263] The methods described herein can also be used to generate isogenic lines of cells (eg, mammalian cells) to study genetic variants.
[0264] Also provided are engineered plants or non-human animals comprising the engineered cells described herein. In some embodiments, the engineered plants or non-human animals are genome-edited non-human animals. The engineered plants or non-human animals can be used as disease models.
[0265] The technology of producing non-human through genome editing or transgenic animals is well known in the art, and includes but is not limited to pronuclear microinjection, viral infection, and the transformation of embryonic stem cells and induced pluripotent stem (iPS) cells. The detailed method that can be used includes but is not limited to those described in the following two: Sundberg and Ichiki (2006 " Genetically Engineered Mice Handbook " CRC Press) and Gibson (2004 " Genome Science Brief Talk (A Primer Of Genome Science) " 2nd edition, Sunderland, Mass., Sinauer Press).
[0266] The engineered animal can be any suitable species, including but not limited to species such as bovines, equines, ovines, canines, cervids, felines, goats, porcines, primates, and less commonly known mammals such as elephants, deer, zebras, or camels.
[0267] Treatment
[0268] Also provided are therapeutic methods using any of the methods described herein for modifying a target nucleic acid in a cell, and diagnostic methods using any of the methods described herein for detecting a target nucleic acid.
[0269] In some embodiments, the application provides a method for treating a disease or condition associated with a target nucleic acid in a cell of an individual, comprising contacting the target nucleic acid with any one of the engineered CRISPR-Cas12b systems described herein, wherein the guide sequence of the guide RNA is complementary to the target sequence of the target nucleic acid, wherein the Cas12b nuclease (e.g., engineered) or its effector protein (e.g., including any of SEQ ID NO: 1 to 22 and 79 to 81) and the guide RNA are associated with each other to bind to the target nucleic acid so as to modify the target nucleic acid, thereby treating a disease or condition. In some embodiments, a mutation (e.g., knockout or knock-in mutation) is introduced into the target nucleic acid. In some embodiments, the expression of the target nucleic acid is enhanced. In some embodiments, the expression of the target nucleic acid is suppressed.
[0270] In some embodiments, the present application provides a method of treating a disease or condition in an individual, comprising administering to the individual an effective amount of any one of the engineered CRISPR-Cas12b systems described herein and a donor DNA encoding a therapeutic agent, wherein the guide sequence of the guide RNA is complementary to the target sequence of the target nucleic acid of the individual, wherein the Cas12b nuclease (e.g., engineered) or its effector protein (e.g., including any of SEQ ID NOs: 1 to 22 and 79 to 81) and the guide RNA associate with each other to bind to the target nucleic acid and insert the donor DNA into the target sequence, thereby treating the disease or condition.
[0271] In some embodiments, the application provides a method for treating a disease or disorder of an individual, comprising administering to the individual an effective amount of an engineered cell comprising a modified target nucleic acid, wherein the engineered cell is prepared by contacting the cell with any one of the engineered CRISPR-Cas12b systems described herein, wherein the guide sequence of the guide RNA is complementary to the target sequence of the target nucleic acid, wherein the engineered Cas12b nuclease (e.g., engineered) or its effector protein (e.g., including SEQ ID NO: 1 to 22 and 79 to 81) and the guide RNA are associated with each other to bind to the target nucleic acid to modify the target nucleic acid. In some embodiments, the engineered cell is an immune cell.
[0272] In some embodiments, a method of treating a disease or condition associated with a target nucleic acid in a cell of an individual (e.g., a human) is provided, comprising contacting the target nucleic acid with the individual (e.g., in vitro or in vivo) or administering to the individual an effective amount of a CRISPR-Cas12b system (e.g., engineered, non-naturally occurring), wherein the CRISPR-Cas12b system comprises: (1) replacing one or more amino acid residues in a reference Cas12b nuclease that interacts with a PAM with a positively charged amino acid residue (e.g., R, H, K), such as one or more of the following positions: 116, 123, 130, 132, 144, 145, 153, 173, 222, 395, 400, and 475; and / or (2) replacing one or more amino acid residues in the reference Cas12b nuclease that are involved in opening DNA with an amino acid residue having an aromatic ring (e.g., F, Y, W). (i) one or more amino acid residues in the RuvC domain of the reference Cas12b nuclease that interacts with the ssDNA substrate (e.g., one or more of the following positions: 300, 301, 304, 329, 636, 639, 647, 682, 757, 758, 761, 764, 768, 852, 854, 856, 857, 858, 860, 862, 863, 865, 866, 867, 869, 938, 956, 957, 958, 994, 1093, and 1097) are replaced by a positively charged amino acid residue (e.g., R, H, K) or a hydrophobic amino acid residue (e.g., F, Y, W, M), wherein the reference Cas12b nuclease comprises SEQ ID NO:1 amino acid sequence, or a nucleic acid encoding an engineered Cas12b nuclease or its effector protein; and (b) a gRNA comprising a guide sequence complementary to the target sequence of the target nucleic acid or a nucleic acid encoding the gRNA, resulting in the target nucleic acid being modified by the engineered Cas12b nuclease or its effector protein, thereby treating a disease or condition. In some embodiments, the gRNA comprises a backbone comprising any one of SEQ ID NOs: 23 and 25 to 53.In some embodiments, a method of treating a disease or condition associated with a target nucleic acid in a cell of an individual (e.g., a human) is provided, comprising contacting the target nucleic acid with the individual (e.g., in vitro or in vivo) or administering to the individual an effective amount of a CRISPR-Cas12b system (e.g., engineered, non-naturally occurring), wherein the CRISPR-Cas12b system comprises one, two, or three types of mutations relative to a reference Cas12b nuclease, wherein the mutations comprise: (1) substitution of one or more amino acid residues in the reference Cas12b nuclease that interacts with the PAM with a positively charged amino acid residue (e.g., R, H, K), such as one or more of the following positions: 116, 123, 130, 132, 144, 145, 153, 173, 222, 395, 400, and 475; and / or (2) substitution of an amino acid residue having an aromatic ring (e.g., F, Y, W) with an amino acid residue having an aromatic ring (e.g., F, Y, W). (i) replacing one or more amino acid residues in the RuvC domain of the reference Cas12b nuclease that interacts with the ssDNA substrate (e.g., one or more of the following positions: 118 and 119) with a positively charged amino acid residue (e.g., R, H, K) or a hydrophobic amino acid residue (e.g., F, Y, W, M) than one or more amino acid residues in the Cas12b nuclease that participate in opening the DNA double strand (e.g., one or more of the following positions: 118 and 119); and / or (ii) replacing one or more amino acid residues in the RuvC domain of the reference Cas12b nuclease that interacts with the ssDNA substrate (e.g., one or more of the following positions: one or more: 300, 301, 304, 329, 636, 639, 647, 682, 757, 758, 761, 764, 768, 852, 854, 856, 857, 858, 860, 862, 863, 865, 866, 867, 869, 938, 956, 957, 958, 994, 1093, and 1097), wherein the reference Cas12b nuclease comprises SEQ ID NO:1 amino acid sequence, or a nucleic acid encoding a Cas12b nuclease (e.g., engineered) or its effector protein; and (b) a gRNA comprising a guide sequence complementary to the target sequence of the target nucleic acid or a nucleic acid encoding the gRNA, wherein the gRNA comprises an engineered scaffold comprising a sequence of any one of SEQ ID NOs: 25 to 53; wherein hybridization of the guide sequence and the target sequence of the target nucleic acid mediates contact of the Cas12b nuclease (e.g., engineered) or its effector protein with the target sequence of the target nucleic acid, resulting in modification of the target nucleic acid by the Cas12b nuclease (e.g., engineered) or its effector protein, thereby treating a disease or disorder. In some embodiments, the engineered Cas12b nuclease or its effector protein comprises an amino acid sequence of any one of SEQ ID NOs: 2 to 22 and 79 to 81.In some embodiments, the method further comprises contacting the target nucleic acid (e.g., in vitro or in vivo) with an effective amount of repair / donor nucleic acid or administering an effective amount of repair / donor nucleic acid to an individual, wherein the nucleic acid of the repair / donor can be incorporated into the modified target nucleic acid at the target sequence (e.g., by homologous recombination). In some embodiments, the modification of the target nucleic acid repairs the mutation (e.g., loss of function mutation) in the target nucleic acid to a wild-type (or non-detrimental) sequence. In some embodiments, the modification of the target nucleic acid introduces an exogenous sequence.
[0273] In some embodiments, the subject is a human. In some embodiments, the subject is an animal, e.g., a model animal such as a rodent (e.g., mouse, rat, hamster), a pet (e.g., cat, dog, rabbit), or a farm animal (e.g., horse, cow, sheep, goat, donkey, pig). In some embodiments, the subject is a mammal.
[0274] In some embodiments, the disease or condition is associated with an abnormality (e.g., pathogenic point mutation) in the target nucleic acid of an individual (e.g., a person). In some embodiments, the disease or condition is treated due to modification (e.g., cutting, base editing, or repair) of the target nucleic acid by the CRISPR-Cas12b system or complex (e.g., repair abnormality). In some embodiments, the disease is caused by overexpression or misexpression (e.g., missense mutation, frameshift mutation, nonsense mutation) of one or more target genes, wherein the CRISPR-Cas12b system or complex can target the one or more target genes to target modification, such as cutting, base editing, or sequence repair (e.g., by further introducing repair / donor template to repair the target gene cut by the CRISPR-Cas12b system or complex by homologous recombination).
[0275] In some embodiments, the disease or condition is selected from the group consisting of cancer, cardiovascular disease, genetic disease, autoimmune disease, metabolic disease, neurodegenerative disease, eye disease, bacterial infection, and viral infection.
[0276] In some embodiments, the disease or disorder is selected from transthyretin amyloidosis (ATTR) (e.g., transthyretin-related wild-type amyloidosis (ATTRwt), transthyretin-related hereditary amyloidosis (ATTRm), familial amyloid polyneuropathy (FAP, ATTR-PN), or familial amyloid cardiomyopathy (FAC, ATTR-CM)), cystic fibrosis, hereditary angioedema (HAE), diabetes, pseudohypertrophic muscular dystrophy, Becker muscular dystrophy (BMD), alpha-1 antitrypsin deficiency (AAT deficiency), Pompe disease, myotonic dystrophy, Huntington's disease, fragile X syndrome (FXS), Friedreich's ataxia (FRDA), amyotrophic lateral sclerosis (ALS), frontotemporal dementia (FTD), hereditary chronic kidney disease, hyperlipidemia, hypercholesterolemia (e.g., familial hypercholesterolemia), Leber congenital amaurosis (LCA), sickle cell disease (SCD), and beta-thalassemia. In some embodiments, the CRISPR-Cas12b system or complex is packaged and delivered via lipid nanoparticles. In some embodiments, the lipid nanoparticles are administered to an individual by intravenous injection or infusion.
[0277] In some embodiments, the target nucleic acid is PCSK9. In some embodiments, the disease or condition is cardiovascular disease. In some embodiments, the disease or condition is coronary artery disease. In some embodiments, the method reduces cholesterol levels in an individual. In some embodiments, the method treats diabetes in an individual. In some embodiments, the disease or condition is hypercholesterolemia, such as familial hypercholesterolemia.
[0278] In some embodiments, the target nucleic acid is HBG1 and / or HBG2. In some embodiments, the disease or condition is sickle cell disease or beta-thalassemia. In some embodiments, the disease or condition is a hereditary persistence of fetal hemoglobin (HPFH), HbS gene deletion HPFH, or HbSHPFH due to a point mutation.
[0279] In some embodiments, the target nucleic acid is CC chemokine receptor (CCR) 5 (CCR5), which encodes the primary HIV-1 coreceptor. In some embodiments, the disease or condition is an infectious disease, such as AIDS. In some embodiments, the disease or condition is a non-infectious disease, such as cancer (e.g., breast cancer or prostate cancer), atherosclerosis, stroke, or inflammatory bowel disease (IBD).
[0280] In some embodiments, the target nucleic acid is CD34. In some embodiments, the disease or disorder is cancer.
[0281] In some embodiments, the target nucleic acid is RING finger protein 2 (RNF2).In some embodiments, the disease or condition is a neurological disorder, such as Luo-Schoch Yamamoto syndrome or nonspecific syndrome intellectual disability.
[0282] Detection method
[0283] The application also provides a method for detecting a target nucleic acid using an engineered Cas12b nuclease or its effector protein (e.g., comprising SEQ ID NO: 2 to 22 and 79 to 81) or any one of the CRISPR-Cas12b systems with improved activity. The use of Cas12b effector protein as a detection agent utilizes such a discovery that V-type CRISPR / Cas12 protein (e.g., Cas12a, Cas12b, Cas12c, Cas12d, Cas12e (CasX), and Cas12i) can promiscuously cut non-targeted single-stranded DNA (ssDNA) once activated by detection of target DNA. The method using Cas12b protein as a detection agent has been described in, for example, US10253365 and WO2020 / 056924, the contents of which are incorporated herein by reference in their entirety. In some embodiments, the target nucleic acid in the detection sample is used to diagnose a disease or illness.
[0284] In some embodiments, once the Cas12b effector protein is activated by the guide RNA (which occurs when the sample includes target DNA that hybridizes to the guide RNA (i.e., the sample includes targeted DNA)), the Cas12b nuclease or its effector protein becomes a nuclease that promiscuously cleaves single-stranded nucleic acids (e.g., non-target ssDNA or RNA, i.e., single-stranded nucleic acids that do not hybridize to the guide sequence of the guide RNA). Thus, when the target DNA (double-stranded or single-stranded) is present in the sample (e.g., in some cases above a threshold amount), the result is cleavage of single-stranded nucleic acids in the sample, which can be detected using any convenient detection method (e.g., using a labeled single-stranded detection nucleic acid, such as DNA or RNA). Cas12b can cleave both ssDNA and ssRNA.
[0285] In some embodiments, a method of detecting target DNA (e.g., double-stranded or single-stranded) in a sample is provided, comprising: (a) contacting the sample with: (i) any of the engineered Cas12b nucleases described herein, or effector proteins thereof (e.g., comprising any of SEQ ID NOs: 2 to 22 and 79 to 81); (ii) a guide RNA comprising a guide sequence that hybridizes to the target DNA; and (iii) a detection nucleic acid that is single-stranded (i.e., a "single-stranded detection nucleic acid") and does not hybridize to the guide sequence of the guide RNA; and (b) measuring a detectable signal produced by cleavage of the single-stranded detection nucleic acid by the engineered Cas12b effector protein. In some embodiments, a method of detecting target DNA (e.g., double-stranded or single-stranded) in a sample is provided, comprising: (a) contacting the sample with (i) any one of the Cas12b nucleases (e.g., engineered or wild-type) described herein or its effector proteins (e.g., comprising any one of SEQ ID NOs: 1 to 22 and 79 to 81); (ii) a guide RNA comprising a guide sequence that hybridizes to the target DNA, and an engineered scaffold comprising any one of SEQ ID NOs: 25 to 53; and (iii) a detector nucleic acid that is single-stranded (i.e., a "single-stranded detector nucleic acid") and does not hybridize to the guide sequence of the guide RNA; and (b) measuring a detectable signal generated by cleavage of the single-stranded detector nucleic acid by the engineered Cas12b effector protein. In some embodiments, a method for detecting a target nucleic acid in a sample is provided, comprising: (a) contacting the sample with any engineered CRISPR-Cas12b system described herein and a labeled detector nucleic acid, wherein the gRNA comprises a guide sequence complementary to the target sequence of the target nucleic acid, and wherein the labeled detector nucleic acid is single-stranded and does not hybridize with the guide sequence of the gRNA; and (b) measuring the detectable signal generated by cutting the labeled detector nucleic acid by the engineered CRISPR-Cas12b system, thereby detecting the target nucleic acid. In some cases, the single-stranded detection nucleic acid includes a fluorescent emission dye pair (e.g., the fluorescent emission dye pair is a fluorescence resonance energy transfer (FRET) pair, a quencher / fluorescent agent pair). In some cases, the target DNA is viral DNA (e.g., papovavirus, hepadnavirus, herpes virus, adenovirus, poxvirus, parvovirus, etc.). In some embodiments, the single-stranded detection nucleic acid is DNA. In some embodiments, the single-stranded detection nucleic acid is RNA. In some embodiments, the engineered Cas12b effector protein is an engineered Cas12b nuclease. In some embodiments, the method is performed in vitro. In some embodiments, the target nucleic acid is present in a cell, such as a bacterial cell, a yeast cell, a plant cell, or an animal cell. In some embodiments, the method is performed in vitro. In some embodiments, the method is performed in vivo. In some embodiments, the target nucleic acid is genomic DNA.In some embodiments, the target sequence is associated with a disease or disorder.
[0286] The method for detecting a target DNA (single-stranded or double-stranded) in a sample disclosed herein can detect the target DNA with high sensitivity. In some cases, the method disclosed herein can be used to detect a target DNA present in a sample containing a plurality of DNAs (including a target DNA and a plurality of non-target DNAs), wherein the target DNA is present at a concentration of 10 7 The non-target DNA is present in one or more copies (e.g., 6 There are one or more copies of non-target DNA, and every 10 5 There are one or more copies of non-target DNA, and every 10 4 There are one or more copies of non-target DNA, and every 10 3 There are one or more copies of non-target DNA. 2 non-target DNA, one or more copies per 50 non-target DNAs, one or more copies per 20 non-target DNAs, one or more copies per 10 non-target DNAs, or one or more copies per 5 non-target DNAs).
[0287] In some embodiments, the engineered Cas12b nuclease or its effector protein described herein (e.g., comprising any of SEQ ID NOs: 2 to 22) can detect target DNA with greater sensitivity compared to a reference Cas12b nuclease (e.g., SEQ ID NO: 1). In some embodiments, the engineered Cas12b effector protein can detect target DNA with 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or greater sensitivity compared to a reference Cas12b nuclease.
[0288] Delivery Method
[0289] In some embodiments, the engineered CRISPR-Cas12b system described herein, or its components, its nucleic acid molecules, or nucleic acid molecules encoding or providing its components, can be delivered to a host cell by various delivery systems such as plasmids or viral vectors (e.g., any of the vectors described in the "Constructs and Vectors" section above). In some embodiments or methods, the engineered CRISPR-Cas12b system can be delivered by other methods, such as nuclear transfection or electroporation of a ribonucleoprotein complex consisting of an engineered Cas12b nuclease or its effector protein and their cognate RNA guides.
[0290] In some embodiments, delivery is via nanoparticles or exosomes.
[0291] In some embodiments, nanoparticles or other direct protein delivery methods can be used to directly deliver paired Cas12b nickase complexes so that the complex containing two paired crRNA elements is delivered together. In addition, protein delivery can be directly delivered to cells by viral vectors or directly, and then the CRISPR array is directly delivered, and the CRISPR array contains two paired spacers for double nicks. In some cases, for direct RNA delivery, RNA can be conjugated to at least one sugar moiety, such as N-acetylgalactosamine (GalNAc) (particularly triantennary GalNAc). In some embodiments, CRISPR-Cas12b system or its components are packaged and delivered via lipid nanoparticles. In some embodiments, lipid nanoparticles are administered to an individual by intravenous injection or infusion.
[0292] IV. Kits and Articles
[0293] Also provided are compositions, kits, unit doses, and articles of manufacture comprising one or more components of any of the engineered Cas12b nucleases or effector proteins thereof described herein, sgRNAs comprising engineered scaffolds (e.g., any of SEQ ID NOs: 25 to 53), or engineered CRISPR-Cas12b systems.
[0294] In some embodiments, a kit is provided, comprising: one or more AAV vectors encoding any of the engineered Cas12b nucleases or their effector proteins or engineered CRISPR-Cas12b systems described herein. In some embodiments, the kit further comprises one or more guide RNAs, such as sgRNAs comprising engineered scaffolds (e.g., any of SEQ ID NOs: 25 to 53). In some embodiments, the kit further comprises donor DNA. In some embodiments, the kit further comprises cells, such as human cells.
[0295] The kit may contain one or more additional components, such as containers, reagents, culture media, cytokines, buffers, antibodies, etc., to allow the propagation of engineered cells. The kit may also contain a device for administering the composition.
[0296] The kit may also include instructions for using the engineered CRISPR-Cas12b system described herein, such as methods for detecting or modifying target nucleic acids. In some embodiments, the kit includes instructions for treating or diagnosing a disease or condition. Instructions related to the use of the kit components typically include information about dosage, dosing regimen, and the route of administration for the intended treatment. The container may be a unit dose, bulk packaging (e.g., multi-dose packaging), or subunit dose. For example, a kit containing a sufficient dose of the composition disclosed herein may be provided to provide effective treatment for an individual over a long period of time. The kit may also include multiple unit doses of the composition and instructions for use, and the amount of packaging is sufficient to store and use in a pharmacy (e.g., a hospital pharmacy and a compounding pharmacy).
[0297] The test kit of the present application is in suitable packaging. Suitable packaging includes but is not limited to vials, bottles, cans, flexible packaging (e.g., sealed Mylar polyester film or plastic bags) etc. The test kit can optionally provide additional components, such as buffer and explanatory information. Therefore, the application also provides products, which include vials (such as sealed vials), bottles, cans, flexible packaging etc.
[0298] The article of manufacture may comprise a container and a label or package insert on or associated with the container. Suitable containers include, for example, bottles, vials, syringes, etc. The container may be formed from a variety of materials (such as glass or plastic). Typically, the container holds a composition effective for treating a disease or condition described herein and may have a sterile access port (e.g., a container may be an intravenous solution bag or a vial with a stopper pierceable by a hypodermic needle). The label or package insert indicates that the composition is used to treat a specific condition of an individual. The label or package insert will also include instructions for administering the composition to the individual.
[0299] Package insert refers to instructions customarily included in commercial packages of therapeutic products, that contain information about the indications, usage, dosage, administration, contraindications, and / or precautions concerning the use of such therapeutic products.
[0300] Additionally, the article of manufacture may further comprise a second container comprising a pharmaceutically acceptable buffer, such as bacteriostatic water for injection (BWFI), phosphate-buffered saline, Ringer's solution, and dextrose solution. It may also include other materials desirable from a commercial and user perspective, including other buffers, diluents, filters, needles, and syringes.
[0301] Exemplary embodiments
[0302] Embodiment 1. An engineered Cas12b nuclease comprising one, two or three types of mutations relative to a reference Cas12b nuclease, wherein the mutations comprise: (1) substitution of one or more amino acid residues in the reference Cas12b nuclease that interact with a protospacer adjacent motif (PAM) with a positively charged amino acid residue; and / or (2) substitution of one or more amino acid residues in the reference Cas12b nuclease that are involved in opening a double-stranded DNA strand with an amino acid residue having an aromatic ring; and / or (3) substitution of one or more amino acid residues in the RuvC domain of the reference Cas12b nuclease that interacts with a single-stranded DNA substrate with a positively charged amino acid residue or a hydrophobic amino acid residue.
[0303] Embodiment 2. The engineered Cas12b nuclease of embodiment 1, wherein the reference Cas12b nuclease is a wild-type Cas12b nuclease.
[0304] Embodiment 3. The engineered Cas12b nuclease of embodiment 1 or 2, wherein the reference Cas12b nuclease comprises the amino acid sequence of SEQ ID NO: 1.
[0305] Embodiment 4. The engineered Cas12b nuclease of any one of embodiments 1-3, comprising replacing one or more amino acid residues in the reference Cas12b nuclease that interacts with the PAM with a positively charged amino acid residue.
[0306] Embodiment 5. The engineered Cas12b nuclease of embodiment 4, wherein the one or more amino acid residues that interact with the PAM are within 9 angstroms of the PAM in the three-dimensional structure.
[0307] Embodiment 6. The engineered Cas12b nuclease of embodiment 4 or 5, wherein the one or more amino acid residues that interact with the PAM are located at one or more of the following positions: 116, 123, 130, 132, 144, 145, 153, 173, 222, 395, 400 and / or 475; wherein the amino acid residues are numbered according to SEQ ID NO: 1.
[0308] Embodiment 7. The engineered Cas12b nuclease of embodiment 6, wherein the one or more amino acid residues that interact with the PAM comprise one or more of the following amino acid residues: D116, K123, D130, D132, N144, K145, E153, D173, Q222, D395, N400 and / or E475; wherein the amino acid residues are numbered according to SEQ ID NO: 1.
[0309] Embodiment 8. The engineered Cas12b nuclease of embodiment 7, wherein the one or more amino acid residues that interact with the PAM comprise one or more of the following amino acid residues: D116 and / or E475; wherein the amino acid residues are numbered according to SEQ ID NO: 1.
[0310] Embodiment 9. The engineered Cas12b nuclease of any one of Embodiments 4-8, wherein the positively charged amino acid residue is R or K.
[0311] Embodiment 10. The engineered Cas12b nuclease of embodiment 9, wherein the substitution of one or more amino acid residues that interact with the PAM in the reference Cas12b nuclease with a positively charged amino acid residue is one or more of the following substitutions: D116R and / or E475R; wherein the amino acid residues are numbered according to SEQ ID NO: 1.
[0312] Embodiment 11. The engineered Cas12b nuclease according to any one of embodiments 1-10, comprising replacing one or more amino acid residues involved in opening the DNA double strand in the reference Cas12b nuclease with an amino acid residue having an aromatic ring.
[0313] Embodiment 12. The engineered Cas12b nuclease of embodiment 11, wherein the one or more amino acid residues involved in opening the DNA double strand interact with the last base pair in the PAM relative to the 3' end of the target strand.
[0314] Embodiment 13. The engineered Cas12b nuclease of embodiment 11 or 12, wherein the one or more amino acid residues involved in opening the DNA double strand are located at one or more of the following positions: 118 and / or 119; wherein the amino acid residues are numbered according to SEQ ID NO: 1.
[0315] Embodiment 14. The engineered Cas12b nuclease of any one of Embodiments 11-13, wherein the amino acid residue having an aromatic ring is Y, F or W.
[0316] Embodiment 15. The engineered Cas12b nuclease of embodiment 14, wherein one or more amino acid residues involved in opening the DNA double strand in the reference Cas12b nuclease are replaced with an amino acid residue having an aromatic ring, which is Q119Y, Q119F or Q119W; wherein the amino acid residues are numbered according to SEQ ID NO: 1.
[0317] Embodiment 16. The engineered Cas12b nuclease of any one of embodiments 1-16, comprising replacing one or more amino acid residues in the reference Cas12b nuclease that are located in the RuvC domain and interact with the single-stranded DNA substrate with a positively charged amino acid residue or a hydrophobic amino acid residue.
[0318] Embodiment 17. The engineered Cas12b nuclease of embodiment 16, wherein the one or more amino acid residues in the RuvC domain that interact with the single-stranded DNA substrate are within 9 angstroms of the single-stranded DNA substrate in the three-dimensional structure.
[0319] Embodiment 18. The engineered Cas12b nuclease of any embodiment 17, wherein the one or more amino acid residues in the RuvC domain that interact with the single-stranded DNA substrate are located at one or more of the following positions: 300, 301, 304, 329, 636, 639, 647, 682, 757, 758, 761, 764, 768, 852, 854, 856, 857, 858, 860, 862, 863, 865, 866, 867, 869, 938, 956, 957, 958, 994, 1093, and / or 1097; wherein the amino acid residues are numbered according to SEQ ID NO: 1.
[0320] Embodiment 19. The engineered Cas12b nuclease of embodiment 18, wherein the one or more amino acid residues in the RuvC domain and interacting with the single-stranded DNA substrate comprise one or more of the following amino acid residues: D300, K301, E304, N329, E636, Q639, T647, Q682, I757, E758, E761, E764, K768, E852, Q854, N856, N857, D858, P860, S862, E863, N865, Q866, L867, Q869, E938, E956, G957, E958, I994, Q1093 and / or W1097; wherein the amino acid residues are numbered according to SEQ ID NO: 1.
[0321] Embodiment 20. The engineered Cas12b nuclease of embodiment 19, comprising replacing one or more of the following amino acid residues: E636, I757, E758, E761, Q854, N857, N865, Q866, Q869 and / or Q1093 with a positively charged amino acid residue; wherein the amino acid residues are numbered according to SEQ ID NO: 1.
[0322] Embodiment 21. The engineered Cas12b nuclease of embodiment 20, wherein the positively charged amino acid residue is R or K.
[0323] Embodiment 22. The engineered Cas12b nuclease of embodiment 21, wherein the substitution of one or more amino acid residues in the reference Cas12b nuclease that are located in the RuvC domain and interact with the single-stranded DNA substrate is one or more of the following substitutions: E636R, I757R, E758R, E761R, Q854R and / or N857K; wherein the amino acid residues are numbered according to SEQ ID NO: 1.
[0324] Embodiment 23. The engineered Cas12b nuclease of embodiment 19, comprising replacing one or more of the following amino acid residues: E758, E761, E863, N865, Q866, Q869, Q956 and / or Q1093 with a hydrophobic amino acid residue; wherein the amino acid residues are numbered according to SEQ ID NO: 1.
[0325] Embodiment 24. The engineered Cas12b nuclease of embodiment 23, wherein the hydrophobic amino acid residue is W, Y, F or M.
[0326] Embodiment 25. The engineered Cas12b nuclease of embodiment 24, wherein the substitution of one or more amino acid residues in the reference Cas12b nuclease located in the RuvC domain and interacting with the single-stranded DNA substrate is one or more of the following substitutions: N865W, N865Y, Q866M, Q869M, Q1093W and / or Q1093Y; wherein the amino acid residues are numbered according to SEQ ID NO: 1.
[0327] Embodiment 26. The engineered Cas12b nuclease of any one of embodiments 1-3, wherein the engineered Cas12b nuclease comprises any one or a combination of the following substitutions: (1) D116R; (2) E475R; (3) Q119F and E475R; (4) Q119F, E475R, and E758R; (5) Q119Y; (6) Q119F; (7) Q119W; (8) I757R; (9) E758R; (10) E761R; (11) K768R; (12) I757R and E758R; (13) I757R and E761R; (14) I757 R and K768R; (15) E758R and E761R; (16) E758R and K768R; (17) E761R and K768R; (18) I757R, E758R and E761R; (19) I757R, E758R and K768R; (20) I757R, E761R and K768R; (21) E758R, E761R and K768R; (22) I757R, E758R, E761R and K768R; (23) Q866M; (24) Q869M; and (25) Q866M and Q869M; wherein the amino acid residues are numbered according to SEQ ID NO:1.
[0328] Embodiment 27. The engineered Cas12b nuclease of any one of embodiments 1-26, comprising an amino acid sequence having at least 85% sequence identity to any one of SEQ ID NOs: 2 to 22.
[0329] Embodiment 28. The engineered Cas12b nuclease of any one of embodiments 1-27, further comprising one or more mutations that increase the flexibility of a flexible region comprising amino acid residues 855-859; wherein the amino acid residue positions are numbered according to SEQ ID NO: 1.
[0330] Embodiment 29. The engineered Cas12b nuclease of embodiment 28, wherein the one or more mutations that increase flexibility comprise N856G.
[0331] Embodiment 30. An engineered Cas12b nuclease comprising any one or more of the following mutations: (1) D116R; (2) E475R; (3) Q119F and E475R; (4) Q119F, E475R, and E758R; (5) Q119Y; (6) Q119F; (7) Q119W; (8) Q119F and E475R; (9) Q 119F, E475R and E758R (10) E636R; (11) I757R; (12) E758R; (13) E761R; (14) Q854R; (15) N857K; (16) Q119F, E475R and E758R; (17) K768R; (18) I757R and E758R; (19) I757R and E761R ; (20) I757R and K768R; (21) E758R and E761R; (22) E758R and K768R; (23) E761R and K768R; (24) I757R, E758R and E761R; (25) I757R, E758R and K768R; (26) I757R, E761R and K768R; (27) E75 8R, E761R and K768R; (28) I757R, E758R, E761R and K768R (29) N865W; (30) N865Y; (31) Q866M; (32) Q869M; (33) Q1093W; (34) Q1093Y; and / or (35) Q866M and Q869M; wherein the amino acid residue positions are numbered according to SEQ ID NO:1.
[0332] Embodiment 31. An engineered Cas12b nuclease comprising the amino acid sequence of any one of SEQ ID NOs: 2 to 22.
[0333] Embodiment 32. An engineered Cas12b effector protein comprising an engineered Cas12b nuclease or a functional derivative thereof according to any one of embodiments 1-31.
[0334] Embodiment 33. The engineered Cas12b effector protein of embodiment 32, wherein the engineered Cas12b nuclease or a functional derivative thereof has enzymatic activity.
[0335] Embodiment 34. The engineered Cas12b effector protein of embodiment 32 or 33, wherein the engineered Cas12b effector protein is capable of inducing double-strand breaks in DNA molecules.
[0336] Embodiment 35. The engineered Cas12b effector protein of embodiment 32 or 33, wherein the engineered Cas12b effector protein is capable of inducing single-strand breaks in DNA molecules.
[0337] Embodiment 36. The engineered Cas12b effector protein of embodiment 32, wherein the engineered Cas12b effector protein comprises an enzymatically inactive mutant of the engineered Cas12b nuclease.
[0338] Embodiment 37. The engineered Cas12b effector protein of embodiment 36, wherein the enzyme-inactive mutant comprises D570A, R785A, E848A, R911A and / or D977A.
[0339] Embodiment 38. The engineered Cas12b effector protein of any one of embodiments 32-37, further comprising a functional domain fused to the engineered Cas12b nuclease or a functional derivative thereof.
[0340] Embodiment 39. The engineered Cas12b effector protein of embodiment 38, wherein the functional domain is selected from a translation initiator domain, a transcription repressor domain, a transactivation domain, an epigenetic modification domain, a nucleobase editing domain, a reverse transcriptase domain, a reporter domain, and a nuclease domain.
[0341] Embodiment 40. The engineered Cas12b effector protein of any one of embodiments 32-37, comprising a first polypeptide of the N-terminal portion of the engineered Cas nuclease or a functional derivative thereof and a second polypeptide comprising the C-terminal portion of the engineered Cas nuclease or a functional derivative thereof, wherein the first polypeptide and the second polypeptide are capable of binding to each other in the presence of a guide RNA comprising a guide sequence to form a clustered regularly interspaced short palindromic repeats (CRISPR) complex that specifically binds to a target nucleic acid comprising a target sequence complementary to the guide sequence.
[0342] Embodiment 41. The engineered Cas12b effector protein of embodiment 40, comprising a first polypeptide and a second polypeptide, wherein the first polypeptide comprises 1 to X amino acid residues from the N-terminus of the engineered Cas12b nuclease or a functional derivative thereof, wherein the second polypeptide comprises X+1 residues from the C-terminus of the engineered Cas12b nuclease or a functional derivative thereof, wherein the first polypeptide and the second polypeptide are capable of binding to each other in the presence of a guide RNA comprising a guide sequence to form a clustered regularly interspaced short palindromic repeats (CRISPR) complex that specifically binds to a target nucleic acid comprising a target sequence complementary to the guide sequence.
[0343] Embodiment 42. The engineered Cas12b effector protein of embodiment 40 or 41, wherein the first polypeptide and the second polypeptide each comprise a dimerization domain.
[0344] Embodiment 43. The engineered Cas12b effector protein of embodiment 42, wherein the first dimerization domain and the second dimerization domain bind to each other in the presence of an inducing agent.
[0345] Embodiment 44. The engineered Cas12b effector protein of embodiment 40 or 41, wherein the first polypeptide and the second polypeptide do not comprise a dimerization domain.
[0346] Embodiment 45. An engineered CRISPR-Cas12b system comprising: (a) the engineered Cas12b effector protein of any one of embodiments 32-44, or a nucleic acid encoding the engineered Cas12b effector protein; and (b) a guide RNA comprising a guide sequence complementary to a target sequence, or a nucleic acid encoding the guide RNA, wherein the engineered Cas12b effector protein and the guide RNA are capable of forming a CRISPR complex that specifically binds to a target nucleic acid comprising the target sequence and induces modification of the target nucleic acid.
[0347] Embodiment 46. The engineered CRISPR-Cas12b system of embodiment 45, wherein the guide RNA comprises crRNA and tracrRNA.
[0348] Embodiment 47. The engineered CRISPR-Cas12b system of embodiment 45 or 46, comprising a precursor guide RNA array encoding multiple crRNAs.
[0349] Embodiment 48. An engineered CRISPR-Cas12b system according to any one of embodiments 45-47, wherein the guide RNA is a single guide RNA (sgRNA).
[0350] Embodiment 49. The engineered CRISPR-Cas12b system of any one of embodiments 45-48, comprising one or more vectors encoding an engineered Cas12b effector protein.
[0351] Embodiment 50. The engineered CRISPR-Cas12b system of embodiment 49, wherein the one or more vectors are adeno-associated virus (AAV) vectors.
[0352] Embodiment 51. The engineered CRISPR-Cas12b system of embodiment 50, wherein the AAV vector further encodes the guide RNA.
[0353] Embodiment 52. A method for detecting a target nucleic acid in a sample, comprising: (a) contacting the sample with the engineered CRISPR-Cas12b system of any one of embodiments 45-51 and a labeled detection nucleic acid, wherein the labeled detection nucleic acid is single-stranded and does not hybridize to the guide sequence of the guide RNA; and (b) measuring a detectable signal produced by cleavage of the labeled detection nucleic acid by the engineered Cas12b effector protein, thereby detecting the target nucleic acid.
[0354] Embodiment 53. A method of modifying a target nucleic acid comprising a target sequence, comprising contacting the target nucleic acid with the engineered CRISPR-Cas12b system of any one of embodiments 45-51.
[0355] Embodiment 54. The method of embodiment 53, wherein the method is performed in vitro.
[0356] Embodiment 55. The method of embodiment 53, wherein the target nucleic acid is present in a cell.
[0357] Embodiment 56. The method of embodiment 55, wherein the cell is a bacterial cell, a yeast cell, a mammalian cell, a plant cell, or an animal cell.
[0358] Embodiment 57. The method of embodiment 53, wherein the method is performed ex vivo.
[0359] Embodiment 58. The method of embodiment 53, wherein the method is performed in vivo.
[0360] Embodiment 59. The method of any one of embodiments 53-58, wherein the target nucleic acid is cut or the target sequence in the target nucleic acid is altered by the engineered CRISPR-Cas12b system.
[0361] Embodiment 60. The method of any one of embodiments 53-58, wherein the expression of the target nucleic acid is altered by the engineered CRISPR-Cas12b system.
[0362] Embodiment 61. The method of any one of Embodiments 53-60, wherein the target nucleic acid is genomic DNA.
[0363] Embodiment 62. The method of any one of embodiments 53-61, wherein the target sequence is associated with a disease or disorder.
[0364] Embodiment 63. The method of any one of Embodiments 53-62, wherein the engineered CRISPR-Cas12b system comprises an array of precursor guide RNAs encoding multiple crRNAs, wherein each crRNA comprises a different guide sequence.
[0365] Embodiment 64. A method of treating a disease or condition associated with a target nucleic acid in a cell of an individual, comprising modifying the target nucleic acid in a cell of the individual using the engineered CRISPR-Cas12b system of any one of embodiments 45-51, thereby treating the disease or condition.
[0366] Embodiment 65. The method of embodiment 64, wherein the disease or condition is selected from the group consisting of cancer, cardiovascular disease, genetic disease, autoimmune disease, metabolic disease, neurodegenerative disease, eye disease, bacterial infection, and viral infection.
[0367] Embodiment 66. An engineered cell comprising a modified target nucleic acid, wherein the target nucleic acid is modified using the method of any one of embodiments 53-63.
[0368] Embodiment 67. An engineered non-human animal comprising one or more engineered cells of embodiment 66.
[0369] Example
[0370] The following examples are merely exemplary embodiments of the present application and therefore should not be considered to limit the present application in any way.The following examples and detailed description are provided by way of illustration and not limitation.
[0371] method
[0372] Plasmid construction
[0373] The coding sequence of AaCas12b is codon optimized for expression in human cells and is synthesized. The nucleic acid sequence encoding the engineered AaCas12b protein mutant is produced by PCR-based directed site mutagenesis. Specifically, the DNA sequence encoding the reference AaCas12b protein is divided into two parts centered on the mutation site. Two pairs of primers are designed to amplify the two parts of the DNA sequence and assembled into a single DNA fragment by Gibson cloning, and the fragment is integrated into the pCAG-2A-eGFP vector. By dividing the DNA encoding the reference AaCas12b protein into multiple sections, and amplifying and assembling using PCR and Gibson cloning, a mutation combination is constructed. The DNA encoding the engineered AaCas12b protein is inserted between the XmaI and NheI sites of the pCAG-2A-eGFP vector. Using protein structure visualization software commonly used in the art (for example, PyMol or Chimera), based on the analysis of the crystal structure of AaCas12b, the position of the mutation in the design AaCas12b protein variant is determined. The crystal structure of AaCas12b is available in the RCSBPDB database under the accession numbers 6LTU, 6LTR, 6LU0, and 6LTP. AaCas12b variants were expressed in human 293T cells using the pCAG-2A-eGFP vector. The DNA sequence encoding the sgRNA scaffold was synthesized de novo and assembled into the pUC19-U6 framework by Gibson cloning. The nucleic acid encoding the spacer sequence was also ligated into the same pUC19-U6 framework.
[0374] Cell culture, transfection, and fluorescence-activated cell sorting (FACS)
[0375] HEK293T cells were cultured in DMEM (Gibco) containing 1% penicillin-streptomycin (Gibco) and 10% fetal bovine serum (Gibco). Cells were seeded in 24-well culture dishes (Corning) for 16 hours until cell coverage reached 70%. Using Lipofectamine 3000 (Invitrogen), 600ng of the plasmid encoding AaCas12b protein and different amounts of the plasmid encoding sgRNA were transfected into cells in each well of the 24-well culture dishes. After 68 hours of transfection, HEK293T cells were digested with trypsin-EDTA (0.05%) (Gibco), and FACS sorting was performed using MoFloXDP (BeckmanCoulter) based on GFP signal (illustrating successful transfection).
[0376] Targeted deep sequencing analysis for genome modifications
[0377] GFP-positive HEK293T cells sorted by FACS were dissolved with buffer L and incubated at 55 ° C for 3 hours, then incubated at 95 ° C for 10 minutes. Corresponding primers were used to PCR amplify dsDNA fragments containing target sites at different genomic loci. For deep sequencing of the target region, cell lysates were used directly as template DNA to be amplified by barcode PCR. The PCR products were purified and pooled into several libraries for high-throughput sequencing. The frequency (%) of indel was analyzed using CRISPResso2 software by calculating the ratio of reads containing insertions or deletions. In this application, the index of indel frequency (%) was used to compare and analyze the gene editing efficiency of different engineered Cas12b proteins and / or in the presence of different sgRNA scaffolds. Any number of reads less than 0.05% of the total reads was discarded.
[0378] Example 1: Substituting one or more amino acid residues that interact with PAM in the reference AaCas12b nuclease with positively charged amino acid residues.
[0379] According to the above method, an engineered AaCas12b enzyme with a single mutation in the amino acid residue that interacts with PAM was designed and expressed. Ten amino acids were selected: D116, K123, D130, D132, N144, K145, E153, D173, Q222, D395, N400, and E475, and each amino acid residue was substituted with arginine (R). Nucleic acids encoding sgRNAs for the target sites CCR5-11 (SEQ ID NO: 63), CD34-7 (SEQ ID NO: 64), and RNF2-1 (SEQ ID NO: 65) were designed, which contained from 5' to 3': DNA encoding the Aa-sg-sgRNA scaffold sequence (SEQ ID NO: 23) - DNA encoding the spacer sequence and cloned into the pUC19-U6 backbone. Using Lipofectamine 3000 (Invitrogen), as described above, 600 ng of plasmid encoding AaCas12b protein and 300 ng of plasmid encoding sgRNA were transfected into HEK293T cells in each well of a 24-well culture dish. Wild-type AaCas12b (SEQ ID NO: 1) was used as a control. The amino acid substitutions in the AaCas12b enzyme and the corresponding gene editing efficiency are shown in Figure 1 and Table 1. Compared with wild-type AaCas12b, AaCas12b variants with amino acid substitutions D116R (SEQ ID NO: 2) or E475R (SEQ ID NO: 3) showed improved gene editing efficiency. Figure 1As shown, AaCas12b-D116R (SEQ ID NO: 2) and AaCas12b-E475R (SEQ ID NO: 3) have an average editing efficiency of more than about 20% at three genomic sites, while the average gene editing efficiency of the reference wild-type AaCas12b nuclease is about 6%. The indel frequency of AaCas12b-D116R (SEQ ID NO: 2) and AaCas12b-E475R (SEQ ID NO: 3) mutants is significantly higher than that of mutants using other AaCas12b in this category. Compared with wild-type AaCas12b, AaCas12b-D395R achieved much higher gene editing efficiency at the CD34-7 site, but not at other tested sites.
[0380] Table 1. Gene editing efficiency of different AaCas12b at different loci
[0381]
[0382] Example 2: Substituting one or more amino acid residues involved in opening the DNA double strand in the reference AaCas12b nuclease with amino acid residues having an aromatic ring.
[0383] According to the above method, an engineered AaCas12b nuclease with a single substitution in the amino acid residue involved in opening the double-stranded DNA was designed and expressed. In short, amino acid residue Q118 or Q119 was replaced by an aromatic amino acid residue (eg, Y, F, or W). The same sgRNA encoding plasmid as in Example 1 was used here. Using Lipofectamine 3000 (Invitrogen), as described above, 600 ng of plasmid encoding AaCas12b protein and 300 ng of plasmid encoding sgRNA were transfected into HE K293T cells in each well of a 24-well culture dish. Wild-type AaCas12b (SEQ ID NO: 1) was used as a control. Amino acid substitutions in the AaCas12b enzyme and the corresponding gene editing efficiency are shown in Figure 2 and Table 2. Compared with wild-type AaCas12b, AaCas12b with amino acid substitutions Q119Y, Q119F, or Q119W showed improved gene editing efficiency at all tested sites. The indel frequencies of AaCas12b-Q119Y, AaCas12b-Q119F, and AaCas12b-Q119W mutants were significantly higher than those of other AaCas12b mutants (Q118Y, Q118F, Q118W) tested in this category.
[0384] Table 2. Gene editing efficiency of different AaCas12b at different loci
[0385]
[0386] Example 3: Substitution of one or more amino acid residues interacting with single-stranded DNA substrates in the RuvC domain of the reference AaCas12b nuclease with positively charged amino acid residues or hydrophobic amino acid residues.
[0387] According to the above method, an engineered AaCas12b nuclease with a single amino acid substitution in the amino acid residue that interacts with a single-stranded DNA substrate in the RuvC domain was designed and expressed. Nucleic acids encoding sgRNAs for the target sites CCR5-3 (SEQ ID NO: 66) and RNF2-5 (SEQ ID NO: 67) were designed, which included from 5' to 3': DNA encoding an Aa-sg-sgRNA scaffold sequence (SEQ ID NO: 23)-encoding spacer sequence DNA and cloned into a pUC19-U6 framework. Using Lipofectamine 3000 (Invitr ogen), 600 ng of a plasmid encoding AaCas12b protein and 300 ng of a plasmid encoding sgRNA were transfected into HEK293T cells in each well of a 24-well culture dish. Wild-type AaCas12b (SEQ ID NO: 1) was used as a control.
[0388] In the first group of AaCas12b mutants, each of the following amino acid residues was replaced with the positively charged amino acid residue arginine (R) in Table 3: E636, I757, E758, E761, Q854, N857, N865, Q866, Q869, and Q1093. The amino acid substitutions and corresponding gene editing efficiencies in the AaCas12b enzyme are shown in Figure 3-4B and Table 3.
[0389] Table 3. Gene editing efficiency of different AaCas12b at different loci
[0390]
[0391]
[0392] In the second group of AaCas12b mutants, each of the amino acid residues in Table 4 was replaced with the positively charged amino acid residue lysine (K). The amino acid substitutions in the AaCas12b enzyme and the corresponding gene editing efficiencies are shown in Tables 4 and Figures 4A to 4B middle.
[0393] As shown in Table 3-4 and Figure 3-4BAs shown, AaCas12b mutants with amino acid substitutions D300R, K301R, E636R, Q639R, T647R, Q682R, I757R, E758R, E761R, K768R, Q854R, N857R, D858R, N865R, Q866R, I994R, Q1093R, W1097R, E636K, Q639K, T647K, Q682K, I757K, E758K, E761K, Q854K, N857K, D858K, N865K, I994K, Q1093K or W1097K have improved gene editing efficiency compared to wild-type AaCas12b. The indel frequencies of E636R, I757R, E758R, E761R, Q854R, D858R, E758K, N857K, I994R, and D858K mutants of AaCas12b were significantly higher than those tested with other AaCas12b mutants in this category (substituted with positively charged amino acids).
[0394] Table 4. Gene editing efficiency of different AaCas12b at different loci
[0395]
[0396]
[0397] In the third group of AaCas12b mutants, each of the following amino acid residues was substituted with a hydrophobic amino acid residue (e.g., Y, F, M, or W): E758, E761, E863, N865, Q866, Q869, E956, and Q1093. The amino acid substitutions and corresponding gene editing efficiencies in the AaCas12b enzyme are shown in Figure 5 And in Table 5. Compared with wild-type AaCas12b, AaCas12b mutants with amino acid substitutions E758W, E758Y, E758M, E761Y, N865W, N865Y, N865F, Q866M, Q869M, Q1093W, Q1093Y, Q1093F, or Q1093M showed improved gene editing efficiency. The indel frequency of N865W, N865Y, Q866M, Q869M, Q1093W, and Q1093Y mutants of AaCas12b was significantly higher than that of other AaCas12b mutants in this category (substituted with hydrophobic amino acid residues).
[0398] Table 5. Gene editing efficiency of different AaCas12b at different loci
[0399]
[0400]
[0401] Example 4. Characterization of the mutation combinations of Examples 1 to 3 and their gene editing efficiency.
[0402] The amino acids with the desired gene editing efficiency screened in Example 1, Example 2, and Example 3 were substituted: Q866M, Q869M, I757R, E758R, E761R, K768R, and I757R to produce AaCas12b proteins with multiple mutations, i.e., Q866M+Q869M, I757R+E758R, I757R+E761R, I 757R+K768R, E758R+E761R, E758R+K768R, E761R+K768R, I757R+E758R+E761R, I757R+E758R+K768R, I757R+E761R+K768R, E758R+E761R+K768R, and I757R+E758R+E761R+K768R. The coding nucleic acids for sgRNAs targeting the target sites CCR5-3 (SEQ ID NO: 66), CCR5-11 (SEQ ID NO: 63), CD34-1 (SEQ ID NO: 68) and RNF2-5 (SEQ ID NO.: 67) were designed, which contained from 5' to 3': DNA encoding the Aa-sg-sgRNA scaffold sequence (SEQ ID NO: 23) - DNA encoding the spacer sequence and cloned into the pUC19-U6 framework. Wild-type AaCas12b (SEQ ID NO: 1) was used as a control. Using Lipofectamine 3000 (Invitrogen), 600 ng of the plasmid encoding the above-mentioned AaCas12b protein and 300 ng of the plasmid encoding sgRNA were transfected into HEK293T cells in each well of a 24-well culture dish. Their gene editing efficiencies are shown in Figure 6 And in Table 6. Compared with wild-type AaCas12b, AaCas12b mutants with amino acid substitution combinations all show significantly improved gene editing efficiency at all test sites. Compared with corresponding single mutants at some test sites, some AaCas12b combination mutants (such as Q866M+Q869M, E758R+E761R, E758R+E768R, I757R+E758R+K768R, and E758R+E761R+K768R) have improved gene editing efficiency.
[0403] Table 6. Gene editing efficiency of different AaCas12b at different loci
[0404]
[0405]
[0406] AaCas12b-Q119F+E475R and AaCas12b-Q119F+A475R+E758R were generated as described above. The same sgRNA encoding plasmid as in Example 1 was used here. Wild-type AaCas12b (SEQ ID NO: 1) was used as a control. Using Lipofectamine 3000 (Invitrogen), 600 ng of the plasmid encoding the above-mentioned AaCas12b protein and 300 ng of the plasmid encoding sgRNA were transfected into HEK293T cells in each well of a 24-well culture dish. Their gene editing efficiency is shown in Figure 2. Figure 7 As shown in Table 7. The results showed that compared with wild-type AaCas12b, AaCas12b-Q119F+E475R and AaCas12b-Q119F+E475R+E758R significantly improved gene editing efficiency at all test sites. AaCas12b-Q119F+E475R+E758R showed the most significant improvement in gene editing efficiency relative to wild-type AaCas12b or the corresponding AaCas12b variant with a single substitution at all test sites (CCR5-11, CD34-7, and RNF2-1).
[0407] Table 7. Gene editing efficiency of different AaCas12b at different loci
[0408]
[0409] Example 5: Enhancing the gene editing activity of engineered AaCas12b using sgRNA with an engineered scaffold.
[0410] In this example, the gene editing activity of various sgRNAs with engineered scaffolds was tested using the AaCas12b mutant (Q119F+E475R+E758R) from Example 4. The encoding nucleic acid for the sgRNA targeting the target site CCR5-11 (SEQ ID NO: 63) was designed, which contained from 5' to 3': DNA encoding the sgRNA scaffold sequence-DNA encoding the spacer sequence and cloned into the pUC19-U6 frame. 600 ng of plasmid encoding AaCas12b variant protein and 300 ng of plasmid encoding sgRNA with engineered scaffold (SEQ ID NO: 25 to 53; modified based on AacCas12b-sgRNA scaffold V0), AacCas12b-sgRNA scaffold (SEQ ID NO: 24; V0, control; H. Yang et al., Cell. 2016; 167(7): 1814-1828.e12), or AaCas12b-Aa-sg scaffold (SEQ ID NO: 23; control) were transfected into HEK293T cells using Lipofectamine 3000 (Invitrogen) in each well of a 24-well culture dish. Their gene editing efficiencies are shown in Figure 9 . Figure 9 The data showed that all sgRNA engineering scaffolds significantly improved the gene editing efficiency of the AaCas12b (Q119F + E475R + E758R) variant compared to the AacCas12b-sgRNA scaffold (V0). Compared to the Aa-sg scaffold, all sgRNA engineering scaffolds (except V1 and V8) also significantly improved the gene editing efficiency of the AaCas12b (Q119F + E475R + E758R) variant.
[0411] Example 6: Engineered AaCas12b with inactivated nuclease activity.
[0412] To generate an inactivated AaCas12b protein, the AaCas12b (Q119F+E475R+E758R) variant from Example 4 (SEQ ID NO: 22) was further modified to include an additional single point mutation (D570A) in the nucleolytic domain ( Figure 10APlasmids encoding i) AaCas12b (Q119F + E475R + E758R) or AaCas12b (Q119F + E475R + E758R + D570A) (SEQ ID NO: 79) under the control of the CMV promoter, and ii) a control sgRNA (not targeting hemoglobin subunit γ1 / 2 (HBG1 / 2), sgRNA1 (target sequence SEQ ID NO: 70 targeting HBG1 / 2), or sgRNA2 (target sequence SEQ ID NO: 71 targeting HBG1 / 2) under the control of the U6 promoter were transfected into HEK293 cells using a method similar to that described above (Table 8; see plasmid construction for details). Figure 10A ). These sgRNAs were constructed using the sgRNA scaffold (V9) (SEQ ID NO: 53). Three days after transfection, genomic DNA was extracted from the transfected cells. T7 endonuclease I (T7EI) mismatch detection assay was performed to determine cleavage efficiency (M. Crisso et al., PLoS One. 2015; 10(8): e013690). Table 9 lists the primer sequences used in the T7EI assay.
[0413] like Figure 10B As shown, compared with AaCas12b(Q119F+E475R+E758R), the catalytic activity of AaCas12b(Q119F+E475R+E758R) was significantly reduced in the cleavage of two different target sites of HBG1 / 2 guided by sgRNA1 or sgRNA2.
[0414] Table 8. PAMs and target sites of sgRNAs targeting HBG1 / 2
[0415] sgRNA PAM Target sequence sgRNA1 TTG AGATAGTTGTGGGGAAGGGGC(SEQ ID NO:70) sgRNA2 TTT GCATTGAGATAGTGTGGGGA(SEQ ID NO:71)
[0416] Table 9. Primer sequences used in T7EI assay
[0417] SEQ ID NO Primer sequences 69 TCCTGCACTGAAACTGTTGC 78 TCCTGAGAAGCGACCTGGA
[0418] To further reduce the nuclease activity of the engineered AaCas12b, additional point mutations were introduced into AaCas12a (Q119F + E475R + E758R + D570A) to generate AaCas12b (Q119F + E475R + E758R + D570A + E848A) (SEQ ID NO: 80) or AaCas12c (Q119F / E475R + E758R + D570A + D977A) (SEQ ID NO: 81). Figure 11APlasmids co-encoding i) AaCas12b (Q119F + E475R + E758R), AaCas12b (Q119F + E475R + E758R + D570A + E848A) or AaCas12a (Q119F / E475R + E758R + D570A + D977A) under the control of CMV promoter, and ii) sgRNA1 (target sequence for HBG1 / 2 SEQ ID NO: 70) or sgRNA2 (target sequence for HBG1 / 2 SEQ ID NO: 11) under the control of U6 promoter were transfected into HEK293 cells using a method similar to that described above (see Figure 11A As negative controls, plasmids encoding AaCas12b(Q119F+E475R+E758R), AaCas12b(Q119F+E475R+E758R+D570A+E848A), or AaCas12b(Q119F+E475R+E758R+D570A+D977A) and a control sgRNA (not targeting any sequence within hemoglobin subunit γ1 / 2 (HBG1 / 2)) were similarly transfected into HEK293 cells without any sequence encoding sgRNA. Figure 11B As shown, AaCas12b(Q119F+E475R+E758R+D570A+E848A) and AaCas12b(Q119F+E475R+E758R+D570A+D977A) completely abolished the nuclease activity of AaCas12b(Q119F+E475R+E758R).
[0419] Example 7: Transcriptional repression using engineered AaCas12b fusion proteins.
[0420] The AaCas12b (Q119F + E475R + E758R + D570A + D977A) (SEQ ID NO: 81) of Example 6 was further engineered to produce a fusion protein to silence the transcription of the target gene. AaCas12b (Q119F + E475R + E758R + D570A + D977A) (two copies of the nuclear localization sequence NLS are located on both sides) is combined with the transcription repression module ZIM3 (SEQ ID NO: 72) The KRAB domain is fused to the C-terminus or N-terminus of AaCas12b (Q119F+E475R+E758R+D570A+D977A), and the fusion proteins are named Cd12bk and Nd12bk, respectively. The same plasmid also encodes a gene that specifically recognizes the SCN9A gene (encoding the voltage-gated sodium channel 1.7Na vsgRNAs at different target sites in 1.7) Figure 12A ; Table 10). These sgRNAs were constructed using the sgRNA scaffold (V9) (SEQ ID NO: 53).
[0421] To examine whether Cd12bk and Nd12bk fusion proteins can recruit chromatin modification complexes to silence the transcription of SCN9A, plasmids encoding the fusion proteins and sgRNA were transfected into Neuro2A (N2a; a mouse neural crest-derived cell line) cells. As a control, a plasmid encoding Cd12bk was similarly transfected into N2a cells with a control sgRNA (not targeting any sequence within SCN9A). Three days after transfection, the transfected cells were collected and RNA was extracted using an RNA extraction kit (Vazyme, catalog number RC112-01). The mRNA level of Nav1.7 in each sample was determined by qPCR. The data were normalized using control sgRNA and Cd12bk ("Cd12bk non-target"). As Figure 12B As shown, Cd12bk or Nd12bk together with sgRNA-msg6, sgRNA-msg8, sgRNA-msg13 or sgRNA-mSG 18 can greatly inhibit the transcription of SCN9A, among which sgRNA-mmsg8 and sgRNA-mssg13 show the strongest inhibition. Nd12bk together with sgRNA-msg11 can also significantly inhibit the transcription of SCN9A. These results show that dAaCas12b fused to KRAB, such as AaCas12b (Q119F + E475R + E758R + D570A + D977A), can be used as a targeted transcriptional regulation tool in eukaryotic cells.
[0422] Table 10. PAMs and target sites of sgRNAs targeting SCN9A
[0423] sgRNA PAM Target site msg6 TTA GCTGCCCGCCACACTGGCGC(SEQ ID NO:73) msg8 TTG GGCGTGGTGATGCTAGGGAT(SEQ ID NO:74) msg11 TTC TAGTCTGCTCAGGATGAAGC(SEQ ID NO:75) msg13 TTC AATCCTGCCCACTGTGCAGG(SEQ ID NO:76) msg18 TTC CCTTGGATCAGAATCCGCAG(SEQ ID NO:77)
[0424] Although the embodiments of the present application have been described above with reference to the accompanying drawings, the present application is not limited to the specific embodiments and application areas described above. The specific embodiments described above are merely illustrative and instructive, and are not restrictive. In light of this specification and without departing from the scope of protection of the claims of this application, those skilled in the art may devise various other embodiments, all of which fall within the scope of protection of this application.
[0425] Exemplary sequences
[0426]
[0427]
[0428]
[0429]
[0430]
[0431]
[0432]
[0433]
[0434]
[0435]
[0436]
[0437]
[0438]
[0439]
[0440]
[0441]
[0442]
[0443]
[0444]
[0445]
[0446]
[0447]
[0448]
[0449]
[0450]
[0451]
[0452]
[0453]
Claims
1. An engineered Cas12b nuclease having amino acid substitutions relative to a reference Cas12b nuclease as follows: Replacing an amino acid residue in the reference Cas12b nuclease that interacts with a protospacer adjacent motif (PAM) with a positively charged amino acid residue, wherein the substitution of an amino acid residue in the reference Cas12b nuclease that interacts with the PAM with the positively charged amino acid residue is one of the following substitutions: D116R, E475R, and D395R; The amino acid sequence of the reference Cas12b nuclease is shown in SEQ ID NO: 1, and the amino acid residues are numbered according to SEQ ID NO:
1.
2. engineered Cas12b nuclease according to claim 1, wherein the amino acid sequence of the engineered Cas12b nuclease is as shown in SEQ ID NO:2 or 3.
3. An engineered Cas12b nuclease, wherein the amino acid substitutions relative to a reference Cas12b nuclease are any one of the following substitution combinations: (1) Q119F+E475R; (2) Q119F+E475R+E758R; and the amino acid sequence of the reference Cas12b nuclease is as shown in SEQ ID NO: 1, wherein the amino acid residues are numbered according to SEQ ID NO:
1.
4. engineered Cas12b nuclease according to claim 3, the amino acid sequence of the wherein said engineered Cas12b nuclease is as shown in any one of SEQ ID NO:21,22.
5. An engineered Cas12b effector protein comprising the engineered Cas12b nuclease according to any one of claims 1 to 4.
6. The engineered Cas12b effector protein of claim 5, wherein the engineered Cas12b nuclease has enzymatic activity.
7. The engineered Cas12b effector protein according to claim 5 or 6, wherein the engineered Cas12b effector protein is capable of: i) inducing double-strand breaks in DNA molecules, and / or ii) inducing single-strand breaks in DNA molecules.
8. The engineered Cas12b effector protein of claim 5, wherein the engineered Cas12b effector protein comprises an enzymatically inactive mutant of the engineered Cas12b nuclease; The enzyme-inactive mutant of the engineered Cas12b nuclease has an amino acid substitution relative to the reference Cas12b nuclease of any one of the following substitution combinations: Q119F+E475R+E758R+D570A, or Q119F+E475R+E758R+D570A+E848A, or Q119F+E475R+E758R+D570A+D977A; and wherein, The amino acid sequence of the reference Cas12b nuclease is shown in SEQ ID NO: 1, wherein the amino acid residues are numbered according to SEQ ID NO:
1.
9. The engineered Cas12b effector protein according to claim 8, wherein the amino acid sequence of the enzymatically inactive mutant of the engineered Cas12b nuclease is shown in any one of SEQ ID NOs: 79 to 81.
10. The engineered Cas12b effector protein of claim 8, wherein the engineered Cas12b effector protein further comprises a Krüppel-associated box (KRA B) domain fused to the engineered Cas12b nuclease, wherein the amino acid sequence of the Krüppel-associated box (KRAB) domain is as shown in SEQ ID NO: 72, wherein the amino acid sequence of the engineered Cas12b nuclease relative to the reference Cas12b nuclease is as shown in SEQ ID NO:
81.
11. An engineered CRISPR-Cas12b system comprising: (a) the engineered Cas12b nuclease according to any one of claims 1 to 4, or the engineered Cas12b effector protein according to any one of claims 5 to 10, or their encoding nucleic acids; and (b) a guide RNA (gRNA) comprising a guide sequence complementary to a target sequence of a target nucleic acid, or a nucleic acid encoding the guide RNA (gRNA), wherein the engineered Cas12b nuclease or the engineered Cas12b effector protein and the guide RNA (gRNA) are capable of forming a CRISPR complex that specifically binds to the target nucleic acid and induces modification of the target nucleic acid.
12. The engineered CRISPR-Cas12b system of claim 11, wherein the guide RNA (gRNA) comprises crRNA and tracrRNA.
13. The engineered CRISPR-Cas12b system of claim 11 or 12, wherein the engineered CRISPR-Cas12b system comprises a precursor gRNA array encoding multiple crRNAs.
14. The engineered CRISPR-Cas12b system of claim 11, wherein the guide RNA (gRNA) is an sgRNA.
15. The engineered CRISPR-Cas12b system of claim 14, wherein the sgRNA comprises the sequence of any one of SEQ ID NOs: 23 to 53.
16. An engineered CRISPR-Cas12b system comprising: (a) a Cas12b nuclease or a Cas12b effector protein having an amino acid sequence of any one of SEQ ID NOs: 2 to 3, 21 to 22, and 79 to 81, or a nucleic acid encoding the same; and (b) a gRNA comprising a guide sequence complementary to a target sequence of a target nucleic acid, or a nucleic acid encoding the gRNA, wherein the gRNA comprises an engineered scaffold comprising the sequence of any one of SEQ ID NOs: 25 to 53; Among them, Cas12b nuclease or Cas12b effector protein and gRNA can form a CRISPR complex that specifically binds to the target nucleic acid and induces modification of the target nucleic acid.
17. The engineered CRISPR-Cas12b system of claim 16, wherein the gRNA comprises crRNA and tracrRNA, and wherein the tracrRNA comprises an engineered scaffold or portion thereof.
18. The engineered CRISPR-Cas12b system of claim 16 or 17, wherein the engineered CRISPR-Cas12b system comprises a precursor gRNA array encoding multiple crRNAs.
19. The engineered CRISPR-Cas12b system of claim 16, wherein the gRNA is an sgRNA.
20. The engineered CRISPR-Cas12b system of any one of claims 11 to 12, 14 to 16, or 19, wherein the engineered CRISPR-Cas12b system comprises one or more vectors encoding the engineered Cas12b nuclease, the engineered Cas12b effector protein, the Cas12b nuclease, or the Cas12b effector protein.
21. The engineered CRISPR-Cas12b system of claim 20, wherein the one or more vectors are adeno-associated virus (AAV) vectors.
22. The engineered CRISPR-Cas12b system of claim 20, wherein the one or more vectors further encode the gRNA.
23. Use of the engineered CRISPR-Cas12b system of any one of claims 11 to 22 in the preparation of a medicament for detecting a target nucleic acid in a sample, wherein The use comprises: (a) contacting the sample with the drug and a labeled detection nucleic acid, wherein the gRNA comprises a guide sequence complementary to the target sequence of the target nucleic acid, and wherein the labeled detection nucleic acid is single-stranded and does not hybridize with the guide sequence of the gRNA; and (b) measuring a detectable signal generated by cleavage of the labeled detection nucleic acid by the engineered Cas12b nuclease or its effector protein, thereby detecting the target nucleic acid; Wherein, when the amino acid substitution of the engineered Cas12b nuclease relative to the reference Cas12b nuclease is as follows: replacing an amino acid residue in the reference Cas12b nuclease that interacts with a protospacer adjacent motif (PAM) with a positively charged amino acid residue, wherein the replacement of the amino acid residue in the reference Cas12b nuclease that interacts with the PAM with the positively charged amino acid residue is a D395R substitution, the target nucleic acid is as shown in SEQ ID NO: 64; The amino acid sequence of the reference Cas12b nuclease is shown in SEQ ID NO: 1, and the amino acid residues are numbered according to SEQ ID NO:
1.
24. Use of the engineered CRISPR-Cas12b system of any one of claims 11 to 22 in the preparation of a medicament for modifying a target nucleic acid comprising a target sequence, wherein the use comprises contacting the target nucleic acid with the medicament; in, When the amino acid substitutions of the engineered Cas12b nuclease relative to the reference Cas12b nuclease are as follows: replacing an amino acid residue in the reference Cas12b nuclease that interacts with a protospacer adjacent motif (PAM) with a positively charged amino acid residue, wherein the replacement of the amino acid residue in the reference Cas12b nuclease that interacts with the PAM with the positively charged amino acid residue is a D395R substitution, the target nucleic acid is as shown in SEQ ID NO: 64; The amino acid sequence of the reference Cas12b nuclease is shown in SEQ ID NO: 1, and the amino acid residues are numbered according to SEQ ID NO:
1.
25. The use according to claim 23 or 24, which is performed in vitro.
26. The use according to claim 23 or 24, wherein the target nucleic acid is present in a cell.
27. The use according to claim 26, wherein the cell is a bacterial cell, a yeast cell, a plant cell or an animal cell.
28. The use according to claim 23 or 24, which is performed ex vivo.
29. The method of claim 24, wherein the target nucleic acid is cut or the target sequence in the target nucleic acid is altered by the engineered CRISPR-Cas12b system.
30. The method of claim 24, wherein the expression of the target nucleic acid is altered by the engineered CRISPR-Cas12b system.
31. The use according to claim 23 or 24, wherein the target nucleic acid is genomic DNA.
32. The method of claim 23 or 24, wherein the engineered CRISPR-Cas12b system comprises a precursor gRNA array encoding multiple crRNAs, and wherein each crRNA comprises a different guide sequence.
33. A pharmaceutical composition for treating a disease or condition associated with a target nucleic acid in a cell of an individual, comprising the engineered CRISPR-Cas12b system of any one of claims 11 to 22; The method of using the pharmaceutical composition includes using the pharmaceutical composition to modify the target nucleic acid in the cells of the individual, thereby treating the disease or condition.
34. The pharmaceutical composition of claim 33, wherein the disease or condition is selected from the group consisting of cancer, cardiovascular disease, genetic disease, autoimmune disease, metabolic disease, neurodegenerative disease, eye disease, bacterial infection, and viral infection.
Citation Information
Patent Citations
Engineered Cas effector protein and using method thereof
CN112195164A
Type V CRISPR / Cas effector proteins for cleaving ssDNAs and detecting target DNAs
US10253365B1
Paint composition
US325160A
Adeno-associated virus as eukaryotic expression vector
US4797368A
Production of recombinant adeno-associated virus vectors
US5173414A