Constructs and their use for efficient and specific genome editing
Engineered Cas12a-based nucleases with modified PAM recognition and split gRNAs improve the precision and efficiency of genome editing, addressing the limitations of existing CRISPR-Cas systems by reducing off-target effects and enhancing targeted gene manipulation.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- ARTISAN DEV LABS INC
- Filing Date
- 2026-02-02
- Publication Date
- 2026-05-19
AI Technical Summary
Existing CRISPR-Cas genome editing systems face challenges in efficiency and precision, particularly in modifying specific genes without unintended off-target effects, limiting their applicability in complex biological systems.
Development of engineered nucleic acid-inducible nucleases, such as Cas12a-based systems, with modified PAM recognition and gRNA sequences, allowing for precise and efficient genome editing by creating staggered DNA overhangs and utilizing split gRNAs for improved targeting.
Enhances the accuracy and efficiency of genome editing by reducing off-target effects and enabling more precise manipulation of target sequences, facilitating applications in various species including human cells and plants.
Smart Images

Figure 2026082943000197 
Figure 2026082943000198 
Figure 2026082943000199
Abstract
Description
Technical Field
[0001] Cross - reference to Related Applications This application claims priority to U.S. Provisional Application No. 63 / 080,552, filed on September 18, 2020, and U.S. Provisional Application No. 63 / 185,315, filed on May 6, 2021, the disclosures of which are hereby incorporated by reference in their entireties.
Background Art
[0002] CRISPR stands for Clustered Regularly Interspaced Short Palindromic Repeats. In palindromic repeats, the nucleotide sequence is the same in both directions. Each of these palindromic repeats is followed by a short segment of spacer DNA. Small clusters of Cas (CRISPR-related system) genes are located adjacent to the CRISPR sequence. The CRISPR / Cas system is a prokaryotic immune system that can confer resistance to exogenous genetic elements, such as those present in plasmids and phages, providing a type of adaptive immunity to prokaryotes. RNA with spacer sequences helps Cas (CRISPR-related) proteins recognize and cleave exogenous DNA. CRISPR sequences are found in approximately 50% of bacterial genomes, and nearly 90% of sequenced archaea have selected efficient and robust metabolic and regulatory networks that prevent the biosynthesis of unwanted metabolites and optimally allocate resources to maximize overall cellular fitness. Progress in this field is challenging due to the complexity of these networks, which limits approaches to understanding their structure and function, as well as the ability to reprogram cellular networks and modify these systems for various applications. Certain approaches to reprogram cellular networks aim to modify a single gene in a complex pathway, but modifying a single gene can result in undesirable modifications to that gene or other genes, hindering the identification of the changes necessary to achieve the desired endpoint, and potentially making the desired endpoint more difficult to attain.
[0003] CRISPR-Cas genome editing and manipulation have had a dramatic impact on biology and biotechnology in general. CRISPR-Cas editing systems require a polynucleotide-induced nuclease, a guide polynucleotide (e.g., guide RNA (gRNA)) that induces the nuclease to cleave specific regions of the genome, and a donor DNA cassette that can be optionally used to incorporate programmable editing at the target site by repairing the cleaved dsDNA. Early demonstrations and applications of CRISPR-Cas editing utilized the Cas9 nuclease and associated gRNA. These systems have been used for gene editing in a wide range of species, from bacteria to animals and, in some cases, higher-order mammalian systems such as humans. However, it has been well demonstrated that key editing parameters, particularly protospacer-adjacent motif (PAM) specificity, editing efficiency, and off-target rates, are species, locus, and nuclease-dependent. There is growing interest in identifying and rapidly characterizing novel nuclease systems that can be used to extend and improve overall editing capabilities.
[0004] Cas12a is known to be a single RNA-inducible CRISPR / Cas endonuclease capable of genome editing with characteristics distinct from those of Cas9. In certain embodiments, Cas12a-based systems enable the rapid and reliable introduction of donor DNA into the genome. Furthermore, Cas12a extends genome editing capabilities. CRISPR / Cas12a genome editing has been evaluated not only in human cells but also in other organisms, including plants. Some features of the CRISPR / Cas12a system differ from those of CRISPR / Cas9.
[0005] The Cas12a nuclease recognizes T-rich protospacer adjacent motif (PAM) sequences (e.g., 5'-TTTN-3' (AsCas12a, LbCas12a) and 5'-TTN-3' (FnCas12a)), while the comparable sequence for SpCas9 is known to be NGG. The Cas12a PAM sequence is located at the 5' end of the target DNA sequence, whereas in Cas9 it is at the 3' end. Furthermore, Cas12a can cleave DNA distal to its PAM near the +18 / +23 position of the protospacer. This cleavage produces staggered DNA overhangs (e.g., sticky ends), while Cas9 cleaves near its PAM after the 3' position of the protospacer on both strands, creating blunt ends. In certain ways, modifying nuclease recognition can lead to improvements over Cas9 or Cas12a, resulting in increased accuracy. Furthermore, since Cas12a is induced by a single crRNA and does not require tracrRNA, it yields a shorter gRNA sequence than the sgRNA used by Cas9.
[0006] Cas12a is also known to exhibit additional ribonuclease activity that functions during crRNA processing. Cas12a can be used as an editing tool for various species (e.g., S. cerevisiae) and enables the use of alternative PAM sequences compared to those recognized by CRISPR / Cas9.
[0007] The well-known Cas12a protein-RNA complex recognizes T-rich PAMs and, upon cleavage, results in staggered DNA double-strand breaks. Cas12a-type nucleases interact with the pseudoknot structure formed by the 5' handle of crRNA. The guide RNA segment, consisting of a seed region and a 3' end, has a binding sequence complementary to the target DNA sequence. To date, Cas12a-type nucleases have been characterized to function with a single gRNA and have been demonstrated to process gRNA arrays. While the Cas12a-type and Cas9 nuclease systems have proven highly influential, neither system has been demonstrated to function as predictably as required to enable all conceivable applications in gene editing technologies.
[0008] Currently, various efforts are being made to manipulate improved CRISPR editing systems with increased efficiency and precision, including manipulation of PAM specificity, stability, and gRNA and / or nuclease sequences. For example, chemical modification of CRISPR / Cas9 gRNA, which is expected to increase gRNA stability, has been shown to result in a 3.8-fold higher indel frequency in human cells. Furthermore, other studies include structure-induced mutagenesis of Cas12a, which is being screened to identify variants with an increased range of recognized PAM sequences. These manipulated AsCas12a recognize TYCV and TATV PAMs in addition to the established TTTV sequence, and their activity has been enhanced in in vitro and tested human cells.
[0009] CRISPR / Cas9, a version of the CRISPR / Cas system, has been modified to provide a useful tool for targeted genome editing. By delivering Cas9 nuclease, which forms a complex with synthetic guide RNA (gRNA), into cells, it becomes possible to cleave / edit the cell's genome at predetermined locations to delete existing genes and / or add new ones. While these systems are useful, they have several significant limitations regarding the efficiency and precision of targeted editing, problems with inaccurate editing, and obstacles when used in commercially relevant situations such as gene replacement. Therefore, there is a need for improved nucleic acid-inducible nuclease systems for more efficient, directed, and precise editing. [Overview of the project]
[0010] Embodiment 1 provides a composition comprising an engineered nucleic acid-inducible nuclease, comprising (i) an engineered nuclease polypeptide comprising an amino acid sequence having at least 60% sequence identity with any one of SEQ ID NOs: 143-177 and 229, or one or more polynucleotides encoding an amino acid sequence having at least 60% sequence identity with any one of SEQ ID NOs: 143-177 and 229. Embodiment 2. The composition according to Embodiment 1, comprising an engineered nuclease polypeptide comprising an amino acid sequence having at least 60% sequence identity with any one of SEQ ID NOs: 144, 153 and 229, or one or more polynucleotides encoding an amino acid sequence having at least 60% sequence identity with any one of SEQ ID NOs: 144, 153 and 229. Embodiment 3. A composition according to Embodiment 1 or 2, comprising an engineered nuclease polypeptide containing an amino acid sequence having at least 60% sequence identity with SEQ ID NO: 144, or one or more polynucleotides encoding an amino acid sequence having at least 60% sequence identity with SEQ ID NO: 144. Embodiment 4. A composition according to any one of the prior embodiments, comprising an engineered nuclease polypeptide containing an amino acid sequence having at least 60% sequence identity with SEQ ID NO: 153, or one or more polynucleotides encoding an amino acid sequence having at least 60% sequence identity with SEQ ID NO: 153. Embodiment 5. A composition according to any one of the prior embodiments, comprising an engineered nuclease polypeptide containing an amino acid sequence having at least 60% sequence identity with SEQ ID NO: 229, or one or more polynucleotides encoding an amino acid sequence having at least 60% sequence identity with SEQ ID NO: 229. Embodiment 6. A composition according to any one of the prior embodiments, wherein the sequence identity is at least 80%. Embodiment 7. A composition according to any one of the prior embodiments, wherein the sequence identity is at least 95%. Embodiment 8. The composition according to any one of the prior embodiments, wherein the sequence identity is 100%.Embodiment 9. The composition according to Embodiment 1, wherein the manipulated nuclease polypeptide does not contain the peptide motif YLFQIYNKDF (SEQ ID NO: 224), or one or more polynucleotides encoding the manipulated nuclease polypeptide do not contain the peptide motif YLFQIYNKDF (SEQ ID NO: 224). Embodiment 10. The composition according to Embodiment 9, comprising one or more polynucleotides encoding an amino acid sequence having at least 60% sequence identity with any one of SEQ ID NOs: 143-151, 161-163, 165, 166, 169, 171-175, 177, and 229, or an amino acid sequence having at least 60% sequence identity with any one of SEQ ID NOs: 143-151, 161-163, 165, 166, 169, 171-175, 177, and 229. Embodiment 11. The composition according to Embodiment 10, comprising one or more polynucleotides encoding an amino acid sequence having at least 60% sequence identity with any one of SEQ ID NOs: 149, 151, 175, and 177, or an amino acid sequence having at least 60% sequence identity with any one of SEQ ID NOs: 149, 151, 175, and 177. Embodiment 12. A composition comprising a targetable guide nucleic acid-induced nuclease complex, comprising an engineered guide nucleic acid nuclease described in any one of the prior embodiments, and further comprising (ii) a compatible guide nucleic acid. Embodiment 13. The composition according to Embodiment 12, wherein the guide nucleic acid is a gRNA and the complex is an RNP. Embodiment 14. The composition according to Embodiment 12 or 13, wherein the guide nucleic acid is a split guide nucleic acid. Embodiment 15. The composition according to Embodiment 13 or 14, wherein the gRNA is an engineered gRNA. Embodiment 16. The composition according to Embodiment 15, wherein the engineered gRNA comprises a conserved gRNA. Embodiment 17. The composition according to Embodiment 16, wherein the preserved gRNA comprises one of sequence numbers 291 to 325, or a portion thereof. Embodiment 18. The composition according to Embodiment 17, wherein the preserved gRNA comprises one of sequence numbers 291 to 325.Embodiment 19. The composition according to Embodiment 18, wherein the portion is a highly conserved portion containing the nucleotide sequence of the secondary structure of the RNA. Embodiment 20. The composition according to Embodiment 18, wherein the secondary structure contains a pseudoknot. Embodiment 21. The composition according to any one of Embodiments 13 to 20, wherein the gRNA is a synthetic gRNA. Embodiment 22. The composition according to Embodiment 21, wherein the gRNA contains one or more chemical modifications.
[0011] Embodiment 23. A method for generating a strand break in or near a target sequence in a target polynucleotide, comprising contacting the target sequence with a targetable nucleic acid-induced nuclease complex described in any one of Embodiments 12 to 22, wherein the compatible guide nucleic acid of the complex targets the target sequence, and the targetable guide nucleic acid-induced nuclease complex generates the strand break. Embodiment 24. The method according to Embodiment 23, wherein the target polynucleotide is in a cell genome. Embodiment 25. The method according to Embodiment 23 or 24, further comprising providing an editing template to be inserted into the target sequence. Embodiment 26. The method according to Embodiment 25, wherein the editing template comprises a transgene. Embodiment 27. The method according to any one of Embodiments 23 to 27, wherein the target polynucleotide is a safe harbor site. Embodiment 28. A cell produced by the embodiment described in Embodiment 23. Embodiment 29. An organism produced by the method according to Embodiment 23. Embodiment 30. A composition comprising one or more manipulated polynucleotides, each comprising one or more polynucleotides, each comprising one or more polynucleotides, each comprising one or more polynucleotides, each comprising a sequence having at least 60% sequence identity with any one of SEQ ID NOs: 1-142 and 225-228. Embodiment 31. The composition according to Embodiment 30, comprising one or more polynucleotides, each comprising one or more polynucleotides, each comprising a sequence, each comprising a sequence, each comprising one or more polynucleotides, each comprising a sequence having at least 60% sequence identity with any one of SEQ ID NOs: 1, 5, 11, 15, 19, 23, 27, 31, 35, 39, 43, 47, 51, 55, 59, 63, 67, 71, 75, 79, 83, 87, 91, 95, 99, 103, 107, 111, 115, 119, 123, 127, 131, 135, 139, and 225. Embodiment 32. The composition according to Embodiment 30 or 31, wherein the polynucleotide encodes one or more additional amino acid sequences at either the N-terminus, C-terminus, or both of the polypeptide encoded by the polynucleotide. Embodiment 33. The composition according to Embodiment 32, wherein the additional amino acid sequence comprises (i) one or more NLSs, (ii) one or more purified tags, (iii) one or a cleavage sequence, and (iv) at least one of FLAG or 3XFLAG.Embodiment 34. The composition according to Embodiment 32, wherein the additional amino acid sequence comprises (i) one or more NLSs, (ii) one or more purified tags, (iii) one or more cleavage sequences, and (iv) at least two of FLAGs or 3XFLAGs. Embodiment 35. The composition according to Embodiment 32, wherein the additional amino acid sequence comprises (i) one or more NLSs, (ii) one or more purified tags, (iii) one or more cleavage sequences, and (iv) at least three of FLAGs or 3XFLAGs. Embodiment 36. The composition according to Embodiment 32, wherein the additional amino acid sequence comprises (i) one or more NLSs, (ii) one or more purified tags, (iii) one or more cleavage sequences, and (iv) FLAGs or 3XFLAGs. Embodiment 37. The composition according to any one of Embodiments 30 to 36, wherein the one or more polynucleotides are codon-optimized. Embodiment 38. The composition according to Embodiment 37, wherein the one or more polynucleotides are codon-optimized for E. coli. Embodiment 39. The composition according to Embodiment 39, comprising one or more polynucleotides containing a sequence corresponding to a sequence having at least 60% sequence identity with any of SEQ ID NOs: 2, 6, 12, 16, 20, 24, 28, 32, 36, 40, 44, 48, 52, 56, 60, 64, 68, 72, 76, 80, 84, 88, 92, 96, 100, 104, 108, 112, 116, 120, 124, 128, 132, 136, 140, 226, and 330. Embodiment 40. The composition according to Embodiment 37, wherein the one or more polynucleotides are codon-optimized for S. cerivisiae. Embodiment 41. The composition according to Embodiment 40, comprising one or more polynucleotides, each containing a sequence corresponding to a sequence having at least 60% sequence identity with any of SEQ ID NOs: 3, 7, 9, 13, 17, 21, 25, 29, 33, 37, 41, 45, 49, 53, 57, 61, 65, 69, 73, 77, 81, 85, 89, 93, 97, 101, 105, 109, 113, 117, 121, 125, 129, 133, 137, 141, and 227. Embodiment 42. The composition according to Embodiment 37, wherein the one or more polynucleotides are codon-optimized for humans.Embodiment 43. The composition according to Embodiment 42, comprising one or more polynucleotides containing a sequence corresponding to a sequence having at least 60% sequence identity with any of SEQ ID NOs: 3, 7, 9, 13, 17, 21, 25, 29, 33, 37, 41, 45, 49, 53, 57, 61, 65, 69, 73, 77, 81, 85, 89, 93, 97, 101, 105, 109, 113, 117, 121, 125, 129, 133, 137, 141, and 227. Embodiment 44. The composition according to any one of Embodiments 30 to 43, wherein the sequence identity is at least 80%. Embodiment 45. The composition according to any one of Embodiments 30 to 43, wherein the sequence identity is at least 95%. Embodiment 46. The composition according to any one of Embodiments 30 to 43, wherein the sequence identity is 100%.
[0012] Built-in by reference All publications, patents, and patent applications referenced herein are incorporated by reference to the same extent as each individual publication, patent, or patent application is incorporated by reference specifically and individually as indicated.
[0013] The following drawings form part of this specification and are included to further illustrate certain embodiments of this disclosure. These particular embodiments may be better understood by referring to one or more of these drawings in combination with the detailed description of the specific embodiments presented herein. [Brief explanation of the drawing]
[0014] [Figure 1] This is an exemplary graph illustrating a depletion assay for evaluating the cleavage efficiency of ART1 nucleic acid-inducible nucleases in several embodiments disclosed herein. [Figure 2] This is an exemplary graph showing a depletion assay for evaluating the cleavage efficiency of ART2 nucleic acid-inducible nucleases in several embodiments disclosed herein. [Figure 3]This is an exemplary graph illustrating a depletion assay for evaluating the cleavage efficiency of ART5 nucleic acid-inducible nucleases in several embodiments disclosed herein. [Figure 4] This is an exemplary graph illustrating a depletion assay for evaluating the cleavage efficiency of ART6 nucleic acid-inducible nucleases in several embodiments disclosed herein. [Figure 5] This is an exemplary graph showing a depletion assay for evaluating the cleavage efficiency of ART8 nucleic acid-inducible nucleases in several embodiments disclosed herein. [Figure 6] This is an exemplary graph showing a depletion assay for evaluating the cleavage efficiency of ART9 nucleic acid-inducible nucleases in several embodiments disclosed herein. [Figure 7] This is an exemplary graph showing a depletion assay for evaluating the cleavage efficiency of ART10 nucleic acid-inducible nucleases in several embodiments disclosed herein. [Figure 8] This is an exemplary graph showing a depletion assay for evaluating the cleavage efficiency of ART11 nucleic acid-inducible nucleases in several embodiments disclosed herein. [Figure 9] This is an exemplary graph showing a depletion assay for evaluating the cleavage efficiency of ART11_L679F(ART11*) nucleic acid-induced nuclease in several embodiments disclosed herein. [Figure 10] This is an exemplary histogram plot showing an experimental GalK editing assay for evaluating the gene editing efficiency of ART2 nucleic acid-inducible nucleases in several embodiments disclosed herein. [Figure 11] This is an exemplary histogram plot showing an experimental GalK editing assay for evaluating the gene editing efficiency of ART11 nucleic acid-inducible nucleases in several embodiments disclosed herein. [Figure 12] This is an exemplary histogram plot showing the enrichment of various PAM sites of ART11 nucleic acid-inducible nucleases in several embodiments disclosed herein. [Figure 13]An exemplary histogram plot showing the enrichment of various PAM sites of the ART11_L679F nucleic acid-guided nuclease of several embodiments disclosed herein. [Figure 14] Shows the %INDEL of guides arranged across the TRAC gene of ART11 in Jurkat cells.
Mode for Carrying Out the Invention
[0015] Several embodiments disclosed herein relate to novel nucleic acid-guided nucleases, guide nucleic acids (e.g., gRNA), and targetable nuclease systems, as well as methods of use. In other embodiments, methods of making and using engineered non-natural nucleic acid-guided nucleases, guide nucleic acids, and targetable nuclease systems are disclosed. In some embodiments, a targetable nuclease system can be used to edit the human genome or the genomes of other species. In some embodiments, the nucleic acid-guided nuclease can include a polypeptide having an amino acid sequence, e.g., a sequence represented by SEQ ID NOs: 143-177 and 229. In an embodiment, the nucleic acid-guided nuclease can include a polynucleotide encoding a nuclease, e.g., a polynucleotide having a nucleic acid sequence represented by SEQ ID NOs: 1-142 and 225-228. In an embodiment, the gRNA can include a gRNA represented by one or more of SEQ ID NOs: 178-188. In other embodiments, the gRNA can be represented by SEQ ID NOs: 178-188. In other embodiments, examples of the gRNA can include split gRNAs that serve as synthetic tracrRNA and cfRNA for the methods and systems disclosed herein. Other sequences useful in the embodiments disclosed herein are provided below.
[0016] In the following sections, various exemplary compositions and methods will be described in order to detail various embodiments of the present disclosure. It will be apparent to those skilled in the art that not all or part of the details outlined herein need to be employed to implement the various embodiments, and that concentrations, times, and other details can be varied through routine experimentation. In some cases, well-known methods or components are not included in the description.
[0017] As used herein, the terms "regulation" and "manipulation" of genome editing can mean an increase, decrease, upregulation, downregulation, induction, change in editing activity, change in binding, change in cleavage, etc. of one or more of the targeted genes or gene clusters of the specific embodiments disclosed herein.
[0018] In certain embodiments of the present disclosure, conventional molecular biology, microbiology, and recombinant DNA techniques within the scope of the art can be employed. Such techniques are well described in the literature and are understood by those skilled in the art.
[0019] In other embodiments, the primers used herein for conventional techniques can include sequencing primers and amplification primers. In some embodiments, the plasmids and oligomers used in conventional techniques can include synthetic oligomers and oligomer cassettes.
[0020] In some embodiments disclosed herein, nucleic acid-inducible nuclease systems and methods of use are provided. The nuclease system may include transcripts and other elements involved in the expression of the engineered nuclease disclosed herein, such as a novel engineered nucleic acid-inducible nuclease protein and a guide sequence (gNA, e.g., gRNA), or a novel gNA, e.g., a sequence encoding a novel gRNA disclosed herein. In some embodiments, the nucleic acid-inducible nuclease system may include at least one CRISPR-related nucleic acid-inducible nuclease construct, which is disclosed herein. In other embodiments, the nucleic acid-inducible nuclease system may include gNA (e.g., gRNA) or at least one novel gNA, e.g., gRNA, with at least one known sequence, e.g., at least one known guide sequence or at least one known scaffold sequence. In some embodiments, the engineered nucleic acid-inducible nucleases of the present invention can be used in systems for editing genes of interest in humans or other species.
[0021] Targetable nuclease systems from bacteria and archaea have emerged as powerful tools for precise genome editing. However, natural nucleases have several limitations, including expression and delivery challenges due to nucleic acid sequence and protein size. In certain embodiments, the novel engineered nucleic acid-inducible nuclease constructs disclosed herein can be constructed to modify the targeting of targeted genes in a subject and / or to improve the efficiency and / or precision of targeted gene editing. Other applications of the novel engineered nucleic acid-inducible nuclease constructs disclosed herein may be, for example, those disclosed herein.
[0022] According to these embodiments, Cas12a is known to be a single RNA-induced CRISPR / Cas endonuclease capable of genome editing with features distinct from those of Cas9. In certain embodiments, Cas12a-based systems enable the rapid and reliable introduction of donor DNA into the genome. Furthermore, Cas12a extends genome editing capabilities. CRISPR / Cas12a genome editing has been evaluated not only in human cells but also in other organisms, including plants. Some features of the CRISPR / Cas12a system differ from those of CRISPR / Cas9.
[0023] The Cas12a nuclease recognizes T-rich protospacer adjacent motif (PAM) sequences (e.g., 5'-TTTN-3' (AsCas12a, LbCas12a) and 5'-TTN-3' (FnCas12a)), while the comparable sequence for SpCas9 is known to be NGG. The Cas12a PAM sequence is located at the 5' end of the target DNA sequence, whereas in Cas9 it is at the 3' end. Furthermore, Cas12a can cleave DNA distal to its PAM near the +18 / +23 position of the protospacer. This cleavage produces staggered DNA overhangs (e.g., sticky ends), while Cas9 cleaves near its PAM after the 3' position of the protospacer on both strands, creating blunt ends. In certain ways, modifying nuclease recognition can lead to improvements over Cas9 or Cas12a, resulting in increased accuracy. Furthermore, in certain embodiments herein, Cas12a may be induced by a single crRNA (gRNA) and does not require tracrRNA, thus yielding a shorter gRNA sequence than the sgRNA used by Cas9. In certain embodiments herein, Cas12a may be induced not by a single gRNA as seen in Cas12a, but by a split gRNA, such as a split gRNA that is a synthetic gRNA, a split gRNA containing a regulated nucleotide, or other manipulated split gRNA, or other manipulated split gRNAs disclosed herein.
[0024] Cas12a is also known to exhibit additional ribonuclease activity that functions in crRNA (gRNA, e.g., split gRNA such as engineered split gRNA) processing. Cas12a can be used as an editing tool for various species (e.g., S. cerevisiae) and enables the use of alternative PAM sequences compared to those recognized by CRISPR / Cas9. Novel nucleases disclosed herein can further recognize the same or alternative PAM sequences. These novel nucleases can provide an alternative system to multiplex genome editing compared to known multiplex approaches and can be used as an improved system in mammalian gene editing. Other meanings of gRNA processing are as discussed herein, e.g., the production of gRNA, e.g., split gRNA such as engineered split gRNA containing conserved or highly conserved RNA sequences that may be important sequences for secondary structures in RNA, e.g., pseudoknot regions, for a particular nuclease, and / or other meanings discussed herein.
[0025] The well-known Cas12a protein-RNA complex recognizes T-rich PAMs and, upon cleavage, results in staggered DNA double-strand breaks. Cas12a-type nucleases interact with the pseudoknot structure formed by the 5' handle of crRNA. The guide RNA segment, consisting of a seed region and a 3' end, has a binding sequence complementary to the target DNA sequence. To date, Cas12a-type nucleases have been characterized to function with a single gRNA and have been demonstrated to process gRNA arrays. While the Cas12a-type and Cas9 nuclease systems have proven highly influential, neither system has been demonstrated to function as predictably as required to enable all conceivable applications in gene editing technologies.
[0026] Currently, various efforts are being made to manipulate improved CRISPR editing systems with increased efficiency and precision, including manipulation of PAM specificity, stability, and gRNA and / or nuclease sequences. For example, chemical modification of CRISPR / Cas9 gRNA, which is expected to increase gRNA stability, has been shown to result in a 3.8-fold higher indel frequency in human cells. Furthermore, other studies include structure-induced mutagenesis of Cas12a, which is being screened to identify variants with an increased range of recognized PAM sequences. These manipulated AsCas12a recognize TYCV and TATV PAMs in addition to the established TTTV sequence, and their activity has been enhanced in in vitro and tested human cells.
[0027] Cas12a-like nucleases and engineered Cas12a-like nucleases (engineered designer nucleases), as well as gRNAs such as gNAs, engineered gNAs, and gRNAs such as those disclosed herein, are intended for use in bacteria and other prokaryotes. In other embodiments, engineered designer nucleases are intended for use in single-celled eukaryotes, such as yeast, mammals, and birds and fish. In certain embodiments, engineered designer nucleases are intended for use in human cells. According to these embodiments, these constructs are generated, for example, to modify specific features of a wild-type gRNA sequence while retaining other desirable features compared to a control from which the gRNA is derived.
[0028] In certain embodiments, the manipulated gRNA constructs disclosed herein can be generated from Cas12as known or yet undiscovered in the art, such as Acidaminococcus massiliensis sp. (e.g., AM_Cas12a strain Marseille-P2828), Sedimentisphaera cyanobacteriorum sp. (SC_Cas12a, strain L21-RPul-D3), Barnesiella sp. An22 (B_Cas12a; An22 An22), Bacteroidetes bacterium HGW-Bacteroidetes-6 sp.XS5 (BB_Cas12a, 08E140C01), Parabacteroides distasonis sp. (PD_Cas12a, strain 8-P5), Collinsella tanakaei sp. (CT_Cas12a, isolated CIM:MAG 294), and Lachnospiraceae bacterium. Examples include, but are not limited to, MC2017 sp. (LB_Cas12a, T350), Coprococcus sp. AF16-5 (Co_Cas12a, AF16-5 AF16-5.Scaf1), or Catenovolum sp. CCB-QB4 (Ca_Cas12a, CCB-QB4 species), Eubacterium rectale (the positive control is a derivative of this Cas12a), Flavobacterium branchiophilum (FB_Cas12a), and / or synthetic constructs (SC_Cas12a), or analogues. In certain embodiments, the construct may contain less than 60% identity to known Cas12as in order to create a novel nuclease. In certain embodiments, novel Cas12a-derived constructs may include constructs that exhibit reduced off-targeting rates and / or improved editing capabilities compared to control or wild-type Cas12a nucleases.
[0029] In some embodiments, the off-targeting rate of the nuclease constructs disclosed herein can be reduced compared to a control for improved editing. For example, the off-targeting rate can be easily tested. According to these embodiments, baseline off-target editing can be evaluated using a wild-type gRNA plasmid compared to an experimentally designed gRNA, and the precision of a novel nuclease can be evaluated compared to a control Cas12a nuclease or other nucleases known in the art as positive controls (e.g., MAD7). In some embodiments, the nuclease constructs disclosed herein may share a conserved encoding motif of a known nuclease. In other embodiments, the nuclease constructs disclosed herein do not share a conserved encoding peptide motif with a known nuclease. In certain embodiments, the nuclease constructs disclosed herein do not encode the peptide motif YLFQIYNKDF (SEQ ID NO: 224) within the encoded nuclease. In certain embodiments, the nucleic acid-derived nuclease constructs disclosed herein may include polypeptides that do not encode the peptide motif YLFQIYNKDF and are represented by SEQ ID NOs. 143-151, 161-163, 165, 166, 169, 171-175, 177, and 229. In other embodiments, the nucleic acid-derived nuclease construct polypeptides disclosed herein include the peptide motif YLFQIYNKDF (SEQ ID NO: 224). In some embodiments, the nucleic acid-derived nuclease construct polypeptides disclosed herein may include the peptide motif YLFQIYNKDF and be represented by polypeptides represented by SEQ ID NOs. 152-160, 164, 167, 168, 170, and 176.
[0030] In certain methods, spacer mutations can be introduced into plasmids to test when a substituted gRNA sequence is created, or to test deletion or insertion variants. Each of these plasmid constructs can then be used to test the accuracy and efficiency of genome editing, such as deletion, substitution, or insertion.
[0031] Alternatively, nuclease constructs prepared by the compositions and methods disclosed herein can be tested for optimal genome editing time for a selected target by observing the editing efficiency over a predetermined period.
[0032] Examples of target polynucleotides for the use of manipulated nucleic acid-inducible nucleases disclosed herein may include sequences / genes or gene segments related to signaling biochemical pathways, e.g., signaling biochemical pathway-related genes or polynucleotides. Other embodiments contemplated herein relate to examples of target polynucleotides related to disease-related genes or polynucleotides.
[0033] A “disease-related” or “disorder-related” gene or polynucleotide can refer to any gene or polynucleotide that results in abnormal levels of transcripts or translations compared to a control, or that results in abnormal morphology in cells derived from diseased tissue compared to tissues or cells from a non-disease control. This may be a gene that becomes expressed at abnormally high levels. This may be a gene that becomes expressed at abnormally low levels, or a gene that contains one or more mutations and whose altered expression or expression is directly correlated with the onset and / or progression of a health condition or disorder. A disease-related gene can refer to a gene that has mutations or genetic diversity that are in linkage disequilibrium with a gene(s) that is the direct cause or progression of a disease or disorder, or the gene(s) that cause or progression of a disease or disorder. The transcribed or translated products may be known or unknown and may be at normal or abnormal levels.
[0034] Those skilled in the art will understand that examples of disease-related genes and polynucleotides are available from: the McKusick-Nathans Institute of Genetic Medicine, Johns Hopkins University (Baltimore, Md.), and the National Center for Biotechnology Information, National Library of Medicine (Bethesda, Md.) (available on the World Wide Web). Hereditary disorders and disorder-related conditions, and further examples of hereditary disorders, are disclosed herein.
[0035] The genetic disorders discussed herein may include, but are not limited to, the following:
[0036] Neoplasm formation: Genes associated with this disorder: PTEN; ATM; ATR; EGFR; ERBB2; ERBB3; ERBB4; Notch1; Notch2; Notch3; Notch4; AKT; AKT2; AKT3; HIF; HIFI a; HIF3a; Met; HRG; Bc12; PPAR alpha; PPAR gamma; WT1 (Wilms tumor); FGF receptor family members (5 members: 1, 2, 3, 4, 5): CDKN2a; APC; RB (retinoblastoma); MEN1; VHL; BRCA1; BRCA2; AR (androgen receptor); TSG101; IGF; IGF receptor; Igfl (4 variants); Igf2 (3 variants); Igf1 receptor; Igf2 receptor; Bax; Bc12; Caspase family (9 members: 1, 2, 3, 4, 6, 7, 8, 9, 12); Kras; Apc.
[0037] Age-related macular degeneration: Genes associated with these diseases include Abcr;Cc12;Cc2;cp (semloplasmin);Timp3;Cathepsin D;VIdlr;Ccr2.
[0038] Schizophrenic disorder: Genes associated with this disorder: Neureglin 1 (Nrg1); Erb4 (neureglin receptor); Complexin 1 (Cp1x1); Tph1 tryptophan hydroxylase; Tph2 tryptophan hydroxylase 2; Neurexin 1; GSK3; GSK3a; GSK3b.
[0039] Trinucleotide repeat disorders: Genes associated with this disorder: 5HTT (Huntington Dx); SBMA / SMAX1 / AR (Kennedy Dx); FXN / X25 (Friedreich's ataxia); ATX3 (Machado-Joseph Dx); ATXN1 and ATXN2 (Spinocerebellar ataxia); DMPK (Myotonic dystrophy); Atrophine 1 and Atn1 (DRPLA Dx); CBP (Creb-BP-global instability); VLDLR (Alzheimer's disease); Atxn7; Atxn10.
[0040] Fragile X syndrome: Genes associated with this disorder: FMR2; FXR1; FXR2; mGLURS.
[0041] Secretase-related disorders: Genes associated with this disorder: APH-1 (alpha and beta); presenyl n (Psenl); nicatrin (Ncstn); PEN-2.
[0042] Other: Genes associated with this disorder: Nosl; Paipl; Nati; Nat2
[0043] Prion-related disorders: Genes associated with this disorder: Prp.
[0044] ALS: Genes associated with this disorder: SOD1; ALS2; STEX; FUS; TARDBP; VEGF (VEGF-a; VEGF-b; VEGF-c).
[0045] Drug addiction: Genes associated with this disorder: Prkce (alcohol); Drd2; Drd4; ABAT (alcohol); GRIA2; GrmS; Grinl; Htrlb; Grin2a; Drd3; Pdyn; Grial (alcohol).
[0046] Autism: Genes associated with this disorder: Mecp2; BZRAP1; MDGA2; SemaSA; Neurexin 1; Fragile X (FMR2 (AFF2); FXR1; FXR2; MglurS).
[0047] Alzheimer's disease: Genes associated with this disorder: El;CHIP;UCH;UBB;Tau;LRP;PICALM;Crastelin;PS1;SORL1;CR1;VIdlr;Ubal;Uba3;CHIP28(Aqp1,Aquaporin 1);Uchll;Uch13;APP.
[0048] Inflammation and immune-related disorders: Genes associated with these disorders: IL-10; IL-1 (IL-la; IL-1b); IL-13; IL-17 (IL-17a (CTLA8); IL-17b; IL-17c; IL-17d; IL-17f); 11-23; Cx3crl; ptpn22; TNFa; NOD2 / CARD15 for IBD; IL-6; IL-12 (IL-12a; IL-12b); CTLA4; Cx3c11, AAT deficiency / mutation, AIDS (KIR3 DL1, NKAT3, NKB1, ANIB11, KIR3DS1, IFNG, CXCL12, SDF1); Autoimmune lymphoproliferative syndrome (TNFRSF6, APT1, FAS, CD95, ALPS1A); Combined immunodeficiency (IL2RG, SCIDX1, SCIDX, IMD4); HIV-1 (CCL5, SCYA5, D17S136E, TCP228), susceptible or infected (IL10, CSIF, CMKBR2, CCR2, CMKBR5, CCCKR5(CC R5)); Immunodeficiency (CD3E, CD3G, AICDA, AID, HIGM2, TNFRSF5, CD40, UNG, DGU, HIGM4, TNFSF5, CD4OLG, HIGM1, IGM, FOXP3, IPEX, AIID, XPID, PIDX, TNFRSF14B, TACI); inflammation (IL-10, IL-1 (IL-la, IL-1b), IL-13, IL-17 (IL-17a (CTLA8), IL-17b, IL-17c, IL-17d, IL- 17f), 11-23, Cx3crl, ptpn22, TNFa, NOD2 / CARD15 for IBD, IL-6, IL-12 (IL-12a, IL-12b), CTLA4, Cx3c11); severe combined immunodeficiency (SCID) s)(JAK3, JAKL, DCLRE1C, ARTEMIS, SCIDA, RAG1, RAG2, ADA, PTPRC, CD45, LCA, IL7R, CD3D, T3D, IL2RG, SCIDX1, SCIDX, IMD4).
[0049] Parkinson's disease: Genes associated with this disorder: x-synuclein; DJ-1; LRRK2; parkin; PINK1.
[0050] Blood and coagulation disorders: Genes associated with these disorders: Anemia (CDAN1, CDA1, RPS19, DBA, PKLR, PK1, NT5C3, UMPH I, PSN1, RHAG, RH50A, NRAMP2, SPTB, ALAS2, ANH I, ASB, ABCB7, ABC7, ASAT); Incomplete lymphocytic syndromes (TAPBP, TPSN, TAP2, ABCB3, PSF2, RINGI 1, MHC2TA, C2TA, RFX5, RFXAP, RFX5), Hemorrhagic disorders (TBXA2R, P2RX I, P2X I); Factor H and Factor H-like-1 (HF1, CFH, HUS); Factor V and Factor VIII (MCFD2); Factor VII deficiency (F7); Factor X deficiency (F10); Factor XI deficiency (F11); Factor XII deficiency ((F12, HAF); Factor XIIIA deficiency (F13A1, F13A); Factor deficiency (F13B); Fanconi anemia (FANCA, FACA, FA1, FA, FAA, FAAP95, FAAP90, FLJ34064, FANCB, FANCC, FACC, BRCA2, FANCD1, FANCD2, FANCD, FACD, FAD, FANCE, FACE, FANCF, XRCC9, FANCG, BRIP1, BACH1, FANCJ, PHF9, FANCL, FANCM, ICIAA1596); Hemophagocytic lymphohistiocytosis syndrome (PRF1, HPLH2, UNC13D, MUNC13-4, HPLH3, HLH3, FHL3); Hemophilia A (F8, F8C, HEMA); Hemophilia B (F9, HEMB), Hemorrhagic disorders (PI, ATT, F5); Leukocyte deficiency and disorders (ITGB2, CD18, LCAMB, LAD, EIF2B1, EIF2BA, EIF2B2, EIF2B3, EIF2B5, LVWM, CACH, CLE, EIF2B4); Sickle cell anemia (HBB); Thalassemia (HBA2, HBB, HBD, LCRB, HBA1).
[0051] Cellular dysregulation and tumor disorders: Genes associated with these disorders: B-cell non-Hodgkin lymphoma (BCL7A, BCL7); leukemia (TALI TCL5, SCL, TAL2, FLT3, NBS 1, NBS, ZNFNIAI, IK1, LYF1, HOXD4, HOX4B, BCR, CML, PHL, ALL, ARNT, KRAS2, RASK2, GMPS, AFIO, ARHGEFI2, LARG, KIAA0382, CALM, CLTH, CEBPA, CEBP, CHIC2, BTL, FLT3, KIT, PBT, LPP, NPM1, NUP214, D9S46E, CAN, CAIN, RUNX 1, CBFA2, AML1, WHSC 1 LI, NSD3, FLT3, AF1Q, NPM 1) , NUMA1, ZNF145, PLZF, PML, MYL, STAT5B, AFI 0, CALM, CLTH, ARLI 1, ARLTS1, P2RX7, P2X7, BCR, CML, PHL, ALL, GRAF, NFI, VRNF, WSS, NFNS, PTPNI 1, PTP2C, SHP2, NS 1 , BCL2, CCND1, PRAD1, BCL1, TCRA, GATA1, GF1, ERYF1, NFE1, ABL1, NQO1, DIA4, NMOR1, NUP2I4, D9S46E, CAN, CAIN).
[0052] Metabolic, hepatic, and renal disorders: Genes associated with these disorders: Amyloid neuropathy (TTR, PALS); Amyloidosis (APOA1, APP, AAA, CVAP, AD1, GSN, FGA, LYZ, UR, PALS); Cirrhosis (KATI 8, KRT8, CaHlA, NAIC, TEX292, KIAA1988); Pancreatic cystic fibrosis (CFTR, ABCC7, CF, MRP7); Glycogen storage disorders (SLC2A2, GLUT2, G6PC, G6PT, G6PT1, GAA, LAMP2, LAMPS, AGL, GDE, GBE1, GYS2, PYGL, PFKM); Hepatocellular adenoma, 142330 (TCF1, HNF1A, MODY3), early-onset liver failure, and neuropathy (SCOD1, SCO1), hepatic lipase deficiency (LIPC), hepatoblastoma, cancer, and Carcinomas (CTNNB1, PDGFRL, PDGRL, PRLTS, AXIN1, AXIN, CTNNB1, TP53, P53, LFS1, IGF2R, MPRI, MET, CASP8, MCH5; medullary cystic kidney disease (UMOD, HNFJ, FJHN, MCKD2, ADMCKD2); phenylketonuria (PAH, PKU1, QDPR, DHPR, PTS); polycystic kidney and liver disease (FCYT, PKHD1, ARPKD, PKD2, PKD4, PKDTS, PRKCSH, G19P1, PCLD, SEC63).
[0053] Musculoskeletal disorders: Genes associated with these disorders: Becker muscular dystrophy (DMD, BMD, MYF6); Duchenne muscular dystrophy (DMD, BMD); Emery-Dreyfus muscular dystrophy (LMNA, LMN1, EMD2, FPLD, CMD1A, HGPS, LGMD1B, LMNA, LMN1, EMD2, FPLD, CMD1A); Facial-scapulohumeral muscular dystrophy (FSHMD1A, FSHD1A); Muscular dystrophy (FKRP, MDC1C, LGMD2I, LAMA2, LAMM, LARGE, KIAA0609, MDC1D, FCMD, TTID, MYOT, CAPN3, CANP3, DYSF, LGMD2B, SGCG, LGMD2C, DMDA1, SCG3, SGCA, ADL, DAG2, LGMD2D, DMDA2, SGCB, LGM) D2E, SGCD, SGD, LGMD2F, CMD1L, TCAP, LGMD2G, CMD1N, TRIM32, HT2A, LGMD2H, FKRP, MDC1C, LGMD2I, TTN, CM D1G, TMD, LGMD2J, POMT1, CAV3, LGMD1C, SEPN1, SELN, RSMD1, PLEC1, PLTN, EBS1); osteopetrosis (LAPS, BMND1, LRP7) , LR3, OPPG, VBCH2, CLCN7, CLC7, OPTA2, OSTM1, GL, TCIRG1, TIRC7, 0C116, OPTB1); muscular atrophy (VAPB, VAPC, ALS8, SMN1, SMA1, SMA2, SMA3, SMA4, BSCL2, SPG17, GARS, SMAD1, CMT2D, HEXB, IGHMBP2, SMUBP2, CATF1, SMARD1).
[0054] Neurological disorders and neuronal disorders: Genes associated with these disorders: ALS (SOD1, ALS2, STEX, FUS, TARDBP, VEGF (VEGF-a, VEGF-b, VEGF-c)); Alzheimer's disease (APP, AAA, CVAP, AD1, APOE, AD2, PSEN2, AD4, STM2, APBB2, FE65L1, NOS3, PLAU, URK, ACE, DCPI, ACEI, MPO, PACIP1, PAXIPIL, PTIP, A2M, BLMH, BMH, PSEN1, AD3); Autism (Mecp2, BZRAP I, MDGA2, Sema5A, Neurex) 1, GLO1, MECP2, RTT, PPMX, MRX16, MRX79, NLGN3, NLGN4, KIAA1260, AUTSX2); Fragile X syndrome (FMR2, FXR1, FXR2, mGLUR5); Huntington's disease and Huntington's disease-like disorder (HD, IT15, PRNP, PRIP, JPH3, JP3, HDL2, TBP, SCA17); Parkinson's disease (NR4A2, NURR1, NOT, TINUR, SNCAIP, TBP, SCA17, SNCA, NACP, PARK1, PARK4, DJ1, PARK7, LRRK2, PARKS, PINK1, PARK6, UCHL1, PARKS, SNCA, NA CP, PARK1, PARK4, PRKN, PARK-2, PDJ, DBH, NDUFV2); Rett syndrome (MECP2, RTT, PPMX, MRX16, MRX79, CDKL5, STK9, MECP2, RTT, PPMX, MRX16, MRX79, x-synuclein, DJ-1); Schizophrenia: Neureglin 1 (Nrg1), Erb4 (neureglin receptor), Complexin 1 (Cp1x1), Tph1 tryptophan hydroxylase, Tph2, tryptophan hydroxylase 2, Neurexin 1, GSK3, GSK3a, GSK3b, 5-HTT (S1c6a4), COMT, DRD (Drd la), SLC6A3, DAOA, DTNBP1, Dao (Daol)); secretase-related disorders (APH-1 (alpha and beta), presenilin (Psenl), nicatrin (Ncstn), PEN-2, Nosl, Parpl, Natl, Nat2);Trinucleotide repeat disorders (HTT (Huntington's Disease), SBMA / SMAX1 / AR (Kennedy's Disease), FXN / X25 (Friedreich's Ataxia), ATX3 (Machado-Joseph's Disease), ATXN1 and ATXN2 (Spinocerebellar Ataxia), DMPK (Myotonic Dystrophy), Atrophine 1 and Atn1 (DRPLA Disease), CBP (Creb-BP-Global Instability), VLDLR (Alzheimer's Disease), Atxn7, Atxn10).
[0055] Eye-related disorders: Genes associated with these disorders: Age-related macular degeneration (Aber, Cc12, Cc2, cp (ceruloplasmin), Timp3, cathepsinD, Vld1r, Ccr2); Cataracts (CRYAA, CRYA1, CRYBB2, CRYB2, PITX3, BFSP2, CP49, CP47, CRYAA, CRYA1, PAX6, AN2, MGDA, CRYBA1, CRYB1, CRYGC, CRYG3, CCL, LIM2, MP19) CRYGD, CRYG4, BFSP2, CP49, CP47, HSF4, CTM, HSF4, CTM, MIP, AQPO, CRYAB, CRYA2, CTPP2, CRYBB1, CRYGD, CRYG4, CRYBB2, CRYB2, CRYGC, CRYG3, CCL, CRYAA, CRYA1, GJA8, CX50, CAE1, GJA3, CX46, CZP3, CAE3, CCM1, CAM, KRIT1); Corneal opacity and dystrophy (APO A1, TGFBI, CSD2, CDGG1, CSD, BIGH3, CDG2, TACSTD2, TROP2, M1S1, VSX1, RINX, PPCD, PPD, KTCN, COL8A2, FECD, PPCD2, PIP5K 3, CFD); Congenital flat cornea (KERA, CNA2); Glaucoma (MYOC, TIGR, GLC1A, JOAG, GPOA, OPTN, GLC1E, FIP2, HYPL, NRP, CYP1B1, GLC3A, OPAL, NT G, NPG, CYP1B1, GLC3A); Leber congenital amaurosis (CRB1, RP12, CRX, CORD2, CRD, RPGRIP1, LCA6, CORD9, RPE65, RP20, AIPL1, LCA4, GUCY2 D, GUC2D, LCA1, CORD6, RDH12, LCA3); macular dystrophy (ELOVL4, ADMD, STGD2, STGD3, RDS, RP7, PRPH2, PRPH, AVMD, AOFMD, VMD2).
[0056] P13K / AKT cell signaling disorders: Genes associated with these disorders: PRKCE;ITGAM;ITGA5;IRAK1;PRKAA2;EIF2AK2;PTEN;EIF4E;PRKCZ;GRK6;MAPK1;TSC1;PLK1;AKT2;IKBKB;PIK3CA;CDK8;CDKN1B;NFKB2;BCL2;PIK3CB;PPP2R1A;MAPK8;BCL2L1;MAPK3;TSC2;ITGAl;KRAS;EIF4EBP1;RELA;PRKCD;NOS3;PRK AA1;MAPK9;CDK2;PPP2CA;PIM1;ITGB7;YWHAZ;ILK;TP53;RAF1;IKBKG;RELB;DYRK1A;CDKN1A;ITGB1;MAP2K2;JAK1;AKT1;JAK2;PIK3R1;CH UK;PDPK1;PPP2R5C;CTNNB1;MAP2K1;NFKB1;PAK3;ITGB3;CCND1;GSK3A;FRAP1;SFN;ITGA2;TTK;CSNK1A1;BRAF;GSK3B;AKT3;FOXO1;SOK;HS P9OAA1;RP S 6KB1.
[0057] ERK / MAPK cell signaling disorders: Genes associated with these disorders: PRKCE;ITGAM;ITGA5;HSPB1;IRAK1;PRKAA2;EIF2AK2;RAC1;RAP1A;TLN1;EIF4E;ELK1;GRK6;MAPK1;RAC2;PLK1;AKT2;PIK3CA;CDK8;CREB1;PRKCI;PTK2;FOS;RPS6KA4;PIK3CB;PPP2R1A;PIK3C3;MAPK8;MAPK3;ITGAl;ETS1;KRAS;MYCN ;EIF4EBP1;PPARG;PRKCD;PRKAA1;MAPK9;SRC;CDK2;PPP2CA;PIM1;PIK3C2A;ITGB7;YWHAZ;PPP1CC;KSR1;PXN;RAF1;FYN;DYRK1A;ITGB1 ;MAP2K2;PAK4;PIK3R1;STAT3;PPP2R5C;MAP2K1;PAK3;ITGB3;ESR1;ITGA2;MYC;TTK;CSNK1A1;CRKL;BRAE;ATF4;PRKCA;SRF;STAT1;SGK.
[0058] Glucocorticoid receptor cell signaling disorders: Genes associated with these disorders: RAC1;TAF4B;EP300;SMAD2;TRAF6;PCAF;ELK1;MAPK1;SMAD3;AKT2;IKBKB;NCOR2;UBE2I;PIK3CA;CREB1;FOS;HSPA5;NFKB2;BCL2;MAP3K14;STAT5B;PIK3CB;PIK3C3;MAPK8;BCL2L1;MAPK3;TSC22D3;MAPK10;NRIP1;KRAS;MAPK 13;RELA;STAT5A;MAPK9;NOS2A;PBX1;NR3C1;PIK3C2A;CDKN1C;TRAF2;SERPINE1;NCOA3;MAPK14;TNF;RAF1;IKBKG;MAP3K7;CREBBP;CD KN1A;MAP2K2;JAK1;IL8;NCOA2;AKT1;JAK2;PIK3R1;CHUK;STAT3;MAP2K1;NFKB1;TGFBR1;ESR1;SMAD4;CEBPB;JUN;AR;AKT3;CCL2;MMP 1;STAT1;IL6;HSP9OAA1.
[0059] Axon-guided cell signaling disorders: Genes associated with these disorders: PRKCE;ITGAM;ROCK1;ITGA5;CXCR4;ADAM12;IGF1;RAC1;RAP1A;El F4E;PRKCZ;NRP1;NTRK2;ARHGEF7;SMO;ROCK2;MAPK1;PGF;RAC2;PTPN11;GNAS;AKT2;PIK3CA;ERBB2;PRKCI;PTK2;CFL1;GNAQ;PIK3CB;CXCL12;PIK3C3;WNT11;PRKD1;GNB2L1;ABL1 ;MAPK3;ITGA1;KRAS;RHOA;PRKCD;PIK3C2A;ITGB7;GLI2;PXN;VASP;RAF1;FYN;ITGB1;MAP2K2;PAK4;ADAM17;AKT 1;PIK3R1;GUI;WNT5A;ADAM10;MAP2K1;PAK3;ITGB3;CDC42;VEGFA;ITGA2;EPHA8;CRKL;RND1;GSK3B;AKT3;PRKCA.
[0060] Ephrin receptor cell signaling disorders: Genes associated with these disorders: PRKCE;ITGAM;ROCK1;ITGA5;CXCR4;IRAK1;PRKAA2;EIF2AK2;RAC1;RAP1A;GRK6;ROCK2;MAPK1;PGF;RAC2;PTPN11;GNAS;PLK1;AKT2;DOK1;CDK8;CREB1;PTK2;CFL1;GNAQ;MAP3K14;CXCL12;MAPK8;GNB2L1;ABL1; MAPK3;ITGA1;KRAS;RHOA;PRKCD;PRKAA1;MAPK9;SRC;CDK2;PIM1;ITGB7;PXN;RAF1;FYN;DYRK1A;ITGB1;MAP2K2;PAK4, AKT1 ;JAK2;STAT3;ADAM10;MAP2K1;PAK3;ITGB3;CDC42;VEGFA;ITGA2;EPHA8;TTK;CSNK1A1;CRKL;BRAF;PTPN13;ATF4;AKT3;SGK.
[0061] Actin cytoskeletal signaling disorders: Genes associated with these disorders: ACTN4;PRKCE;ITGAM;ROCK1;ITGA5;IRAK1;PRKAA2;EIF2AK2;RAC1;INS;ARHGEF7;GRK6;ROCK2;MAPK1;RAC2;PLK1;AKT2;PIK3CA;CDK8;PTK2;CFL1;PIK3CB;MYH9;DIAPH1;PIK3C3;MAPK8;F2R;MAPK3 ;SLC9A1;ITGA1;KRAS;RHOA;PRKCD;PRKAA1;MAPK9;CDK2;PIM1;PIK3C2A;ITGB7;PPP1CC;PXN;VIL2;RAF1;GSN;DYRK1A ;ITGB1;MAP2K2;PAK4;PIP5K1A;PIK3R1;MAP2K1;PAK3;ITGB3;CDC42;APC;ITGA2;TTK;CSNK1A1;CRKL;BRAF;VAV3;SGK.
[0062] Huntington's disease cell signaling disorders: Genes associated with these disorders: PRKCE; IGF1; EP300; RCOR1; PRKCZ; HDAC4; TGM2; MAPK1; CAPNS1; AKT2; EGFR; NCOR2; SP1; CAPN2; PIK3CA; HDAC5; CREB1; PRKC1; HS PA5 ;REST;GNAQ;PIK3CB;PIK3C3;MAPK8;IGF1R;PRKD1;GNB2L1;BCL2L1;CAPN1;MAPK3;CASP8;HDAC2;HDAC7A;PRKCD;HDAC11;MAPK9;HDAC9;PIK 3C2A;HDAC3;TP53;CASP9;CREBBP;AKT1;PIK3R1;PDPK1;CASP1;APAF1;FRAP1;CASP2;JUN;BAX;ATF4;AKT3;PRKCA;CLTC;SGK;HDAC6;CASP3.
[0063] Apoptotic cell signaling disorders: Genes associated with these disorders: PRKCE;ROCK1;BID;IRAK1;PRKAA2;EIF2AK2;BAK1;BIRC4;GRK6;MAPK1;CAPNS1;PLK1;AKT2;IKBKB;CAPN2;CDK8;FAS;NFKB2;BCL2;MAP3K14;MAPK8;BCL2L1;CAPN1;MAPK3;CASP8;KRAS;RELA;PRKCD;PRKAA1;MAPK9;CDK2;PIM1;TP53;TNF;RAF1;IKBKG;RELB;CASP9;DYRK1A;MAP2K2;CHUK;APAF1;MAP2K1;NFKB1;PAK3;LMNA;CASP2;BIRC2;TTK;CSNK1A1;BRAF;BAX;PRKCA;SGK;CASP3:BTRC3:PARPI.
[0064] B cell receptor signaling disorders: Genes associated with these disorders: RAC1;PTEN;LYN;ELK1;MAPK1;RAC2;PTPN11;AKT2;IKBKB;PIK3CA;CREB1;SYK;NFKB2;CAMK2A;MAP3K14;PIK3CB;PIK3C3;MAPK8;BCL2L1;ABL1;MAPK3;ETS1;KRAS;MAPK13;RELA;PTPN6;MAPK9;EGR1;PIK3C2A;BTK;MAPK14;RAF1;IKBKG;RELB;MAP3K7;MAP2K2;AKT1;PIK3R1;CHUK;MAP2K1;NFKB1;CDC42;GSK3A;FRAP1;BCL6;BCL10;JUN;GSK3B;ATF4;AKT3;VAV3;RPS6KB1.
[0065] Leukocyte extravasation cell signaling disorders: Genes associated with these disorders: ACTN4;CD44;PRKCE;ITGAM;ROCK1;CXCR4;CYBA;RAC1;RAP1A;PRKCZ;ROCK2;RAC2;PTPN11;MMP14;PIK3CA;PRKCI;PTK2;PIK3CB;CXCL12;PIK3C3;MAPK8;PRKD1;ABL1;MAPK10;CYBB;MAPK13;RHOA;PRKCD;MAPK9;SRC;PIK3C2A;BTK;MAPK14;NOX1;PXN;VIL2;VASP;ITGB1;MAP2K2;CTNND1;PIK3R1;CTNNB1;CLDN1;CDC42;FUR;ITK;CRKL;VAV3;CTTN;PRKCA;MMPl;MMP9.
[0066] Integrin cell signaling disorders: Genes associated with these disorders: ACTN4;ITGAM;ROCK1;ITGA5;RAC1;PTEN;RAP1A;TLN1;ARHGEF7;MAPK1;RAC2;CAPNS1;AKT2;CAPN2;PIK3CA;PTK2;PIK3CB;PIK3C3;MAPK8;CAV1;CAPN1;ABL1;MAPK3;ITGAl;KRAS;RHOA;SRC;PIK3C2A;ITGB7;PPP1CC;ILK;PXN;VASP;RAF1;FYN;ITGB1;MAP2K2;PAK4;AKT1;PIK3R1;TNK2;MAP2K1;PAK3;ITGB3;CDC42;RND3;ITGA2;CRKL;BRAF;GSK3B;AKT3.
[0067] Acute-phase response cell signaling disorders: Genes associated with these disorders: IRAK1; SOD2; MYD88; TRAF6; ELK1; MAPK1; PTPN11; AKT2; IKBKB; PIK3CA; FOS; NFKB2; MAP3K14; PIK3CB; MAPK8; RIPK1; MAPK3; IL6ST; KRAS; MAPK13; IL6R; RELA; SOCS1; MAPK9; FTL; NR3C1; TRAF2; SERPINE1; MAPK14; TNF; RAF1; PDK1; IKBKG; RELB; MAP3K7; MAP2K2; AKT1; JAK2; PIK3R1; CHUK; STAT3; MAP2K1; NFKB1; FRAP1; CEBPB; JUN; AKT3; IL1R1; IL6.
[0068] PTEN cell signaling disorders: Genes associated with these disorders: ITGAM;ITGA5;RAC1;PTEN;PRKCZ;BCL2L11;MAPK1;RAC2;AKT2;EGFR;IKBKB;CBL;PIK3CA;CDKN1B;PTK2;NFKB2;BCL2;PIK3CB;BCL2L1;MAPK3;ITGA1;KRAS;ITGB7;ILK;PDGFRB;INSR;RAF1;IKBKG;CASP9;CDKN1A;ITGB1;MAP2K2;AKT1;PIK3R1;CHUK;PDGFRA;PDPK1;MAP2K1;NFKB1;ITGB3;CDC42;CCND1;GSK3A;ITGA2;GSK3B;AKT3;FOXO1;CASP3.
[0069] p53 cell signaling disorders: Genes associated with these disorders: RPS6KB1 PTEN;EP300;BBC3;PCAF;FASN;BRCA1;GADD45A;BIRC5;AKT2;PIK3CA;CHEK1;TP53INP1;BCL2;PIK3CB;PIK3C3;MAPK8;THBS 1;ATR;BCL2L1;E2F1;PMAIP1;CHEK2;TNFASF10B;TP73;RB1;HDAC9;CDK2;PIK3C2A;MAPK14;TP53;LRDD;CDKN1A;HIPK2;AKT1;PIK3R1;RAM2B;APAF1;CTNNB1;SIRT1;CCND1;PRKDC;ATM;SFN;CDKN2A;JUN;SNAI2;GSK3B;BAX;AKT3.
[0070] Aryl hydrocarbon receptor cell signaling disorders: Genes associated with these disorders: HSPB1;EP300;FASN;TGM2;RXRA;MAPK1;NQO1;NCOR2;SP1;ARNT;CDKN1B;FOS;CHEK1;SMARCA4;NFKB2;MAPK8;ALDH1A1;ATR;E2F1;MAPK3;NRIP1;CHEK2;RELA;TP73;GSTP1;RB1;SRC;CDK2;AHR;NFE2L2;NCOA3;TP53;TNF;CDKN1A;NCOA2;APAF1;NFKB1;CCND1;ATM;ESR1;CDKN2A;MYC;JUN;ESR2;BAX;IL6;CYP1B1;HSP9OAA1.
[0071] Disorders in xenobiotic metabolic cell signaling: Genes associated with these disorders: PRKCE;EP300;PRKCZ;RXRA;MAPK1;NQO1;NCOR2;PIK3CA;ARNT;PRKCI;NFKB2;CAMK2A;PIK3CB;PPP2R1A;PIK3C3;MAPK8;PRKD1;ALDH1A1;MAPK3;NRIP1;KRAS;MAPK13;PRKCD;GSTP1;MAPK9;NOS2A;ABCB1;AHR;PPP2CA;FTL;NFE2L2;PIK3C2A;PPARGC1A;MAPK14;TNF;RAF1;CREBBP;MAP2K2;PIK3R1;PPP2R5C;MAP2K1;NFKB1;KEAP1;PRKCA;EIF2AK3;IL6;CYP1B1;HSP9OAA1.
[0072] SAPL / JNK cell signaling disorders: Genes associated with these disorders: PRKCE; IRAK1; PRKAA2; EIF2AK2; RAC1; ELK1; GRK6; MAPK1; GADD45A; RAC2; PLK1; AKT2; PIK3CA; FADD; CDK8; PIK3CB; PIK3C3; MAPK8; RIPK1; GNB2L1; IRS1; MAPK3; MAPK10; DAXX; KRAS; PRKCD; PRKAA1; MAPK9; CDK2; PIM1; PIK3C2A; TRAF2; TP53; LCK; MAP3K7; DYRK1A; MAP2K2; PIK3R1; MAP2K1; PAK3; CDC42; JUN; TTK; CSNK1A1; CRKL; BRAF; SGK.
[0073] PPAr / RXR cell signaling disorders: Genes associated with these disorders: PRKAA2;EP300;INS;SMAD2;TRAF6;PPARA;FASN;RXRA;MAPK1;SMAD3;GNAS;IKBKB;NCOR2;ABCA1;GNAQ;NFKB2;MAP3K14;STAT5B;MAPK8;IASI;MAPK3;KRAS;RELA;PRKAA1;PPARGC1A;NCOA3;MAPK14;INSR;RAF1;IKBKG;RELB;MAP3K7;CREBBP;MAP2K2;JAK2;CHUK;MAP2K1;NFKB1;TGFBAl;SMAD4;JUN;IL1R1;PRKCA;IL6;HSP9OAA1;ADIPOO.
[0074] NF-KB cell signaling disorders: Genes associated with these disorders: IRAK1; EIF2AK2; EP300; INS; MYD88; PRKCZ; TRAF6; TBK1; AKT2; EGFR; IKBKB; PIK3CA; BTRC; NFKB2; MAP3K14; PIK3CB; PIK3C3; MAPK8; RIPK1; HDAC2; KRAS; RELA; PIK3C2A; TRAF2; TLR4: PDGFRB; TNF; INSR; LCK; IKBKG; RELB; MAP3K7; CREBBP; AKT1; PIK3R1; CHUK; PDGFRA; NFKB1; TLR2; BCL10; GSK3B; AKT3; TNFAIP3; IL1R1.
[0075] Neuregulin cell signaling disorders: Genes associated with these disorders: ERBB4; PRKCE; ITGAM; ITGA5; PTEN; PRKCZ; ELK1; MAPK1; PTPN11; AKT2; EGFR; ERBB2; PRKCI; CDKN1B; STAT5B; PRKD1; MAPK3; ITGA1; KRAS; PRKCD; STAT5A; SRC; ITGB7; RAF1; ITGB1; MAP2K2; ADAM17; AKT1; PIK3R1; PDPK1; MAP2K1; ITGB3; EREG; FRAP1; PSEN1; ITGA2; MYC; NRG1; CRKL; AKT3; PRKCA; HS P9OAA1; RPS6KB1.
[0076] Wnt and beta-catenin cell signaling disorders: Genes associated with these disorders: CD44;EP300;LRP6;DVL3;CSNK1E;GJA1;SMO;AKT2;PIN1;CDH1;BTRC;GNAQ;MARK2;PPP2R1A;WNT11;SRC;DKK1;PPP2CA;SOX6;SFRP2;ILK;LEF1;SOX9;TP53;MAP3K7;CREBBP;TCF7L2;AKT1;PPP2R5C;WNT5A;LAPS;CTNNB1;TGFBR1;CCND1;GSK3A;DVL1;APC;CDKN2A;MYC;CSNK1A1;GSK3B;AKT3;SOX2.
[0077] Insulin receptor signaling disorders: Genes associated with these disorders: PTEN; INS; EIF4E; PTPN1; PRKCZ; MAPK1; TSC1; PTPN11; AKT2; CBL; PIK3CA; PRKCI; PIK3CB; PIK3C3; MAPK8; IASI; MAPK3; TSC2; KRAS; EIF4EBP1; SLC2A4; PIK3C2A; PPP1CC; INSR; RAF1; FYN; MAP2K2; JAK1; AKT1; JAK2; PIK3R1; PDPK1; MAP2K1; GSK3A; FRAP1; CRKL; GSK3B; AKT3; FOXO1; SGK; RPS6KB1.
[0078] IL-6 cell signaling disorders: Genes associated with these disorders: HSPB1;TRAF6;MAPKAPK2;ELK1;MAPK1;PTPN11;IKBKB;FOS;NFKB2;MAP3K14;MAPK8;MAPK3;MAPK10;IL6ST;KRAS;MAPK13;IL6R;RELA;SOCS1;MAPK9;ABCB1;TRAF2;MAPK14;TNF;RAF1;IKBKG;RELB;MAP3K7;MAP2K2;IL8;JAK2;CHUK;STAT3;MAP2K1;NFKB1;CEBPB;JUN;IL1R1;SRF;IL6.
[0079] Hepatic cholestasis and cellular signaling disorders: Genes associated with these disorders: PRKCE;IRAK1;INS;MYD88;PRKCZ;TRAF6;PPARA;RXRA;IKBKB;PRKCI;NFKB2;MAP3K14;MAPK8;PRKD1;MAPK10;RELA;PRKCD;MAPK9;ABCB1;TRAF2;TLR4;TNF;INSR;IKBKG;RELB;MAP3K7;IL8;CHUK;NR1H2;TJP2;NFKB1;ESR1;SREBF1;FGFR4;JUN;IL1R1;PRKCA;IL6.
[0080] IGF-1 cell signaling disorders: Genes associated with these disorders: IGF1; PRKCZ; ELK1; MAPK1; PTPN11; NEDD4; AKT2; PIK3CA; PRKCI; PTK2; FOS; PIK3CB; PIK3C3; MAPK8; IGF1R; IRS1; MAPK3; IGFBP7; KRAS; PIK3C2A; YWHAZ; PXN; RAF1; CASP9; MAP2K2; AKT1; PIK3R1; PDPK1; MAP2K1; IGFBP2; SFN; JUN; CYR61; AKT3; FOXO1; SRF; CTGF; RPS6KB1.
[0081] NRF2-mediated oxidative stress response signaling disorders: Genes associated with these disorders: PRKCE;EP300;SOD2;PRKCZ;MAPK1;SQSTM1;NQO1;PIK3CA;PRKCI;FOS;PIK3CB;PIK3C3;MAPK8;PRKD1;MAPK3;KRAS;PRKCD;GSTP1;MAPK9;FTL;NFE2L2;PIK3C2A;MAPK14;RAF1;MAP3K7;CREBBP;MAP2K2;AKT1;PIK3R1;MAP2K1;PPIB;JUN;KEAP1;GSK3B;ATF4;PRKCA;EIF2AK3;HSP9OAA1.
[0082] Hepatic fibrosis / hepatic stellate cell activation signaling disorders: Genes associated with these disorders: EDN1; IGF1; KDR; FLT1; SMAD2; FGFR1; MET; PGF; SMAD3; EGFR; FAS; CSF1; NFKB2; BCL2; MYH9; IGF1R; IL6R; RELA; TLR4; PDGFRB; TNF; RELB; IL8; PDGFRA; NFKB1; TGFBR1; SMAD4; VEGFA; BAX; IL1R1; CCL2; HGF; MMP1; STAT1; IL6; CTGF; MMP9.
[0083] PPAR signaling disorders: Genes associated with these disorders: EP300; INS; TRAF6; PPARA; RXRA; MAPK1; IKBKB; NCOR2; FOS; NFKB2; MAP3K14; STAT5B; MAPK3; NRIP1; KRAS; PPARG; RELA; STAT5A; TRAF2; PPARGC1A; PDGFRB; TNF; INSR; RAF1; IKBKG; RELB; MAP3K7; CREBBP; MAP2K2; CHUK; PDGFRA; MAP2K1; NFKB1; JUN; IL1R1; HSP9OAA1.
[0084] Fc epsilon RI signaling disorders: Genes associated with these disorders: PRKCE; RAC1; PRKCZ; LYN; MAPK1; RAC2; PTPN11; AKT2; PIK3CA; SYK; PRKCI; PIK3CB; PIK3C3; MAPK8; PRKD1; MAPK3; MAPK10; KRAS; MAPK13; PRKCD; MAPK9; PIK3C2A; BTK; MAPK14; TNF; RAF1; FYN; MAP2K2; AKT1; PIK3R1; PDPK1; MAP2K1; AKT3; VAV3; PRKCA.
[0085] G protein-coupled receptor signaling disorders: Genes associated with these disorders: PRKCE;RAP1A;RGS16;MAPK1;GNAS;AKT2;IKBKB;PIK3CA;CREB1;GNAQ;NFKB2;CAMK2A;PIK3CB;PIK3C3;MAPK3;KRAS;RELA;SRC;PIK3C2A;RAF1;IKBKG;RELB;FYN;MAP2K2;AKT1;PIK3R1;CHUK;PDPK1;STAT3;MAP2K1;NFKB1;BRAF;ATF4;AKT3;PRKCA.
[0086] Inositol phosphate metabolism signaling disorders: Genes associated with these disorders: PRKCE; IRAK1; PRKAA2; EIF2AK2; PTEN; GRK6; MAPK1; PLK1; AKT2; PIK3CA; CDK8; PIK3CB; PIK3C3; MAPK8; MAPK3; PRKCD; PRKAA1; MAPK9; CDK2; PIM1; PIK3C2A; DYRK1A; MAP2K2; PIP5K1A; PIK3R1; MAP2K1; PAK3; ATM; TTK; CSNK1A1; BRAF; SGK.
[0087] PDGF signaling disorders: Genes associated with these disorders: EIF2AK2;ELK1;ABL2;MAPK1;PIK3CA;FOS;PIK3CB;P IK3 C3;MAPK8;CAV1;ABL1;MAPK3;KRAS;SRC;PIK3C2A;PDGFRB;RAF1;MAP2K2;JAK1;JAK2;PIK3R1;PDGFRA;STAT3;SPHK1;MAP2K1;MYC;JUN;CRKL;PRKCA;SRF;STAT1;SPHK2 VEGF signaling disorders: Genes associated with these disorders: ACTN4; ROCK1; KDR; FLT1; ROCK2; MAPK1; PGF; AKT2; PIK3CA; ARNT; PTK2; BCL2; PIK3CB; PIK3C3; BCL2L1; MAPK3; KRAS; HIF1A; NOS3; PIK3C2A; PXN; RAF1; MAP2K2; ELAVL1; AKT1; PIK3R1; MAP2K1; SFN; VEGFA; AKT3; FOXO1; PRKCA.
[0088] Natural killer cell signaling disorders: Genes associated with these disorders: PRKCE;RAC1;PRKCZ;MAPK1;RAC2;PTPN11;KIR2DL3;AKT2;PIK3CA;SYK;PRKCI;PIK3CB;PIK3C3;PRKD1;MAPK3;KRAS;PRKCD;PTPN6;PIK3C2A;LCK;RAF1;FYN;MAP2K2;PAK4;AKT1;PIK3R1;MAP2K1;PAK3;AKT3;VAV3;PRKCA.
[0089] Cell cycle: Gl / S checkpoint control signaling disorders: Genes associated with these disorders: HDAC4; SMAD3; SUV39H1; HDAC5; CDKN1B; BTRC; ATR; ABL1; E2F1; HDAC2; HDAC7A; RB1; HDAC11; HDAC9; CDK2; E2F2; HDAC3; TP53; CDKN1A; CCND1; E2F4; ATM; RBL2; SMAD4; CDKN2A; MYC; NRG1; GSK3B; RBL1; HDAC6.
[0090] T cell receptor signaling disorders: Genes associated with these disorders: RAC1;ELK1;MAPK1;IKBKB;CBL;PIK3CA;FOS;NFKB2;PIK3CB;PIK3C3;MAPK8;MAPK3;KRAS;RELA,PIK3C2A;BTK;LCK;RAF1;IKBKG;RELB,FYN;MAP2K2;PIK3R1;CHUK;MAP2K1;NFKB1;ITK;BCL10;JUN;VAV3.
[0091] Death receptor dysfunction: Genes associated with these dysfunctions: CRADD;HSPB1;BID;BIRC4;TBK1;IKBKB;FADD;FAS;NFKB2;BCL2;MAP3K14;MAPK8;RIPK1;CASP8;DAXX;TNFRSF10B;RELA;TRAF2;TNF;IKBKG;RELB;CASP9;CHUK;APAF1;NFKB1;CASP2;BIRC2;CASP3;BIRC3.
[0092] FGF cell signaling disorders: Genes associated with these disorders: RAC1;FGFR1;MET;MAPKAPK2;MAPK1;PTPN11;AKT2;PIK3CA;CREB1;PIK3CB;PIK3C3;MAPK8;MAPK3;MAPK13;PTPN6;PIK3C2A;MAPK14;RAF1;AKT1;PIK3R1;STAT3;MAP2K1;FGFR4;CRKL;ATF4;AKT3;PRKCA;HGF.
[0093] GM-CSF cell signaling disorders: Genes associated with these disorders: LYN;ELK1;MAPK1;PTPN11;AKT2;PIK3CA;CAMK2A;STAT5B;PIK3CB;PIK3C3;GNB2L1;BCL2L1;MAPK3;ETS1;KRAS;RUNX1;PIM1;PIK3C2A;RAF1;MAP2K2;AKT1;JAK2;PIK3R1;STAT3;MAP2K1;CCND1;AKT3;STAT1.
[0094] Amyotrophic lateral sclerosis (ALS) cell signaling disorders: Genes associated with these disorders: BID; IGF1; RAC1; BIRC4; PGF; CAPNS1; CAPN2; PIK3CA; BCL2; PIK3CB; PIK3C3; BCL2L1; CAPN1; PIK3C2A; TP53; CASP9; PIK3R1; RAB5A; CASP1; APAF1; VEGFA; BIRC2; BAX; AKT3; CASP3; BIRC3 PTPN1; MAPK1; PTPN11; AKT2; PIK3CA; STAT5B; PIK3CB; PIK3C3; MAPK3; KRAS; SOCS1; STAT5A; PTPN6; PIK3C2A; RAF1; CDKN1A; MAP2K2; JAK1; AKT1; JAK2; PIK3R1; STAT3; MAP2K1; FRAP1; AKT3; STAT1.
[0095] JAK / Stat cell signaling disorders: Genes associated with these disorders: PTPN1; MAPK1; PTPN11; AKT2; PIK3CA; STAT5B; PIK3CB; PIK3C3; MAPK3; KRAS; SOCS1; STAT5A; PTPN6; PIK3C2A; RAF1; CDKN1A; MAP2K2; JAK1; AKT1; JAK2; PIK3R1; STAT3; MAP2K1; FRAP1; AKT3; STAT1.
[0096] Disorders in nicotinic acid and nicotinamide metabolic cell signaling: Genes associated with these disorders: IRAK1;PRKAA2;EIF2AK2;GRK6;MAPK1;PLK1;AKT2;CDK8;MAPK8;MAPK3;PRKCD;PRKAA1;PBEF1;MAPK9;CDK2;PIM1;DYRK1A;MAP2K2;MAP2K1;PAK3;NT5E;TTK;CSNK1A1;BRAF;SGK.
[0097] Chemokine cell signaling disorders: Genes associated with these disorders: CXCR4; ROCK2; MAPK1; PTK2; FOS; CFL1; GNAQ; CAMK2A; CXCL12; MAPK8; MAPK3; KRAS; MAPK13; RHOA; CCR3; SRC; PPP1CC; MAPK14; NOX1; RAF1; MAP2K2; MAP2K1; JUN; CCL2; PRKCA.
[0098] IL-2 cell signaling disorders: Genes associated with these disorders: ELK1; MAPK1; PTPN11; AKT2; PIK3CA; SYK; FOS; STAT5B; PIK3CB; PIK3C3; MAPK8; MAPK3; KRAS; SOCS1; STAT5A; PIK3C2A; LCK; RAF1; MAP2K2; JAK1; AKT1; PIK3R1; MAP2K1; JUN; AKT3.
[0099] Synaptic long-term inhibitory signaling disorders: Genes associated with these disorders: PRKCE; IGF1; PRKCZ; PRDX6; LYN; MAPK1; GNAS; PRKCI; GNAQ; PPP2R1A; IGF1R; PRKD1; MAPK3; KRAS; GRN; PRKCD; NOS3; NOS2A; PPP2CA; YWHAZ; RAF1; MAP2K2; PPP2R5C; MAP2K1; PRKCA.
[0100] Estrogen receptor cell signaling disorders: Genes associated with these disorders: TAF4B; EP300; CARM1; PCAF; MAPK1; NCOR2; SMARCA4; MAPK3; NRIP1; KRAS; SRC; NR3C1; HDAC3; PPARGC1A; RBM9; NCOA3; RAF1; CREBBP; MAP2K2; NCOA2; MAP2K1; PRKDC; ESR1; ESR2.
[0101] Protein ubiquitination pathway cell signaling disorders: Genes associated with these disorders: TRAF6; SMURF1; BIRC4; BRCAl; UCHL1; NEDD4; CBL; UBE2I; BTRC; HSPA5; USP7; USP10; FBXW7; USP9X; STUB1; USP22; B2M; BIRC2; PARK2; USP8; USP1; VHL; HSP9OAA1; BIRC3.
[0102] IL-10 cell signaling disorders: Genes associated with these disorders: TRAF6; CCR1; ELK1; IKBKB; SP1; FOS; NFKB2; MAP3K14; MAPK8; MAPK13; RELA; MAPK14; TNF; IKBKG; RELB; MAP3K7; JAK1; CHUK; STAT3; NFKB1; JUN; IL1R1; IL6.
[0103] VDR / RXR activation signaling disorders: Genes associated with these disorders: PRKCE; EP300; PRKCZ; RXRA; GADD45A; HES1; NCOR2; SP1; PRKCI; CDKN1B; PRKD1; PRKCD; RUNX2; KLF4; YY1; NCOA3; CDKN1A; NCOA2; SPP1; LAPS; CEBPB; FOXO1; PRKCA.
[0104] TGF-beta cell signaling disorders: Genes associated with these disorders: EP300; SMAD2; SMURF1; MAPK1; SMAD3; SMAD1; FOS; MAPK8; MAPK3; KRAS; MAPK9; RUNX2; SERPINE1; RAF1; MAP3K7; CREBBP; MAP2K2; MAP2K1; TGFBR1; SMAD4; JUN; SMAD5.
[0105] Toll-like receptor cell signaling disorders: Genes associated with these disorders: IRAK1; EIF2AK2; MYD88; TRAF6; PPARA; ELK1; IKBKB; FOS; NFKB2; MAP3K14; MAPK8; MAPK13; RELA; TLR4; MAPK14; IKBKG; RELB; MAP3K7; CHUK; NFKB1; TLR2; JUN.
[0106] p38 MAPK cell signaling disorders: Genes associated with these disorders: HSPB1; IRAK1; TRAF6; MAPKAPK2; ELK1; FADD; FAS; CREB1; DDIT3; RPS6KA4; DAXX; MAPK13; TRAF2; MAPK14; TNF; MAP3K7; TGFBR1; MYC; ATF4; IL1R1; SRF; STAT1.
[0107] Neurolrophin / TRK cell signaling disorders: Genes associated with these disorders: NTRK2; MAPK1; PTPN11; PIK3CA; CREB1; FOS; PIK3CB; PIK3C3; MAPK8; MAPK3; KRAS; PIK3C2A; RAF1; MAP2K2; AKT1; PIK3R1; PDPK1; MAP2K1; CDC42; JUN; ATF4.
[0108] Other cellular dysfunctions related to gene modification are intended herein, including, for example, FXR / RXR activation, synaptic long-term potentiation, calcium signaling, EGF signaling, hypoxic signaling in the cardiovascular system, LPS / IL-1 mediated inhibition of RXR function, LXR / RXR activation, amyloid processing, IL-4 signaling, and cell cycle: G2 / M DNA damage checkpoint control, nitric oxide signaling in the cardiovascular system, purine metabolism, cAMP-mediated signaling, mitochondrial dysfunction, Notch signaling, endoplasmic reticulum stress pathway, pyrimidine metabolism, Parkinsonian signaling, cardiac and beta-adrenergic signaling, glycolysis / gluconeogenesis, interferon signaling, sonic hedgehog signaling, glycerophospholipid metabolism, phospholipid degradation, tryptophan metabolism, lysine degradation, nucleotide excision repair pathway, starch and sucrose metabolism, aminoglycoside metabolism, arachidonic acid metabolism, circadian rhythm signaling, coagulation system, dopamine receptor signaling, glutathione metabolism, glycerolipid metabolism, linoleic acid metabolism, methionine metabolism, pyruvate metabolism, arginine and proline metabolism, eicosanoid signaling, fructose and mannose metabolism, galactose metabolism, stilbene, coumarin and ligni Biosynthesis, antigen presentation pathways, steroid biosynthesis, butanoic acid metabolism, citrate cycle, fatty acid metabolism, glycerophospholipid metabolism, histidine metabolism, inositol metabolism, xenobiotic metabolism by cytochrome P450, methane metabolism, phenylalanine metabolism, propanoic acid metabolism, selenoamino acid metabolism, sphingolipid metabolism, aminophosphonic acid metabolism, androgen and estrogen metabolism, ascorbic acid and aldalic acid metabolism, bile acid biosynthesis, cysteine metabolism, fatty acid biosynthesis, glutamate receptor signaling, NRF2-mediated oxidative stress response, pentose phosphate pathway, interconversion of pentoses and glucuronic acid, retinol metabolism, riboflavin metabolism, tyrosine metabolism, ubiquinone biosynthesis, degradation of valine, leucine, and isoleucine, metabolism of glycine, serine, and threonine, lysine degradation, pain / taste, or developmental neurology of mitochondrial function, or a combination of these.
[0109] Nucleic acid-induced nucleases can include native sequences, engineered sequences, or engineered nucleotide sequences of synthetic variants. Non-limiting examples of the types of manipulations that can be performed to obtain non-native nuclease systems include: Manipulations may include codon optimization to promote or improve expression in host cells, such as heterologous host cells. Manipulations can reduce the size or molecular weight of the nuclease to facilitate expression or delivery. Manipulations can alter the selection of PAMs to change the specificity of PAMs or broaden the range of PAMs that are recognized. Manipulations can alter, increase, or decrease the stability, processing capacity, specificity, or efficiency of a targetable nuclease system. Manipulations can alter, increase, or decrease the stability of a protein. Manipulations can alter, increase, or decrease the processing capacity of nucleic acid scans. Manipulations can alter, increase, or decrease the specificity of a target sequence. Manipulations can alter, increase, or decrease nuclease activity. Manipulations can alter, increase, or decrease editing efficiency. Manipulations can alter, increase, or decrease transformation efficiency. The manipulation can alter, increase, or decrease nucleases or induce nucleic acid expression. As used herein, non-natural nucleic acid sequences may be manipulated sequences of synthetic variants or manipulated nucleotide sequences. Such non-natural nucleic acid sequences can be obtained by amplification, cloning, construction, synthesis, generation from synthetic oligonucleotides or dNTPs, or otherwise, using methods known to those skilled in the art.In certain embodiments, examples of non-natural nucleic acid-inducible nucleases disclosed herein include engineered polypeptide sequences (e.g., SEQ ID NOs. 143-177, 229, 257-262, and 330, which may also include additional amino acid sequences described herein); engineered polynucleotide sequences that thus encode (e.g., SEQ ID NOs. 1-142, 225-228, 230-256, and 230, which may also include additional nucleotide sequences described herein); one or more polynucleotides, including engineered gRNA, that are compatible with these nucleic acid-inducible nucleases, including a synthetic variant or a portion of a nucleotide sequence or a portion of a nucleotide sequence (e.g., SEQ ID NOs. 178-188, sequences described in Table 3, or a portion thereof); and / or other nucleic acid-inducible nucleases that include others described herein.
[0110] Nucleic acid-inducible nucleases are disclosed herein. The disclosed nucleic acid-inducible nucleases are understood to be functional in vitro or in prokaryotic, archaeal, or eukaryotic cells for in vitro, in vivo, or ex vivo application.Additionally, the phytoplankton is Thiomicrospira, Succinivibrio, Candidatus, Porph yromonas、Acidaminococcus、Acidomonococcus、Barnesiella、Prevotel the、Smithella、Moraxella、Synergistes、Francisella、Leptospira、Catenibacterium、Kandleria、Clostridium、Dorea、Coprococcus、Enterococ cus, Fructobacillus, Weissella, Pediococcus, Collinsella, Corynebacter, Sutterella, Legionella, Treponema, Roseburia, Filifactor, Lach nospiraceae、Eubacterium、Sedimentisphaera、Streptococcus、Lactobacillus、Mycoplasma、Bacteroides、Flaviivola、Flavobacterium、Sphae rochaeta、Azospirillum、Gluconacetobacter、Neisseria、Roseburia、Parvibaculum、Parabacteroides、Staphylococcus、Nitratifractor、Myco plasma、Alicyclobacillus、Brevibacillus、Bacillus、Bacteroidetes、Brevibacillus、Carnobacterium、Clostridiaridium、Clostridium、Desulf onatronum、Desulfovibrio、Helcococcus、Leptotrichia、Listeria、Methanomethyophilus、Methylobacterium、Opitutaceae、Paludibacter、Rho dobacter、Sphaerochaeta、Tuberibacillus、Oleiphilus、Omnitrophica Parkubacteria and Campylobacter The rest of the snow is covered with a snowflake.The species of organisms in these genera may be as discussed elsewhere herein. Suitable gRNAs may be derived from organisms of genera or unclassified genera within the kingdom, including but not limited to Firmicute, Actinobacteria, Bacteroidetes, Proteobacteria, Spirochates, and Tenericutes. Suitable gRNAs may be derived from organisms of genera or unclassified genera within the phylum, including but not limited to Erysipelotrichia, Clostridia, Bacilli, Actinobacteria, Bacteroidetes, Catenovolum, Coprococcus, Flavobacteria, Alphaproteobacteria, Betaproteobacteria, Gammaproteobacteria, Deltaproteobacteria, Epsilonproteobacteria, Spirochaetes, and Mollicutes. Suitable gRNAs may originate from organisms of genera or unclassified genera within the order, including but not limited to Clostridiales, Lactobacillales, Actinomycetales, Bacteroidales, Flavobacteriales, Rhizobiales, Rhodospirillales, Burkholderiales, Neisseriales, Legionellales, Nautiliales, Campylobacterales, Spirochaetales, Mycoplasmatales, and Thiotrichales.Suitable gRNAs may originate from organisms of genera within these families or unclassified genera, including but not limited to those of Lachnospiraceae, Enterococcaceae, Leuconostocaceae, Lactobacillaceae, Streptococcaceae, Peptostreptococcaceae, Staphylococcaceae, Eubacteriaceae, Corynebacterineae, Bacteroidaceae, Flavobacterium, Cryomoorphaceae, Rhodobiaceae, Rhodospirillaceae, Acetobacteraceae, Sutterellaceae, Neisseriaceae, Legionellaceae, Nautiliaceae, Campylobacteraceae, Spirochaetaceae, Mycoplasmataceae, Pisciririckettsiaceae, and Francisellaceae. In some embodiments, suitable gRNAs may be derived from organisms of genera within families or unclassified genera, including Acidaminococcus, Sedimentisphaera, Barnesiella sp., Bacteroidetes, Parabacteroides, Lachnospiraceae, Coprococcus sp., Catenovolum sp., and Collinsella. Other nucleic acid-inducible nucleases are described in U.S. Patent Application Publication US20160208243 filed December 18, 2015, U.S. Patent Application Publication US20140068797 filed March 15, 2013, U.S. Patent No. 8,697,359 filed October 15, 2013, and Zetsche et al., Cell 2015 Oct.22;163(3):759-71.
[0111] Some nucleic acid-derived nucleases suitable for use in the methods, systems, and compositions of this disclosure include, but are not limited to, those derived from organisms such as: Thiomicrospira sp.XS5, Eubacterium rectale, Succinivibrio dextrinosolvens, Candidatus Methanoplasma termitum, Candidatus Methanomethylophilus alvus, Porphyromonas crevioricanis, Flavobacterium branchiophilum, Acidaminococcus Sp., Acidomonococcus sp., Lachnospiraceae bacterium COE1, Prevotella brevis ATCC 19188, Smithella sp.SCADC, Moraxella bovoculi, Synergistes jonesii, Bacteroidetes oral taxa 274, Francisella tularensis, Leptospira inadai serovar Lyme str.10, Acidomonococcus sp. Crystal structure (5B43) S.mutans, S.agalactiae, S.equisimilis, S.sanguinis, S.pneumonia; C.jejuni, C.coli; N.salsuginis, N.tergarcus; S.auric ularis, S. carnosus; N. meningitides, N. gonorrhoeae; L. monocytogenes, L. ivanovii; C. botulinum, C. difficile, C. tetani, C.sordellii; Francisella tularensis 1, Prevotella albensis, Lachnospiraceae bacterium MC2017 1, Butyrivibrio proteoclasticus, Butyrivibrio proteoclasticus B316, Peregrinibacteria bacterium GW2011_GWA2_33_10, Parcubacteria bacterium GW2011_GWC2_44_17, Smithella sp.SCADC, Acidaminococcus sp.BV3L6, Lachnospiraceae bacterium MA2020, Candidatus Methanoplasma termitum, Eubacterium eligens, Moraxella bovoculi 237, Leptospira inadai, Lachnospiraceae bacterium ND2006, Porphyromonas crevioricanis 3, Prevotella disiens, Porphyromonas macacae, Catenibacterium sp.CAG:290, Kandleria vitulina, Clostridiales bacterium KA00274, Lachnospiraceae bacterium 3-2, Dorea longicatena, Coprococcus catus GD / 7, Enterococcus columbae DSM 7374, Fructobacillus sp.EFB-N1, Weissella halotolerans, Pediococcus acidilactici, Lactobacillus curvatus, Streptococcus pyogenes, Lactobacillus versmoldensis, Filifactor alocis ATCC 35896, Alicyclobacillus acidoterrestris, Alicyclobacillus acidoterrestris ATCC 49025, Desulfovibrio inopinatus, Desulfovibrio inopinatus DSM 10711, Oleiphilus sp.Oleiphilus sp.HI0009, Candidtus kefeldibacteria, Parcubacteria CasY.4, Omnitrophica WOR 2 bacterium GWF2, Bacillus sp.NSP2.1, Bacillus thermoamylovorans, Catenovulum sp.CCB-QB4, Coprococcus sp.AF16-5, Lachnospiraceae bacterium MC2017, Collinsella, tanakaei, Parabacteroides distasonis, Bacteroidetes bacterium HGW-Bacteroidetes-6, Barnesiella sp.An22, Sedimentesphaera cyanobacteriorum, and Acidaminococcus massiliensis.
[0112] In some embodiments, the nucleic acid-inducible nucleases disclosed herein include polypeptides having an amino acid sequence that is at least 50% identical to any one of SEQ ID NOs: 143-177 and 229. In some embodiments, the nucleic acid-inducible nucleases disclosed herein include polypeptides having an amino acid sequence that is at least 60%, 65%, 75%, 85%, 95%, 99%, or 100% identical to one or more amino acid sequences from SEQ ID NOs: 143-177 and 229. In some embodiments, the nucleic acid-inducible nucleases disclosed herein include polypeptides having an amino acid sequence that is at least 60%, 65%, 75%, 85%, 95%, 99%, or 100% identical to one or more amino acid sequences from SEQ ID NOs: 143-151. In some embodiments, the nucleic acid-derived nucleases disclosed herein include an amino acid sequence having at least 85%, 90%, 95%, 99%, or 100% amino acid identity with any one of SEQ ID NOs: 143, 144, 147, 148, 150, and 151. In some embodiments, the nucleic acid-derived nucleases disclosed herein include a polypeptide having at least 85%, 90%, 95%, 99%, or 100% amino acid identity with the amino acid sequence represented by SEQ ID NO: 144.
[0113] In certain embodiments of this specification, a nucleic acid-inducible nuclease, such as an engineered nucleic acid-inducible nuclease, includes an amino acid sequence having at least 60%, 65%, 75%, 85%, 90%, 95%, 99%, or 100% sequence identity with the amino acid sequence represented by any one of SEQ ID NOs: 143-177 and 229, for example, at least 85%, possibly at least 90%, possibly at least 95%, possibly at least 99%, or even 100% sequence identity. In certain embodiments, a nucleic acid-inducible nuclease, such as an engineered nucleic acid-inducible nuclease, includes an amino acid sequence having at least 60% sequence identity with the amino acid sequence represented by any one of SEQ ID NOs: 143-177 and 229. In certain embodiments, a nucleic acid-inducible nuclease, such as an engineered nucleic acid-inducible nuclease, includes an amino acid sequence having at least 80% sequence identity with the amino acid sequence represented by any one of SEQ ID NOs: 143-177 and 229. In certain embodiments, the nucleic acid-inducible nuclease, such as the manipulated nucleic acid-inducible nuclease, includes an amino acid sequence having at least 90% sequence identity with the amino acid sequence represented by any one of SEQ ID NOs: 143-177 and 229. In certain embodiments, the nucleic acid-inducible nuclease, such as the manipulated nucleic acid-inducible nuclease, includes an amino acid sequence having at least 95% sequence identity with the amino acid sequence represented by any one of SEQ ID NOs: 143-177 and 229. In certain embodiments, the nucleic acid-inducible nuclease, such as the manipulated nucleic acid-inducible nuclease, includes an amino acid sequence having 100% sequence identity with the amino acid sequence represented by any one of SEQ ID NOs: 143-177 and 229. These amino acid sequences may be the entire amino acid sequence of the nuclease polypeptide, or they may be the original nuclease polypeptide with the additional amino acid sequences added as described above.
[0114] In certain embodiments, the nucleases disclosed herein do not share a conserved peptide motif, or a polynucleotide encoding the nuclease, or a nucleotide sequence encoding such a motif, with any known nuclease. In certain embodiments, the nucleases disclosed herein do not contain the peptide motif YLFQIYNKDF (SEQ ID NO: 224). In certain embodiments, one or more polynucleotides encoding the nucleases disclosed herein do not encode the peptide motif YLFQIYNKDF (SEQ ID NO: 224) within the encoded nuclease. The motif of SEQ ID NO: 224 may be absent entirely, or it may have 1, 2, 3, 4, 5, or more than 5 substituted amino acids compared to SEQ ID NO: 224. The substitutions may be conserved, radical, or any combination thereof. In certain embodiments, sequences other than SEQ ID NO: 224 may have at least one radical substitution. In certain embodiments, sequences other than sequence number 224 may have at least one, two, three, or four substitutions having a Sneath's index value of at least 10, 15, 20, or 25. In certain embodiments, sequences other than sequence number 224 may have at least one substitution having a Sneath's index value of at least 25.
[0115] In certain embodiments of this specification, a nucleic acid-inducible nuclease, such as an engineered nucleic acid-inducible nuclease, comprises an amino acid sequence having at least 60%, 65%, 75%, 85%, 90%, 95%, 99%, or 100% sequence identity with any one of the amino acid sequences represented by SEQ ID NOs: 143-151, 161-163, 165, 166, 169, 171-175, 177, and 229, for example, at least 85%, possibly at least 90%, possibly at least 95%, possibly at least 99%, or even 100% sequence identity. In certain embodiments, the nucleic acid-inducible nuclease, such as the manipulated nucleic acid-inducible nuclease, includes an amino acid sequence having at least 60% sequence identity with the amino acid sequence represented by any one of SEQ ID NOs: 143-151, 161-163, 165, 166, 169, 171-175, 177, and 229. In certain embodiments, the nucleic acid-inducible nuclease, such as the manipulated nucleic acid-inducible nuclease, includes an amino acid sequence having at least 80% sequence identity with the amino acid sequence represented by any one of SEQ ID NOs: 143-151, 161-163, 165, 166, 169, 171-175, 177, and 229. In certain embodiments, the nucleic acid-inducible nuclease, such as the manipulated nucleic acid-inducible nuclease, contains an amino acid sequence having at least 90% sequence identity with the amino acid sequence represented by any one of SEQ ID NOs: 143-151, 161-163, 165, 166, 169, 171-175, 177, and 229. In certain embodiments, the nucleic acid-inducible nuclease, such as the manipulated nucleic acid-inducible nuclease, contains an amino acid sequence having at least 95% sequence identity with the amino acid sequence represented by any one of SEQ ID NOs: 143-151, 161-163, 165, 166, 169, 171-175, 177, and 229. In certain embodiments, the manipulated nucleic acid-inducible nuclease, such as the nucleic acid-inducible nuclease, contains an amino acid sequence that has 100% sequence identity with any one of the amino acid sequences represented by SEQ ID NOs: 143-151, 161-163, 165, 166, 169, 171-175, 177, and 229.These amino acid sequences may be the entire amino acid sequence of the nuclease polypeptide, or they may be the original nuclease polypeptide with the additional amino acid sequences added as described above.
[0116] In certain embodiments of this specification, a nucleic acid-inducible nuclease, such as an engineered nucleic acid-inducible nuclease, includes an amino acid sequence having at least 60%, 65%, 75%, 85%, 90%, 95%, 99%, or 100% sequence identity with the amino acid sequence represented by any one of SEQ ID NOs: 149, 151, 175, and 177, for example, at least 85%, possibly at least 90%, possibly at least 95%, possibly at least 99%, or even 100% sequence identity. In certain embodiments, a nucleic acid-inducible nuclease, such as an engineered nucleic acid-inducible nuclease, includes an amino acid sequence having at least 60% sequence identity with the amino acid sequence represented by any one of SEQ ID NOs: 149, 151, 175, and 177. In certain embodiments, a nucleic acid-inducible nuclease, such as a manipulated nucleic acid-inducible nuclease, includes an amino acid sequence having at least 80% sequence identity with the amino acid sequence represented by any one of SEQ ID NOs: 149, 151, 175, and 177. In certain embodiments, a nucleic acid-inducible nuclease, such as a manipulated nucleic acid-inducible nuclease, includes an amino acid sequence having at least 90% sequence identity with the amino acid sequence represented by any one of SEQ ID NOs: 149, 151, 175, and 177. In certain embodiments, a nucleic acid-inducible nuclease, such as a manipulated nucleic acid-inducible nuclease, includes an amino acid sequence having at least 95% sequence identity with the amino acid sequence represented by any one of SEQ ID NOs: 149, 151, 175, and 177. In certain embodiments, the manipulated nucleic acid-inducible nuclease, such as the nucleic acid-inducible nuclease, comprises an amino acid sequence having 100% sequence identity with the amino acid sequence represented by any one of SEQ ID NOs: 149, 151, 175, and 177. These amino acid sequences may be the entire amino acid sequence of the nuclease polypeptide, or they may be the original nuclease polypeptide with the additional amino acid sequences added as described above.
[0117] In certain embodiments of this specification, a nucleic acid-inducible nuclease, such as an engineered nucleic acid-inducible nuclease, includes an amino acid sequence having at least 60%, 65%, 75%, 85%, 90%, 95%, 99%, or 100% sequence identity with the amino acid sequence represented by any one of SEQ ID NOs: 144, 153, and 229, for example, at least 85%, possibly at least 90%, possibly at least 95%, possibly at least 99%, or even 100% sequence identity. In certain embodiments, a nucleic acid-inducible nuclease, such as an engineered nucleic acid-inducible nuclease, includes an amino acid sequence having at least 60% sequence identity with the amino acid sequence represented by any one of SEQ ID NOs: 144, 153, and 229. In certain embodiments, a nucleic acid-inducible nuclease, such as an engineered nucleic acid-inducible nuclease, includes an amino acid sequence having at least 80% sequence identity with the amino acid sequence represented by any one of SEQ ID NOs: 144, 153, and 229. In certain embodiments, the nucleic acid-inducible nuclease, such as the manipulated nucleic acid-inducible nuclease, includes an amino acid sequence having at least 90% sequence identity with the amino acid sequence represented by any one of SEQ ID NOs: 144, 153, and 229. In certain embodiments, the nucleic acid-inducible nuclease, such as the manipulated nucleic acid-inducible nuclease, includes an amino acid sequence having at least 95% sequence identity with the amino acid sequence represented by any one of SEQ ID NOs: 144, 153, and 229. In certain embodiments, the nucleic acid-inducible nuclease, such as the manipulated nucleic acid-inducible nuclease, includes an amino acid sequence having 100% sequence identity with the amino acid sequence represented by any one of SEQ ID NOs: 144, 153, and 229. These amino acid sequences may be the entire amino acid sequence of the nuclease polypeptide, or they may be the original nuclease polypeptide with the additional amino acid sequences added as described above.
[0118] In certain embodiments of this specification, a nucleic acid-inducible nuclease, such as an engineered nucleic acid-inducible nuclease, includes an amino acid sequence having at least 60%, 65%, 75%, 85%, 90%, 95%, 99%, or 100% sequence identity with the amino acid sequence represented by SEQ ID NO: 144, for example, at least 85%, possibly at least 90%, possibly at least 95%, possibly at least 99%, or even 100% sequence identity. In certain embodiments, a nucleic acid-inducible nuclease, such as an engineered nucleic acid-inducible nuclease, includes an amino acid sequence having at least 60% sequence identity with the amino acid sequence represented by SEQ ID NO: 144. In certain embodiments, a nucleic acid-inducible nuclease, such as an engineered nucleic acid-inducible nuclease, includes an amino acid sequence having at least 80% sequence identity with the amino acid sequence represented by SEQ ID NO: 144. In certain embodiments, the nucleic acid-inducible nuclease, such as the manipulated nucleic acid-inducible nuclease, includes an amino acid sequence having at least 90% sequence identity with the amino acid sequence represented by SEQ ID NO: 144. In certain embodiments, the nucleic acid-inducible nuclease, such as the manipulated nucleic acid-inducible nuclease, includes an amino acid sequence having at least 95% sequence identity with the amino acid sequence represented by SEQ ID NO: 144. In certain embodiments, the nucleic acid-inducible nuclease, such as the manipulated nucleic acid-inducible nuclease, includes an amino acid sequence having 100% sequence identity with the amino acid sequence represented by any one of SEQ ID NOs: 144. These amino acid sequences may be the entire amino acid sequence of the nuclease polypeptide, or they may be the original nuclease polypeptide with the additional amino acid sequences added as described above.
[0119] In certain embodiments of this specification, a nucleic acid-inducible nuclease, such as an engineered nucleic acid-inducible nuclease, includes an amino acid sequence having at least 60%, 65%, 75%, 85%, 90%, 95%, 99%, or 100% sequence identity with the amino acid sequence represented by SEQ ID NO: 153, for example, at least 85%, possibly at least 90%, possibly at least 95%, possibly at least 99%, or even 100% sequence identity. In certain embodiments, a nucleic acid-inducible nuclease, such as an engineered nucleic acid-inducible nuclease, includes an amino acid sequence having at least 60% sequence identity with the amino acid sequence represented by SEQ ID NO: 153. In certain embodiments, a nucleic acid-inducible nuclease, such as an engineered nucleic acid-inducible nuclease, includes an amino acid sequence having at least 80% sequence identity with the amino acid sequence represented by SEQ ID NO: 153. In certain embodiments, the nucleic acid-inducible nuclease, such as the manipulated nucleic acid-inducible nuclease, includes an amino acid sequence having at least 90% sequence identity with the amino acid sequence represented by SEQ ID NO: 153. In certain embodiments, the nucleic acid-inducible nuclease, such as the manipulated nucleic acid-inducible nuclease, includes an amino acid sequence having at least 95% sequence identity with the amino acid sequence represented by SEQ ID NO: 153. In certain embodiments, the nucleic acid-inducible nuclease, such as the manipulated nucleic acid-inducible nuclease, includes an amino acid sequence having 100% sequence identity with the amino acid sequence represented by any one of SEQ ID NOs: 153. These amino acid sequences may be the entire amino acid sequence of the nuclease polypeptide, or they may be the original nuclease polypeptide with the additional amino acid sequences added as described above.
[0120] Nucleic acid-inducible nucleases may include an amino acid sequence that is an engineered sequence, i.e., an amino acid sequence that does not match any known native sequence, even if it does not include additional amino acid sequences such as those described herein. In certain embodiments herein, a nucleic acid-inducible nuclease, such as an engineered nucleic acid-inducible nuclease, includes an amino acid sequence having at least 60%, 65%, 75%, 85%, 90%, 95%, 99%, or 100% sequence identity with the amino acid sequence represented by SEQ ID NO: 229, for example, at least 85%, possibly at least 90%, possibly at least 95%, possibly at least 99%, or even 100% sequence identity. In certain embodiments, a nucleic acid-inducible nuclease, such as an engineered nucleic acid-inducible nuclease, includes an amino acid sequence having at least 60% sequence identity with the amino acid sequence represented by SEQ ID NO: 229. In certain embodiments, the nucleic acid-inducible nuclease, such as the manipulated nucleic acid-inducible nuclease, includes an amino acid sequence having at least 80% sequence identity with the amino acid sequence represented by SEQ ID NO: 229. In certain embodiments, the nucleic acid-inducible nuclease, such as the manipulated nucleic acid-inducible nuclease, includes an amino acid sequence having at least 90% sequence identity with the amino acid sequence represented by SEQ ID NO: 229. In certain embodiments, the nucleic acid-inducible nuclease, such as the manipulated nucleic acid-inducible nuclease, includes an amino acid sequence having at least 95% sequence identity with the amino acid sequence represented by SEQ ID NO: 229. In certain embodiments, the nucleic acid-inducible nuclease, such as the manipulated nucleic acid-inducible nuclease, includes an amino acid sequence having 100% sequence identity with the amino acid sequence represented by any one of SEQ ID NOs: 229. These amino acid sequences may be the entire amino acid sequence of the nuclease polypeptide, or they may be the original nuclease polypeptide with the additional amino acid sequences added as described above.
[0121] Accordingly, compositions and methods providing and / or utilizing engineered nucleic acid-inducible nuclease systems, their components and products, as well as other compositions and methods, are disclosed herein. As used herein, “engineered nucleic acid-inducible nuclease system” may also be referred herein to as a novel engineered nucleic acid-inducible nuclease construct, and non-natural nucleic acid-inducible nuclease systems, for example, may also be referred to as a nucleic acid-inducible nuclease system in which the system is non-natural. The system may include a) one or more components in its final form, used in one or more ways, for example in a host cell, after further processing at one or more sites, and b) one or more polynucleotides encoding one or more components, and the system may include a) and b). The components include one or more of the following: 1) an engineered nucleic acid-inducible nuclease or a part thereof, such as its active moiety; 2) an engineered guide nucleic acid compatible with the engineered nucleic acid-inducible nuclease or a part thereof, such as gRNA, and / or encoding or other polynucleotides; and 3) one or more engineered polynucleotides. This system can also include other components such as editing templates.
[0122] As used herein, "operated," "operated," etc., which are also referred to herein as "novel," may refer to a non-natural composition or method.
[0123] In certain embodiments, examples of non-natural nucleic acid-inducible nucleases disclosed herein include nucleic acid-inducible nucleases produced using polynucleotide sequences, such as manipulated polynucleotide sequences (e.g., SEQ ID NOs. 1-142, 225-228, and 330, or their subgroups, as described in more detail herein), and synthetic variant gNAs, such as gRNA sequences (e.g., SEQ ID NOs. 178-188). In certain embodiments, synthetic variants including gNAs, such as those shown in Table 3 and described in more detail herein, may be used.
[0124] In certain embodiments, engineered nucleic acid-inducible nucleases are provided herein. As used herein, “engineered nucleic acid-inducible nuclease” or similar terms means a non-natural nucleic acid-inducible nuclease, and the nuclease may be non-natural for any reason, such as comprising an engineered nuclease polypeptide and / or one or more engineered polynucleotides encoding it.
[0125] As used herein, “nucleic acid-induced nuclease” may also be simply referred to as nuclease, CRISPR-associated (Cas) nuclease, Cas12a-like, etc., and may refer to a nuclease that can bind to and cleave at or near a target sequence in a target polynucleotide together with a compatible guide nucleic acid, such as a compatible gRNA. “Target sequence,” as also referred herein as target nucleic acid, target polynucleotide sequence, etc., may refer to a sequence to which the guide sequence is complementary, where hybridization between the target sequence and the guide sequence enables the activity of the nuclease complex, such as the manipulated nuclease complex. The target polynucleotide of a targetable nuclease complex may be any polynucleotide that is endogenous or exogenous to the host cell. “Target polynucleotide” may refer to the polynucleotide on which the target sequence is located, as the term is used herein.
[0126] Guide nucleic acids (gNAs), such as gRNAs, and polynucleotides encoding gNAs or gRNAs, or parts of gNAs or gRNAs, are disclosed herein. In certain embodiments, the gNAs, such as gRNAs, are engineered gNAs, such as engineered gRNAs.
[0127] As used herein, “guide nucleic acid” or “guide polynucleotide” (gNA) may refer to one or more polynucleotides, and a gNA comprises 1) a guide sequence that can hybridize to a target sequence, and 2) a scaffold sequence that can interact with or complex with a nucleic acid-inducible nuclease. A gNA, such as a gRNA, needs to complex with a compatible nucleic acid-inducible nuclease in order for the nuclease complex to be located at or near the target sequence and cleave it. “Guide RNA (gRNA),” also referred herein as an RNA guide polynucleotide, is a gNA whose nucleotide is either a native or modified ribonucleotide.
[0128] The target polynucleotide of a targetable nuclease complex may be any polynucleotide that is endogenous or exogenous to the host cell. “Target polynucleotide” may refer to the polynucleotide in which the target sequence is located, as used herein. The target polynucleotide may include coding nucleotides or non-coding nucleotides. In certain embodiments, the target sequence is located within the target polynucleotide, which is a safe harbor site (SHS).
[0129] The guide nucleic acid can be provided as one or more nucleic acids.
[0130] In certain embodiments, guide nucleic acids, e.g., gRNAs, are provided as two distinct polynucleotides that associate to form a functional guide nucleic acid, e.g., gRNA (split or dual guide nucleic acid, e.g., split or dual gRNA). In certain embodiments, nucleic acid-inducible nucleases, such as the engineered nucleic acid-inducible nucleases disclosed herein, bind to compatible gNAs, e.g., compatible gRNAs, including split gNAs, e.g., split gRNAs, where the nuclease, or the nuclease sequence from which it is derived, does not bind to split gNAs, e.g., split gRNAs in its native state, but binds to single gNAs, e.g., single gRNAs. In at least some of these native nucleases, e.g., Cas12a, tracrRNA is absent in the native gRNA. In certain embodiments herein, gRNAs such as split gRNAs include tracrRNA. For a further discussion of these non-native gNAs, see PCT Publication WO2021067788.
[0131] In certain embodiments, the guide sequence and scaffold sequence are provided as a single polynucleotide (a single guide nucleic acid, e.g., a single gRNA).
[0132] Nucleic acid-inducible nucleases, such as the manipulated nucleic acid-inducible nucleases disclosed herein, when bound to a compatible gNA, such as a compatible gRNA, form a targetable nuclease complex, also referred herein as ribonucleoprotein (RNP) (in the case of gRNA), a complexed nucleic acid-inducible nuclease, which can bind to and cleave at or near the target sequence in a target polynucleotide determined by the guide sequence of the guide nucleic acid. The guide polynucleotide may be DNA. The guide polynucleotide may be RNA. The guide polynucleotide may contain both DNA and RNA. The guide polynucleotide may contain modified nucleotides or non-native nucleotides. If the guide polynucleotide contains RNA, the RNA guide polynucleotide may be encoded by a DNA sequence on a polynucleotide molecule such as a plasmid, linear construct, or edited cassette disclosed herein.
[0133] Generally, guide polynucleotides can form complexes with compatible nucleic acid-inducible nucleases and hybridize with target sequences, thereby directing the nucleases toward the target sequences. A target nucleic acid-inducible nuclease capable of forming complexes with a guide polynucleotide may be referred to as a guide polynucleotide-compatible nucleic acid-inducible nuclease. Furthermore, a guide polynucleotide capable of forming complexes with a nucleic acid-inducible nuclease may be referred to as a guide polynucleotide or guide nucleic acid-compatible nucleic acid.
[0134] gNA, e.g., gRNA, may be a natural gNA, e.g., a natural gRNA. In certain embodiments, gNA, e.g., gRNA, may be an engineered gNA, e.g., an engineered gRNA. When the term is used herein, “engineered guide nucleic acid,” e.g., “engineered gRNA,” may also be called a novel guide nucleic acid, e.g., a novel gRNA, and may include a non-natural guide nucleic acid, e.g., a non-natural gRNA, or an orthogonal gNA, e.g., an orthogonal gRNA.
[0135] Nucleic acid-induced nucleases can be fitted with guide nucleic acids not found in the nuclease's endogenous host. These orthogonal guide nucleic acids can be determined by empirical testing. These orthogonal guide nucleic acids may originate from different bacterial species, be synthesized, or be manipulated to be unnatural.
[0136] Orthogonal guide nucleic acids compatible with nucleic acid-induced nucleases may contain one or more common features. These common features may include sequences outside the pseudoknot region, the pseudoknot region, or the primary or secondary structure.
[0137] Targetable nucleic acid-inducible nuclease complexes are disclosed herein. When the term “targetable nucleic acid-inducible nuclease complex” is used herein, it may refer to a nucleic acid-inducible nuclease bound to a compatible gNA, the complex having the function of binding to or near a target sequence in a target polynucleotide and generating at least one strand break there. In certain embodiments, the targetable nucleic acid-inducible complex comprises a gNA which is a gRNA, and such complexes may be referred to as “ribonucleoprotein” or “RNP”. In certain embodiments, the targetable nucleic acid-inducible nuclease complex, e.g., RNP, is an engineered targetable nucleic acid-inducible nuclease complex, e.g., an engineered RNP. When these terms are used herein, “manipulated targetable nucleic acid-inducible nuclease complex,” for example, “manipulated RNP,” etc., may refer to a targetable nucleic acid-inducible nuclease complex, for example, an RNP, where the nucleic acid-inducible nuclease comprises a manipulated nucleic acid-inducible nuclease, the guide nucleic acid, for example, gRNA comprises a manipulated guide nucleic acid, for example, manipulated gRNA, or both. In certain embodiments, both the nuclease and gNA, for example, gRNA, are manipulated. In embodiments in which a manipulated nucleic acid-inducible nuclease is used, any suitable nucleic acid-inducible nuclease, such as the nucleic acid-inducible nucleases disclosed herein, may be used. In embodiments in which a manipulated gNA, for example, manipulated gRNA, is used, any suitable manipulated gNA, for example, manipulated gRNA, such as the manipulated gNA, for example, manipulated gRNA disclosed herein, may be used.
[0138] A targetable nucleic acid-inducible nuclease complex, e.g., RNP, can be produced by any preferred method known in the art. At one pole, both a nucleic acid-inducible nuclease and its compatible gNA, e.g., gRNA, are produced synthetically and subsequently bound to form a targetable nucleic acid-inducible nuclease complex, e.g., RNP. The complex can be introduced into a host cell by any preferred method, e.g., electroporation. At the other pole, a targetable nucleic acid-inducible nuclease complex, e.g., RNP, is produced in the host cell by transcription and / or translation of one or more polynucleotides introduced into the host cell, where one or more polynucleotides include portions encoding one or more components of the targetable nucleic acid-inducible nuclease complex, e.g., RNP, e.g., one or more portions encoding a nuclease, one or more portions encoding one or more gNA, e.g., gRNA, and one or more portions encoding one or more editing templates. Modulatory elements and others can be added to make the polynucleotide manipulable and to produce one or more vectors, as discussed herein and known in the art. One or more vectors are introduced into a host cell by any preferred method. Various components are produced by the cell and assemble within a targetable nucleic acid-inducible nuclease complex, e.g., RNP, within the cell. Any of these, or any variation between these two poles, are available and have been widely described in the art. See, for example, U.S. Patent No. 10,337,028. In certain embodiments, the targetable nucleic acid-inducible nuclease is produced within a first host cell by introducing a suitable one or more polynucleotides packaged in one or more suitable vectors into a first host cell that produces the nuclease, followed by extraction and purification to a preferred degree. It will be apparent that the various purification tags, cleavage sequences, FLAGs, and 3XFLAGs described herein are useful in assisting the isolation and purification of nucleases.In certain embodiments, one or more compatible gNAs, e.g., one or more compatible gRNAs, e.g., one or more modified nucleotides, e.g., split gNAs, e.g., split gRNAs, or single gNAs, e.g., single gRNAs (split gNAs in certain embodiments), comprising one or more chemically modified nucleotides, are synthesized as complete gNAs or gRNAs. The synthesized gNAs, e.g., gRNAs, may be introduced into a host cell, where, upon encountering a nuclease, the compatible gNAs, e.g., compatible gRNAs bind to a targetable nucleic acid-inducible nuclease complex, e.g., RNP. The synthesized gNAs, e.g., gRNAs, may be exposed to a suitable extracellular nuclease for a sufficient time to allow for the formation of a targetable nucleic acid-inducible nuclease complex, e.g., RNP, which is subsequently introduced into the host cell.
[0139] In certain embodiments, the sequence to be incorporated includes a transgene.
[0140] In certain embodiments, the compositions and methods disclosed herein utilize nucleic acid-inducible nucleases comprising an engineered nuclease polypeptide. As used herein, the term "engineered nuclease polypeptide" may also be referred herein to as an engineered sequence, and may refer to a nuclease polypeptide comprising a non-natural amino acid sequence, where the polypeptide functions as a nucleic acid-inducible nuclease either as is or in combination with a compatible gNA, such as gRNA, with additional processing. The non-natural amino acid sequence may be an amino acid sequence that differs from the natural sequence, for example, by the substitution and / or addition of non-natural amino acids in the sequence, for example, at the N-terminus, C-terminus, or a combination thereof.
[0141] As used herein, “engineered nucleic acid-inducible nuclease” is also referred to herein as Cas12a-like nuclease and is a non-natural nuclease. Nucleases can be non-natural for reasons including, but not limited to, the inclusion of engineered nuclease polypeptides.
[0142] In certain embodiments, the manipulated nuclease polypeptide may include a nuclease polypeptide having at least 60%, 65%, 75%, 85%, 90%, 95%, 99%, or 100% sequence identity, e.g., at least 85%, possibly at least 90%, possibly at least 95%, possibly at least 99%, or even 100% sequence identity, with any one amino acid sequence represented by SEQ ID NOs: 143-177 and 229, for example, with additional amino acids added at the amino terminus, carboxy terminus, or both. Such additions may include any preferred additions. Exemplary additions include one or more nuclear localization sequences (NLS), one or more purified tags, one or more cleavage sequences, one or more markers, one or more FLAG or 3XFLAG sequences, or a combination thereof, where each addition may occur at the amino terminus or carboxy terminus of the core amino acid sequence, as required. The term "at the amino terminus," as used herein, includes amino acid additions that are added before the amino terminus and are directly or indirectly linked to the amino terminus. The term "at the amino terminus," as used herein, also includes amino acid additions that are added after the amino terminus and are directly or indirectly linked to the amino terminus. One or more additional amino acids may be cleaved during the preparation and / or processing of the nuclease polypeptide.
[0143] When used herein, terms such as "nuclease polypeptide" may refer to a polypeptide having an amino acid sequence that functions as a nucleic acid-inducible nuclease either in its original form or, with additional processing, in combination with a compatible gNA, such as a compatible gRNA. The term "natural nuclease polypeptide," also referred to herein as a natural nuclease polypeptide sequence, may refer, when used herein, to a nuclease polypeptide found in nature, such as a nucleic acid-inducible nuclease found in prokaryotes.
[0144] As used herein, terms such as "original nuclease polynucleotide" may refer to the nuclease polypeptide from which the manipulated nuclease polypeptide is derived. In some cases, the original nuclease polypeptide may be a naturally occurring nuclease polypeptide.
[0145] As used herein, the term “manipulated nuclease polypeptide” may also be referred to herein as “manipulated sequence,” and may refer to a non-natural nuclease polypeptide. A nuclease polypeptide may be non-natural for any reason, including having an amino acid sequence different from known natural nuclease polypeptides, or having a nuclease polypeptide that includes the original nuclease polypeptide (which may be a natural nuclease polypeptide) to which one or more additional amino acid sequences are added at the N-terminus, C-terminus, or both.
[0146] The additional amino acid sequences that can be added to the original nuclease polypeptide can be any preferred amino acid sequence. In certain embodiments, the additional amino acid sequences may include one or more nuclear localization sequences (NLS), one or more purification tags, one or more cleavage sequences, one or more FLAGs or 3XFLAGs, and / or one or more markers. If multiple types of amino acid sequences, for example multiple NLSs, are used, each amino acid sequence may be the same as or different from the others, and each may be added to the N-terminus or C-terminus of the original nuclease polypeptide. In certain embodiments, the order and / or type of the additional amino acid sequences may be a specific order and / or type. The order is read from the N-terminus to the C-terminus.
[0147] NLS
[0148] In certain embodiments, the additional amino acid sequence may include one or more nuclear localization sequences (NLS), e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more NLS, which may be added to the original nuclease polypeptide. In some embodiments, the manipulated nuclease polypeptide includes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more NLS at the amino terminus of the original nuclease polypeptide, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more NLS at the carboxy terminus of the original nuclease polypeptide, or a combination thereof (e.g., one or more NLS at the amino terminus and one or more NLS at the carboxy terminus). In certain embodiments, the manipulated nuclease polypeptide includes 1 to 3, optionally 1 to 2, e.g., 1 NLS at the amino terminus. In certain embodiments, the manipulated nuclease polypeptide includes 3 to 5, optionally 3 to 4, e.g., 3 NLS at the carboxy terminus. When multiple NLSs exist, each can be selected independently of others so that a single NLS may exist in multiple copies and / or in combination with one or more other NLSs present in two or more copies. In certain embodiments, four NLSs are added to the original nuclease polypeptide. In some of these embodiments, one NLS is at the N-terminus and three are at the C-terminus. In certain embodiments, the manipulated nuclease polypeptide provided herein includes at least one myc-associated NLS containing the sequence PAAKKKKLD (SEQ ID NO: 279) or an amino acid sequence having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity thereto, and in certain embodiments, the myc-associated NLS is at the N-terminus of the original nuclease polypeptide. In certain embodiments, the manipulated nuclease polypeptide provided herein comprises at least one nucleoplasmin NLS having the sequence KRPAATKKAGQAKKKK (SEQ ID NO: 264) or an amino acid sequence having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity thereto, wherein in certain embodiments, the nucleoplasmin NLS is C-terminal.In certain embodiments, the manipulated nuclease polypeptide provided herein comprises at least one or at least two SV40NLS sequences having sequence PKKKRKV (SEQ ID NO: 263) or an amino acid sequence having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity therewith, wherein in certain embodiments the SV40NLS is C-terminal. In certain embodiments, the manipulated nuclease polypeptide provided herein comprises one NLS at the N-terminus and three NLS at the C-terminus, for example, one myc-associated NLS at the N-terminus and one nucleoplasmin NLS and two SV40NLS at the C-terminus. In certain embodiments, the manipulated nuclease polypeptide provided herein comprises one myc-associated NLS having the sequence PAAKKKKLD (SEQ ID NO: 279) or an amino acid sequence having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity with it at its N-terminus; one nucleoplasmin NLS containing the sequence KRPAATKKAGQAKKKK (SEQ ID NO: 264) or an amino acid sequence having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity with it at its C-terminus; and two SV40 NLS containing the sequence PKKKRKV (SEQ ID NO: 263) or an amino acid sequence having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity with it. Generally, one or more NLSs are adjacent to the original nuclease polypeptide on its N-side, C-side, or both. For example, in an embodiment having four NLSs, the order could be NLS-original nuclease polypeptide-NLS-NLS-NLS. Therefore, if some or all of the other tags are removed, for example, by cleavage in the cleavage sequence, the remaining portion of the manipulated nuclease polypeptide retains the NLSs.
[0149] Non-limiting examples of NLS include NLS of the SV40 virus large T antigen having the amino acid sequence PKKKRKV (SEQ ID NO: 263) or an amino acid sequence having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity with it; and NLS derived from nucleoplasm (e.g., sequence KRPAATKKAGQAKKKK (SEQ ID NO: 264) or at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity with it). A single nucleoplasmin bifid NLS containing an amino acid sequence with sequence identity; an amino acid sequence PAAKRVKLD (SEQ ID NO: 265) or an amino acid sequence having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity with it; or a c-myc NLS having an amino acid sequence RQRRNELKRSP (SEQ ID NO: 266) or an amino acid sequence having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity with it; or an hRNPA1 M9 having an amino acid sequence NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 267) or an amino acid sequence having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity with it. NLS; the sequence of the IBB domain derived from importin-alpha RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 268) or an amino acid sequence having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity thereto; the sequence of the myoma T protein VSRKRPRP (SEQ ID NO: 269) or at least 50, 60, 70, 75, 80, 85, 90 Amino acid sequences having 95% or 98% sequence identity, and PPKKARED (SEQ ID NO: 270) or amino acid sequences having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity with it; human p53 sequence PQPKKKPL (SEQ ID NO: 271) or amino acid sequences having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity with it; mouse c-abl IV sequence SALIKKKKKMAP (SEQ ID NO: 272) or sequences having at least 50, 60, 70, 75, 80, 85, 90, 95,or an amino acid sequence having 98% sequence identity; the influenza virus NS1 sequence DRLRR (SEQ ID NO: 273) or an amino acid sequence having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity with it, and PKQKKRK (SEQ ID NO: 274) or an amino acid sequence having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity with it; the hepatitis virus delta antigen sequence RKLKKKIKKL (SEQ ID NO: 275) or an amino acid sequence having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity with it; the mouse Mx1 protein sequence REKKKFLKRR (SEQ ID NO: 275) Examples include amino acid sequences having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity with sequence number 276; the human poly(ADP-ribose) polymerase sequence KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 277) or amino acid sequences having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity with it; and NLS sequences derived from the steroid hormone receptor (human) glucocorticoid sequence RKCLQAGMNLEARKTKK (SEQ ID NO: 278) or amino acid sequences having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity with it. In certain embodiments, the manipulated nuclease polypeptide disclosed herein comprises an original nuclease polypeptide having at least 50, 60%, 65%, 75%, 85%, 95%, 99%, or 100% identity with one or more amino acid sequences from SEQ ID NOs: 143-177 or 229, for example, at least 60%, possibly at least 85%, and in certain embodiments at least 95%, or even 100% identity with an amino acid sequence, and at least one myc-associated NLS comprising the sequence PAAKKKKLD (SEQ ID NO: 279) or an amino acid sequence having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity with it, wherein in certain embodiments the myc-associated NLS is at the N-terminus. Additionally or alternatively, the manipulated nuclease polypeptide comprises:The manipulated nuclease polypeptide may contain at least one nucleoplasmin NLS having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity with the sequence KRPAATKKAGQAKKKK (SEQ ID NO: 264), in which case the nucleoplasmin NLS is C-terminal. Additionally or alternatively, the manipulated nuclease polypeptide may contain at least one, or at least two, for example one, or in which case two SV40NLS sequences having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity with the sequence PKKKRKV (SEQ ID NO: 263), in which case the SV40NLS is C-terminal. In certain embodiments, the manipulated nuclease polypeptides disclosed herein include an original nuclease polypeptide having an amino acid sequence that is at least 50, 60%, 65%, 75%, 85%, 95%, 99%, or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 143-177 or 229, for example, at least 60%, optionally at least 85%, and in certain embodiments at least 95%, or even 100% identical, and comprising one NLS at the N-terminus and three NLSs at the C-terminus, for example, one myc-associated NLS as described above at the N-terminus, and one nucleoplasmin NLS and two SV40 NLS as described above at the C-terminus. In certain embodiments, the manipulated nuclease polypeptide disclosed herein comprises an original nuclease polypeptide having an amino acid sequence that is at least 50, 60%, 65%, 75%, 85%, 95%, 99%, or 100% identical to any one of SEQ ID NOs: 143-177 or 229, for example, at least 60%, possibly at least 85%, and in certain embodiments at least 95%, or even 100% identical, and at the N-terminus, a myc-associated NLS comprising the sequence PAAKKKKLD (SEQ ID NO: 279) or an amino acid sequence having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity thereto, and at the C-terminus, the sequence KRPAATKKAGQAKKKK (SEQ ID NO: 264) or at least 50, 60,It comprises one nucleoplasmin NLS containing an amino acid sequence having 70, 75, 80, 85, 90, 95, or 98% sequence identity, and two SV40NLS containing the sequence PKKKRKV (SEQ ID NO: 263) or an amino acid sequence having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity with it.
[0150] In certain embodiments, the manipulated nuclease polypeptide disclosed herein comprises an original nuclease polypeptide that is at least 50, 60%, 65%, 75%, 85%, 95%, 99%, or 100% identical to the amino acid sequence of any of SEQ ID NOs: 144, 153, and 229 (specifically SEQ ID NOs: 144, specifically SEQ ID NOs: 153, and specifically SEQ ID NOs: 229), e.g., at least 60%, specifically at least 85%, and in certain embodiments at least 95%, or even 100% identical, and at least one myc-associated NLS comprising the sequence PAAKKKKLD (SEQ ID NOs: 279) or an amino acid sequence having at least 50, 60, 70, 75, 80, 85, 90, 95%, or 98% sequence identity thereto, wherein in certain embodiments the myc-associated NLS is N-. Additionally or alternatively, the manipulated nuclease polypeptide may contain at least one nucleoplasmin NLS having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity with the sequence KRPAATKKAGQAKKKK (SEQ ID NO: 264), wherein in certain embodiments the nucleoplasmin NLS is C-terminal. Additionally or alternatively, the manipulated nuclease polypeptide may contain at least one, or at least two, for example one, or two in certain embodiments, SV40NLS sequences having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity with the sequence PKKKRKV (SEQ ID NO: 263), wherein in certain embodiments the SV40NLS is C-terminal.In certain embodiments, the manipulated nuclease polypeptide disclosed herein comprises an original nuclease polypeptide having an amino acid sequence that is at least 50, 60%, 65%, 75%, 85%, 95%, 99%, or 100% identical to the amino acid sequence of any of SEQ ID NOs: 144, 153, and 229 (specifically SEQ ID NOs: 144, specifically SEQ ID NOs: 153, and specifically SEQ ID NOs: 229), for example, at least 60%, specifically at least 85%, and in certain embodiments at least 95%, or even 100% identical, and comprising one NLS at the N-terminus and three NLSs at the C-terminus, for example, one myc-associated NLS as described above at the N-terminus, and one nucleoplasmin NLS and two SV40 NLS as described above at the C-terminus. In certain embodiments, the manipulated nuclease polypeptide disclosed herein is an original nuclease polypeptide having an amino acid sequence that is at least 50, 60%, 65%, 75%, 85%, 95%, 99%, or 100% identical to the amino acid sequence of any of SEQ ID NOs: 144, 153, and 229 (specifically SEQ ID NOs: 144, specifically SEQ ID NOs: 153, and specifically SEQ ID NOs: 229), for example at least 60%, specifically at least 85%, and in certain embodiments at least 95%, or even 100% identical, and at the N-terminus, the sequence PAAKKKKLD (SEQ ID NOs: 279) or The present invention comprises one myc-related NLS containing an amino acid sequence having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity, one nucleoplasmin NLS containing the sequence KRPAATKKAGQAKKKK (SEQ ID NO: 264) or an amino acid sequence having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity with it at its C-terminus, and two SV40 NLS containing the sequence PKKKRKV (SEQ ID NO: 263) or an amino acid sequence having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity with it.
[0151] In certain embodiments, the manipulated nuclease polypeptide disclosed herein comprises an original nuclease polypeptide that is at least 50, 60%, 65%, 75%, 85%, 95%, 99%, or 100% identical to the amino acid sequence of SEQ ID NO: 153, for example, at least 60%, possibly at least 85%, and in certain embodiments at least 95%, or even 100% identical, and at least one myc-associated NLS comprising the sequence PAAKKKKLD (SEQ ID NO: 279) or an amino acid sequence having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity thereto, wherein in certain embodiments the myc-associated NLS is at the N-terminus. Additionally or alternatively, the manipulated nuclease polypeptide may contain at least one nucleoplasmin NLS having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity with the sequence KRPAATKKAGQAKKKK (SEQ ID NO: 264), wherein in certain embodiments the nucleoplasmin NLS is C-terminal. Additionally or alternatively, the manipulated nuclease polypeptide may contain at least one, or at least two, for example one, or two in certain embodiments, SV40NLS sequences having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity with the sequence PKKKRKV (SEQ ID NO: 263), wherein in certain embodiments the SV40NLS is C-terminal. In certain embodiments, the manipulated nuclease polypeptide disclosed herein comprises an original nuclease polypeptide having an amino acid sequence that is at least 50, 60%, 65%, 75%, 85%, 95%, 99%, or 100% identical to the amino acid sequence of SEQ ID NO: 153, for example, at least 60%, optionally at least 85%, and in certain embodiments at least 95%, or even 100% identical, and comprising one NLS at the N-terminus and three NLSs at the C-terminus, for example, one myc-associated NLS as described above at the N-terminus, and one nucleoplasmin NLS and two SV40 NLS as described above at the C-terminus.In certain embodiments, the manipulated nuclease polypeptide disclosed herein comprises an original nuclease polypeptide having an amino acid sequence that is at least 50, 60%, 65%, 75%, 85%, 95%, 99%, or 100% identical to the amino acid sequence of SEQ ID NO: 153, for example at least 60%, possibly at least 85%, and in certain embodiments at least 95%, or even 100% identical, and at the N-terminus, the sequence PAAKKKKLD (SEQ ID NO: 279) or at least 50, 60, 70, 75, 80, 85, 90, 95, or It comprises at least one myc-related NLS containing an amino acid sequence having 98% sequence identity, one nucleoplasmin NLS containing the sequence KRPAATKKAGQAKKKK (SEQ ID NO: 264) or an amino acid sequence having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity with it at its C-terminus, and two SV40 NLS containing the sequence PKKKRKV (SEQ ID NO: 263) or an amino acid sequence having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity with it.
[0152] In certain embodiments, the manipulated nuclease polypeptide disclosed herein comprises an original nuclease polypeptide that is at least 50, 60%, 65%, 75%, 85%, 95%, 99%, or 100% identical to the amino acid sequence of SEQ ID NO: 229, for example, at least 60%, possibly at least 85%, and in certain embodiments at least 95%, or even 100% identical, and at least one myc-associated NLS comprising the sequence PAAKKKKLD (SEQ ID NO: 279) or an amino acid sequence having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity thereto, wherein in certain embodiments the myc-associated NLS is N-. Additionally or alternatively, the manipulated nuclease polypeptide may contain at least one nucleoplasmin NLS having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity with the sequence KRPAATKKAGQAKKKK (SEQ ID NO: 264), wherein in certain embodiments the nucleoplasmin NLS is C-terminal. Additionally or alternatively, the manipulated nuclease polypeptide may contain at least one, or at least two, for example one, or two in certain embodiments, SV40NLS sequences having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity with the sequence PKKKRKV (SEQ ID NO: 263), wherein in certain embodiments the SV40NLS is C-terminal. In certain embodiments, the manipulated nuclease polypeptide disclosed herein comprises an original nuclease polypeptide having an amino acid sequence that is at least 50, 60%, 65%, 75%, 85%, 95%, 99%, or 100% identical to the amino acid sequence of SEQ ID NO: 229, for example, at least 60%, optionally at least 85%, and in certain embodiments at least 95%, or even 100% identical, and comprising one NLS at the N-terminus and three NLSs at the C-terminus, for example, one myc-associated NLS as described above at the N-terminus, and one nucleoplasmin NLS and two SV40 NLS as described above at the C-terminus.In certain embodiments, the manipulated nuclease polypeptide disclosed herein comprises an original nuclease polypeptide having an amino acid sequence that is at least 50, 60%, 65%, 75%, 85%, 95%, 99%, or 100% identical to the amino acid sequence of SEQ ID NO: 229, for example at least 60%, possibly at least 85%, and in certain embodiments at least 95%, or even 100% identical, and at the N-terminus, the sequence PAAKKKKLD (SEQ ID NO: 279) or at least 50, 60, 70, 75, 80, 85, 90, 95, or It comprises at least one myc-related NLS containing an amino acid sequence having 98% sequence identity, one nucleoplasmin NLS containing the sequence KRPAATKKAGQAKKKK (SEQ ID NO: 264) or an amino acid sequence having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity with it at its C-terminus, and two SV40 NLS containing the sequence PKKKRKV (SEQ ID NO: 263) or an amino acid sequence having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity with it. b) Refined tags
[0153] In addition to comprising one or more NLSs and / or other additional amino acid sequences described herein, or instead, the manipulated nuclease polypeptides disclosed herein may include one or more purification tags, which may be at the N-terminus or C-terminus of the original nuclease polypeptide. Any preferred purification tag(s) may be used. Exemplary purification tags include poly-his tags that can contain gly at their N-terminus, such as the Gly-6xHis tag (SEQ ID NO: 332) or the Gly-8xHis tag (SEQ ID NO: 333). Further exemplary purification tags include hemagglutinin (HA), c-myc, T7, and Glu-Glu; maltose-binding protein (mbp); N-terminal glutathione S-transferase (GST); and calmodulin-binding peptide (CBP). In certain embodiments, the manipulated nuclease polypeptide may, in addition to the one or more NLSs described above, include, for example, the Gly-6xhis tag at its N-terminus. In certain embodiments, the manipulated nuclease polypeptide may, for example, include the Gly-8xhis tag at its N-terminus. Generally, when a Gly-polyhis tag is used, it is the most N-terminal sequence to which it is attached.
[0154] In certain embodiments, the manipulated nuclease polypeptides disclosed herein may contain a poly-his tag, or a gly-polyhis tag such as a Gly-6xHis tag or a Gly-8xHis tag, for example at the N-terminus. These Gly-6xHis or Gly-8xHis tags are applicable for several reasons, including: 1) the 6xHis or 8xHis tag can be used in protein purification to enable binding to a chromatography column for purification; and 2) the N-terminal glycine further enables site-specific chemical modification, allowing for more advanced protein manipulation. Furthermore, the Gly-6xHis or Gly-8xHis are designed to be readily removed by digestion with tobacco ecchi disease virus (TEV) protease, if necessary. The Gly-6xHis or Gly-8xHis tag can be positioned at the N-terminus. The Gly-6xHis tag is described in more detail in Martos-Maldonado et al., Nat Commun. (2018) 17;9(1):3307, the disclosure of which is incorporated herein.
[0155] In certain embodiments, the manipulated nuclease polypeptide disclosed herein comprises an original nuclease polypeptide that is at least 50, 60%, 65%, 75%, 85%, 95%, 99%, or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 143-177 or 229, or any one of the amino acid sequences of SEQ ID NOs: 144, 153, and 229 (specifically SEQ ID NOs: 144, specifically SEQ ID NOs: 153, specifically SEQ ID NOs: 229), for example at least 60%, specifically at least 85%, and in certain embodiments at least 95%, or even 100% identical, and a Gly-poly-His tag at the N-terminus, for example, a Gly-6xHis tag (SEQ ID NOs: 332) or a Gly-8xHis tag (SEQ ID NOs: 333). In certain embodiments, the tag is a Gly-6xHis tag (SEQ ID NOs: 332). In certain embodiments, the tag is a Gly-8xHis tag (SEQ ID NOs: 333). In addition, or alternatively, the manipulated nuclease polypeptide may contain a FLAG (SEQ ID NO: 281) or 3XFLAG (SEQ ID NO: 280) at either the carboxyl or amino terminus. In certain embodiments, the tag is an amino-terminated 3XFLAG (SEQ ID NO: 280). If a gly-polyhis tag is also present, the 3XFLAG may be located inside the gly-polyhis tag. In certain embodiments, the tag is a carboxyl-terminated 3XFLAG (SEQ ID NO: 280).
[0156] In certain embodiments, the manipulated nuclease polypeptide disclosed herein comprises an original nuclease polypeptide that is at least 50, 60%, 65%, 75%, 85%, 95%, 99%, or 100% identical to the amino acid sequence of any of SEQ ID NOs: 144, 153, and 229 (specifically SEQ ID NOs: 144, specifically SEQ ID NOs: 153, and specifically SEQ ID NOs: 229), for example at least 60%, specifically at least 85%, and in certain embodiments at least 95%, or even 100% identical, and a Gly-poly-His tag at the N-terminus, for example, a Gly-6xHis tag (SEQ ID NOs: 332) or a Gly-8xHis tag (SEQ ID NOs: 333). In certain embodiments, the tag is a Gly-6xHis tag (SEQ ID NOs: 332). In certain embodiments, the tag is a Gly-8xHis tag (SEQ ID NOs: 333). In addition, or instead, the manipulated nuclease polypeptide may contain a FLAG (SEQ ID NO: 281) or 3XFLAG (SEQ ID NO: 280) at either the carboxyl or amino terminus. In certain embodiments, the tag is an amino-terminated 3XFLAG (SEQ ID NO: 280). If a gly-polyhis tag is also present, the 3XFLAG may be located inside the gly-polyhis tag. In certain embodiments, the tag is an amino-terminated 3XFLAG (SEQ ID NO: 280).
[0157] In certain embodiments, the manipulated nuclease polypeptide disclosed herein comprises the original nuclease polypeptide having at least 50, 60%, 65%, 75%, 85%, 95%, 99%, or 100% identity with the amino acid sequence of SEQ ID NO: 153, for example, at least 60%, possibly at least 85%, and in certain embodiments at least 95%, or even 100% identity, and a Gly-poly-His tag at the N-terminus, for example, a Gly-6xHis tag (SEQ ID NO: 332) or a Gly-8xHis tag (SEQ ID NO: 333). In certain embodiments, the tag is a Gly-6xHis tag (SEQ ID NO: 332). In certain embodiments, the tag is a Gly-8xHis tag (SEQ ID NO: 333). In addition, or instead, the manipulated nuclease polypeptide may contain a FLAG (SEQ ID NO: 281) or a 3XFLAG (SEQ ID NO: 280) at either the carboxyl terminus or the amino terminus. In certain embodiments, the tag is an amino-terminated 3XFLAG (SEQ ID NO: 280). If a gly-polyhis tag is also present, the 3XFLAG may be located inside the gly-polyhis tag. In certain embodiments, the tag is a carboxy-terminated 3XFLAG (SEQ ID NO: 280).
[0158] In certain embodiments, the manipulated nuclease polypeptide disclosed herein comprises the original nuclease polypeptide having at least 50, 60%, 65%, 75%, 85%, 95%, 99%, or 100% identity with the amino acid sequence of SEQ ID NO: 229, for example, at least 60%, possibly at least 85%, and in certain embodiments at least 95%, or even 100% identity, and a Gly-poly-His tag at the N-terminus, for example, a Gly-6xHis tag (SEQ ID NO: 332) or a Gly-8xHis tag (SEQ ID NO: 333). In certain embodiments, the tag is a Gly-6xHis tag (SEQ ID NO: 332). In certain embodiments, the tag is a Gly-8xHis tag (SEQ ID NO: 333). FLAG or 3XFLAG
[0159] In addition, or alternatively, the manipulated nuclease polypeptide may contain a FLAG (SEQ ID NO: 281) or 3XFLAG (SEQ ID NO: 280) at either the carboxyl or amino terminus. In certain embodiments, the tag is an amino-terminated 3XFLAG (SEQ ID NO: 280). If a gly-polyhis tag is also present, the 3XFLAG may be located inside the gly-polyhis tag. In certain embodiments, the tag is a carboxyl-terminated 3XFLAG (SEQ ID NO: 280). Cutting sequence
[0160] In addition to comprising one or more NLSs, purified tags, and / or other additional amino acid sequences described herein, or instead, the manipulated nuclease polypeptides disclosed herein may include one or more cleavage sequences, which may be at the N-terminus or C-terminus. Any preferred cleavage sequence may be used, and if multiple cleavage sequences are used, they may be identical or different. In certain embodiments, the cleavage sequence includes a tobacco eczema virus protease cleavage sequence, which is referred herein to as the "TEV sequence" (SEQ ID NO: 331). The TEV sequence may be at the amino terminus. Generally, the cleavage sequence, e.g., the TEV sequence, is positioned such that the cleavage at the cleavage sequence leaves other additional amino acid sequences, in particular any NLS attached to the original nuclease polypeptide, intact. combination
[0161] This specification discloses engineered nuclease polypeptides comprising two or more additional amino acid sequences added to the original nuclease polypeptide. In addition to such engineered nuclease polynucleotides disclosed above, additional engineered nuclease polynucleotides may include:
[0162] In certain embodiments, the manipulated nuclease polypeptide or its active portion disclosed herein is (i) Refinement tag; (ii) cleavage site; (iii) NLS; (iv) the original nuclease polypeptide; and (v) Three NLS Contains ingredients that include [specific ingredients].
[0163] In certain embodiments, the manipulated nuclease polypeptide further comprises (vi)3XFLAG at the C-terminus. In certain embodiments, the manipulated polypeptide further comprises v(3x)FLAG at the N-terminus. In certain embodiments, the N-terminal 3XFLAG is located between the purification tag and the cleavage site. In certain embodiments, the components are in order, i.e., from the amino terminus to the carboxyl terminus of the nuclease polypeptide. In certain embodiments, the purification tag comprises a Gly-polyhis tag such as a Gly-6xHis or Gly-8xHis tag. In certain embodiments, the purification tag comprises a Gly-6xHis tag. In certain embodiments, the cleavage site comprises a TEV. In certain embodiments, the N-terminal NLS includes a myc-related NLS, such as the amino acid sequence PAAKRVKLD (SEQ ID NO: 265) or an amino acid sequence having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity thereto, or a c-myc NLS having the amino acid sequence RQRRNELKRSP (SEQ ID NO: 266) or an amino acid sequence having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity thereto. In certain embodiments, the original nuclease polypeptide has an amino acid sequence that is at least 50, 60%, 65%, 75%, 85%, 95%, 99%, or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 143-177 and 229, or any one of the amino acid sequences of SEQ ID NOs: 144, 153, and 229 (specifically SEQ ID NOs: 144, 153, and 229), for example, at least 60%, specifically at least 85%, and in certain embodiments at least 95%, or even 100% identical. In certain embodiments, the original nuclease polypeptide does not contain the peptide motif of SEQ ID NO: 224.In certain embodiments, the original nuclease polypeptide has an amino acid sequence that is at least 50, 60%, 65%, 75%, 85%, 95%, 99%, or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 143-151, 161-163, 165, 166, 169, 171-175, 177, or 229, for example at least 60%, possibly at least 85%, and in certain embodiments at least 95%, or even 100% identical. In certain embodiments, the original nuclease polypeptide has an amino acid sequence that is at least 50, 60%, 65%, 75%, 85%, 95%, 99%, or 100% identical to the amino acid sequence of SEQ ID NO: 144, 153, or 229, for example, at least 60%, possibly at least 85%, and in certain embodiments at least 95%, or even 100% identical. In certain embodiments, the original nuclease polypeptide has an amino acid sequence that is at least 50, 60%, 65%, 75%, 85%, 95%, 99%, or 100% identical to the amino acid sequence of SEQ ID NO: 153, for example, at least 60%, possibly at least 85%, and in certain embodiments at least 95%, or even 100% identical.In certain embodiments, the original nuclease polypeptide has an amino acid sequence that is at least 50, 60%, 65%, 75%, 85%, 95%, 99%, or 100% identical to the amino acid sequence of SEQ ID NO: 229, for example, at least 60%, possibly at least 85%, and in certain embodiments at least 95%, or even 100% identical. In a particular embodiment, the three N-terminal NLSs include, at their C-terminuses, one nucleoplasmin NLS, for example, the sequence KRPAATKKAGQAKKKK (SEQ ID NO: 264) or an amino acid sequence having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity with it, and two SV40NLSs, for example, the sequence PKKKRKV (SEQ ID NO: 263) or an amino acid sequence having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity with it. In certain embodiments, the manipulated nuclease polypeptide comprises an amino acid sequence that is at least 50, 60%, 65%, 75%, 85%, 95%, 99%, or 100% identical to the amino acid sequence of SEQ ID NO: 257, for example, at least 60%, optionally at least 85%, and in certain embodiments at least 95%, or even 100% identical. In certain embodiments, the manipulated nuclease polypeptide comprises an amino acid sequence that is at least 50, 60%, 65%, 75%, 85%, 95%, 99%, or 100% identical to the amino acid sequence of SEQ ID NO: 260, for example, at least 60%, optionally at least 85%, and in certain embodiments at least 95%, or even 100% identical. In certain embodiments, the manipulated nuclease polypeptide comprises an amino acid sequence that is at least 50, 60%, 65%, 75%, 85%, 95%, 99%, or 100% identical to the amino acid sequence of SEQ ID NO: 261, for example, at least 60%, possibly at least 85%, and in certain embodiments at least 95%, or even 100% identical. In certain embodiments, the manipulated nuclease polynucleotide further comprises 3XFLAG (SEQ ID NO: 280).In certain embodiments, the manipulated nuclease polypeptide comprises an amino acid sequence that is at least 50, 60%, 65%, 75%, 85%, 95%, 99%, or 100% identical to the amino acid sequence of SEQ ID NO: 258, for example, at least 60%, possibly at least 85%, and in certain embodiments at least 95%, or even 100% identical.
[0164] In certain embodiments, the manipulated nuclease polypeptide or its active portion disclosed herein is (i) Refinement tag; (ii) cleavage site; (iii) NLS; (iv) Original nuclease polypeptide; (v) Three NLS Includes.
[0165] In certain embodiments, the manipulated nuclease polypeptide further comprises (vi)3XFLAG at the C-terminus. In certain embodiments, the manipulated polypeptide further comprises (vi)(3x)FLAG at the N-terminus. In certain embodiments, the N-terminal 3XFLAG is located between the purification tag and the cleavage site. In certain embodiments, the components are in order, i.e., from the amino terminus to the carboxyl terminus of the nuclease polypeptide. In certain embodiments, the purification tag comprises a Gly-polyhis tag such as a Gly-6xHis or Gly-8xHis tag. In certain embodiments, the purification tag comprises a Gly-8xHis tag. In certain embodiments, the cleavage site comprises a TEV. In certain embodiments, the N-terminal NLS includes a myc-related NLS, such as the amino acid sequence PAAKRVKLD (SEQ ID NO: 265) or an amino acid sequence having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity thereto, or a c-myc NLS having the amino acid sequence RQRRNELKRSP (SEQ ID NO: 266) or an amino acid sequence having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity thereto. In certain embodiments, the original nuclease polypeptide has an amino acid sequence that is at least 50, 60%, 65%, 75%, 85%, 95%, 99%, or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 143-177 and 229, or any one of the amino acid sequences of SEQ ID NOs: 144, 153, and 229 (specifically SEQ ID NOs: 144, 153, and 229), for example, at least 60%, specifically at least 85%, and in certain embodiments at least 95%, or even 100% identical. In certain embodiments, the original nuclease polypeptide does not contain the peptide motif of SEQ ID NO: 224.In certain embodiments, the original nuclease polypeptide has an amino acid sequence that is at least 50, 60%, 65%, 75%, 85%, 95%, 99%, or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 143-151, 161-163, 165, 166, 169, 171-175, 177, or 229, for example at least 60%, possibly at least 85%, and in certain embodiments at least 95%, or even 100% identical. In certain embodiments, the original nuclease polypeptide has an amino acid sequence that is at least 50, 60%, 65%, 75%, 85%, 95%, 99%, or 100% identical to the amino acid sequence of SEQ ID NO: 144, 153, or 229, for example, at least 60%, possibly at least 85%, and in certain embodiments at least 95%, or even 100% identical. In certain embodiments, the original nuclease polypeptide has an amino acid sequence that is at least 50, 60%, 65%, 75%, 85%, 95%, 99%, or 100% identical to the amino acid sequence of SEQ ID NO: 153, for example, at least 60%, possibly at least 85%, and in certain embodiments at least 95%, or even 100% identical.In certain embodiments, the original nuclease polypeptide has an amino acid sequence that is at least 50, 60%, 65%, 75%, 85%, 95%, 99%, or 100% identical to the amino acid sequence of SEQ ID NO: 229, for example, at least 60%, possibly at least 85%, and in certain embodiments at least 95%, or even 100% identical. In a particular embodiment, the three N-terminal NLSs include, at their C-terminuses, one nucleoplasmin NLS, for example, the sequence KRPAATKKAGQAKKKK (SEQ ID NO: 264) or an amino acid sequence having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity with it, and two SV40NLSs, for example, the sequence PKKKRKV (SEQ ID NO: 263) or an amino acid sequence having at least 50, 60, 70, 75, 80, 85, 90, 95, or 98% sequence identity with it. In certain embodiments, the manipulated nuclease polypeptide comprises an amino acid sequence that is at least 50, 60%, 65%, 75%, 85%, 95%, 99%, or 100% identical to the amino acid sequence of SEQ ID NO: 259, for example, at least 60%, optionally at least 85%, and in certain embodiments at least 95%, or even 100% identical. In certain embodiments, the manipulated nuclease polypeptide comprises an amino acid sequence that is at least 50, 60%, 65%, 75%, 85%, 95%, 99%, or 100% identical to the amino acid sequence of SEQ ID NO: 262, for example, at least 60%, optionally at least 85%, and in certain embodiments at least 95%, or even 100% identical.
[0166] Nucleic acid-inducible nucleases may be encoded by one or more polynucleotides. One or more polynucleotides may be natural. In certain embodiments, one or more polynucleotides include engineered polynucleotides. An engineered polynucleotide is a non-natural polynucleotide, e.g., a polynucleotide encoding an engineered nucleic acid-inducible nuclease, wherein the encoded amino acid sequence is modified from the natural sequence by one or more substitutions in the natural nuclease polypeptide, the addition of one or more amino acid sequences to the C and / or N terminus of the nuclease polypeptide, or both. Engineered polynucleotides may additionally or alternatively be produced by codon optimization, for example, a natural polynucleotide in one species being optimized for transcription and / or translation in the other species, where at least 1, 2, 5, 10, 20, 50, 100, 200, or 500 codons in the polynucleotide differ between the two. Engineered polynucleotides encoding nucleic acid-inducible nucleases, such as nucleic acid-inducible nucleases disclosed herein, can be codon-optimized for prokaryotes, such as E. coli. Engineered polynucleotides encoding nucleic acid-inducible nucleases, such as nucleic acid-inducible nucleases disclosed herein, can be codon-optimized for unicellular eukaryotes, such as yeasts, such as S. cerivisae. Engineered polynucleotides encoding nucleic acid-inducible nucleases, such as nucleic acid-inducible nucleases disclosed herein, can be codon-optimized for multicellular eukaryotes, such as humans.
[0167] Polynucleotides encoding nucleic acid-inducible nucleases are disclosed herein. In certain embodiments, the polynucleotides are natural. In certain embodiments, the polynucleotides are manipulated, for example, so that the encoded nuclease polypeptide comprises a manipulated nuclease polypeptide, or the polynucleotide is codon-optimized, or both.
[0168] In certain embodiments, one or more polynucleotides are provided that encode one or more amino acid sequences corresponding to any one of SEQ ID NOs: 143-177 and 229, or at least 50, 60, 70, 80, 90, 95, or 100% identical amino acid sequences to one or more amino acid sequences corresponding to any one of SEQ ID NOs: 143-177 and 229. In certain embodiments, the encoded polypeptide does not contain the peptide motif YLFQIYNKDF (SEQ ID NO: 224). Accordingly, in certain embodiments, polynucleotides are provided that encode one or more amino acid sequences corresponding to any one of the amino acid sequences of SEQ ID NOs: 143-151, 161-163, 165, 166, 169, 171-175, 177, and 229, or an amino acid sequence that is at least 50, 60, 70, 80, 90, 95, or 100% identical to one or more amino acid sequences corresponding to any one of the amino acid sequences of SEQ ID NOs: 143-151, 161-163, 165, 166, 169, 171-175, 177, and 229. In certain embodiments, the encoded polypeptide, which does not contain the peptide motif YLFQIYNKDF (SEQ ID NO: 224), comprises at least one amino acid substitution that is a radical amino acid substitution and / or has a Snees index value greater than 20. In certain embodiments, a polynucleotide is provided that codes for one or more amino acid sequences corresponding to any one of sequence numbers 149, 151, 175, and 177, or a polynucleotide that codes for at least 50, 60, 70, 80, 90, 95, or 100% identical amino acid sequences to one or more amino acid sequences corresponding to any one of sequence numbers 149, 151, 175, and 177. In certain embodiments, a polynucleotide is provided that codes for one or more amino acid sequences corresponding to any one of sequence numbers 144, 153, and 229, or a polynucleotide that codes for at least 50, 60, 70, 80, 90, 95, or 100% identical amino acid sequences to one or more amino acid sequences corresponding to any one of sequence numbers 144, 153, and 229.In certain embodiments, a polynucleotide is provided that encodes an amino acid sequence corresponding to SEQ ID NO: 144, or an amino acid sequence that is at least 50, 60, 70, 80, 90, 95, or 100% identical to the amino acid sequence corresponding to SEQ ID NO: 144. In certain embodiments, a polynucleotide is provided that encodes an amino acid sequence corresponding to SEQ ID NO: 153, or an amino acid sequence that is at least 50, 60, 70, 80, 90, 95, or 100% identical to the amino acid sequence corresponding to SEQ ID NO: 153. In certain embodiments, a polynucleotide is provided that encodes an amino acid sequence corresponding to SEQ ID NO: 229, or an amino acid sequence that is at least 50, 60, 70, 80, 90, 95, or 100% identical to the amino acid sequence corresponding to SEQ ID NO: 229. In certain embodiments, the sequence encodes one or more additional amino acid sequences at either the N-terminus, C-terminus, or both of the polypeptide encoded by the polynucleotide. The type, combination, N-terminus or C-terminus, and / or order of the additional amino acid sequences may be any of those disclosed herein. In certain embodiments of the latter, a polynucleotide is provided that encodes an amino acid sequence corresponding to SEQ ID NOs. 257-262, or an amino acid sequence that is at least 50, 60, 70, 80, 90, 95, or 100% identical to the amino acid sequences corresponding to SEQ ID NOs. 257-262. In certain embodiments of the above, the polynucleotide is an engineered polynucleotide.
[0169] In certain embodiments, one or more polynucleotides containing sequences corresponding to any one of SEQ ID NOs: 1-142 and 225-228, or polynucleotide sequences that are at least 50, 60, 70, 80, 90, 95, or 100% identical to one or more polynucleotide sequences containing sequences corresponding to any one of SEQ ID NOs: 1-142 and 225-228, are provided herein. In certain embodiments, the sequence encodes one or more additional amino acid sequences at either the N-terminus, C-terminus, or both of the polypeptide encoded by the polynucleotide. The type, combination, N-terminus or C-terminus, and / or order of the additional amino acid sequences may be any of those disclosed herein. The sequence can be codon-optimized for, for example, E. coli, S. cerevisiae, or human codon optimization. In certain embodiments of the above, the polynucleotide is an engineered polynucleotide. In certain embodiments of the latter, the polynucleotide is codon-optimized for E. coli, S. cerevisiae, or human.
[0170] In a particular embodiment, the sequence corresponds to any one of the sequence numbers 1, 5, 11, 15, 19, 23, 27, 31, 35, 39, 43, 47, 51, 55, 59, 63, 67, 71, 75, 79, 83, 87, 91, 95, 99, 103, 107, 111, 115, 119, 123, 127, 131, 135, 139, and 225, or the sequence numbers 1, 5, 11, 15, 19, 23, 27, 31, 35, 39, 43, 47 Provided herein are one or more polynucleotides comprising one or more polynucleotide sequences that are at least 50, 60, 70, 80, 90, 95, or 100% identical to one of the polynucleotide sequences corresponding to any one of 51, 55, 59, 63, 67, 71, 75, 79, 83, 87, 91, 95, 99, 103, 107, 111, 115, 119, 123, 127, 131, 135, 139, and 225. In certain embodiments, the sequence encodes one or more additional amino acid sequences at either the N-terminus, C-terminus, or both of the polypeptide encoded by the polynucleotide. The type, combination, N-terminus or C-terminus, and / or order of the additional amino acid sequences may be any of those disclosed herein. The sequence can be codon-optimized for, for example, E. coli, S. cerevisiae, or human codon optimization. In certain embodiments, a polynucleotide corresponding to any one of SEQ ID NOs. 230-256 and 330 is provided, or one or more polynucleotide sequences corresponding to any one of SEQ ID NOs. 230-256 and 330 are provided to be at least 50, 60, 70, 80, 90, 95, or 100% identical to one or more of the polynucleotide sequences. In certain embodiments of the above, the polynucleotide is an engineered polynucleotide. In certain embodiments of the latter, the polynucleotide is codon-optimized for E. coli, S. cerevisiae, or human.
[0171] In some embodiments, the guide RNA (gRNA) disclosed herein may be any gRNA. In other embodiments, the gRNA disclosed herein may contain a nucleic acid sequence with at least 50% nucleic acid identity to any one of SEQ ID NOs. 178-188. In some embodiments, the gRNA disclosed herein contains a nucleic acid sequence having about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, over 95%, or 100% nucleic acid identity to any one of SEQ ID NOs. 178-188. In some embodiments, the gRNA disclosed herein contains a nucleic acid sequence with at least 50%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or over 95% nucleic acid identity to any one of SEQ ID NOs. 178-188. In some embodiments, the manipulated polynucleotide (gRNA) may be split into fragments containing synthetic tracrRNA and crRNA.
[0172] In some embodiments, the gRNAs disclosed herein contain a nucleic acid sequence that is at least 50% nucleic acid identical to SEQ ID NO: 188. In some embodiments, the gRNAs disclosed herein contain a nucleic acid sequence that is about 10%, 20%, 30%, 40%, 50%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, over 95%, or 100% nucleic acid identical to SEQ ID NO: 188. In some embodiments, the gRNAs disclosed herein contain a nucleic acid sequence that is at least 50%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or over 95% nucleic acid identical to SEQ ID NO: 188.
[0173] In some embodiments, the polynucleotides encoding the nucleic acid-inducible nucleases disclosed herein include nucleic acid sequences having at least 50% nucleic acid identity with any nucleic acid represented by SEQ ID NOs: 1-142 or 225. In some embodiments, the nucleic acid-inducible nucleases disclosed herein include nucleic acid sequences having about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, over 95%, or 100% polynucleotide identity with any one of SEQ ID NOs: 1-142 or 225.
[0174] In some cases, the nucleic acid-induced nucleases disclosed herein are encoded from nucleic acid sequences. Such nucleic acids can be codon-optimized for expression in desired host cells. Suitable host cells, in non-limiting examples, may include prokaryotic cells such as E. coli, P. aeruginosa, B. subtilus, and V. natriegens, eukaryotic cells such as S. cerevisiae, plant cells, insect cells, nematode cells, amphibian cells, fish cells, or mammalian cells including human cells.
[0175] Nucleic acid sequences encoding nucleic acid-induced nucleases can be operably ligated to a promoter. Such nucleic acid sequences may be linear or circular. The nucleic acid sequence may be contained within a larger linear or circular nucleic acid sequence that includes an origin of replication, a selection marker or selectable marker, a terminator, other components of a targetable nuclease system, such as a guide nucleic acid, or additional elements such as an edit or recorder cassette as disclosed herein. In some embodiments, the nucleic acid sequence may include at least one glycine, at least one 6X histidine tag, and / or at least one 3X nuclear localization signal tag. The larger nucleic acid sequence may be a recombinant expression vector, as described in more detail below.
[0176] Generally, guide polynucleotides can form complexes with compatible nucleic acid-inducible nucleases and can direct the nucleases toward the target sequence by hybridizing with the target sequence. A target nucleic acid-inducible nuclease that can form complexes with a guide polynucleotide may be referred to as a nucleic acid-inducible nuclease compatible with the guide polynucleotide. Furthermore, a guide polynucleotide capable of forming complexes with a nucleic acid-inducible nuclease may be referred to as a guide polynucleotide or guide nucleic acid compatible with the nucleic acid-inducible nuclease. In some embodiments, the polynucleotides (gRNAs) disclosed herein can be split into fragments containing synthetic tracrRNA and crRNA. Examples of gRNAs are, but are not limited to, those shown in Table 1.
[0177] [Table 1]
[0178] The guide polynucleotide may be DNA. The guide polynucleotide may be RNA. The guide polynucleotide may contain both DNA and RNA. The guide polynucleotide may contain modified nucleotides or non-native nucleotides. If the guide polynucleotide contains RNA, the RNA guide polynucleotide may be encoded by a DNA sequence on a polynucleotide molecule such as a plasmid, linear construct, or edited cassette disclosed herein.
[0179] A guide polynucleotide may include a guide sequence. The guide sequence is a polynucleotide sequence that hybridizes with the target sequence and has sufficient complementarity with the target polynucleotide sequence to induce sequence-specific binding of a complex nucleic acid-induced nuclease to the target sequence. The degree of complementarity between the guide sequence and its corresponding target sequence is approximately 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99%, or higher, or approximately greater than 50%, greater than 60%, greater than 75%, greater than 80%, greater than 85%, greater than 90%, greater than 95%, greater than 97.5%, greater than 99%, or higher, when optimally aligned using a preferred alignment algorithm. Optimal alignment can be determined using any preferred algorithm for aligning sequences. In some embodiments, the guide sequence may have a nucleotide length of approximately 5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 75 or more, or more than approximately 5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 75 or more. In other embodiments, the guide sequence may have a nucleotide length of less than approximately 75, 50, 45, 40, 35, 30, 25, 20. Preferably, the guide sequence is 10 to 30 nucleotides long. The guide sequence may be 15 to 20 nucleotides long. The guide sequence may be 15 nucleotides long. The guide sequence may be 16 nucleotides long. The guide sequence may be 17 nucleotides long. The guide sequence may be 18 nucleotides long. The guide sequence may be 19 nucleotides long. The guide sequence may be 20 nucleotides long.
[0180] Guide polynucleotides may include scaffold sequences. Generally, a “scaffold sequence” can include any sequence having enough sequences to facilitate the formation of a targetable nuclease complex, and while targetable nuclease complexes include, but are not limited to, nucleic acid-induced nucleases, guide polynucleotides may include both scaffold sequences and guide sequences. Sufficient sequences within a scaffold sequence to facilitate the formation of a targetable nuclease complex may include some degree of complementarity along the lengths of two sequence regions within the scaffold sequence, such as one or two sequence regions involved in the formation of a secondary structure. In some cases, one or two sequence regions may be contained on or encoded on the same polynucleotide. In some cases, one or two sequence regions may be contained on or encoded on a different polynucleotide. The optimal alignment can be determined by any preferred alignment algorithm, which can further describe secondary structures such as self-complementarity within either of the one or two sequence regions. In some embodiments, the degree of complementarity between one or two sequence regions along the shorter of the two lengths is, when optimally aligned, about 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97.5%, 99%, or more, or about 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97.5%, 99%, or more. In some embodiments, at least one of the two sequence regions may have a nucleotide length of about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50 or more, or a nucleotide length of about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50 or more.
[0181] The scaffold sequence of the subject guide polynucleotide may include a secondary structure. The secondary structure may include a pseudoknot region. In some cases, the binding dynamics of the guide polynucleotide to the nucleic acid-inducible nuclease are partially determined by the secondary structure within the scaffold sequence. In some cases, the binding dynamics of the guide polynucleotide to the nucleic acid-inducible nuclease are partially determined by the nucleic acid sequence including the scaffold sequence. In some embodiments, the present invention provides a nuclease that binds to a guide polynucleotide which may include a conserved scaffold sequence. For example, a nucleic acid-inducible nuclease for use in this disclosure may bind to a conserved pseudoknot region. Thus, the scaffold sequence may include a secondary structure. The secondary structure may include a pseudoknot region. In some cases, the binding dynamics of the guide polynucleotide to the nucleic acid-inducible nuclease are partially determined by the secondary structure within the scaffold sequence. In some cases, the binding dynamics of the guide polynucleotide to the nucleic acid-inducible nuclease are partially determined by the nucleic acid sequence including the scaffold sequence.
[0182] In certain methods, the compatibility scaffold sequence of a compatibility guide nucleic acid can be found by scanning sequences adjacent to the locus of a native nucleic acid-inducible nuclease gene. For example, a native nucleic acid-inducible nuclease can encode on the genome in close proximity to the corresponding compatibility guide nucleic acid or scaffold sequence. See, for example, Example 3.
[0183] Table 3 below provides the conserved DNA sequences for each of ART1 to ART35. These sequences encode the conserved gRNA sequences for each nuclease, and the RNA sequences can be constructed from these sequences (conserved RNA sequences in Table 3). Furthermore, some or all of the RNA sequences can be subjected to further processing to remove one or more nucleotides from any end. For each specific ART nuclease, the spacer length is a specific number of NTs, and the scaffold sequence length is a specific number of NTs. Thus, in a particular embodiment of the gRNA (before processing) for a particular nuclease, these lengths are as shown in Table 3, and the total length of the gRNA (before potential additional processing to produce the final gRNA) for a particular ART nuclease is the sum of the two. In the discussion of the various embodiments herein, it is understood that ART nucleases and gRNAs containing conserved sequences (or parts thereof, see below) refer to the specific ART nucleases disclosed herein and the corresponding specific gRNAs, or parts thereof.
[0184] Therefore, in each ART nuclease, the conserved RNA sequence may also include a portion of the RNA sequence, for example, a shortened version of the conserved RNA, such as a sequence shortened by one or more nucleotides at either the 5' and / or 3' ends. While not bound by theory, these portions can, at least in some cases, correspond to, for example, the RNA sequence representing the final gRNA after editing, and / or highly conserved sequences present in many or all gRNAs for use with a particular ART nuclease. In the latter case, these may be, for example, RNA sequences necessary to produce important secondary structures in the final gRNA. The gRNA of a particular ART nuclease may include a conserved portion, or a highly conserved portion, such as a highly conserved portion containing the nucleotide sequence of a pseudoknot, which is a secondary structure of the gRNA.
[0185] Therefore, in certain embodiments, a conserved gRNA, i.e., a gRNA containing a conserved portion or part thereof, such as a conserved scaffold portion of a particular ART nuclease, is used with the ART nuclease. In certain embodiments, the conserved gRNA contains one or a part thereof of any of the sequences of SEQ ID NOs. 291–325. That portion is a contiguous sequence within a conserved RNA sequence in which one or more nucleotides have been removed from either the 5' end, the 3' end, or both (in this context, “removed” simply means that that portion of the conserved RNA is not present, regardless of how it was produced). This portion may be any preferred portion, and in a particular embodiment, this portion comprises 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides removed from the 5' end of the gRNA, and in a particular embodiment, this portion comprises 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides removed from the 3' end of the gRNA (as long as at least one nucleotide is removed). In certain embodiments, highly conserved regions are used, such as a secondary structure of gRNA, for example, a highly conserved region containing a nucleotide sequence of a secondary structure including a pseudoknot.
[0186] In certain embodiments, the gRNA for use with a particular ART nuclease is a split gRNA, such as a split gRNA containing modified nucleotides. In certain embodiments, the gRNA for use with a particular ART nuclease is a single gRNA, such as a single gRNA containing modified nucleotides. Suitable gNAs, e.g., gRNAs, may be produced by any suitable method, e.g., in a natural setting, in a host cell, by synthesis (gRNA with modified nucleotides), or by any other suitable method. Such methods are well known in the art. In certain embodiments, the gRNA for use with a particular ART nuclease is produced by synthesis. In certain embodiments, the gRNA for use with a particular ART nuclease is produced by synthesis and contains modified nucleotides, such as chemically modified nucleotides. Synthetic gRNAs containing modified nucleotides are described in more detail in U.S. Patent Application Publication No. 20160289675. In certain embodiments, the gRNA for use with a particular nucleic acid-inducible nuclease includes a conserved gRNA or a portion thereof, such as a conserved gRNA or a portion thereof derived from one of the above SEQ ID NOs: 291-325. In certain embodiments, the conserved gRNA includes highly conserved regions, such as those forming secondary structures like pseudoknots. It is understood that the term "nucleotide" as used herein may be either a native nucleotide or a modified nucleotide. For example, sequences 291–325 described herein should be interpreted, depending on the context, as either sequences containing native ribonucleotides or sequences containing one or more chemically modified ribonucleotides.
[0187] [Table 2-1]
[0188] [Table 2-2]
[0189] [Table 2-3]
[0190] The guide polynucleotide, or "gRNA," can be represented by any one of the sequences represented by SEQ ID NOs. 178–188, or by another suitable gRNA. In some embodiments, the manipulated polynucleotide (gRNA) can be split into fragments containing synthetic tracrRNA and crRNA. In some embodiments, a gRNA represented by having at least 80%, or at least 85%, or at least 90%, or at least 95% identity with any one of the sequences represented by SEQ ID NOs. 178–223 may include synthetic tracrRNA and crRNA.
[0191] As used herein, “guide nucleic acid” or “guide polynucleotide” may refer to one or more polynucleotides and may include 1) a guide sequence that can hybridize to a target sequence, and 2) a scaffold sequence that can interact with or complex with a nucleic acid-induced nuclease described herein. The guide nucleic acid may be provided as one or more nucleic acids. In certain embodiments, the guide sequence and scaffold sequence are provided as a single polynucleotide. In other embodiments, the guide nucleic acid may comprise at least one amplicon-targeting fragment.
[0192] A guide nucleic acid may be compatible with a nucleic acid-inducible nuclease if its two elements can form a functional, targetable nuclease complex capable of cleaving a target sequence. In certain methods, the compatible scaffold sequence of a compatible guide nucleic acid can be found by scanning sequences adjacent to the locus of a native nucleic acid-inducible nuclease gene. For example, a native nucleic acid-inducible nuclease may encode on the genome in close proximity to its corresponding compatible guide nucleic acid or scaffold sequence.
[0193] Nucleic acid-induced nucleases can be fitted with guide nucleic acids not found in the nuclease's endogenous host. These orthogonal guide nucleic acids can be determined by empirical testing. These orthogonal guide nucleic acids may originate from different bacterial species, be synthesized, or be manipulated to be unnatural.
[0194] A common nucleic acid-induced nuclease and a compatible orthogonal guide nucleic acid may contain one or more common features. These common features may include sequences outside the pseudoknot region, the pseudoknot region, or the primary or secondary structure.
[0195] Guide nucleic acids can be manipulated to target a desired target sequence by modifying the guide sequence so that it is complementary to the target sequence, thereby enabling hybridization between the guide and target sequences. A guide nucleic acid having a manipulated guide sequence can be called a manipulated guide nucleic acid. Manipulated guide nucleic acids are often non-natural and not found in nature.
[0196] The manipulated guide nucleic acid can be formed using the Synthetic Tracr RNA (STAR) system. When combined with the Cas12a protein, STAR can form at least one ribonucleoprotein (RNP) complex that targets a specific genomic locus. STAR leverages the innate properties of CRISPR (clustered, regularly arranged short palindromic sequence repeats), and the CRISPR system functions like an immune system against viral and plasmid DNA invasion. Short DNA sequences (spacers) from the invading virus are incorporated into CRISPR loci within the bacterial genome, acting as a "memory" of the previous infection. Reinfection triggers complementary mature CRISPR RNA (crRNA), which detects the matching viral sequence. The crRNA and transactivated crRNA (tracrRNA) together induce CRISPR-associated (Cas) nucleases, causing double-strand breaks in the "foreign" DNA sequence. The prokaryotic CRISPR "immune system" has been engineered to function as a simple, easy, and rapidly implementable RNA-guided mammalian genome editing tool. STAR (including synthetic crRNA and tracrRNA), when combined with the Cas12a protein, can form ribonucleoprotein (RNP) complexes that target specific genomic loci. Manipulated guide nucleic acids formed using the RNA(STAR) system may result in split gRNAs. An example of a split gRNA for use disclosed herein may include the sequence represented by Sequence ID No. 188.
[0197] In some embodiments, the ribonucleoprotein (RNP) complex may comprise at least one nuclease disclosed herein. In some embodiments, the RNP complex may comprise at least one nuclease having an amino acid sequence of about 75%, about 85%, about 95%, about 99%, or identity with SEQ ID NOs: 143-177 or 229. In some examples, the RNP complex comprising the nucleases disclosed herein may further comprise at least one STAR gRNA. In some other examples, the RNP complex comprising the nucleases disclosed herein may further comprise at least one non-STAR gRNA. In some other examples, the RNP complex comprising the nucleases disclosed herein may further comprise at least one polynucleotide. In some embodiments, the polynucleotide comprising the RNP complex disclosed herein may be greater than about 50 nucleotides in length. In some embodiments, the polynucleotide comprising the RNP complex disclosed herein may be greater than about 50, about 150, about 500, about 1000 nucleotides, or 1000 nucleotides in length. In some embodiments, multiple nucleases can be attached to the RNP complex to affect the overall editing efficiency. In other embodiments, two or more gRNAs can be attached to the RNP complex to improve efficiency by enabling multiple editing of two or more sites in a single transfection. In other embodiments, two or more DNA templates can be attached to the RNP to enable multiple editing of one or more sites based on a specific desired repair outcome.
[0198] Nuclease system Other embodiments disclosed herein are targetable nuclease systems. In certain embodiments, the targetable nuclease system may include a nucleic acid-inducible nuclease and a compatible guide nucleic acid (also interchangeably referred to herein as “guide polynucleotide” and “gRNA”). The targetable nuclease system may include a nucleic acid-inducible nuclease or a polynucleotide sequence encoding a nucleic acid-inducible nuclease. The targetable nuclease system may include a guide nucleic acid or a polynucleotide sequence encoding a guide nucleic acid.
[0199] Generally, the targetable nuclease systems disclosed herein may be characterized by elements that facilitate the formation of a targetable nuclease complex at a site of a target sequence, wherein the targetable nuclease complex comprises a nucleic acid-inducible nuclease and a guide nucleic acid.
[0200] The guide nucleic acid, together with the nucleic acid-inducible nuclease, forms a targetable nuclease complex that can bind to a target sequence within the target polynucleotide, as determined by the guide sequence of the guide nucleic acid.
[0201] Generally, generating a double-strand break requires, in most cases, a targetable nuclease complex to bind to a target sequence as determined by a guide nucleic acid, and the nuclease to recognize a protospacer-adjacent motif (PAM) sequence adjacent to the target sequence.
[0202] A targetable nuclease complex may include a nucleic acid-inducible nuclease containing any one of sequence numbers 143-177 and 229, and a compatible guide nucleic acid. A targetable nuclease complex may include a nucleic acid-inducible nuclease containing any one of sequence numbers 143-151, and a compatible guide nucleic acid. A targetable nuclease complex may include a nucleic acid-inducible nuclease containing any one of sequence numbers 143-177, and a compatible guide nucleic acid represented by sequence numbers 178-188. A targetable nuclease complex may include a nucleic acid-inducible nuclease encoding a nuclease represented by any one of sequence numbers 1-142, and a compatible gRNA or a gRNA represented by any one of sequence numbers 178-188. In certain embodiments, the guide nucleic acid may include a scaffold sequence compatible with the selected nucleic acid-inducible nuclease. In any of these embodiments, the guide sequence can be manipulated to be complementary to any desired target sequence. The selected guide sequence can be manipulated to hybridize with any desired target sequence.
[0203] The target sequence of a targetable nuclease complex may be any polynucleotide, endogenous or exogenous, for prokaryotic or eukaryotic cells, or in vitro. For example, the target sequence may be a polynucleotide present in the nucleus of a eukaryotic cell. The target sequence may be a sequence encoding a gene product (e.g., a protein) or a non-coding sequence (e.g., regulatory polynucleotide or junk DNA). In this specification, the target sequence is assumed to be associated with a PAM, i.e., a short sequence recognized by the targetable nuclease complex. The exact sequence and length requirements of the PAM vary depending on the nucleic acid-inducible nuclease used, but the PAM may be a 2-5 base pair sequence adjacent to the target sequence. Examples of PAM sequences are described in the Examples section below, and those skilled in the art can identify further PAM sequences for use with a given nucleic acid-inducible nuclease. Furthermore, manipulation of the PAM interaction (PI) domain may enable programming of PAM specificity, improving the fidelity of target site recognition and enhancing the reusability of the nucleic acid-inducible nuclease genome manipulation platform. Nucleic acid-induced nucleases may be manipulated to alter their PAM specificity, for example, as described in Kleinstiver et al., Nature. 2015 Jul. 23;523(7561):481-5, the entire disclosure of which is incorporated herein.
[0204] The PAM site is a nucleotide sequence adjacent to the target sequence. In most cases, nucleic acid-inducible nucleases can only cleave the target sequence if a suitable PAM is present. PAMs are nucleic acid-inducible nuclease-specific and can differ between two different nucleic acid-inducible nucleases. PAMs can be 5' or 3' of the target sequence. PAMs can be upstream or downstream of the target sequence. PAMs can have a nucleotide length of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more. Often, the nucleotide length of a PAM is between 2 and 6.
[0205] In some embodiments disclosed herein, the PAM may be provided on a separate oligonucleotide. In such cases, since there is no adjacent PAM on the same polynucleotide as the target sequence, providing the PAM on an oligonucleotide enables cleavage of the target sequence that would otherwise be impossible.
[0206] A polynucleotide sequence encoding components of a targetable nuclease system may contain one or more vectors. Generally, as used herein, the term “vector” can refer to a nucleic acid molecule capable of transporting another nucleic acid to which it is linked. Examples of vectors include, but are not limited to, nucleic acid molecules that are single-stranded, double-stranded, or partially double-stranded; nucleic acid molecules containing one or more free ends, or (e.g., circular) without free ends; nucleic acid molecules containing DNA, RNA, or both; and other types of polynucleotides known in the art. One type of vector is a “plasmid,” which refers to a circular double-stranded DNA loop into which additional DNA fragments can be inserted by standard molecular cloning techniques, for example. Another type of vector is a viral vector, in which a viral-derived DNA or RNA sequence is present within the vector and packaged in a virus (e.g., retroviruses, replication-deficient retroviruses, adenoviruses, replication-deficient adenoviruses, and adeno-associated viruses). Other vectors (e.g., non-episomal mammalian vectors) may be incorporated into the host cell's genome upon introduction into the host cell. The recombinant expression vector may contain the nucleic acid of the present invention in a form suitable for nucleic acid expression in host cells, which may mean that the recombinant expression vector contains one or more regulatory elements that can be operably linked to the nucleic acid sequence to be expressed and can be selected based on the host cell used for expression.
[0207] In some embodiments, the regulatory element can be operably linked to one or more elements of a targetable nuclease system in order to drive the expression of one or more components of the targetable nuclease system.
[0208] In some embodiments, the vector may include a regulatory element operably linked to a polynucleotide sequence encoding a nucleic acid-inducible nuclease. The polynucleotide sequence encoding the nucleic acid-inducible nuclease may be codon-optimized for expression in target cells such as prokaryotic or eukaryotic cells. Eukaryotic cells may be cells of yeast, fungi, algae, plants, animals, or humans. Eukaryotic cells may originate from organisms such as mammals, including but not limited to humans, mice, rats, rabbits, dogs, or non-human mammals, including non-human primates.
[0209] Generally, codon optimization can refer to the process of modifying a nucleic acid sequence to enhance expression in a target host cell by replacing at least one codon in the natural sequence with a codon that is more or most frequently used in that host cell gene, while maintaining the natural amino acid sequence. Different species exhibit specific biases to codons of particular amino acids. As intended herein, genes can be tuned for optimal gene expression in a given organism based on codon optimization. Codon usage tables are readily available, such as the "Codon Usage Database" available at www.kazusa.or.jp, and these tables can be adapted in many ways. See Nakamura, Y., et al. "Codon usage tabulated from the international DNA sequence databases: status for the year 2000" Nucl. Acids Res. 28:292 (2000). In certain embodiments, codon-optimized polynucleotides are provided. In certain embodiments, one or more polynucleotides are provided that encode at least 50, 60, 70, 80, 90, 95, or 100% identical amino acid sequences to one or more amino acid sequences corresponding to any one of SEQ ID NOs: 143-177 and 229, or to any one of SEQ ID NOs: 144, 153, and 229 (possibly SEQ ID NOs: 144, 153, and 229), where the polynucleotides are codon-optimized, for example, codon-optimized to E. coli, or codon-optimized to S. cerevisiae, or codon-optimized to human. In certain embodiments, the sequences encode one or more additional amino acid sequences at either the N-terminus, C-terminus, or both of the polypeptide encoded by the polynucleotide.The type, combination, N-terminus or C-terminus, and / or order of the additional amino acid sequences may be any of those disclosed herein.
[0210] In certain embodiments, sequences 2-4, 6-10, 12-14, 16-18, 20-22, 24-26, 28-30, 32-34, 36-38, 40-42, 44-46, 48-50, 52-54, 56-58, 60-62, 64-66, 68-70, 72-74, 76-78, 80-82, 84-86, 88-90, 92-94, 96-98, One or more codon-optimized polynucleotides corresponding to any of the following: 100-102, 104-106, 108-110, 112-114, 116-118, 120-122, 124-126, 128-130, 132-134, 136-138, 140-142, 226-228, and 330, or SEQ ID NOs: 2-4, 6-10, 12-14, 16 ~18, 20~22, 24~26, 28~30, 32~34, 36~38, 40~42, 44~46, 48~50, 52~54, 56~58, 60~62, 64~66, 68~70, 72~74, 76~78, 80~82, 84~86, 88~90, 92~94, 96~98, 100~102, 104~106, 108~110, 112~114, Provided herein are polynucleotide sequences that are at least 50, 60, 70, 80, 90, 95, or 100% identical to one or more polynucleotide sequences corresponding to any one of 116-118, 120-122, 124-126, 128-130, 132-134, 136-138, 140-142, 226-228, and 330. In certain embodiments, the sequences encode one or more additional amino acid sequences at either the N-terminus, C-terminus, or both of the polypeptide encoded by the polynucleotides. The type, combination, N-terminus or C-terminus, and / or order of the additional amino acid sequences may be any of those disclosed herein.In certain embodiments, one or more E. coli codon-optimized polynucleotides corresponding to any one of SEQ ID NOs: 2, 6, 12, 16, 20, 24, 28, 32, 36, 40, 44, 48, 52, 56, 60, 64, 68, 72, 76, 80, 84, 88, 92, 96, 100, 104, 108, 112, 116, 120, 124, 128, 132, 136, 140, 226, and 330, or SEQ ID NOs: 2, 6, 12, 16, 20, Provided herein are polynucleotide sequences that are at least 50, 60, 70, 80, 90, 95, or 100% identical to one or more polynucleotide sequences corresponding to any one of 24, 28, 32, 36, 40, 44, 48, 52, 56, 60, 64, 68, 72, 76, 80, 84, 88, 92, 96, 100, 104, 108, 112, 116, 120, 124, 128, 132, 136, 140, 226, and 330. In certain embodiments, the sequences encode one or more additional amino acid sequences at either the N-terminus, C-terminus, or both of the polypeptide encoded by the polynucleotides. The type, combination, N-terminus or C-terminus, and / or order of the additional amino acid sequences may be any of those disclosed herein. In certain embodiments, one or more S. cerivisiae codon-optimized polynucleotides corresponding to any of SEQ ID NOs: 3, 7, 9, 13, 17, 21, 25, 29, 33, 37, 41, 45, 49, 53, 57, 61, 65, 69, 73, 77, 81, 85, 89, 93, 97, 101, 105, 109, 113, 117, 121, 125, 129, 133, 137, 141, and 227, or SEQ ID NOs: 3, 7, 9, 13, 17 The Specified Specified Polynucleotide Sequences are provided that are at least 50, 60, 70, 80, 90, 95, or 100% identical to one or more polynucleotide sequences corresponding to any one of 21, 25, 29, 33, 37, 41, 45, 49, 53, 57, 61, 65, 69, 73, 77, 81, 85, 89, 93, 97, 101, 105, 109, 113, 117, 121, 125, 129, 133, 137, 141, and 227.In certain embodiments, the sequence encodes one or more additional amino acid sequences at either the N-terminus, C-terminus, or both of the polypeptide encoded by the polynucleotide. The type, combination, N-terminus or C-terminus, and / or order of the additional amino acid sequences may be any of those disclosed herein. In certain embodiments, one or more human codon-optimized polynucleotides corresponding to any of SEQ ID NOs: 4, 10, 14, 18, 22, 26, 30, 34, 38, 42, 46, 50, 54, 58, 62, 66, 70, 74, 78, 82, 86, 90, 94, 98, 102, 106, 110, 114, 118, 122, 126, 130, 134, 138, 142, and 228, or SEQ ID NOs: 4, 10, 14, 18, 22, 26 Provided herein are polynucleotide sequences that are at least 50, 60, 70, 80, 90, 95, or 100% identical to one or more polynucleotide sequences corresponding to any one of 30, 34, 38, 42, 46, 50, 54, 58, 62, 66, 70, 74, 78, 82, 86, 90, 94, 98, 102, 106, 110, 114, 118, 122, 126, 130, 134, 138, 142, and 228. In certain embodiments, the sequences encode one or more additional amino acid sequences at either the N-terminus, C-terminus, or both of the polypeptide encoded by the polynucleotides. The type, combination, N-terminus or C-terminus, and / or order of the additional amino acid sequences may be any of those disclosed herein.
[0211] Nucleic acid-inducible nucleases and one or more guide nucleic acids can be delivered as either DNA or RNA. By delivering both the nucleic acid-inducible nuclease and the guide nucleic acid as RNA molecules (including unmodified or base- or backbone-modified RNAs), the retention time of the nucleic acid-inducible nuclease in cells can be reduced. This can reduce the level of off-target cleavage activity in target cells. Since delivering the nucleic acid-inducible nuclease as mRNA takes time to translate to protein, it may be advantageous to deliver the guide nucleic acid several hours after the delivery of the nucleic acid-inducible nuclease mRNA to maximize the level of guide nucleic acid available for interaction with the nucleic acid-inducible nuclease protein. In other cases, the nucleic acid-inducible nuclease mRNA and the guide nucleic acid are delivered simultaneously. In other examples, the guide nucleic acid is delivered sequentially after the nucleic acid-inducible nuclease mRNA, for example, 0.5, 1, 2, 3, 4 hours or more.
[0212] Guide nucleic acids, either in RNA form or encoded on a DNA expression cassette, can be introduced into host cells that may contain a vector or a nucleic acid-inducible nuclease encoded on a chromosome. The guide nucleic acid may be provided within the cassette as one or more polynucleotides, which may be continuous or discontinuous within the cassette. In certain embodiments, the guide nucleic acid is provided within the cassette as a single continuous polynucleotide.
[0213] Nucleic acid-inducible nucleases (DNA or RNA) and guide nucleic acids (DNA or RNA) can be introduced into host cells using various delivery systems. According to these embodiments, useful systems include, but are not limited to, yeast systems, lipofection systems, microinjection systems, microparticle gun systems, virosomes, liposomes, immunoliposomes, polycations, lipid:nucleic acid conjugates, virions, artificial virions, viral vectors, electroporation, cell-permeable peptides, nanoparticles, nanowires (Shalek et al., Nano Letters, 2012), and exosomes. A molecular Trojan liposome (Pardridge et al., Cold Spring Harb Protoc; 2010; doi:10.1101 / pdb.prot5407) can be used to deliver manipulated nucleases and guide nucleases across the blood-brain barrier.
[0214] In some embodiments, editing templates are also provided. The editing template may be a component of the vector described herein, may be contained in a separate vector, or may be provided as a separate polynucleotide such as an oligonucleotide, a linear polynucleotide, or a synthetic polynucleotide. In some cases, the editing template is on the same polynucleotide as the guide nucleic acid. In some embodiments, the editing template is designed to function as a template in homologous recombination, such as within or near a target sequence to be nicked or cleaved by a nucleic acid-induced nuclease, as part of a complex disclosed herein. The editing template polynucleotide may be of any preferred length, such as about 10, 15, 20, 25, 50, 75, 100, 150, 200, 500, 1000 or more nucleotide lengths, or about greater than 10, greater than 15, greater than 20, greater than 25, greater than 50, greater than 75, greater than 100, greater than 150, greater than 200, greater than 500, greater than 1000 or more nucleotide lengths. In some embodiments, the editing template polynucleotide is complementary to a portion of the polynucleotide that may contain the target sequence. When optimally aligned, the editing template polynucleotide may overlap with one or more nucleotides of the target sequence (e.g., about 1, 5, 10, 15, 20, 25, 30, 35, 40, or more, or about 1, 5, 10, 15, 20, 25, 30, 35, 40, or more nucleotides). In some embodiments, when the polynucleotides, which may include the editing template sequence and the target sequence, are optimally aligned, the nearest nucleotide of the template polynucleotide is within about 1, 5, 10, 15, 20, 25, 50, 75, 100, 200, 300, 400, 500, 1000, 5000, 10000, or more nucleotides from the target sequence.
[0215] In some embodiments, methods are provided for delivering one or more vectors or polynucleotides, such as linear polynucleotides, one or more transcripts thereof, and / or one or more proteins transcribed therefrom, to a host cell. In some embodiments, the present invention further provides cells produced by such methods, and organisms may comprise or be produced from such cells. In some embodiments, an engineered nuclease combined with (and optionally forming a complex with) a guide nucleic acid is delivered to the cell.
[0216] Nucleic acids can be introduced into cells such as prokaryotic cells, eukaryotic cells, mammalian cells, or target tissues using conventional viral and nonviral-based gene transfer methods. These methods can be used to deliver nucleic acids encoding components of engineered nucleic acid-inducible nuclease systems to cells in cultures or host organisms. Nonviral vector delivery systems include DNA plasmids, RNA (e.g., transcripts of vectors described herein), naked nucleic acids, and nucleic acids complexed with delivery vehicles such as liposomes. Viral vector delivery systems include DNA viruses and RNA viruses that, after delivery to cells, have either an episome or an integrated genome. Any gene therapy known in the art is considered useful herein. Nonviral methods of nucleic acid delivery are intended herein. Adeno-associated virus ("AAV") vectors can also be used, for example, in the in vitro production of nucleic acids and peptides, and for in vivo and ex vivo gene therapy procedures, to transduce target nucleic acids into cells.
[0217] In some embodiments, host cells are transiently or nontransiently transfected with one or more vectors, linear polynucleotides, polypeptides, nucleic acid-protein complexes, or any combination thereof as described herein. In some embodiments, cells are transfected in vitro, in culture medium, or ex vivo. In some embodiments, cells are transfected as they occur naturally within the subject. In some embodiments, the cells to be transfected are harvested from the subject. In some embodiments, the cells are derived from cells harvested from the subject, such as a cell line.
[0218] In some embodiments, novel cell lines containing one or more transfection-derived sequences are established using cells transfected with one or more vectors, linear polynucleotides, polypeptides, nucleic acid-protein complexes, or any combination thereof as described herein. In some embodiments, novel cell lines may contain cells containing modifications but lacking any other exogenous sequences, using cells transiently transfected with components of the engineered nucleic acid-induced nuclease system described herein (such as by transient transfection of one or more vectors or transfection with RNA) and modified via the activity of the engineered nuclease complex.
[0219] In some embodiments, one or more vectors described herein are used to produce non-human transgenic cells, organisms, animals, or plants. In some embodiments, the transgenic animals are mammals such as mice, rats, or rabbits. Methods for producing transgenic cells, organisms, plants, and animals are known in the art and generally begin with cell transformation or transfection methods, such as those described herein.
[0220] In certain embodiments, the engineered nuclease complex, the "target sequence," may refer to a sequence in which a guide sequence is designed to be complementary, where hybridization between the target sequence and the guide sequence facilitates the formation of the engineered nuclease complex. The target sequence may include any polynucleotide, such as DNA, RNA, or a DNA-RNA hybrid. The target sequence may be located in the nucleus or cytoplasm of a cell. The target sequence may be placed in vitro or in a cell-free environment.
[0221] In some embodiments, the formation of an engineered nuclease complex may include a guide nucleic acid hybridized to a target sequence and complexed with one or more novel engineered nucleases as disclosed herein, resulting in a break in one or both strands within or near the target sequence (e.g., within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50 or more base pairs therefrom). The break may occur within the target sequence, at 5′ of the target sequence, upstream of the target sequence, at 3′ of the target sequence, or downstream of the target sequence.
[0222] In some embodiments, one or more vectors driving the expression of one or more components of a targetable nuclease system are introduced into a host cell or in vitro, or a targetable nuclease complex is formed at one or more target sites. For example, the nucleic acid-inducible nuclease and the guide nucleic acid may each be operably linked to separate regulatory elements on separate vectors. Alternatively, two or more elements expressed from the same or different regulatory elements may be combined into a single vector with one or more additional vectors providing any components of the targetable nuclease system not included in the first vector. The targetable nuclease system elements combined in a single vector may be positioned in any preferred orientation, such that one element is located 5' ("upstream") or 3' ("downstream") of the second element. The coding sequence of one element may be located on the same or opposite strand of the coding sequence of the second element and oriented in the same or opposite direction. In some embodiments, a single promoter drives the expression of transcripts encoding the nucleic acid-inducible nuclease and one or more guide nucleic acids. In some embodiments, the nucleic acid-inducible nuclease and one or more guide nucleic acids are operably ligated to the same promoter and expressed from there. In other embodiments, one or more guide nucleic acids, or polynucleotides encoding one or more guide nucleic acids, are introduced into a cell or in vitro environment that may already contain the nucleic acid-inducible nuclease or a polynucleotide sequence encoding the nucleic acid-inducible nuclease.
[0223] In some embodiments, when multiple different guide sequences are used, a single expression construct can be used intracellularly or in vitro to target nuclease activity to multiple different corresponding target sequences. For example, a single vector may contain about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or more guide sequences, or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more guide sequences. In other embodiments, vectors containing about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more guide sequences, or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more guide sequences, can be provided and optionally delivered to cells in vivo or in vitro.
[0224] In some embodiments, the methods and compositions disclosed herein may include multiple guide nucleic acids such that each guide nucleic acid has a different guide sequence, thereby targeting different target sequences. According to these embodiments, multiple guide nucleic acids can be used for multiplexing, in which multiple targets are targeted simultaneously. Furthermore, or alternatively, multiple guide nucleic acids may be introduced into a cell population so that each cell in the population receives a different or random guide nucleic acid, thereby targeting multiple different target sequences across the entire cell population. In such cases, the collection of subsequently modified cells may be referred to as a library.
[0225] In other embodiments, the methods and compositions disclosed herein include a plurality of different nucleic acid-inducible nucleases, each having one or more different corresponding guide nucleic acids, thereby enabling targeting of different target sequences by different nucleic acid-inducible nucleases. In some such cases, each nucleic acid-inducible nuclease can enable two or more non-overlapping, partially overlapping, or fully overlapping multiplexing events corresponding to a plurality of distinct guide nucleic acids.
[0226] In some embodiments, nucleic acid-inducible nucleases have DNA cleavage activity or RNA cleavage activity. In some embodiments, nucleic acid-inducible nucleases direct cleavage of one or both strands at a location in the target sequence, such as within the target sequence and / or within the complement of the target sequence. In some embodiments, nucleic acid-inducible nucleases direct cleavage of one or both strands within approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500, or more base pairs from the first or last nucleotide of the target sequence.
[0227] In certain embodiments, the present invention provides a method for modifying a target sequence in vitro, or in vivo, ex vivo, or in vitro, within a prokaryotic or eukaryotic cell. In some embodiments, the method comprises extracting cells or cell populations, such as prokaryotic cells, or cells or cell populations derived from humans, non-human animals, or plants (including microalgae or other organisms), and modifying the cells(s). Culturing can be carried out at any stage in vitro or ex vivo. The cells(s) can even be reintroduced into a host, such as a non-human animal or plant (including microalgae). In the case of reintroduced cells, these may be stem cells.
[0228] In some embodiments, the method comprises binding a targetable nuclease complex to a target sequence to cause cleavage of the target sequence, thereby modifying the target sequence, wherein the targetable nuclease complex comprises a nucleic acid-inducible nuclease complexed with a guide nucleic acid, and the guide sequence of the guide nucleic acid hybridizes to the target sequence in the target polynucleotide. In some embodiments, the present invention provides a method for modifying the expression of a target polynucleotide in vitro or in prokaryotic or eukaryotic cells. In some embodiments, the method comprises binding a targetable nuclease complex to a target sequence having a target polynucleotide such that the binding can result in an increase or decrease in the expression of the target polynucleotide, wherein the targetable nuclease complex comprises a nucleic acid-inducible nuclease complexed with a guide nucleic acid, and the guide sequence of the guide nucleic acid hybridizes to the target sequence in the target polynucleotide.
[0229] In certain embodiments, the present invention provides a kit comprising one or more of the elements disclosed in the methods and compositions described above. The elements may be provided individually or in combination and in any suitable container such as vials, bottles, or tubes. In some embodiments, the kit includes instructions in one or more languages, for example, two or more languages.
[0230] In some embodiments, the kit includes one or more reagents for use in a process utilizing one or more of the elements described herein. The reagents may be provided in any suitable container. For example, the kit may provide one or more reaction buffers or storage buffers. The reagents may be provided in an assay-ready form or in a form requiring the addition of one or more other components before use (e.g., in a concentrate or lyophilized form). The buffer may be any buffer, including but not limited to sodium carbonate buffer, sodium bicarbonate buffer, borate buffer, Tris buffer, MOPS buffer, HEPES buffer, and combinations thereof. In some embodiments, the buffer is alkaline. In some embodiments, the buffer has a pH of about 7 to about 10. In some embodiments, the kit includes one or more oligonucleotides corresponding to a guide sequence for insertion into a vector so as to operably link the guide sequence and the regulatory element. In some embodiments, the kit includes an editing template.
[0231] In some embodiments, the targetable nuclease complex has a wide range of utility, including modification of target sequences (e.g., deletion, insertion, translocation, inactivation, activation) in diverse cell types. Such targetable nuclease complexes of the present invention have a wide range of applications, for example, in biochemical pathway optimization, genome-wide studies, genomic manipulation, gene therapy, drug screening, disease diagnosis, and prognosis. An exemplary targetable nuclease complex comprises a nucleic acid-inducible nuclease disclosed herein, which is complexed with a guide nucleic acid, where the guide sequence of the guide nucleic acid can hybridize to a target sequence within a target polynucleotide. The guide nucleic acid may include a guide sequence ligated to a scaffold sequence. The scaffold sequence may include one or more sequence regions that have a degree of complementarity such that they together form a secondary structure.
[0232] The editing template polynucleotide may contain the sequence to be incorporated (e.g., a mutant gene). The sequence for incorporation may be endogenous or exogenous to the cell. Examples of sequences to be incorporated include protein-coding polynucleotides or non-coding RNAs (e.g., microRNAs). Thus, the sequence for incorporation may be operablely linked to a suitable regulatory sequence(s). Alternatively, the sequence to be incorporated may provide a regulatory function. The sequence to be incorporated may be a mutation or variant of an endogenous wild-type sequence. Alternatively, the sequence to be incorporated may be a wild-type version of an endogenous mutant sequence. Furthermore, or alternatively, the sequencing of the sequence to be incorporated may be a variant or mutant form of an endogenous mutation or variant sequence.
[0233] In certain embodiments, the upstream or downstream sequence may include approximately 20 bp to approximately 2500 bp, for example, approximately 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, or approximately 2500 bp. In some embodiments, exemplary upstream or downstream sequences have approximately 15 bp to approximately 2000 bp, approximately 30 bp to approximately 1000 bp, approximately 50 bp to approximately 750 bp, approximately 600 bp to approximately 1000 bp, or approximately 700 bp to approximately 1000 bp.
[0234] In some embodiments, the editing template polynucleotide may further include markers. In certain embodiments, several markers can facilitate screening for targeted incorporation. Examples of suitable markers may include, but are not limited to, restriction sites, fluorescent proteins, or selection markers. In certain embodiments, the exogenous polynucleotide template may be constructed using recombination techniques.
[0235] In one exemplary method for modifying a target polynucleotide by incorporating an editing template polynucleotide, a double-strand break is introduced into the genomic sequence by an engineered nuclease complex, and the break can be repaired by homologous recombination using the editing template, resulting in the template being incorporated into the target polynucleotide. The presence of the double-strand break can increase the efficiency of the editing template incorporation.
[0236] Methods for modifying the expression of polynucleotides in cells are disclosed herein. Some methods involve increasing or decreasing the expression of a target polynucleotide by using a targetable nuclease complex that binds to the target polynucleotide.
[0237] The detection of gene expression levels can be performed in real time during amplification assays. In one embodiment, the amplified product can be directly visualized using a fluorescent DNA conjugate, which includes, but is not limited to, a DNA intercalator and a DNA groove conjugate. Since the amount of intercalator incorporated into the double-stranded DNA molecule may be proportional to the amount of amplified DNA product, the amount of amplified product can be easily determined by quantifying the fluorescence of the intercalated dye using conventional optics of the art. Suitable DNA conjugate dyes for this application include, but are not limited to, SYBR green, SYBR blue, DAPI, iodine propidium, Hoeste, SYBR gold, ethidium bromide, acridine, proflavin, acridine orange, acriflavin, fluorocoumanin, ellipticin, daunomycin, chloroquine, distamycin D, chromomycin, homidium, mitramycin, ruthenium polypyridyl, anthramycin, and others known to those skilled in the art.
[0238] In some embodiments, other fluorescent labels, such as sequence-specific probes, can be used in the amplification reaction to facilitate the detection and quantification of the amplified product. Probe-based quantitative amplification relies on the sequence-specific detection of the desired amplified product. This is achieved by utilizing fluorescent target-specific probes (e.g., TaqMan® probes) to improve specificity and sensitivity. Methods for carrying out probe-based quantitative amplification are well established in the art.
[0239] In some embodiments, drug-induced changes in the expression of sequences related to signaling biochemical pathways can also be determined by examining the corresponding gene products. Protein-level determination may involve a) contacting proteins in a biological sample with a drug that specifically binds to proteins related to signaling biochemical pathways, and b) identifying any drug-protein complex thus formed. In one aspect of this embodiment, the drug that specifically binds to proteins related to signaling biochemical pathways is an antibody, preferably a monoclonal antibody.
[0240] In some embodiments, the amount of drug:polypeptide complex formed during the binding reaction can be quantified by a standard quantitative assay. As shown above, the formation of the drug:polypeptide complex can be directly measured by the amount of label remaining at the binding site. Alternatively, proteins associated with signal transduction biochemical pathways are tested for their ability to compete with labeled analogs for the binding site of a particular drug. In this competition assay, the amount of captured label is inversely proportional to the amount of signal transduction biochemical pathway-related protein sequences present in the test sample.
[0241] In some embodiments, many techniques for protein analysis based on the general principles outlined above are known in the art and are intended herein. These include, but are not limited to, radioimmunoassays, ELISA (enzyme-linked immunoradiometric assays), "sandwich" immunoassays, immunoradiometric assays, in-sight immunoassays (e.g., using colloidal gold, enzymes, or radioisotope labeling), Western blot analysis, immunoprecipitation assays, immunofluorescence assays, and SDS-PAGE.
[0242] In some embodiments, when carrying out the subject method, it may be desirable to identify the expression patterns of proteins related to signaling biochemical pathways in different body tissues, different cell types, and / or different intracellular structures. These studies can be performed using tissue-specific, cell-specific, or intracellular structure-specific antibodies that can bind to protein markers selectively expressed in specific tissues, cell types, or intracellular structures.
[0243] In other embodiments, altered expression of genes associated with signaling biochemical pathways can also be determined by examining changes in the activity of gene products compared to control cells. Assays for drug-induced changes in the activity of proteins associated with signaling biochemical pathways depend on the biological activity and / or signaling pathway under investigation. For example, if the protein is a kinase, changes in its ability to phosphorylate downstream substrates can be determined by a variety of assays known in the art. Typical assays include, but are not limited to, immunoblotting and immunoprecipitation using antibodies such as anti-phosphotyrosine antibodies that recognize phosphorylated proteins. Furthermore, kinase activity can be detected by high-throughput chemiluminescence assays.
[0244] In certain embodiments where proteins associated with signal transduction biochemical pathways are part of a signal transduction cascade that results in fluctuations in intracellular pH conditions, pH-sensitive molecules such as fluorescent pH dyes can be used as reporter molecules. In another example where proteins associated with signal transduction biochemical pathways are ion channels, fluctuations in membrane potential and / or intracellular ion concentration can be monitored. Many commercially available kits and high-throughput instruments are suitable for rapid and robust screening of ion channel modulators. Representative instruments include FLIPR® (Molecular Devices, Inc.) and VIPR (Aurora Biosciences). These instruments can simultaneously detect reactions in over 1000 sample wells of a microplate and provide real-time measurement and functional data within 1 second or even 1 millisecond.
[0245] When carrying out any of the methods disclosed herein, a suitable vector can be introduced into a cell, tissue, organism, or embryo via one or more methods known in the art, including but not limited to microinjection, electroporation, sonoporation, microparticle gun, calcium phosphate-mediated transfection, cationic transfection, liposome transfection, dendrimer transfection, heat shock transfection, nucleofection transfection, magnetofection, lipofection, impalefection, phototransfection, nucleic acid uptake enhanced by proprietary drugs, and delivery via liposomes, immunoliposomes, virosomes, or artificial virions. In some methods, the vector is introduced into the embryo by microinjection. The vector(s) may be microinjected into the nucleus or cytoplasm of the embryo. In some methods, the vector(s) may be introduced into a cell by nucleofection.
[0246] The target polynucleotide of a targetable nuclease complex can be any polynucleotide that is endogenous or exogenous to the host cell. For example, the target polynucleotide may be a polynucleotide present in the nucleus of a eukaryotic cell, the genome of a prokaryotic cell, or an extrachromosomal vector of the host cell. The target polynucleotide may be a sequence that codes for a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide or junk DNA).
[0247] Some embodiments disclosed herein relate to the use of the engineered nucleic acid-inducible nuclease systems disclosed herein for, for example, targeting and knocking out genes, amplifying genes, and / or repairing certain mutations associated with DNA repeat instability and medical disorders. These nuclease systems can be used to suppress and correct defects in genomic instability. In other embodiments, the engineered nucleic acid-inducible nuclease systems disclosed herein can be used to correct genetic defects associated with Lafora disease. Lafora disease is an autosomal recessive disorder characterized by progressive myoclonus epilepsy, which can begin in adolescence as epileptic seizures. The condition leads to seizures, muscle spasms, difficulty walking, dementia, and ultimately death.
[0248] In yet another aspect of the present invention, a modified / novel nucleic acid-induced nuclease system can be used to correct genetic eye diseases resulting from certain gene mutations.
[0249] Some other embodiments of the present invention relate to the modification of defects associated with a wide range of genetic diseases, which are further described in the subject subsection Genetic Disorders on the National Institutes of Health website. Specific genetic disorders of the brain include, but are not limited to, adrenoleukodystrophy, corpus callosum agenesis, Ecardi syndrome, Alpers disease, gliablastoma, Alzheimer's disease, Barth syndrome, Batten disease, CADASIL, cerebellar degeneration, Fabry disease, Gerstmann-Streusler-Scheinker disease, Huntington's disease, and other triplet repeat disorders, Leigh disease, Lesch-Nyhan syndrome, Menkes disease, mitochondrial myopathy, and NINDS colposephary, or other brain disorders in which a genetically related cause is present. In some embodiments, the genetically related disorder may be neoplasm. In some embodiments where the condition is neoplasm, the target gene may include one or more of the genes listed above. In some embodiments, the health condition contemplated herein may be age-related macular degeneration or schizophrenia-related disorder. In other embodiments, the condition may be a trinucleotide repeat disorder or fragile X syndrome. In other embodiments, the condition may be a secretase-related disorder. In some embodiments, the condition may be a prion-related disorder. In some embodiments, the condition may be ALS. In some embodiments, the condition may be a drug addiction related to prescription drugs or illegal drugs. According to these embodiments, a protein associated with addiction may be, for example, ABAT.
[0250] In some embodiments, the condition may be autism. In some embodiments, the health condition may be an inflammation-related condition, such as overexpression of inflammatory cytokines. Other inflammation-related proteins may include one or more of the following: monocyte chemotactic protein-1 (MCP1) encoded by the Ccr2 gene, CC chemokine receptor type 5 (CCR5) encoded by the Ccr5 gene, IgG receptor IIB (FCGR2b, also known as CD32) encoded by the Fcgr2b gene, or FcεR1g (FCER1g) protein encoded by the Fcer1g gene, or other proteins genetically associated with these conditions. In some embodiments, the condition may be Parkinson's disease. According to these embodiments, Parkinson's disease-related proteins may include, but are not limited to, α-synuclein, DJ-1, LRRK2, PINK1, Parkin, UCHL1, Symphyrin-1, and NURR1.
[0251] Cardiovascular-related proteins that contribute to heart disease may include, but are not limited to, IL1b (interleukin-1-beta), XDH (xanthine dehydrogenase), TP53 (tumor protein p53), PTGIS (prostaglandin-12 (prostacyclin) synthase), MB (myoglobin), IL4 (interleukin-4), ANGPT1 (angiopoietin-1), ABCG8 (ATP-binding cassette, subfamily G (WHITE) member 8), or CTSK (cathepsin K), or other known contributing factors to these conditions.
[0252] In some embodiments, the condition may be Alzheimer's disease. According to these embodiments, proteins associated with Alzheimer's disease may include very low-density lipoprotein receptor protein (VLDLR) encoded by the VLDLR gene, ubiquitin-like modifier activator 1 (UBA1) encoded by the UBA1 gene, or, for example, NEDD8 activator E1 catalytic subunit protein (UBE1C) encoded by the UBA3 gene, or other genetically related contributing factors.
[0253] In some embodiments, the condition may be autism spectrum disorder. According to these embodiments, proteins associated with autism spectrum disorder may include benzodiazapine receptor (peripheral)-associated protein 1 (BZRAP1) encoded by the BZRAP1 gene, AF4 / FMR2 family member 2 protein (AFF2) (also referred to as MFR2) encoded by the AFF2 gene, fragile X psychiatric autosomal homolog 1 protein (FXR1) encoded by the FXR1 gene, or fragile X psychiatric autosomal homolog 2 protein (FXR2) encoded by the FXR2 gene, or other genetically related contributing factors.
[0254] In some embodiments, the condition may be macular degeneration. According to these embodiments, proteins associated with macular degeneration may include, but are not limited to, ATP-binding cassettes, subfamily A (ABC1) member 4 protein (ABCA4) encoded by the ABCR gene, apolipoprotein E protein (APOE) encoded by the APOE gene, or chemokine (CC motif) ligand 2 protein (CCL2) encoded by the CCL2 gene, or other genetically related contributing factors.
[0255] In some embodiments, the condition may be schizophrenia. According to these embodiments, proteins associated with schizophrenia include NRG1, ErbB4, CPLX1, TPH1, TPH2, NRXN1, GSK3A, BDNF, DISCI, GSK3B, and combinations thereof.
[0256] In some embodiments, the state may be tumor suppression. According to these embodiments, proteins associated with tumor suppression may include ATM (ataxia telangiectasia mutation), ATR (ataxia telangiectasia and Rad3-associated), EGFR (epidermal growth factor receptor), ERBB2 (v-erb-b2 erythroblastic leukemia virus oncogene homolog 2), ERBB3 (v-erb-b2 erythroblastic leukemia virus oncogene homolog 3), ERBB4 (v-erb-b2 erythroblastic leukemia virus oncogene homolog 4), Notch 1, Notch 2, Notch 3, or Notch 4, or other genetically related contributing factors.
[0257] In some embodiments, the condition may be secretase dysfunction. According to these embodiments, proteins associated with secretase dysfunction may include PSENEN (presenilin enhancer 2 homolog (C. elegans)), CTSB (cathepsin B), PSEN1 (presenilin 1), APP (amyloid beta (A4) precursor protein), APH1B (anterior pharyngeal deficiency 1 homolog B (C. elegans)), PSEN2 (presenilin 2 (Alzheimer's disease 4)), or BACE1 (beta-site APP cleavage enzyme 1), or other genetically related contributing factors.
[0258] In some embodiments, the condition may be amyotrophic lateral sclerosis (ALS). According to these embodiments, proteins associated with ALS may include SOD1 (superoxide dismutase 1), ALS2 (amyotrophic lateral sclerosis 2), FUS (fused to sarcoma), TARDBP (TARDNA-binding protein), VAGFA (vascular endothelial growth factor A), VAGFB (vascular endothelial growth factor B), and VAGFC (vascular endothelial growth factor C), as well as any combination thereof, or other genetically related contributing factors.
[0259] In some embodiments, the condition may be prion disease disorder. According to these embodiments, proteins associated with prion disease disorder may include SOD1 (superoxide dismutase 1), ALS2 (amyotrophic lateral sclerosis 2), FUS (fused to sarcoma), TARDBP (TARDNA-binding protein), VAGFA (vascular endothelial growth factor A), VAGFB (vascular endothelial growth factor B), and VAGFC (vascular endothelial growth factor C), as well as any combination thereof, or other genetically related contributing factors. Examples of proteins associated with neurodegenerative states in prion disorders include A2M (alpha-2 macroglobulin), AATF (anti-apoptotic transcription factor), ACPP (acid phosphatase prostate), ACTA2 (actin alpha-2 smooth muscle aorta), ADAM22 (ADAM metallopeptidase domain), ADORA3 (adenosine A3 receptor), or ADRA1D (alpha-1D adrenergic receptor for alpha-1D adrenergic receptor), or other genetically related contributing factors.
[0260] In some embodiments, the condition may be an immunodeficiency disorder. According to these embodiments, proteins associated with the immunodeficiency disorder may include A2M [alpha-2-macroglobulin], AANAT [arylalkylamine N-acetyltransferase], ABCA1 [ATP-binding cassette, subfamily A (ABC1), member 1], ABCA2 [ATP-binding cassette, subfamily A (ABC1), member 2], or ABCA3 [ATP-binding cassette, subfamily A (ABC1), member 3], or other genetically related contributing factors.
[0261] In some embodiments, the condition may be an immunodeficiency disorder. According to these embodiments, proteins associated with an immunodeficiency disorder that may include a trinucleotide repeat disorder include AR (androgen receptor), FMR1 (fragile x intellectual disability 1), HTT (huntingtin), or DMPK (myotonic dystrophy protein kinase), FXN (frataxin), ATXN2 (ataxin 2), or other genetically related contributing factors.
[0262] In some embodiments, the condition may be a neurotransmission disorder. According to these embodiments, proteins associated with the neurotransmission disorder may include SST (somatostatin), NOS1 (nitric oxide synthase 1 (neuronal type)), ADRA2A (adrenergic, alpha-2A- receptor), ADRA2C (adrenergic, alpha-2C- receptor), TACR1 (tachykinin receptor 1), or HTR2c (5-hydroxytryptamine (serotonin) receptor 2C), or other genetically related contributing factors. In other embodiments, sequences related to neurodevelopment may include, but are not limited to, A2BP1 [ataxin 2-binding protein 1], AADAT [aminoadipate aminotransferase], AANAT [arylalkylamine N-acetyltransferase], ABAT [4-aminobutyrate aminotrans-ABCA1 [ATP-binding cassette, subfamily A (ABC1), member 1]], or ABCA13 [ATP-binding cassette, subfamily A (ABC1), member 13], or other genetically related contributing factors.
[0263] In further embodiments, genetic health conditions include Ecardi-Goutierre syndrome; Alexander disease; Alan Herndon-Dudley syndrome; POLG-related disorders; alpha-mannosidosis (types II and III); Alström syndrome; Angelman syndrome; ataxia vasodilator; neuronal ceroid lipofuscinosis; beta-thalassemia; bilateral optic atrophy and (infantile) optic atrophy type I; retinoblastoma (bilateral); Kanban's disease; cerebroophthalmofacial skeletal syndrome 1 [COFS1]; cerebral tendon xanthomatous syndrome; Cornelia de Lange syndrome; MAPT-related disorders. Disorders; hereditary prion diseases; Dravet syndrome; early-onset familial Alzheimer's disease; Friedreich's ataxia [FRDA]; Fryns syndrome; fucose storage disease; Fukuyama congenital muscular dystrophy; galactosialidosis; Gaucher disease; organic acidemia; hemophagocytic lymphohistiocytosis; Hutchinson-Gilford progeria syndrome; mucolipidosis II; infantile free sialic acid storage diseases; PLA2G6-associated neurodegeneration; Jarber-Lange-Nielsen syndrome; junctional epidermolysis bullosa; Huntington's disease; Krabbe disease (infantile); mitochondrial DNA-associated Lie syndrome and NARP; Lesch-Nyhan syndrome, LIST-associated lissencephaly; Lowe syndrome; maple syrup urine disease; MECP2 duplication syndrome; ATP7A-associated copper transport disorder; LAMA2-associated muscular dystrophy; arylsulfatase A deficiency; mucopolysaccharidosis type I, II, or III; peroxisome biosynthesis disorder; Zellbegger syndrome spectrum; neurodegenerative diseases with intracerebral iron accumulation; acid sphingomyelinase deficiency; Niemann-Pick disease type C; glycine encephalopathy; ARX-associated disorders; urea cycle disorders; COL1A1 / 2-associated osteogenesis imperfecta; mitochondrial D NA depletion syndrome; PLP1-related disorder; Perry syndrome; Phelan-McDermid syndrome; Glycogen storage disorder type II (Pompe disease) (infantile); MAPT-related disorder; MECP2-related disorder; rhombodysplasia punctate type 1; Robert syndrome; Sandhoff disease; Schindler disease type 1; Adenosine deaminase deficiency; Smith-Lemle-Oppitz syndrome; Spinal muscular atrophy; Infantile spinocerebellar ataxia; Hexosaminidase A deficiency; Fatal dysplasia type 1; Collagen type VI-related disorder; Usher syndrome type I; Congenital muscular dystrophy; Wolff-Hirschorn syndrome;Lysosomal acid lipase deficiency; and xeroderma pigmentosum; are possible but not limited to these.
[0264] In other embodiments, the hereditary disorders of animals covered by the editing systems disclosed herein may include, but are not limited to, hip dysplasia, bladder conditions, epilepsy, cardiac disorders, degenerative myelopathy, brachycephalic syndrome, glycogen branching enzyme deficiency (GBED), equine focal cutaneous asthenia (HERDA), hyperkalemia-induced periodic paralysis (HYPP), malignant hyperthermia (MH), polysaccharide-storing myopathy type 1 (PSSM1), junctional epidermolysis bullosa, cerebellar malnutrition, lavender foal syndrome, fatal familial insomnia, or other animal-related hereditary disorders.
[0265] In some embodiments of the present invention, the nuclease and / or gRNA sequence may include sequences having homologous substitutions, such as basic to basic, acidic to acidic, or polar to polar (for example, both substitution and replacement are used herein to mean the exchange of an existing amino acid residue or nucleotide with an alternative residue or nucleotide). Non-homologous substitutions are also intended, for example, from one class of residue to another, or involving the inclusion of non-natural amino acids such as ornithine (hereinafter referred to herein as Z), ornithine diaminobutyrate (hereinafter referred to herein as B), orleucine ornithine (hereinafter referred to herein as O), pyridylalanine, thienylalanine, naphthylalanine, and phenylglycine.
[0266] In certain embodiments disclosed herein, the engineered nucleic acid-inducible nuclease constructs can recognize protospacer-adjacent motif (PAM) sequences other than TTTN, or in addition to TTTN. In other embodiments, the engineered nucleic acid-inducible nuclease constructs disclosed herein may be further mutated to improve targeting efficiency, or selected from a library for specific targeted features.
[0267] Other embodiments disclosed herein relate to vectors comprising the constructs disclosed herein, which are useful for further analysis and selection of improved genome editing features.
[0268] Other embodiments disclosed herein include a kit for packaging and transporting a nucleic acid-inducible nuclease construct and / or novel gRNA, or a known gRNA, disclosed herein, and further include at least one container. In certain embodiments, several reagents required for the kit may be included for convenience, ease of transport, and efficiency.
[0269] In certain embodiments, a method is provided herein for generating a strand break in or near a target sequence in a target polynucleotide, comprising contacting the target sequence with an engineered targetable nucleic acid-induced nuclease complex disclosed herein, such as an RNP, wherein a guide nucleic acid of the complex's compatibility targets the target sequence, and causing the targetable guide nucleic acid-induced nuclease complex to generate a strand break. The target polynucleotide may be any suitable target polynucleotide, such as a target polynucleotide in a cellular genome. The target polynucleotide may be a safe harbor site. The method may further include providing an editing template to be inserted into the target sequence. The editing template may include any suitable sequence to be inserted into the break, and in certain embodiments, the editing template includes a transgene. In certain embodiments, cells or organisms produced by the method are provided herein.
[0270] In certain embodiments, the present invention provides a method comprising delivering one or more vectors or polynucleotides, such as linear polynucleotides, one or more transcripts thereof, and / or one or more proteins transcribed therefrom, as described herein, to a host cell. In some embodiments, the present invention further provides cells produced by such methods, and organisms comprising or produced from such cells. In some embodiments, an engineered nuclease combined with (and optionally forming a complex with) a guide nucleic acid is delivered to the cell.
[0271] Certain embodiments provide exemplary methods for modifying target polynucleotides by incorporating an editing template, in which a double-strand break is introduced into the genomic sequence by an engineered nuclease complex, and the break can be repaired by homologous recombination using the editing template, thereby incorporating the template into the target polynucleotide. The presence of the double-strand break can enhance the efficiency of editing template incorporation.
[0272] Further objects, advantages, and novel features of the present disclosure will become apparent to those skilled in the art upon consideration of the following examples taken in conjunction with the present disclosure.
[0273] Appendix A, which includes a sequence listing, is attached to this specification and constitutes a part of this application.
[0274] The following examples are not intended to be limiting.
[0275] IV. Examples The following examples are included to demonstrate preferred embodiments of the present disclosure. The techniques disclosed in the examples that follow represent techniques discovered by the inventors to function well in the implementation of the present disclosure and, accordingly, may be considered to constitute a suitable mode of its practice. However, those skilled in the art should understand that, upon consideration of the present disclosure, many changes may be made to the specific embodiments disclosed without departing from the spirit and scope of the present disclosure and still obtain the same or similar results.
Examples
[0276] In one exemplary method, the selection criteria were set to identify sequences having less than 60% AA sequence similarity to Cas12a, less than 60% AA sequence similarity to a positive control nuclease, and a query coverage rate of more than 80%. After several screening rounds, 35 nucleases were identified and designated herein as ART1-35 for further study. An overview of the investigation is shown in Table 2.
[0277]
Table 3-1
[0278]
Table 3-2
Examples
[0279] In some methods, codon optimization can reduce nucleotide sequence similarity in most cases, as described in Example 8, but without altering the amino acid sequence of the protein. Further manipulations were applied to the sequences to improve the activity of the nucleases outside of their native context. The native sequences of 35 ART nucleases were manipulated to include glycine, 6x histidine, and 3x nuclear localization signal tags.
[0280] These Gly-6xHis tags are applicable for several reasons, including: 1) the 6xHis tag can be used in protein purification to enable binding to a chromatography column for purification; and 2) the N-terminal glycine further enables site-specific chemical modification, allowing for advanced protein manipulation. Furthermore, Gly-6xHis is designed to be readily removed by digestion with tobacco ecchi disease virus (TEV) protease if necessary. In these constructs, the Gly-6xHis tag is positioned at the N-terminus. The Gly-6xHis tag is described in more detail in Martos-Maldonado et al., Nat Commun. (2018) 17;9(1):3307, the disclosure of which is incorporated herein.
[0281] NLS (Nuclear Localization Signal) fragments were added to improve transport to the nucleus. The NLS fragments used in these examples were successfully added to the Cas9 construct as already described in Perli et al., Science. (2016) 353(6304); Menoret et al., Sci Rep. (2015) 5:14410; and EnGen® Spy Cas9 NLS product information from New England Biolabs (NEB), the entire disclosure of which is incorporated herein. [Examples]
[0282] In another exemplary method, it is understood that a CRISPR-Cas genome editing system requires at least two components in certain embodiments: a guide RNA (gRNA) and a CRISPR-associated (Cas) nuclease. The guide RNA is a specific RNA sequence that recognizes a target DNA region of interest and guides the Cas nuclease to this region for editing. The gRNA may consist of two parts: a guide sequence, which is a nucleotide sequence of 17–29 or longer complementary to the target DNA, and a scaffold sequence, which functions as a binding scaffold for the Cas nuclease to facilitate editing. In one method, the conserved sequences of the gRNAs of nucleases ART1–ART35 were discovered by searching 5000 bp upstream of the start codon and 1000 bp downstream of the stop codon for each of ART1–ART35, and standard methods were used to determine the conserved sequences of the putative gRNA coding segments. The conserved DNA sequences for each of ART1–ART35 are provided in Table 3 of this application. These sequences encode conserved portions of the gRNA of each nuclease, and RNA sequences can be constructed from these sequences (conserved RNA sequences in Table 3). Furthermore, as is known in the art and described elsewhere in this specification, some or all of the RNA sequences can be subjected to further processing to remove one or more nucleotides from any end. [Examples]
[0283] In another exemplary method, the cleavage efficiency of ART nuclease was tested in vivo. The cleavage efficiency of ART nuclease was tested in vivo in Escherichia coli (E. coli). In these methods, the assay is based on an in vivo depletion assay in E. coli. First, a glycerol stock of E. coli MG1655 containing a plasmid expressing ART nuclease was removed from a -80°C freezer, and 20 μL of cells were placed in two 4 mL LB (Luria-Bertani medium) containers containing 34 μg / mL chloramphenicol in a 15 mL tube. The cells were cultured overnight at 30°C and 200 rpm. Next, 4 mL of the overnight culture was placed in 200 mL of LB medium containing 34 μg / mL chloramphenicol in two 1 L flasks. The cells were then OD 600 The cells were cultured at 30°C and 200 rpm until the saturation reached 0.5–0.6. The flask was placed in a shaking water bath incubator at 42°C and 200 rpm for 15 minutes. Next, the flask was placed in ice while gently shaking by hand and kept in ice for 15 minutes. Then, the cells were transferred from the flask to 50 mL tubes (four tubes for 200 mL of cells) and centrifuged at 8000 rpm and 4°C for 5 minutes, and the supernatant was removed. Next, 50 mL of ice-cold 10% glycerol was added to 200 mL of culture medium, and the cells were resuspended. The resuspended cells were centrifuged at 8000 rpm and 4°C for 5 minutes, the supernatant was removed, and 2 mL of ice-cold 10% glycerol was added. The cells were gently resuspended with a pipette and divided into 50 μL of competent cells. The mixture was then dispensed into 72 chilled 0.1 cm electroporation cuvettes (Bio-rad).
[0284] Plasmids containing 24 gRNAs and one untargeted control gRNA were diluted to 25 ng / µl in nuclease-free water. gRNA_EC1~gRNA_EC23 are 18 targeted loci, which are the galK, lpd, accA, cynT, cynS, adhE, oppA, fabI, ldhA, pntA, pta, accD, pheA, accB, accC, aroE, aroB, and aroK genes. 2 μL (50 ng) of chilled plasmid was placed in an electroporation cuvette and electroporated at 1800 V. Next, 950 μL of LB medium was added to the cuvette and mixed, and the cells were transferred to a 96-deep-well plate (Light Labs). The 96-deep-well plate containing the cells was incubated at 30°C and 200 rpm for 2 hours.
[0285] After culturing for 2 hours, the harvested cells were diluted to 10^0, 10^1, and 10^2. Next, 10 μL of cells were placed in 90 μL of ddH2O and mixed by pipetting. After dilution, 8 μL of cells were taken from each dilution and pipetted onto LB agar plates containing 34 μg / mL chloramphenicol and 100 μg / mL carbenicillin, and allowed to dry for several minutes without cover. Next, the cover was returned to the plate, and the plate was returned to culture overnight at 30°C. The results were confirmed the following day by counting the number of colonies.
[0286] The results of depletion assays using ART1, ART2, ART5, ART6, ART8, ART9, ART10, ART11, and ART11_L679F (also referred to herein as ART11* or ART11 variant) are shown in Figures 1-9, where the data represent percentage cleavage efficiency = 1 - (number of colonies on the plate containing on-target gRNA / number of colonies on the plate containing non-target gRNA) * 100%. [Examples]
[0287] In another exemplary method, the editing efficiency of ART nuclease was tested in vivo in Escherichia coli (E. coli). In these methods, the assays were based on in vivo editing assays in E. coli. First, a glycerol stock of E. coli MG1655 containing a plasmid expressing ART nuclease was removed from a -80°C freezer, 20 μL of stock cells were taken, and placed in 4 mL of LB medium containing 34 μg / mL chloramphenicol in one 15 mL tube. The cells were cultured overnight at 30°C and 200 rpm. Next, 1 mL of the overnight culture was placed in 50 mL of LB medium containing 34 μg / mL chloramphenicol and 0.2% arabinose in a 500 mL flask. The cells were then OD 600 The cells were cultured at 30°C and 200 rpm until the pH reached 0.5–0.6. The flask was then placed in a shaking water bath incubator at 42°C and 200 rpm for 15 minutes. Next, the flask was placed in ice while gently shaking by hand and kept in ice for 15 minutes. Afterward, the cells were transferred from the flask to a 50 mL tube and centrifuged at 8000 rpm and 4°C for 5 minutes, with the supernatant removed. Next, 25 mL of ice-cold 10% glycerol was added to 50 mL of culture medium, and the cells were resuspended. The resuspended cells were centrifuged at 8000 rpm and 4°C for 5 minutes, with the supernatant removed, and 0.5 mL of ice-cold 10% glycerol was added. The cells were gently resuspended using a pipette and then divided into 50 μL portions of competent cells. The mixture was then divided into nine chilled 0.1 cm electroporation cuvettes.
[0288] The plasmid containing 3 gRNAs was diluted to 25 ng / μL with nuclease-free water. gRNA_EC1 to gRNA_EC3 targeted the galK gene. 2 μL (50 ng) of the cooled gRNA plasmid and 2 μL (50 ng) of ssDNA (used as a DNA repair template) were placed into an electroporation cuvette, and electroporation was performed at 1800 V. Next, 950 μL of LB medium was added to the cuvette and mixed, then the cells were removed from the cuvette and placed into a 1.5 mL tube. The tube containing the cells was incubated at 30 °C and 200 rpm for 2 hours.
[0289] After recovery, 5 μL of the cells were pipetted onto a MacConkey agar plate containing 34 μg / mL chloramphenicol, 100 μg / mL carbenicillin, and 1% galactose, and the cells were spread using a sterile plating bead. After removing the plating bead, the cover was put back on the plate, and the plate was returned to incubation at 30 °C overnight. The next day, the editing efficiency was calculated using the following formula.
[0290] [Number]
[0291] The results of a representative editing assay using ART2 are shown in Figure 10, and those regarding ART11 are shown in Figure 11. [Example]
[0292] In another exemplary method, the cleavage efficiency of ART nucleases can be tested in vivo in eukaryotic cells. In these methods, the assay is based on an in vivo DNA cleavage assay. Jurkat cells, an immortalized lineage of human T lymphocytes, are cultured in RPMI1640 medium containing 10% fetal bovine serum (FBS), periodically split, and then harvested for transfection. Two target loci, DNMT1 and TRAC43, are selected in genomic Jurkat DNA as targets. Nuclease ART2 and a control nuclease are diluted to 20 mg / mL in a storage buffer (e.g., NaCl 300 mM, sodium phosphate 50 mM, EDTA 0.1 mM, DTT 1 mM, and glycerol 10%). Similarly, gRNA is diluted to 100 μM in nuclease-free water. RNA-protein complexes (RNPs) are prepared by mixing 1 μL of nuclease solution with 1.5 μL of gRNA solution. The complexes form in a 96-well V-bottom plate during a 10-minute incubation at room temperature.
[0293] Cells were counted in a NucleoCounter NC-200 and their viability was estimated. The collected cells were transfected in transfection buffer (SF Cell Line 96-well Nucleofector Kit, Lonza) at a rate of 100 × 10⁶. 5Resuspend at a concentration of cells / mL. Add 20 μL of this solution to the well containing the formed RNP, mix by pipetting, and transfer to a 96-well Nucleocuvette plate (Lonza). Electroporate the cells. In some cases, before nucleoporation, mix two-component gRNA (split gRNA; STAR) in a 1:1 volume and anneal at 37°C for 30 minutes to form a gRNA solution. Note that STAR gRNA is split gRNA from which crRNA and tracrRNA have been separated. ART2 mRNA and gRNA (either single or STAR) are co-delivered immediately after resuspending in a suitable nucleoporation buffer (Lonza) and delivered via an optimized nucleoporator program (Lonza).
[0294] Following electroporation, 80 μL of fresh RPMI1640 medium containing 10% FBS is added to a Nucleocuvette plate immediately after electroporation. The solution is mixed, and 50 μL is transferred to a 96-well flat-bottom culture plate containing 150 μL of fresh medium. Cells are cultured for 72 hours and then harvested for DNA extraction. Cells are harvested by centrifugation at 1000 × g for 10 minutes and washed with buffer (PBS). The supernatant is carefully removed, and the cell pellet is treated with 20 μL of preheated QuickExtract DNA Extraction Solution (Lucigen). The plate is placed in a thermocycler (Biorad) and subjected to temperature treatment (e.g., 65°C for 15 minutes, 68°C for 15 minutes, 95°C for 10 minutes, and cooled to 4°C). Cell fragments are harvested by centrifugation, and the supernatant containing genomic DNA is collected. DNA fragments containing the target site are amplified by PCR, and the DNA is prepared for sequencing. The Illumina-compatible adapter sequence and index sequence for sample identification are added to the PCR product at the target site during the second round of PCR. The PCR products from the second round are pooled and fed into an Illumina MiSeq sequencing instrument for 2x150 paired-end sequencing. The editing frequency was determined using the Crispresso2 analysis package. [Examples]
[0295] In another exemplary method, the ART nuclease RNP was introduced into mammalian cells for gene editing.
[0296] 1.1 Cells and Cultures Jurkat, clone E6-1, acute T-cell leukemia cells were purchased (ATCC) and cultured in RPMI-1640 (Thermo Fisher) supplemented with 10% fetal bovine serum (FBS, Thermo Fisher) according to the manufacturer's instructions. All cell cultures were grown and maintained in a humidified incubator (Heracell VIOS 160i, Thermo Fisher) with 5% CO2 at 37°C.
[0297] 1.2 Introduction of ART11 RNP into mammalian cells for gene editing Ribonucleoprotein (RNP) was generated by complex formation between a single gRNA and ART11 nuclease. A single gRNA was synthesized (IDT), and recombinant ART11 was produced and purified (Aldevron). Recombinant ART11 nuclease was stored at -80°C in 25 mM Tris-HCl pH 7.4, 300 mM NaCl, 0.1 mM EDTA, 1 mM DTT, and 50% (v / v) glycerol buffer before use. A single gRNA was resuspended in IDTE pH 7.5 buffer (IDT) to produce a 100 M stock, which was stored at -80°C before use. μART11 nuclease and gRNA were mixed at room temperature for 10 minutes to form RNP. After complex formation, the RNPs were resuspended in a suitable nucleoporation buffer (e.g., SF buffer, Lonza) and delivered to mammalian cells via an optimized nucleoporator program (e.g., CA-137, Lonza).
[0298] 1.3 DNA collection for sequencing of amplicons Cells were cultured for 48 hours before being harvested for DNA extraction. Cells were separated by centrifugation (200 × g for 5 minutes), pelletized, and washed with buffer (PBS). After carefully removing the supernatant, the cell pellet was treated with 20 μL of QuickExtract DNA Extraction Solution (Lucigen). The sample was placed in a thermal cycler and subjected to temperature treatment (e.g., 65°C for 15 minutes, 68°C for 15 minutes, 95°C for 10 minutes, and then cooled to 4°C). DNA fragments containing the target site were amplified by PCR, and the DNA was prepared for sequencing. Illumina-compatible adapter sequences and index sequences for sample identification were added to the PCR product of the target site during the second round of PCR. The PCR product from the second round was pooled and loaded into an Illumina MiSeq sequencing instrument for 2x150 pair-end sequencing. Editing frequencies were determined using the Crispresso2 analysis package. The results are shown in Figure 14. [Examples]
[0299] In another exemplary method, codon optimization of non-natural nucleic acid sequences disclosed herein was performed using the Codon Optimization Tool (Integrated DNA technologies). In this method, an organism (such as bacteria (e.g., Escherichia coli K12), yeast (e.g., Saccharomyces cerevisiae), or multicellular eukaryote (e.g., Homo sapiens (human))) was selected from the "Organisms" section to express a nuclease for a wide range of applications. Next, the DNA base or amino acid sequence was loaded into the blank field of the Codon Optimization Tool. Then, the sequence was optimized. The resulting DNA sequence was codon-optimized for bacteria (e.g., Escherichia coli K12), yeast (e.g., Saccharomyces cerevisiae), or multicellular eukaryote (e.g., Homo sapiens (human)).
[0300] Examples of non-natural nucleic acid sequences disclosed herein include codon-optimized ART2 sequences for expression in bacteria such as Escherichia coli (e.g., SEQ ID NO: 6), codon-optimized sequences for expression in yeast such as Saccharomyces cerevisiae (e.g., SEQ ID NO: 7), and codon-optimized sequences for expression in multicellular eukaryotes such as Homo sapiens (human) (e.g., SEQ ID NO: 8). These non-natural nucleic acid sequences were obtained by amplification, cloning, construction, synthesis, generation from synthetic oligonucleotides or dNTPs, or by other means using methods known to those skilled in the art. Codon-optimized nucleases have been used to edit cell lines by expression from cell line plasmids, or have been expressed at high levels in protein-producing cell lines and subsequently purified for editing with RNPs. [Examples]
[0301] In another exemplary method, two or more ART nucleases (RNPs) are introduced into mammalian cells for gene editing. Ribonucleoproteins (RNPs) are produced by complex formation between a single gRNA or STAR gRNA and each ART nuclease, and a mixture of multiple ART nuclease RNPs is used for transfection. As in Example 6, a single or STAR gRNA is synthesized, and recombinant ART is produced and purified. The recombinant ART nucleases are then stored at -80°C in 25 mM Tris-HCl pH 7.4, 300 mM NaCl, 0.1 mM EDTA, 1 mM DTT, and 50% (v / v) glycerol buffer before use. The single or STAR gRNA is resuspended in IDTE buffer (10 mM Tris, 0.1 mM EDTA) pH 7.5 buffer to produce a 100 μM stock, which is stored at -80°C before use. Immediately before nucleoporation, recombinant ART was diluted in a working buffer consisting of 20 mM HEPES and 150 mM KCl pH 7.5, and the gRNA was diluted to the final working concentration in IDTE pH 7.5 buffer (annealing was performed first in the case of STAR; see Section 1.4). After diluting the ART nuclease and gRNA, both were mixed in a 1:1 volume (2:1 ratio of gRNA to nuclease) at 37°C for 10 minutes to form RNP. After complex formation, the RNP was resuspended in a suitable nucleoporation buffer (Lonza) and delivered via an optimized nucleoporator program (Lonza). [Examples]
[0302] In another exemplary method, ART nucleases are combined with multiple gRNAs to construct multi-target RNPs, which are then introduced into mammalian cells for gene editing. Ribonucleoproteins (RNPs) are produced by complexing either a single or multiple ART nucleases with either a single or multiple STAR gRNAs that target different sites within the genome. As in Example 6, single or STAR gRNAs are synthesized, and recombinant ARTs are produced and purified. The recombinant ART nucleases are stored at -80°C in 25 mM Tris-HCl pH 7.4, 300 mM NaCl, 0.1 mM EDTA, 1 mM DTT, and 50% (v / v) glycerol buffer before use. The single or STAR gRNAs are resuspended in IDTE buffer (10 mM Tris, 0.1 mM EDTA) pH 7.5 buffer to produce a 100 μM stock, which is stored at -80°C before use. Immediately before nucleoporation, recombinant ART was diluted in a working buffer consisting of 20 mM HEPES and 150 mM KCl pH 7.5, and the gRNA was diluted to the final working concentration in IDTE pH 7.5 buffer (annealing was performed first in the case of STAR; see Section 1.4). After diluting the ART nuclease and gRNA, both were mixed in a 1:1 volume (2:1 ratio of gRNA to nuclease) at 37°C for 10 minutes to form RNP. After complex formation, the RNP was resuspended in a suitable nucleoporation buffer (Lonza) and delivered via an optimized nucleoporator program (Lonza). [Examples]
[0303] In another exemplary method, the ART nuclease RNP is introduced into mammalian cells for gene editing using a polynucleotide DNA repair template used for directional repair or editing of a target genome. The ribonucleoprotein (RNP) is produced by complex formation with a single or STAR gRNA and the ART nuclease, and a polynucleotide DNA repair template in the size range of 20 bp to 20 kbp is added to the RNP. As in Example 6, a single or STAR gRNA is synthesized, and recombinant ART is produced and purified. The recombinant ART nuclease is stored at -80°C in 25 mM Tris-HCl pH 7.4, 300 mM NaCl, 0.1 mM EDTA, 1 mM DTT, and 50% (v / v) glycerol buffer before use. The polynucleotide DNA template is either synthesized from a plasmid stock containing source DNA material or is a commercially synthesized synthetic DNA material (IDT, Genewiz). The DNA template is either single-stranded DNA (ssDNA) or double-stranded DNA (dsDNA) and has homology arms adjacent to the insertion or editing region proximal to the gRNA cleavage site. Resuspend a single or STAR gRNA in IDTE buffer (10 mM Tris, 0.1 mM EDTA) pH 7.5 to produce a 100 μM stock, which should be stored at -80°C before use. Immediately before nucleoporation, dilute the recombinant ART in a working buffer consisting of 20 mM HEPES and 150 mM KCl pH 7.5, and dilute the gRNA to the final working concentration in IDTE pH 7.5 buffer (anneal first in the case of STAR; see Section 1.4). After diluting the ART nuclease and gRNA, mix both in a 1:1 volume (2:1 ratio of gRNA to nuclease) at 37°C for 10 minutes, and add the DNA template at the optimal concentration to form the RNP. After complex formation, the RNPs were resuspended in a suitable nucleoporation buffer (e.g., Lonza) and delivered via an optimized nucleoporator program (e.g., Lonza). [Examples]
[0304] In another exemplary method, the ART nuclease RNP is introduced into mammalian cells for gene editing using a mixture of polynucleotide DNA repair templates used for directional repair or editing of multiple target genomes. Ribonucleoprotein (RNP) is produced by complex formation with a single or STAR gRNA and the ART nuclease, and a mixture of polynucleotide DNA repair templates ranging in size from 20 bp to 20 kbp is added to the RNP. As in Example 6, a single or STAR gRNA is synthesized, and recombinant ART is produced and purified. The recombinant ART nuclease is stored at -80°C in 25 mM Tris-HCl pH 7.4, 300 mM NaCl, 0.1 mM EDTA, 1 mM DTT, and 50% (v / v) glycerol buffer before use. The polynucleotide DNA templates are either synthesized from plasmid stocks containing source DNA material or are commercially synthesized synthetic DNA material (IDT, Genewiz). The DNA template is either single-stranded DNA (ssDNA) or double-stranded DNA (dsDNA) and has homology arms adjacent to the insertion or editing region proximal to the gRNA cleavage site. Resuspend a single or STAR gRNA in IDTE buffer (10 mM Tris, 0.1 mM EDTA) pH 7.5 buffer to produce a 100 μM stock, which should be stored at -80°C before use. Immediately before nucleoporation, dilute the recombinant ART in a working buffer consisting of 20 mM HEPES and 150 mM KCl pH 7.5, and dilute the gRNA to the final working concentration in IDTE pH 7.5 buffer (anneal first in the case of STAR; see Section 1.4). After diluting the ART nuclease and gRNA, mix both in a 1:1 volume (2:1 ratio of gRNA to nuclease) at 37°C for 10 minutes, and add the DNA template to the optimal concentration to form the RNP. After complex formation, the RNPs were resuspended in a suitable nucleoporation buffer (e.g., Lonza) and delivered via an optimized nucleoporator program (e.g., Lonza). [Examples]
[0305] In this example, the PAM sequences of representative nucleases disclosed herein were evaluated.
[0306] Using a glycerol stock of E. coli MG1655 containing a plasmid expressing ART nuclease, and a 100 μL cell stock, seeded in 4 mL of LB medium containing 34 mg / mL chloramphenicol in a 15 mL tube. The cells were cultured overnight (12-16 hours) in a 30°C, 200 rpm shaking incubator. After overnight growth, 1 mL of the overnight cell culture was added to a 250 mL flask containing 25 mL of LB medium containing 34 mg / mL chloramphenicol. The cells were cultured overnight in a 30°C, 200 rpm shaking incubator. 600 The cells were cultured until the pH reached 0.5-0.6. The cells were transferred from the flask to a 50 mL tube and centrifuged at 8000 rpm and 4°C for 5 minutes, and the supernatant was removed. Next, 25 mL of ice-cold 10% glycerol was added and the cells were resuspended. The resuspended cells were centrifuged at 8000 rpm and 4°C for 5 minutes, the supernatant was removed, and 2 mL of ice-cold 10% glycerol was added. The cells were gently resuspended with a pipette and divided into 50 μL aliquots of competent cells.
[0307] Prepare for electroporation by transferring 50 μL of prepared competent cells to an electroporation cuvette with a 0.1 cm gap on ice. Add 200 ng of PAM plasmid library (carrying on-target sites with various PAM sequences) to the electroporation cuvette and electroporate at 1800 V. Add 950 μL of Super Optimal Broth (SOB) medium, mix gently, and transfer the entire volume to a 1.5 mL tube. Incubate at 30°C and 200 rpm for 2 hours. From each tube, transfer the cells to another 15 mL plastic tube containing 4 mL of LB medium containing 34 mg / mL chloramphenicol and 50 mg / mL kanamycin, and incubate overnight. For each culture, add 1 mL of the overnight cell culture to a 250 mL flask containing 25 mL of LB medium containing 34 mg / mL chloramphenicol, and incubate in a shaking incubator at 30°C and 200 rpm. 600 The cells were incubated until the pH reached 0.5-0.6. The flask was placed in a 42°C, 200 rpm shaking water bath incubator for 15 minutes. The flask was then placed in ice while gently shaking by hand and kept in ice for 15 minutes. The cells were transferred from the flask to a 50 mL tube and centrifuged at 8000 rpm and 4°C for 5 minutes, with the supernatant removed. Next, 25 mL of ice-cold 10% glycerol was added to resuspend the cells. The resuspended cells were centrifuged at 8000 rpm and 4°C for 5 minutes, with the supernatant removed, and 2 mL of ice-cold 10% glycerol was added. The cells were gently resuspended with a pipette and divided into 50 μL aliquots of competent cells.
[0308] Prepare a new round of electroporation by transferring 50 μL of prepared competent cells to an electroporation cuvette with a 0.1 cm gap on ice. Add 100 μg of either a non-targeted control or on-targeting gRNA plasmid and electroporate at 1800 V. Add 950 μL of SOP medium, mix gently, and transfer the entire volume to a 1.5 mL tube. Incubate at 30°C and 200 rpm for 2 hours. Seed aliquots of the harvested cells onto LB agar plates containing 50 mg / mL kanamycin and 100 mg / mL carbenicillin, and incubate overnight at 30°C. Harvest the cells and purify the plasmid. The plasmid is used as template DNA for a PCR reaction using primers miniseq_galKOFF_T225-F2(TCGTCGGCAGCGTCAGATGTGTATAAGAGACAGcgtaccctggttggcagcgaatac) (SEQ ID NO: 326) and miniseq_galKOFF_T225-R2(GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAGacgcacgcgttttgccacgatc) (SEQ ID NO: 327). Illumina-compatible adapter sequences and index sequences for sample identification are added to the PCR product during the second round of PCR. The PCR product from the second round is pooled and loaded into an Illumina MiSeq sequencing instrument for 2x150 paired-end sequencing.
[0309] The NGS data was aligned with a reference PAM library template using the VSEARCH tool. The alignment threshold was 0.9. The aligned data was filtered using a pandas program. The data threshold was 100%. The normalized readings for each PAM were calculated using the following formula, where PAM hits are the sum of PAM hits after running pandas, and total hits are the total hits before running VSEARCH.
[0310]
number
[0311] The degree of enrichment was calculated using the following formula, where normalized readings are used. y This refers to the normalized readings of each PAM in the non-targeted controlled experiment, and the normalized readings x This represents the normalized readings of each PAM in on-targeting gRNA experiments.
[0312]
number
[0313] Figure 11 shows the results for the top hits in the PAM region of ART11. Figure 13 shows the results for the top hits in the PAM region of ART11_L679F. [Examples]
[0314] This example describes site-directed mutagenesis (SDM) for developing the mutant nuclease ART11_L679F.
[0315] [Table 4]
[0316] Exponential amplification (PCR) using NEB SDM kits
[0317] [Table 5]
[0318] b. Run PCR on 5 μl of the PCR product in a 1% agarose gel at 120V for 30 minutes.
[0319] Kinase, ligase, and DpnI (KLD) treatment Combine the following reagents:
[0320] [Table 6]
[0321] Mix thoroughly by pipetting up and down, then incubate at room temperature for 30 minutes.
[0322] Transformation 1. Thaw the tube containing NEB 5-alpha Competent E. coli cells on ice.
[0323] 2. Add 5 μl of the KLD mix from step II to the tube containing the thawed cells. Gently tap the tube 4-5 times to mix. Do not vortex.
[0324] 3. Place the mixture on ice for 30 minutes.
[0325] 4. Administer a heat shock at 4.42°C for 30 seconds.
[0326] 5. Place on ice for 5 minutes.
[0327] Transfer 6,950 μl of room temperature SOC to the mixture using a pipette.
[0328] 7. Incubate at 30°C for 2 hours while shaking (200 rpm).
[0329] 8. Gently tap and invert the tube to thoroughly mix the cells, then spread 10 and 100 μl onto a selection plate and incubate overnight at 30°C.
[0330] 9. Select several colonies for Sanger sequencing to find the modified plasmid.
[0331] Appendix A is incorporated herein by reference in its entirety for all purposes.
[0332] The above statements in this disclosure are provided for illustrative and explanatory purposes only. They are not intended to limit this disclosure to any or any of the forms disclosed herein. While the descriptions in this disclosure include descriptions of one or more embodiments and specific variations and modifications, other variations and modifications that may, for example, be within the scope of the disclosure and may be within the scope of the art and knowledge of those skilled in the art after understanding this disclosure are also included. It is intended that rights be obtained including, to the extent permitted, alternative embodiments, including alternative, interchangeable, and / or equivalent structures, functions, scopes, or processes, whether or not such alternative, interchangeable, and / or equivalent structures, functions, scopes, or processes are disclosed herein, and without the intent to disclose patentable subject matter.
[0333] Preferred embodiments of the present invention have been illustrated and described herein, but it will be apparent to those skilled in the art that these embodiments are provided merely as examples. Those skilled in the art will anticipate numerous variations, modifications, and substitutions without departing from the present invention. It should be understood that various alternatives to the embodiments of the present invention described herein may be employed in the practice of the present invention. The following claims define the scope of the present invention and are intended to encompass the methods and structures included in these claims, as well as their equivalents.
[0334] [Table 7] TIFF2026082943000014.tif223158TIFF2026082943000015.tif223158TIFF2026082943000016.tif223158TIFF2026082943000017.tif223158TIFF2026082943000018.tif223158TIFF2026082943000019.tif223158TIFF2026082943000020.tif223158TIFF2026082943000021.tif223158TIFF2026082943000022.tif223158TIFF2026082943000023.tif223158TIFF2026082943000024.tif223158TIFF2026082943000025.tif223158TIFF2026082943000026.tif223158TIFF2026082943000027.tif223158TIFF2026082943000028.tif223158TIFF2026082943000029.tif223158TIFF2026082943000030.tif223158TIFF2026082943000031.tif223158TIFF2026082943000032.tif223158TIFF2026082943000033.tif223158TIFF2026082943000034.tif223158TIFF2026082943000035.tif223158TIFF2026082943000036.tif223158TIFF2026082943000037.tif223158TIFF2026082943000038.tif223158TIFF2026082943000039.tif223158TIFF2026082943000040.tif223158TIFF2026082943000041.tif223158TIFF2026082943000042.tif223158TIFF2026082943000043.tif223158TIFF2026082943000044.tif223158TIFF2026082943000045.tif223158TIFF2026082943000046.tif223158TIFF2026082943000047.tif223158TIFF2026082943000048.tif223158TIFF2026082943000049.tif223158TIFF2026082943000050.tif223158TIFF2026082943000051.tif223158TIFF2026082943000052.tif223158TIFF2026082943000053.tif223158TIFF2026082943000054.tif223158TIFF2026082943000055.tif223158TIFF2026082943000056.tif223158TIFF2026082943000057.tif223158TIFF2026082943000058.tif223158TIFF2026082943000059.tif223158TIFF2026082943000060.tif223158TIFF2026082943000061.tif223158TIFF2026082943000062.tif223158TIFF2026082943000063.tif223158TIFF2026082943000064.tif223158TIFF2026082943000065.tif223158TIFF2026082943000066.tif223158TIFF2026082943000067.tif223158TIFF2026082943000068.tif223158TIFF2026082943000069.tif223158TIFF2026082943000070.tif223158TIFF2026082943000071.tif223158TIFF2026082943000072.tif223158TIFF2026082943000073.tif223158TIFF2026082943000074.tif223158TIFF2026082943000075.tif223158TIFF2026082943000076.tif223158TIFF2026082943000077.tif223158TIFF2026082943000078.tif223158TIFF2026082943000079.tif223158TIFF2026082943000080.tif223158TIFF2026082943000081.tif223158TIFF2026082943000082.tif223158TIFF2026082943000083.tif223158TIFF2026082943000084.tif223158TIFF2026082943000085.tif223158TIFF2026082943000086.tif223158TIFF2026082943000087.tif223158TIFF2026082943000088.tif223158TIFF2026082943000089.tif223158TIFF2026082943000090.tif223158TIFF2026082943000091.tif223158TIFF2026082943000092.tif223158TIFF2026082943000093.tif223158TIFF2026082943000094.tif223158TIFF2026082943000095.tif223158TIFF2026082943000096.tif223158TIFF2026082943000097.tif223158TIFF2026082943000098.tif223158TIFF2026082943000099.tif223158TIFF2026082943000100.tif223158TIFF2026082943000101.tif223158TIFF2026082943000102.tif223158TIFF2026082943000103.tif223158TIFF2026082943000104.tif223158TIFF2026082943000105.tif223158TIFF2026082943000106.tif223158TIFF2026082943000107.tif223158TIFF2026082943000108.tif223158TIFF2026082943000109.tif223158TIFF2026082943000110.tif223158TIFF2026082943000111.tif223158TIFF2026082943000112.tif223158TIFF2026082943000113.tif223158TIFF2026082943000114.tif223158TIFF2026082943000115.tif223158TIFF2026082943000116.tif223158TIFF2026082943000117.tif223158TIFF2026082943000118.tif223158TIFF2026082943000119.tif223158TIFF2026082943000120.tif223158TIFF2026082943000121.tif223158TIFF2026082943000122.tif223158TIFF2026082943000123.tif223158TIFF2026082943000124.tif223158TIFF2026082943000125.tif223158TIFF2026082943000126.tif223158TIFF2026082943000127.tif223158TIFF2026082943000128.tif223158TIFF2026082943000129.tif223158TIFF2026082943000130.tif223158TIFF2026082943000131.tif223158TIFF2026082943000132.tif223158TIFF2026082943000133.tif223158TIFF2026082943000134.tif223158TIFF2026082943000135.tif223158TIFF2026082943000136.tif223158TIFF2026082943000137.tif223158TIFF2026082943000138.tif223158TIFF2026082943000139.tif223158TIFF2026082943000140.tif223158TIFF2026082943000141.tif223158TIFF2026082943000142.tif223158TIFF2026082943000143.tif223158TIFF2026082943000144.tif223158TIFF2026082943000145.tif223158TIFF2026082943000146.tif223158TIFF2026082943000147.tif223158TIFF2026082943000148.tif223158TIFF2026082943000149.tif223158TIFF2026082943000150.tif223158TIFF2026082943000151.tif223158TIFF2026082943000152.tif223158TIFF2026082943000153.tif223158TIFF2026082943000154.tif223158TIFF2026082943000155.tif223158TIFF2026082943000156.tif223158TIFF2026082943000157.tif223158TIFF2026082943000158.tif223158TIFF2026082943000159.tif223158TIFF2026082943000160.tif223158TIFF2026082943000161.tif223158TIFF2026082943000162.tif223158TIFF2026082943000163.tif223158TIFF2026082943000164.tif223158TIFF2026082943000165.tif223158TIFF2026082943000166.tif223158TIFF2026082943000167.tif223158TIFF2026082943000168.tif223158TIFF2026082943000169.tif223158TIFF2026082943000170.tif223158TIFF2026082943000171.tif223158TIFF2026082943000172.tif223158TIFF2026082943000173.tif223158TIFF2026082943000174.tif223158TIFF2026082943000175.tif223158TIFF2026082943000176.tif223158TIFF2026082943000177.tif223158TIFF2026082943000178.tif223158TIFF2026082943000179.tif223158TIFF2026082943000180.tif22315 8TIFF2026082943000181.tif223158TIFF2026082943000182.tif223158TIFF2026 082943000183.tif223158TIFF2026082943000184.tif223158TIFF2026082943000 185.tif223158TIFF2026082943000186.tif223158TIFF2026082943000187.tif22 3158TIFF2026082943000188.tif223158TIFF2026082943000189.tif223158TIFF 2026082943000190.tif223158TIFF2026082943000191.tif223158TIFF202608294 3000192.tif223158TIFF2026082943000193.tif223158TIFF2026082943000194.t if223158TIFF2026082943000195.tif223158TIFF2026082943000196.tif223158. [Sequence Listing Free Text]
[0335] Sequence ID 2: Description of artificial sequence: Synthetic polynucleotide Sequence ID 3: Description of artificial sequence: Synthetic polynucleotide Sequence ID 4: Description of artificial sequence: Synthetic polynucleotide Sequence ID 6: Description of artificial sequence: Synthetic polynucleotide Sequence ID 7: Description of artificial sequence: Synthetic polynucleotide Sequence ID 8: Description of artificial sequence: Synthetic polynucleotide Sequence ID 9: Description of artificial sequence: Synthetic polynucleotide Sequence ID 10: Description of artificial sequence: Synthetic polynucleotide Sequence ID 12: Description of artificial sequence: Synthetic polynucleotide Sequence ID 13: Description of artificial sequence: Synthetic polynucleotide Sequence ID 14: Description of artificial sequence: Synthetic polynucleotide Sequence ID 16: Description of artificial sequence: Synthetic polynucleotide Sequence ID 17: Description of artificial sequence: Synthetic polynucleotide Sequence ID 18: Description of artificial sequence: Synthetic polynucleotide Sequence ID 20: Description of artificial sequence: Synthetic polynucleotide Sequence ID 21: Description of artificial sequence: Synthetic polynucleotide Sequence ID 22: Description of artificial sequence: Synthetic polynucleotide Sequence ID 24: Description of artificial sequence: Synthetic polynucleotide Sequence ID 25: Description of artificial sequence: Synthetic polynucleotide Sequence ID 26: Description of artificial sequence: Synthetic polynucleotide Sequence ID 28: Description of artificial sequence: Synthetic polynucleotide Sequence ID 29: Description of artificial sequence: Synthetic polynucleotide Sequence ID 30: Description of artificial sequence: Synthetic polynucleotide Sequence ID 31: Description of unknown element: Parcubacteria bacterial sequence Sequence ID 32: Description of artificial sequence: Synthetic polynucleotide Sequence ID 33: Description of artificial sequence: Synthetic polynucleotide Sequence ID 34: Description of artificial sequence: Synthetic polynucleotide Sequence ID 36: Description of artificial sequence: Synthetic polynucleotide Sequence ID 37: Description of artificial sequence: Synthetic polynucleotide Sequence ID 38: Description of artificial sequence: Synthetic polynucleotide Sequence ID 40: Description of artificial sequence: Synthetic polynucleotide Sequence ID 41: Description of artificial sequence: Synthetic polynucleotide Sequence ID 42: Description of artificial sequence: Synthetic polynucleotide Sequence ID 44: Description of artificial sequence: Synthetic polynucleotide Sequence ID 45: Description of artificial sequence: Synthetic polynucleotide Sequence ID 46: Description of artificial sequence: Synthetic polynucleotide Sequence ID 48: Description of artificial sequence: Synthetic polynucleotide Sequence ID 49: Description of artificial sequence: Synthetic polynucleotide Sequence ID 50: Description of artificial sequence: Synthetic polynucleotide Sequence ID 52: Description of artificial sequence: Synthetic polynucleotide Sequence ID 53: Description of artificial sequence: Synthetic polynucleotide Sequence ID 54: Description of artificial sequence: Synthetic polynucleotide Sequence ID 56: Description of artificial sequence: Synthetic polynucleotide Sequence ID 57: Description of artificial sequence: Synthetic polynucleotide Sequence ID 58: Description of artificial sequence: Synthetic polynucleotide Sequence ID 60: Description of artificial sequence: Synthetic polynucleotide Sequence ID 61: Description of artificial sequence: Synthetic polynucleotide Sequence ID 62: Description of artificial sequence: Synthetic polynucleotide Sequence ID 64: Description of artificial sequence: Synthetic polynucleotide Sequence ID 65: Description of artificial sequence: Synthetic polynucleotide Sequence ID 66: Description of artificial sequence: Synthetic polynucleotide Sequence ID 68: Description of artificial sequence: Synthetic polynucleotide Sequence ID 69: Description of artificial sequence: Synthetic polynucleotide Sequence ID 70: Description of artificial sequence: Synthetic polynucleotide Sequence ID 71: Description of unknown element: Sodaliphilus pleomorphus sequence Sequence ID 72: Description of artificial sequence: Synthetic polynucleotide Sequence ID 73: Description of artificial sequence: Synthetic polynucleotide Sequence ID 74: Description of artificial sequence: Synthetic polynucleotide Sequence ID 76: Description of artificial sequence: Synthetic polynucleotide Sequence ID 77: Description of artificial sequence: Synthetic polynucleotide Sequence ID 78: Description of artificial sequence: Synthetic polynucleotide Sequence ID 80: Description of artificial sequence: Synthetic polynucleotide Sequence ID 81: Description of artificial sequence: Synthetic polynucleotide Sequence ID 82: Description of artificial sequence: Synthetic polynucleotide Sequence ID 84: Description of artificial sequence: Synthetic polynucleotide Sequence ID 85: Description of artificial sequence: Synthetic polynucleotide Sequence ID 86: Description of artificial sequence: Synthetic polynucleotide Sequence ID 88: Description of artificial sequence: Synthetic polynucleotide Sequence ID 89: Description of artificial sequence: Synthetic polynucleotide Sequence ID 90: Description of artificial sequence: Synthetic polynucleotide Sequence ID 91: Description of unknown element: Lachnospiraceae bacterial sequence Sequence ID 92: Description of artificial sequence: Synthetic polynucleotide Sequence ID 93: Description of artificial sequence: Synthetic polynucleotide Sequence ID 94: Description of artificial sequence: Synthetic polynucleotide Sequence ID 96: Description of artificial sequence: Synthetic polynucleotide Sequence ID 97: Description of artificial sequence: Synthetic polynucleotide Sequence ID 98: Description of artificial sequence: Synthetic polynucleotide Sequence ID 100: Description of artificial sequence: Synthetic polynucleotide Sequence ID 101: Description of artificial sequence: Synthetic polynucleotide Sequence ID 102: Description of artificial sequence: Synthetic polynucleotide Sequence ID 104: Description of artificial sequence: Synthetic polynucleotide Sequence ID 105: Description of artificial sequence: Synthetic polynucleotide Sequence ID 106: Description of artificial sequence: Synthetic polynucleotide Sequence ID 108: Description of artificial sequence: Synthetic polynucleotide Sequence ID 109: Description of artificial sequence: Synthetic polynucleotide Sequence ID 110: Description of artificial sequence: Synthetic polynucleotide Sequence ID 112: Description of artificial sequence: Synthetic polynucleotide Sequence ID 113: Description of artificial sequence: Synthetic polynucleotide Sequence ID 114: Description of artificial sequence: Synthetic polynucleotide Sequence ID 116: Description of artificial sequence: Synthetic polynucleotide Sequence ID 117: Description of artificial sequence: Synthetic polynucleotide Sequence ID 118: Description of artificial sequence: Synthetic polynucleotide Sequence ID 120: Description of artificial sequence: Synthetic polynucleotide Sequence ID 121: Description of artificial sequence: Synthetic polynucleotide Sequence ID 122: Description of artificial sequence: Synthetic polynucleotide Sequence ID 124: Description of artificial sequence: Synthetic polynucleotide Sequence ID 125: Description of artificial sequence: Synthetic polynucleotide Sequence ID 126: Description of artificial sequence: Synthetic polynucleotide Sequence ID 128: Description of artificial sequence: Synthetic polynucleotide Sequence ID 129: Description of artificial sequence: Synthetic polynucleotide Sequence ID 130: Description of artificial sequence: Synthetic polynucleotide Sequence ID 132: Description of artificial sequence: Synthetic polynucleotide Sequence ID 133: Description of artificial sequence: Synthetic polynucleotide Sequence ID 134: Description of artificial sequence: Synthetic polynucleotide Sequence ID 136: Description of artificial sequence: Synthetic polynucleotide Sequence ID 137: Description of artificial sequence: Synthetic polynucleotide Sequence ID 138: Description of artificial sequence: Synthetic polynucleotide Sequence ID 140: Description of artificial sequence: Synthetic polynucleotide Sequence ID 141: Description of artificial sequence: Synthetic polynucleotide Sequence ID 142: Description of artificial sequence: Synthetic polynucleotide Sequence ID 150: Description of unknown element: Parcubacteria bacterial sequence Sequence ID 160: Description of unknown element: Sodaliphilus pleomorphus sequence Sequence ID 165: Description of unknown element: Lachnospiraceae bacterial sequence Sequence ID 178: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 179: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 180: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 181: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 182: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 183: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 184: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 185: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 186: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 187: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 188: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 189: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 190: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 191: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 192: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 193: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 194: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 195: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 196: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 197: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 198: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 199: Description of artificial sequence: Synthetic oligonucleotide Sequence ID No. 200: Description of artificial sequence: Synthetic oligonucleotide Sequence ID No. 201: Description of artificial sequence: Synthetic oligonucleotide Sequence ID No. 202: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 203: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 204: Description of artificial sequence: Synthetic oligonucleotide Sequence ID No. 205: Description of artificial sequence: Synthetic oligonucleotide Sequence ID No. 206: Description of artificial sequence: Synthetic oligonucleotide Sequence ID No. 207: Description of artificial sequence: Synthetic oligonucleotide Sequence ID No. 208: Description of artificial sequence: Synthetic oligonucleotide Sequence ID No. 209: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 210: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 211: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 212: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 213: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 214: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 215: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 216: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 217: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 218: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 219: Description of artificial sequence: Synthetic oligonucleotide Sequence ID No. 220: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 221: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 222: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 223: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 224: Description of unknown element: ART nuclease motif array Sequence ID 225: Description of artificial sequence: Synthetic polynucleotide Sequence ID 226: Description of artificial sequence: Synthetic polynucleotide Sequence ID 227: Description of artificial sequence: Synthetic polynucleotide Sequence ID 228: Description of artificial sequence: Synthetic polynucleotide Sequence ID 229: Description of artificial sequence: Synthetic polypeptide Sequence ID 230: Description of artificial sequence: Synthetic polynucleotide Sequence ID 231: Description of artificial sequence: Synthetic polynucleotide Sequence ID 232: Description of artificial sequence: Synthetic polynucleotide Sequence ID 233: Description of artificial sequence: Synthetic polynucleotide Sequence ID 234: Description of artificial sequence: Synthetic polynucleotide Sequence ID 235: Description of artificial sequence: Synthetic polynucleotide Sequence ID 236: Description of artificial sequence: Synthetic polynucleotide Sequence ID 237: Description of artificial sequence: Synthetic polynucleotide Sequence ID 238: Description of artificial sequence: Synthetic polynucleotide Sequence ID 239: Description of artificial sequence: Synthetic polynucleotide Sequence ID 240: Description of artificial sequence: Synthetic polynucleotide Sequence ID 241: Description of artificial sequence: Synthetic polynucleotide Sequence ID 242: Description of artificial sequence: Synthetic polynucleotide Sequence ID 243: Description of artificial sequence: Synthetic polynucleotide Sequence ID 244: Description of artificial sequence: Synthetic polynucleotide Sequence ID 245: Description of artificial sequence: Synthetic polynucleotide Sequence ID 246: Description of artificial sequence: Synthetic polynucleotide Sequence ID 247: Description of artificial sequence: Synthetic polynucleotide Sequence ID 248: Description of artificial sequence: Synthetic polynucleotide Sequence ID 249: Description of artificial sequence: Synthetic polynucleotide Sequence ID 250: Description of artificial sequence: Synthetic polynucleotide Sequence ID 251: Description of artificial sequence: Synthetic polynucleotide Sequence ID 252: Description of artificial sequence: Synthetic polynucleotide Sequence ID 253: Description of artificial sequence: Synthetic polynucleotide Sequence ID 254: Description of artificial sequence: Synthetic polynucleotide Sequence ID 255: Description of artificial sequence: Synthetic polynucleotide Sequence ID 256: Description of artificial sequence: Synthetic polynucleotide Sequence ID 257: Description of artificial sequence: Synthetic polypeptide Sequence ID 258: Description of artificial sequence: Synthetic polypeptide Sequence ID 259: Description of artificial sequence: Synthetic polypeptide Sequence ID 260: Description of artificial sequence: Synthetic polypeptide Sequence ID 261: Description of artificial sequence: Synthetic polypeptide Sequence ID 262: Description of artificial sequence: Synthetic polypeptide Sequence ID 264: Description of unknown element: Nucleoplasmin bipartite NLS sequence Sequence ID 265: Description of unknown element: C-myc NLS sequence Sequence ID 266: Description of unknown element: C-myc NLS sequence Sequence ID 268: Description of unknown element: IBB domain derived from importin α sequence Sequence ID 269: Description of unknown element: Myoma T protein sequence Sequence ID 270: Description of unknown element: Myoma T protein sequence Sequence ID 279: Description of unknown element: Nuclear localization sequence Sequence ID 280: Description of artificial sequence: Synthetic peptide Sequence ID 281: Description of artificial sequence: Synthetic peptide Sequence ID 282: Description of unknown element: ART nuclease motif array Sequence ID 283: Description of unknown element: ART nuclease motif array Sequence ID 284: Description of unknown element: ART nuclease motif array Sequence ID 285: Description of unknown element: ART nuclease motif array Sequence ID 286: Description of unknown element: ART nuclease motif array Sequence ID 287: Description of unknown element: ART nuclease motif array Sequence ID 288: Description of unknown element: ART nuclease motif array Sequence ID 289: Description of unknown element: ART nuclease motif array Sequence ID 290: Description of unknown element: ART nuclease motif array Sequence ID 291: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 292: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 293: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 294: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 295: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 296: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 297: Description of artificial sequence: Synthetic oligonucleotide Sequence ID No. 298: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 299: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 300: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 301: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 302: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 303: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 304: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 305: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 306: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 307: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 308: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 309: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 310: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 311: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 312: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 313: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 314: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 315: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 316: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 317: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 318: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 319: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 320: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 321: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 322: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 323: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 324: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 325: Description of artificial sequence: Synthetic oligonucleotide Sequence ID 326: Description of artificial sequence: Synthetic primer Sequence ID 327: Description of artificial sequence: Synthetic primer Sequence ID 328: Description of artificial sequence: Synthetic primer Sequence ID 329: Description of artificial sequence: Synthetic primer Sequence ID 330: Description of artificial sequence: Synthetic polynucleotide Sequence ID 332: Description of artificial sequence: Synthetic peptide Sequence ID 333: Description of artificial sequence: Synthetic peptide Sequence ID 334: Description of artificial sequence: Synthetic 6xHis tag Sequence ID 335: Description of artificial sequence: Synthetic 8xHis tag
Claims
1. (i) an engineered nuclease polypeptide comprising an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NOs: 144, 153, and 229, or an engineered nucleic acid-derived nuclease comprising one or more polynucleotides encoding an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NOs: 144, 153, and 229, (ii) A compatible guide nucleic acid that forms a targetable nucleic acid-induced nuclease complex with the nucleic acid-induced nuclease and can induce the nuclease to target a sequence in a human cell. A composition containing the following:
2. The composition according to claim 1, wherein the manipulated nuclease polypeptide does not contain the peptide motif YLFQIYNKDF (SEQ ID NO: 224), or one or more polynucleotides encoding the manipulated nuclease polypeptide do not contain the peptide motif YLFQIYNKDF (SEQ ID NO: 224).
3. The composition according to claim 1, wherein the guide nucleic acid is gRNA and the complex is RNP.
4. The composition according to claim 1 or 3, wherein the guide nucleic acid is a split guide nucleic acid.
5. The composition according to claim 3 or 4, wherein the gRNA is a manipulated gRNA.
6. The composition according to claim 5, wherein the manipulated gRNA includes conserved gRNA.
7. The composition according to claim 6, wherein the conserved gRNA comprises one of sequence numbers 291 to 325, or a portion thereof.
8. The composition according to claim 7, wherein the portion is a highly conserved portion containing the nucleotide sequence of the secondary structure of the RNA.
9. The composition according to claim 8, wherein the secondary structure includes a pseudoknot.
10. The composition according to any one of claims 3 to 9, wherein the gRNA comprises one or more chemical modifications.
11. A method for generating a chain break in or near a target sequence within a target polynucleotide in a human cell, comprising: contacting the target polynucleotide with a targetable nucleic acid-induced nuclease complex as defined in any one of claims 1 to 10, wherein the compatible guide nucleic acid of the complex targets the target sequence, and the targetable guide nucleic acid-induced nuclease complex generates the chain break.
12. The method according to claim 11, wherein the target polynucleotide is located in the cell genome.
13. The method according to claim 11 or 12, further comprising providing an editing template to be inserted into the target sequence.
14. The method according to claim 13, wherein the editing template includes an introduced gene.
15. The method according to any one of claims 11 to 14, wherein the target polynucleotide is a safe harbor site.
16. A method for producing manipulated cells by the method described in claim 11.
17. A method for producing an manipulated organism by the method of claim 11.
18. A method for modifying the expression of a target polynucleotide in a human cell, comprising contacting a targetable nucleic acid-induced nuclease complex as defined in any one of claims 1 to 15 with a target sequence, wherein a compatibility guide nucleic acid of the complex targets the target sequence, and the targetable nucleic acid-induced nuclease complex binds to the target sequence in the target polynucleotide, thereby enabling the binding to result in an increase or decrease in the expression of the target polynucleotide.
19. A composition comprising one or more manipulated polynucleotides, each containing one or more polynucleotides, each containing one or more polynucleotides, each containing a sequence corresponding to a sequence having at least 90% sequence identity with any one of sequence numbers 5-10, 43-46, 225-228, 230-248, 253-256, and 330.
20. The composition according to claim 19, wherein the polynucleotide encodes one or more additional amino acid sequences at either the N-terminus, the C-terminus, or both of the polypeptide encoded by the polynucleotide.
21. The aforementioned additional amino acid sequence is (i) One or more NLS, (ii) One or more refined tags, (iii) One or more cleavage sequences, and (iv) FLAG or 3XFLAG The composition according to claim 20, comprising at least one of the following.
22. The composition according to any one of claims 19 to 21, wherein the one or more polynucleotides are codon-optimized for E. coli, S. cerevisiae, or human.
23. The composition according to claim 22, comprising one or more polynucleotides including a sequence corresponding to a sequence having at least 90% sequence identity with any of sequence numbers 6, 7, 9, 44, 45, 226, 227, and 330.