In vitro and in situ circularized rnas
DNA endonucleases and zinc finger proteins with specific sequences address the immune response issue, enabling efficient genetic editing and increased RNA expression in cells.
Patent Information
- Authority / Receiving Office
- AU · AU
- Patent Type
- Applications
- Current Assignee / Owner
- RGT UNIV OF CALIFORNIA
- Filing Date
- 2025-02-28
- Publication Date
- 2026-07-16
AI Technical Summary
Existing DNA endonucleases provoke an immune response, such as interferon secretion, limiting their use in genetic engineering applications.
Development of DNA endonucleases with specific amino acid sequences, such as SEQ ID NO:33-53, that do not provoke an immune response, including deimmunized enzymes and zinc finger proteins with novel alpha helix domains for genetic engineering.
The DNA endonucleases and zinc finger proteins enable efficient genetic editing without immune activation, facilitating the use of large protein-encoding nucleic acid sequences in cells like cardiomyocytes and neurons, with increased expression and persistence of circularized RNA.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
DNA ENDONUCLEASE COMPOSITIONS
[0240] The compositions and methods provided herein include DNA endonuclease compositions and methods for designing, generating, and / or using the same. The DNA endonuclease compositions provided herein including embodiments thereof may be used, inter alia, in genetic engineering methods. The DNA endonuclease enzyme provided herein, including embodiments thereof do not provoke or produce an immune response (e.g., interferon secretion). Thus, in an aspect is provided a DNA endonuclease enzyme including the amino acid sequence of any one of SEQ ID NO:33-53.
[0241] In embodiments, the DNA endonuclease enzy me includes the amino acid sequence of SEQ ID NO:33. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33. wherein the amino acid residue corresponding to position 28 is proline or leucine; the amino acid residue corresponding to position 237 is leucine or cysteine; the amino acid residue corresponding to position 286 is tyrosine or glutamine; the amino acid residue corresponding to position 318 is serine, histidine, or cysteine; the amino acid residue corresponding to position 368 is serine or cysteine; the amino acid residue corresponding to position 498 is phenylalanine or threonine; the amino acid residue corresponding to position 514 is leucine, glycine, or threonine; the amino acid residue corresponding to position 616 is leucine or glycine; the amino acid residue corresponding to position 623 is leucine or glutamine; the amino acid residue corresponding to position 636 is leucine or aspartic acid; the amino acid residue corresponding to position 704 is phenylalanine or alanine; the amino acid residue corresponding to position 727 is leucine, glycine, or proline; the amino acid residue corresponding to position 816 is leucine or aspartic acid; the amino acid residue corresponding to position 1016 is tyrosine, glycine, or lysine; the amino acid residue corresponding to position 1245 is leucine or glycine; the amino acid residue corresponding to position 1273 is isoleucine or glutamine; the amino acid residue corresponding to position 1282 is leucine, alanine, or glutamic acid; and the amino acid residue corresponding to position 1294 is ty rosine or glutamine.
[0242] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 28 is proline or leucine; the amino acid residue corresponding to position 237 is leucine or cysteine; the amino acid residue corresponding to position 286 is tyrosine or glutamine; the amino acid residue corresponding to position 318 is serine, histidine, or cysteine; the amino acid residue corresponding to position 368 is serine or cysteine; the amino acid residue corresponding to position 498 is phenylalanine or threonine; the amino acid residue corresponding to position 514 is leucine, glycine, or threonine; the amino acid residue corresponding to position 616 is leucine or glycine; the amino acid residue corresponding to position 623 is leucine or glutamine; the amino acid residue corresponding to position 636 is leucine or aspartic acid; the amino acid residue corresponding to position 704 is phenylalanine or alanine; the amino acid residue corresponding to position 727 is leucine, glycine, or proline; the amino acid residue corresponding to position 816 is leucine or aspartic acid; the amino acid residue corresponding to position 1016 is tyrosine, glycine, or lysine; the amino acid residue corresponding to position 1245 is leucine or glycine; the amino acid residue corresponding to position 1273 is isoleucine or glutamine; the amino acid residue corresponding to position 1282 is leucine, alanine, or glutamic acid; or the amino acid residue corresponding to position 1294 is tyrosine or glutamine.
[0243] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 28 is proline or leucine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 28 is proline. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 28 is leucine.
[0244] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 237 is leucine or cysteine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 237 is leucine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 237 is cysteine.
[0245] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 286 is tyrosine or glutamine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33. wherein the amino acid residue corresponding to position 286 is tyrosine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 286 is glutamine.
[0246] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 318 is serine, histidine, or cysteine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 318 is serine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 318 is histidine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 318 is cysteine.
[0247] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 368 is serine or cysteine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 368 is serine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 368 is cysteine.
[0248] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 498 is phenylalanine or threonine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 498 is phenylalanine. In embodiments, the DNA endonuclease enzy me includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 498 is threonine.
[0249] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 514 is leucine, glycine, or threonine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 514 is leucine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 514 is glycine. In embodiments, the DNA endonuclease enzy me includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 514 is threonine.
[0250] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 616 is leucine or glycine. In embodiments, the DNA endonuclease enzy me includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 616 is leucine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 616 is glycine.
[0251] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 623 is leucine or glutamine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 623 is leucine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 623 is glutamine.
[0252] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 636 is leucine or aspartic acid. In embodiments, the DNA endonuclease enzy me includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 636 is leucine. In embodiments, the DNA endonuclease enzy me includes the amino acid sequence of SEQ ID NO:33. wherein the amino acid residue corresponding to position 636 is aspartic acid.
[0253] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 704 is phenylalanine or alanine. In embodiments, the DNA endonuclease enzy me includes the amino acid sequence of SEQ ID NO:33. wherein the amino acid residue corresponding to position 704 is phenylalanine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 704 is alanine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 727 is leucine, glycine, or proline.
[0254] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 816 is leucine or aspartic acid. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33. wherein the amino acid residue corresponding to position 816 is leucine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 816 is aspartic acid.
[0255] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33. wherein the amino acid residue corresponding to position 1016 is tyrosine, glycine, or lysine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 1016 is ty rosine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 1016 is glycine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 1016 is lysine.
[0256] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 1245 is leucine or glycine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 1245 is leucine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 1245 is glycine.
[0257] In embodiments, the DNA endonuclease enzy me includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 1273 is isoleucine or glutamine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 1273 is isoleucine. In embodiments, the DNA endonuclease enzy me includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 1273 is glutamine.
[0258] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 1282 is leucine, alanine, or glutamic acid. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 1282 is leucine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 1282 is alanine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 1282 is glutamic acid.
[0259] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 1294 is tyrosine or glutamine. In embodiments, the DNA endonuclease enzy me includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 1294 is tyrosine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 1294 is glutamine.
[0260] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:34. In embodiments, the DNA endonuclease enzyme is the amino acid sequence of SEQ ID NO:34. In one further embodiment, the DNA endonuclease enzyme is referred to herein as Cas9V4.
[0261] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:35. In embodiments, the DNA endonuclease enzy me is the amino acid sequence of SEQ ID NO:35. In one further embodiment, the DNA endonuclease enzy me is referred to herein as Cas9Vl.
[0262] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:36. In embodiments, the DNA endonuclease enzyme is the amino acid sequence of SEQ ID NO:36. In one further embodiment, the DNA endonuclease enzyme is referred to herein as Cas9V2.
[0263] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:37. In embodiments, the DNA endonuclease enzyme is the amino acid sequence of SEQ ID NO:37. In one further embodiment, the DNA endonuclease enzyme is referred to herein as Cas9V3.
[0264] In embodiments, the DNA endonuclease enzy me includes the amino acid sequence of SEQ ID NO:38. In embodiments, the DNA endonuclease enzyme is the amino acid sequence of SEQ ID NO:38. In one further embodiment, the DNA endonuclease enzyme is referred to herein as Cas9V5.
[0265] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:39. In embodiments, the DNA endonuclease enzy me is the amino acid sequence of SEQ ID NO:39. In one further embodiment, the DNA endonuclease enzyme is referred to herein as Cas9V6.
[0266] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:40. In embodiments, the DNA endonuclease enzyme is the amino acid sequence of SEQ ID NO:40. In one further embodiment, the DNA endonuclease enzy me is referred to herein as Cas9V7.
[0267] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:41. In embodiments, the DNA endonuclease enzyme is the amino acid sequence of SEQ ID NO:41. In one further embodiment, the DNA endonuclease enzyme is referred to herein as Cas9V8.
[0268] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:42. In embodiments, the DNA endonuclease enzy me is the amino acid sequence of SEQ ID NO:42. In one further embodiment, the DNA endonuclease enzy me is referred to herein as Cas9V9.
[0269] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:43. In embodiments, the DNA endonuclease enzyme is the amino acid sequence of SEQ ID NO:43. In one further embodiment, the DNA endonuclease enzyme is referred to herein as Cas9V10.
[0270] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:44. In embodiments, the DNA endonuclease enzyme is the amino acid sequence of SEQ ID NO:44. In one further embodiment, the DNA endonuclease enzyme is referred to herein as Cas9Vl 1.
[0271] In embodiments, the DNA endonuclease enzy me includes the amino acid sequence of SEQ ID NO:45. In embodiments, the DNA endonuclease enzyme is the amino acid sequence of SEQ ID NO:45. In one further embodiment, the DNA endonuclease enzyme is referred to herein as Cas9Vl 2.
[0272] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:46. In embodiments, the DNA endonuclease enzy me is the amino acid sequence of SEQ ID NO:46. In one further embodiment, the DNA endonuclease enzyme is referred to herein as Cas9V13.
[0273] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:47. In embodiments, the DNA endonuclease enzyme is the amino acid sequence of SEQ ID NO:47. In one further embodiment, the DNA endonuclease enzy me is referred to herein as Cas9V14.
[0274] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:48. In embodiments, the DNA endonuclease enzyme is the amino acid sequence of SEQ ID NO:48. In one further embodiment, the DNA endonuclease enzyme is referred to herein as Cas9V15.
[0275] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:49. In embodiments, the DNA endonuclease enzy me is the amino acid sequence of SEQ ID NO:49. In one further embodiment, the DNA endonuclease enzy me is referred to herein as Cas9V16.
[0276] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:50. In embodiments, the DNA endonuclease enzyme is the amino acid sequence of SEQ ID NO:50. In one further embodiment, the DNA endonuclease enzyme is referred to herein as Cas9V17.
[0277] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:51. In embodiments, the DNA endonuclease enzyme is the amino acid sequence of SEQ ID NO:51. In one further embodiment, the DNA endonuclease enzyme is referred to herein as Cas9V18.
[0278] In embodiments, the DNA endonuclease enzy me includes the amino acid sequence of SEQ ID NO:52. In embodiments, the DNA endonuclease enzyme is the amino acid sequence of SEQ ID NO:52. In one further embodiment, the DNA endonuclease enzyme is referred to herein as Cas9Vl 9.
[0279] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:53. In embodiments, the DNA endonuclease enzy me is the amino acid sequence of SEQ ID NO:53. In one further embodiment, the DNA endonuclease enzyme is referred to herein as Cas9V20.
[0280] In embodiments, the DNA endonuclease enzyme includes an amino acid sequence having 95% sequence identity to SEQ ID NO:34. In embodiments, the DNA endonuclease enzy me an amino acid sequence having 95% sequence identity7 to SEQ ID NO:35. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 95% sequence identity to SEQ ID NO:36. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 95% sequence identity to SEQ ID NO:37. In embodiments, the DNA endonuclease enzy me an amino acid sequence having 95% sequence identity7 to SEQ ID NO:38. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 95% sequence identity to SEQ ID NO:39. In embodiments, the DNA endonuclease enzyme 111 an amino acid sequence having 95% sequence identity to SEQ ID NO:40. In embodiments, the DNA endonuclease enzy me an amino acid sequence having 95% sequence identity' to SEQ ID NO: 1. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 95% sequence identity to SEQ ID NO:42. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 95% sequence identity to SEQ ID NO:43. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 95% sequence identity to SEQ ID NO:44. In embodiments, the DNA endonuclease enzy me an amino acid sequence having 95% sequence identity to SEQ ID NO:45. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 95% sequence identity to SEQ ID NO:46. In embodiments, the DNA endonuclease enzy me an amino acid sequence having 95% sequence identity7 to SEQ ID NO:47. In embodiments, the DNA endonuclease enzy me an amino acid sequence having 95% sequence identity7 to SEQ ID NO:48. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 95% sequence identity to SEQ ID NO:49. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 95% sequence identity to SEQ ID NO:50. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 95% sequence identity to SEQ ID N0:51. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 95% sequence identity7 to SEQ ID NO:52. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 95% sequence identity to SEQ ID NO:53.
[0281] In embodiments, the DNA endonuclease enzyme includes an amino acid sequence having 96% sequence identity to SEQ ID NO:34. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 96% sequence identity to SEQ ID NO:35. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 96% sequence identity7 to SEQ ID NO:36. In embodiments, the DNA endonuclease enzy me an amino acid sequence having 96% sequence identity to SEQ ID NO:37. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 96% sequence identity to SEQ ID NO:38. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 96% sequence identity to SEQ ID NO:39. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 96% sequence identity7 to SEQ ID NO:40. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 96% sequence identity to SEQ ID NO: 1. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 96% sequence identity to SEQ ID NO:42. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 96% sequence identity to SEQ ID NO:43. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 96% sequence identity to SEQ ID NO:44. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 96% sequence identity to SEQ ID NO:45. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 96% sequence identity to SEQ ID NO:46. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 96% sequence identity to SEQ ID NO:47. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 96% sequence identity to SEQ ID NO:48. In embodiments, the DNA endonuclease enzy me an amino acid sequence having 96% sequence identity to SEQ ID NO:49. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 96% sequence identity to SEQ ID NO:50. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 96% sequence identity to SEQ ID N0:51. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 96% sequence identity to SEQ ID NO: 52. In embodiments, the DNA endonuclease enzy me an amino acid sequence having 96% sequence identity to SEQ ID NO:53.
[0282] In embodiments, the DNA endonuclease enzyme includes an amino acid sequence having 97% sequence identity to SEQ ID NO:34. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 97% sequence identity to SEQ ID NO:35. In embodiments, the DNA endonuclease enzy me an amino acid sequence having 97% sequence identity to SEQ ID NO:36. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 97% sequence identity to SEQ ID NO:37. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 97% sequence identity to SEQ ID NO:38. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 97% sequence identity to SEQ ID NO:39. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 97% sequence identity to SEQ ID NO:40. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 97% sequence identity to SEQ ID NO:1. In embodiments, the DNA endonuclease enzy me an amino acid sequence having 97% sequence identity to SEQ ID NO:42. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 97% sequence identity to SEQ ID NO:43. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 97% sequence identity to SEQ ID NO:44. In embodiments, the DNA endonuclease enzy me an amino acid sequence having 97% sequence identity to SEQ ID NO:45. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 97% sequence identity to SEQ ID NO:46. In embodiments, the DNA endonuclease enzy me an amino acid sequence having 97% sequence identity to SEQ ID NO:47. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 97% sequence identity’ to SEQ ID NO:48. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 97% sequence identity to SEQ ID NO:49. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 97% sequence identity’ to SEQ ID NO:50. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 97% sequence identity to SEQ ID N0:51. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 97% sequence identity7 to SEQ ID NO:52. In embodiments, the DNA endonuclease enzy me an amino acid sequence having 97% sequence identity’ to SEQ ID NO:53.
[0283] In embodiments, the DNA endonuclease enzy me includes an amino acid sequence having 98% sequence identity to SEQ ID NO:34. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 98% sequence identity to SEQ ID NO:35. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 98% sequence identity7 to SEQ ID NO:36. In embodiments, the DNA endonuclease enzy me an amino acid sequence having 98% sequence identity to SEQ ID NO:37. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 98% sequence identity to SEQ ID NO:38. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 98% sequence identity’ to SEQ ID NO:39. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 98% sequence identity' to SEQ ID NO:40. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 98% sequence identity to SEQ ID NO: 1. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 98% sequence identity to SEQ ID NO:42. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 98% sequence identity to SEQ ID NO:43. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 98% sequence identity to SEQ ID NO:44. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 98% sequence identity to SEQ ID NO:45. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 98% sequence identity to SEQ ID NO:46. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 98% sequence identity to SEQ ID NO:47. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 98% sequence identity to SEQ ID NO:48. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 98% sequence identity to SEQ ID NO:49. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 98% sequence identity to SEQ ID NO:50. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 98% sequence identity to SEQ ID N0:51. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 98% sequence identity to SEQ ID NO: 52. In embodiments, the DNA endonuclease enzy me an amino acid sequence having 98% sequence identity to SEQ ID NO:53.
[0284] In embodiments, the DNA endonuclease enzyme includes an amino acid sequence having 99% sequence identity to SEQ ID NO:34. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 99% sequence identity to SEQ ID NO:35. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 99% sequence identity to SEQ ID NO:36. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 99% sequence identity to SEQ ID NO:37. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 99% sequence identity to SEQ ID NO:38. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 99% sequence identity to SEQ ID NO:39. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 99% sequence identity to SEQ ID NO:40. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 99% sequence identity7 to SEQ ID NO: 1. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 99% sequence identity to SEQ ID NO:42. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 99% sequence identity to SEQ ID NO:43. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 99% sequence identity to SEQ ID NO:44. In embodiments, the DNA endonuclease enzy me an amino acid sequence having 99% sequence identity to SEQ ID NO:45. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 99% sequence identity to SEQ ID NO:46. In embodiments, the DNA endonuclease enzy me an amino acid sequence having 99% sequence identity to SEQ ID NO:47. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 99% sequence identity to SEQ ID NO:48. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 99% sequence identity to SEQ ID NO:49. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 99% sequence identity to SEQ ID NO:50. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 99% sequence identity to SEQ ID N0:51. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 99% sequence identity to SEQ ID NO:52. In embodiments, the DNA endonuclease enzy me an amino acid sequence having 99% sequence identity7 to SEQ ID NO:53.
[0285] In embodiments, the DNA endonuclease enzyme is a deimmunized DNA endonuclease enzyme. In embodiments, the DNA endonuclease enzyme does not produce or provoke an immune or inflammatory' response relative to a non-deimmunized DNA endonuclease enzyme. In embodiments, the DNA endonuclease enzyme does not increase secretion of an immune cytokine relative to a non-deimmunized DNA endonuclease enzy me. In embodiments, the DNA endonuclease enzyme does not increase IFNy secretion. IFNB secretion, RIG-I secretion, or IL6 secretion relative to anon-deimmunized DNA endonuclease enzyme. In embodiments, the DNA endonuclease enzyme does not increase IFNy secretion relative to a non-deimmunized DNA endonuclease enzy me. In embodiments, the DNA endonuclease enzyme does not increase IFNB secretion relative to a nondeimmunized DNA endonuclease enzyme. In embodiments, the DNA endonuclease enzyme does not increase RIG-I secretion relative to anon-deimmunized DNA endonuclease enzyme. In embodiments, the DNA endonuclease enzyme does not increase IL6 secretion relative to a non-deimmunized DNA endonuclease enzyme.
[0286] In embodiments, the DNA endonuclease enzyme provided herein including embodiments thereof is useful in a method of editing a target gene. ZINC FINGER COMPOSITIONS
[0287] The compositions and methods provided herein include zinc finger protein compositions including novel alpha helix domain sequences and methods of for designing, generating, and / or using the same. The zinc finger protein compositions provided herein including embodiments thereof may be used, inter alia, in genetic engineering methods. Thus, in an aspect is provided a poly dactyl zinc finger protein including a first alpha helix domain including the amino acid sequence of SEQ ID NO: 55 a second alpha helix domain including the amino acid sequence of SEQ ID NO:56, a third alpha helix domain including the amino acid sequence of SEQ ID NO:57, a fourth alpha helix domain including the amino acid sequence of SEQ ID NO: 58, a fifth alpha helix domain including the amino acid sequence of SEQ ID NO:59, and a sixth alpha helix domain including the amino acid sequence of SEQ ID NO:60.
[0288] In embodiments, the first alpha helix domain includes a first alpha helix including the amino acid sequence of SEQ ID NO:55. In embodiments, the second alpha helix domain includes a second alpha helix including the amino acid sequence of SEQ ID NO:56. In embodiments, the third alpha helix domain includes a third alpha helix including the amino acid sequence of SEQ ID NO: 57. In embodiments, the fourth alpha helix domain includes a fourth alpha helix including the amino acid sequence of SEQ ID NO:58. In embodiments, the fifth alpha helix domain includes a fifth alpha helix including the amino acid sequence of SEQ ID NO:59. In embodiments, the sixth alpha helix domain includes a sixth alpha helix including the amino acid sequence of SEQ ID NO:60.
[0289] In embodiments, the poly dactyl zinc finger protein includes from N-terminus to C-terminus: the first alpha helix domain, the second alpha helix domain, the third alpha helix domain, the fourth alpha helix domain, the fifth alpha helix domain, and the sixth alpha helix domain.
[0290] In embodiments, the first alpha helix domain is bound to the second alpha helix domain. In embodiments, the second alpha helix domain is bound to the third alpha helix domain. In embodiments, the third alpha helix domain is bound to the fourth alpha helix domain. In embodiments, the fourth alpha helix domain is bound to the fifth alpha helix domain. In embodiments, the fifth alpha helix domain is bound to the sixth alpha helix. In embodiments, the first alpha helix domain is bound to the second alpha helix domain; the second alpha helix domain is bound to the third alpha helix domain; the third alpha helix domain is bound to the fourth alpha helix domain: the fourth alpha helix domain is bound to the fifth alpha helix domain; and the fifth alpha helix domain is bound to the sixth alpha helix.
[0291] In embodiments, the first alpha helix domain, the second alpha helix domain, the third alpha helix domain, the fourth alpha helix domain, the fifth alpha helix domain, and the sixth alpha helix domain are bound together. In embodiments, the first alpha helix domain, the second alpha helix domain, the third alpha helix domain, the fourth alpha helix domain, the fifth alpha helix domain, and the sixth alpha helix domain are bound together in a single, continuous amino acid sequence.
[0292] In embodiments, the poly dactyl zinc finger protein includes an amino acid sequence having at least 80% sequence identity to the sequence of SEQ ID NO:61. In embodiments, the poly dactyl zinc finger protein includes an amino acid sequence having at least 85% sequence identity to the sequence of SEQ ID NO:61. In embodiments, the poly dactyl zinc finger protein includes an amino acid sequence having at least 90% sequence identity to the sequence of SEQ ID NO:61. In embodiments, the poly dactyl zinc finger protein includes an amino acid sequence having at least 95% sequence identity to the sequence of SEQ ID NO:61. In embodiments, the poly dactyl zinc finger protein includes an amino acid sequence having at least 96% sequence identity to the sequence of SEQ ID NO:61. In embodiments, the poly dactyl zinc finger protein includes an amino acid sequence having at least 97% sequence identity to the sequence of SEQ ID NO:61. In embodiments, the poly dactyl zinc finger protein includes an amino acid sequence having at least 98% sequence identity to the sequence of SEQ ID NO:61. In embodiments, the poly dactyl zinc finger protein includes an amino acid sequence having at least 99% sequence identity to the sequence of SEQ ID NO:61. In embodiments, the poly dactyl zinc finger protein includes an amino acid sequence having at least 100% sequence identity to the sequence of SEQ ID NO:61.
[0293] In embodiments, the poly dactyl zinc finger protein is an amino acid sequence having at least 80% sequence identity to the sequence of SEQ ID NO:61. In embodiments, the poly dactyl zinc finger protein is an amino acid sequence having at least 85% sequence identity to the sequence of SEQ ID NO:61. In embodiments, the polydactyl zinc finger protein is an amino acid sequence having at least 90% sequence identity to the sequence of SEQ ID NO:61. In embodiments, the poly dactyl zinc finger protein is an amino acid sequence having at least 95% sequence identity to the sequence of SEQ ID NO:61. In embodiments, the polydactyl zinc finger protein is an amino acid sequence having at least 96% sequence identity to the sequence of SEQ ID NO:61. In embodiments, the polydactyl zinc finger protein is an amino acid sequence having at least 97% sequence identity to the sequence of SEQ ID NO:61. In embodiments, the polydactyl zinc finger protein is an amino acid sequence having at least 98% sequence identity to the sequence of SEQ ID NO:61. In embodiments, the polydactyl zinc finger protein is an amino acid sequence having at least 99% sequence identity to the sequence of SEQ ID NO:61. In embodiments, the polydactyl zinc finger protein is an amino acid sequence having at least 100% sequence identity to the sequence of SEQ ID NO:61.
[0294] In embodiments, the polydactyl zinc finger protein includes the amino acid sequence of SEQ ID NO:61. In embodiments, the polydactyl zinc finger protein is the amino acid sequence of SEQ ID NO:61.
[0295] In embodiments, the polydacty l zinc finger protein is capable of binding to a nucleic acid encoding a PCSK9 gene. In embodiments, the nucleic acid is DNA. In embodiments, the nucleic acid is RNA. In embodiments, the PCSK9 gene is a human PCSK9 gene.
[0296] In embodiments, the polydactyl zinc finger protein provided herein including embodiments thereof is useful in a method of editing a target nucleic acid. In embodiments, the target nucleic acid encodes a PCSK9 gene. In embodiments the PCSK9 gene is a human PCSK9 gene. METHODS
[0297] The compositions provided herein including embodiments thereof may be used, inter alia, to circularize an RNA molecule in situ or in vitro. The inventors surprisingly found that large protein-encoding nucleic acid sequences (e.g., greater than 750 nucleotides) could be efficiently circularized using the methods and compositions provided herein including embodiments thereof. The circularized RNA molecules including a protein-encoding nucleic acid sequence have increased expression in cells (e.g., cardiomyocytes and neurons) and increased RNA persistence. The methods provided herein including embodiments thereof allowed for efficient generation of circularized RNA and delivery of large constructs (e.g., protein-encoding nucleic acid sequences). Thus, in an aspect is provided a method of forming a circularized ribonucleic acid (RNA) in a cell, the method including transfecting a cell with a circularizable linear RNA compound that is capable of circularizing within the cell, thereby forming a circularized RNA, wherein the linear RNA compound is a nucleic acid including from 5' to 3': a first member of a split ligation stem, an internal ribosome entry site (IRES) domain, a 5' untranslated region (UTR), a protein-encoding nucleic acid sequence at least 750 nucleotides in length, a 3' UTR, a poly-adenine (poly A) domain, and a second member of the split ligation stem.
[0298] In embodiments, the method further includes allowing the cell to translate the protein-encoding nucleic acid sequence, thereby forming a protein within the cell
[0299] In embodiments, the method includes cleaving a ribozyme-cleavable linear RNA compound to form the circularizable linear RNA compound, wherein the ribozyme-cleavable linear RNA compound includes from 5' to 3': a first ribozyme domain, an internal ribosome entry site (IRES) domain, a 5' untranslated region (UTR), a protein-encoding nucleic acid sequence at least 750 nucleotides in length, a 3' UTR, a poly-adenine (polyA) domain, and a second ribozy me domain, wherein the cleaving includes allowing the first ribozyme domain to cleave the ribozyme-cleavable linear RNA compound in the 3' direction, thereby forming a 5' end including the first member of the split ligation stem; and the second ribozyme domain to cleave the ribozyme-cleavable linear RNA compound in the 5' direction, thereby forming a 3' end including the second member of the split ligation stem.
[0300] In embodiments, the first ribozyme domain is a first twister ribozyme domain and the second ribozyme domain is a second twister ribozyme domain.
[0301] In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently are a Sma-1-402 twister ribozyme domain, an env-112 twister ribozyme domain, a Slyc-2-1 twister ribozyme domain, an Osin-1-1 twister ribozyme domain, an env-270 twister ribozyme domain, an env-94 twister ribozyme domain, an env-935 twister ribozyme domain, a Sma-1-66 twister ribozyme domain, a Dre-1-3 twister ribozy me domain, an Osa-1-4 twister ribozyme domain, an Eana-1-1 twister ribozy me domain, an Osa-1-8 twister ribozyme domain, an Osa-1-3 twister ribozyme domain, an env-13 twister ribozyme domain, a Dre-1-4 twister ribozyme domain, an Aage-1-1 twister ribozyme domain, a Spol-1-1 twister ribozyme domain, aTtru-1-1 twister ribozyme domain, an Hsap-1-1 twister ribozyme domain, or an Hsap-1-2 twister ribozy me domain. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently are a Sma-1-402 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently are an env-112 twister ribozyme domain. In embodiments, the first twister ribozy me domain and the second twister ribozy me domain independently are a Slyc-2-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently are an Osin-1-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently are an env-270 twister ribozy me domain. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently are an env-94 twister ribozyme domain. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently are an env-935 twister ribozyme domain. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently are a Sma-1-66 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently are a Dre-1-3 twister ribozyme domain. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently are an Osa-1-4 twister ribozy me domain. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently are an Eana-1-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently are an Osa-1-8 twister ribozy me domain. In embodiments, the first twister ribozy me domain and the second twister ribozy me domain independently are an Osa-1-3 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently are an env-13 twister ribozyme domain. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently are a Dre-1-4 twister ribozyme domain. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently are an Aage-1-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently are a Spol-1-1 twister ribozyme domain. In embodiments, the first twister ribozy me domain and the second twister ribozy me domain independently are a Ttru-1-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently are an Hsap-1-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently are an Hsap-1-2 twister ribozyme domain.
[0302] In embodiments, the first twister ribozyme domain and the second twister ribozyme domain are a Sma-1-402 twister ribozyme domain, an env-112 twister ribozyme domain, a Slyc-2-1 twister ribozyme domain, an Osin-1-1 twister ribozy me domain, an env-270 twister ribozyme domain, an env-94 twister ribozyme domain, an env-935 twister ribozy me domain, a Sma-1-66 twister ribozyme domain, a Dre-1-3 twister ribozy me domain, an Osa-1-4 twister ribozyme domain, an Eana-1-1 twister ribozyme domain, an Osa-1-8 twister ribozy me domain, an Osa-1-3 twister ribozyme domain, an env-13 twister ribozyme domain, a Dre-1-4 twister ribozyme domain, an Aage-1-1 twister ribozyme domain, a Spol-1-1 twister ribozyme domain, a Ttru-1-1 twister ribozyme domain, an Hsap-1-1 twister ribozyme domain, or an Hsap-1-2 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain are a Sma-1-402 twister ribozyme domain. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain are an env-112 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain are a Slyc-2-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain are an Osin-1-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second tw ister ribozyme domain are an env-270 twister ribozy me domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain are an env-94 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozy me domain are an env-935 twister ribozyme domain. In embodiments, the first twister ribozy me domain and the second twister ribozy me domain are a Sma-1-66 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain are a Dre-1-3 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain are an Osa-1-4 twister ribozyme domain. In embodiments, the first twister ribozy me domain and the second twister ribozy me domain are an Eana-1-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain are an Osa-1-8 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain are an Osa-1-3 twister ribozyme domain. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain are an env-13 twister ribozy me domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain are a Dre-1-4 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain are an Aage-1-1 twister ribozyme domain. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain are a Spol-1-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain are a Ttru-1-1 twister ribozyme domain. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain are an Hsap-1-1 twister ribozyme domain. In embodiments, the first twister ribozy me domain and the second twister ribozy me domain are an Hsap-1-2 twister ribozyme domain.
[0303] In embodiments, the first twister ribozyme domain or the second twister ribozyme domain is a Sma-1-402 twister ribozyme domain, an env-112 twister ribozyme domain, a Slyc-2-1 twister ribozyme domain, an Osin-1-1 twister ribozy me domain, an env-270 twister ribozyme domain, an env-94 twister ribozyme domain, an env-935 twister ribozy me domain, a Sma-1-66 twister ribozyme domain, a Dre-1-3 twister ribozy me domain, an Osa-1-4 twister ribozyme domain, an Eana-1 -1 twister ribozyme domain, an Osa-1 -8 twister ribozyme domain, an Osa-1-3 twister ribozy me domain, an env-13 twister ribozyme domain, a Dre-1-4 twister ribozy me domain, an Aage-1-1 twister ribozyme domain, a Spol-1-1 twister ribozyme domain, a Ttru-1-1 twister ribozyme domain, an Hsap-1-1 twister ribozyme domain, or an Hsap-1-2 twister ribozyme domain. In embodiments, the first twister ribozyme domain or the second twister ribozy me domain is a Sma-1-402 twister ribozyme domain. In embodiments, the first twister ribozy me domain or the second twister ribozyme domain is an env-112 twister ribozyme domain. In embodiments, the first twister ribozyme domain or the second twister ribozyme domain is a Slyc-2-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain or the second twister ribozyme domain is an Osin-1-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain or the second twister ribozyme domain is an env-270 twister ribozyme domain. In embodiments, the first twister ribozyme domain or the second twister ribozyme domain is an env-94 twister ribozyme domain. In embodiments, the first twister ribozyme domain or the second twister ribozyme domain is an env-935 twister ribozyme domain. In embodiments, the first twister ribozyme domain or the second twister ribozyme domain is a Sma-1-66 twister ribozyme domain. In embodiments, the first twister ribozyme domain or the second twister ribozyme domain is a Dre-1-3 twister ribozyme domain. In embodiments, the first twister ribozyme domain or the second twister ribozy me domain is an Osa-1-4 twister ribozyme domain. In embodiments, the first twister ribozyme domain or the second twister ribozyme domain is an Eana-1-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain or the second twister ribozyme domain is an Osa-1-8 twister ribozyme domain. In embodiments, the first twister ribozyme domain or the second twister ribozyme domain is an Osa-1-3 twister ribozyme domain. In embodiments, the first twister ribozy me domain or the second twister ribozyme domain is an env-13 twister ribozy me domain. In embodiments, the first twister ribozyme domain or the second twister ribozyme domain is a Dre-1-4 twister ribozyme domain. In embodiments, the first twister ribozy me domain or the second twister ribozyme domain is an Aage-1-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain or the second twister ribozy me domain is a Spol-1-1 twister ribozy me domain. In embodiments, the first twister ribozyme domain or the second twister ribozyme domain is a Ttru-1-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain or the second twister ribozyme domain is an Hsap-1-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain or the second twister ribozy me domain is an Hsap-1-2 twister ribozy me domain.
[0304] In embodiments, the first twister ribozyme domain is a Sma-1-402 twister ribozyme domain, an env-112 twister ribozyme domain, a Slyc-2-1 twister ribozyme domain, an Osin-1-1 twister ribozyme domain, an env-270 twister ribozyme domain, an env-94 twister ribozyme domain, an env-935 twister ribozy me domain, a Sma-1-66 twister ribozy me domain, a Dre-1-3 twister ribozyme domain, an Osa-1-4 twister ribozyme domain, an Eana-1-1 twister ribozyme domain, an Osa-1-8 twister ribozyme domain, an Osa-1-3 twister ribozyme domain, an env-13 twister ribozyme domain, a Dre-1-4 twister ribozyme domain, an Aage-1-1 twister ribozy me domain, a Spol-1-1 twister ribozyme domain, a Ttru-1-1 twister ribozyme domain, an Hsap-1-1 twister ribozy me domain, or an Hsap-1-2 twister ribozyme domain. In embodiments, the first twister ribozyme domain is a Sma-1-402 twister ribozyme domain. In embodiments, the first twister ribozyme domain is an env-112 twister ribozyme domain. In embodiments, the first twister ribozyme domain is a Slyc-2-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain is an Osin-1-1 twister ribozyme domain. In embodiments, the first twister ribozy me domain is an env-270 twister ribozyme domain. In embodiments, the first twister ribozyme domain is an env-94 twister ribozyme domain. In embodiments, the first twister ribozyme domain is an env-935 twister ribozyme domain. In embodiments, the first twister ribozyme domain is a Sma-1-66 twister ribozy me domain. In embodiments, the first twister ribozy me domain is a Dre-1-3 twister ribozyme domain. In embodiments, the first twister ribozyme domain is an Osa-1-4 twister ribozyme domain. In embodiments, the first twister ribozyme domain is an Eana-1-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain is an Osa-1-8 twister ribozy me domain. In embodiments, the first twister ribozy me domain is an Osa-1-3 twister ribozyme domain. In embodiments, the first twister ribozyme domain is an env-13 twister ribozyme domain. In embodiments, the first twister ribozyme domain is a Dre-1-4 twister ribozyme domain. In embodiments, the first twister ribozyme domain is an Aage-1-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain is a Spol-1-1 twister ribozy me domain. In embodiments, the first twister ribozy me domain is a Ttru-1-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain is an Hsap-1 -1 twister ribozyme domain. In embodiments, the first twister ribozyme domain is an Hsap-1-2 twister ribozy me domain.
[0305] In embodiments, the second twister ribozyme domain is a Sma-1-402 twister ribozyme domain, an env-112 twister ribozyme domain, a Slyc-2-1 twister ribozyme domain, an Osin-1-1 twister ribozyme domain, an env-270 twister ribozyme domain, an env-94 twister ribozyme domain, an env-935 twister ribozy me domain, a Sma-1-66 twister ribozyme domain, a Dre-1-3 twister ribozyme domain, an Osa-1-4 twister ribozyme domain, an Eana-1-1 twister ribozyme domain, an Osa-1-8 twister ribozyme domain, an Osa-1-3 twister ribozyme domain, an env-13 twister ribozyme domain, a Dre-1-4 twister ribozyme domain, an Aage-1-1 twister ribozy me domain, a Spol-1-1 twister ribozyme domain, a Ttru-1-1 twister ribozy me domain, an Hsap-1-1 twister ribozy me domain, or an Hsap-1-2 twister ribozyme domain. . In embodiments, the second twister ribozyme domain is a Sma-1-402 twister ribozyme domain. In embodiments, the second twister ribozyme domain is an env-112 twister ribozyme domain. In embodiments, the second twister ribozyme domain is a Slyc-2-1 twister ribozyme domain. In embodiments, the second twister ribozyme domain is an Osin-1 -1 twister ribozy me domain. In embodiments, the second twister ribozyme domain is an env-270 twister ribozyme domain. In embodiments, the second twister ribozyme domain is an env-94 twister ribozy me domain. In embodiments, the second twister ribozyme domain is an env-935 twister ribozyme domain. In embodiments, the second twister ribozyme domain is a Sma-1-66 twister ribozyme domain. In embodiments, the second twister ribozyme domain is a Dre-1-3 twister ribozyme domain. In embodiments, the second twister ribozyme domain is an Osa-1-4 twister ribozyme domain. In embodiments, the second twister ribozyme domain is an Eana-1-1 twister ribozyme domain. In embodiments, the second twister ribozyme domain is an Osa-1-8 twister ribozy me domain. In embodiments, the second twister ribozy me domain is an Osa-1-3 twister ribozyme domain. In embodiments, the second twister ribozyme domain is an env-13 twister ribozyme domain. In embodiments, the second twister ribozyme domain is a Dre-1-4 twister ribozyme domain. In embodiments, the second twister ribozy me domain is an Aage-1-1 twister ribozyme domain. In embodiments, the second twister ribozyme domain is a Spol-1-1 twister ribozy me domain. In embodiments, the second twister ribozyme domain is a Ttru-1 -1 twister ribozyme domain. In embodiments, the second twister ribozyme domain is an Hsap-1-1 twister ribozyme domain. In embodiments, the second twister ribozy me domain is an Hsap-1-2 twister ribozyme domain.
[0306] In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include the nucleotide sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include the nucleotide sequence of SEQ ID NO:1. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include the nucleotide sequence of SEQ ID NO:2. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include the nucleotide sequence of SEQ ID NO:3. In embodiments, the first twister ribozyme domain and the second twister ribozy me domain independently include the nucleotide sequence of SEQ ID NO:4. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include the nucleotide sequence of SEQ ID NO:5. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include the nucleotide sequence of SEQ ID NO:6. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include the nucleotide sequence of SEQ ID NO:7. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include the nucleotide sequence of SEQ ID NO:8. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include the nucleotide sequence of SEQ ID NO:9. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include the nucleotide sequence of SEQ ID NO: 10. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently7 include the nucleotide sequence of SEQ ID NO:11. In embodiments, the first twister ribozy me domain and the second twister ribozy me domain independently include the nucleotide sequence of SEQ ID NO: 12. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include the nucleotide sequence of SEQ ID NO: 13. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include the nucleotide sequence of SEQ ID NO: 14. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include the nucleotide sequence of SEQ ID NO:15. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include the nucleotide sequence of SEQ ID NO: 16. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include the nucleotide sequence of SEQ ID NO: 17. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include the nucleotide sequence of SEQ ID NO: 18. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include the nucleotide sequence of SEQ ID NO: 19. In embodiments, the first twister nbozyme domain and the second twister ribozyme domain independently include the nucleotide sequence of SEQ ID NO:20. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include the nucleotide sequence of SEQ ID N0:21. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include the nucleotide sequence of SEQ ID NO:22.
[0307] In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 70% sequence identity7 to 10, 15, 20, 25, 30, 35, 40, 45, 50, 51, 52, or 53 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 70% sequence identity to 10 contiguous nucleotides of the sequence of any one of SEQ ID NOs :1-22. In embodiments, the first twister ribozy me domain and the second twister ribozy me domain independently include a nucleotide sequence having 70% sequence identity to 15 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently7 include a nucleotide sequence having 70% sequence identity7 to 20 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 70% sequence identity to 25 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozy me domain independently include a nucleotide sequence having 70% sequence identity to 30 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1 -22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include a nucleotide sequence having 70% sequence identity to 35 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 70% sequence identity to 40 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 70% sequence identity to 45 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 70% sequence identity to 50 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 70% sequence identity' to 51 contiguous nucleotides of the sequence of any one of SEQ ID NOs :1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 70% sequence identity to 52 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 70% sequence identity to 53 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22.
[0308] In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 75% sequence identity to 10, 15, 20, 25, 30, 35, 40, 45, 50, 51, 52, or 53 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 75% sequence identity to 10 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 75% sequence identity to 15 contiguous nucleotides of the sequence of any one of SEQ ID NOs :1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 75% sequence identity7 to 20 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 75% sequence identity to 25 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozy me domain independently include a nucleotide sequence having 75% sequence identity7 to 30 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister nbozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 75% sequence identity7 to 35 contiguous nucleotides of the sequence of any one of SEQ ID NOs :1-22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include a nucleotide sequence having 75% sequence identity to 40 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 75% sequence identity to 45 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 75% sequence identity to 50 contiguous nucleotides of the sequence of any7 one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozy me domain independently include a nucleotide sequence having 75% sequence identity to 51 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently7 include a nucleotide sequence having 75% sequence identity7 to 52 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 75% sequence identity to 53 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22.
[0309] In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 80% sequence identity7 to 10, 15, 20, 25, 30, 35, 40, 45, 50, 51, 52, or 53 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozy me domain independently include a nucleotide sequence having 80% sequence identity to 10 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 80% sequence identity7 to 15 contiguous nucleotides of the sequence of any one of SEQ ID NOs :1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 80% sequence identity to 20 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 80% sequence identity to 25 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 80% sequence identity to 30 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozy me domain independently include a nucleotide sequence having 80% sequence identity to 35 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently7 include a nucleotide sequence having 80% sequence identity7 to 40 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 80% sequence identity to 45 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozy me domain independently include a nucleotide sequence having 80% sequence identity to 50 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include a nucleotide sequence having 80% sequence identity7 to 51 contiguous nucleotides of the sequence of any one of SEQ ID NOs :1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 80% sequence identity to 52 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1 -22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 80% sequence identity to 53 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22.
[0310] In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 85% sequence identity to 10, 15, 20, 25, 30, 35, 40, 45, 50, 51, 52, or 53 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 85% sequence identity to 10 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include a nucleotide sequence having 85% sequence identity to 15 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 85% sequence identity to 20 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include a nucleotide sequence having 85% sequence identity to 25 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 85% sequence identity to 30 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 85% sequence identity7 to 35 contiguous nucleotides of the sequence of any7 one of SEQ ID NOs :1-22. In embodiments, the first twister ribozy me domain and the second twister ribozy me domain independently include a nucleotide sequence having 85% sequence identity to 40 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 85% sequence identity to 45 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 85% sequence identity to 50 contiguous nucleotides of the sequence of any one of SEQ ID NOs:l-22. In embodiments, the first twister ribozy me domain and the second twister ribozy me domain independently include a nucleotide sequence having 85% sequence identity to 51 contiguous nucleotides of the sequence of any one of SEQ ID NOs:l-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 85% sequence identity to 52 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 85% sequence identity to 53 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22.
[0311] In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 90% sequence identity to 10, 15, 20. 25. 30, 35, 40, 45, 50, 51, 52, or 53 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 90% sequence identity to 10 contiguous nucleotides of the sequence of any one of SEQ ID NOs :1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 90% sequence identity' to 15 contiguous nucleotides of the sequence of any one of SEQ ID NOs :1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 90% sequence identity to 20 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 90% sequence identity to 25 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 90% sequence identity to 30 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 90% sequence identity to 35 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1 -22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 90% sequence identity' to 40 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 90% sequence identity to 45 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 90% sequence identity to 50 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include a nucleotide sequence having 90% sequence identity' to 51 contiguous nucleotides of the sequence of any one of SEQ ID NOs :1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 90% sequence identity to 52 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozy me domain independently include a nucleotide sequence having 90% sequence identity to 53 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22.
[0312] In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 95% sequence identity to 10, 15, 20, 25, 30, 35, 40, 45, 50, 51, 52, or 53 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 95% sequence identity to 10 contiguous nucleotides of the sequence of any7 one of SEQ ID NOs :1-22. In embodiments, the first twister ribozy me domain and the second twister ribozy me domain independently include a nucleotide sequence having 95% sequence identity to 15 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently7 include a nucleotide sequence having 95% sequence identity7 to 20 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 95% sequence identity to 25 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 95% sequence identity to 30 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include a nucleotide sequence having 95% sequence identity to 35 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 95% sequence identity to 40 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include a nucleotide sequence having 95% sequence identity to 45 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 95% sequence identity to 50 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 95% sequence identity to 51 contiguous nucleotides of the sequence of any one of SEQ ID NOs :1-22. In embodiments, the first twister ribozy me domain and the second twister ribozy me domain independently include a nucleotide sequence having 95% sequence identity to 52 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently' include a nucleotide sequence having 95% sequence identity to 53 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22.
[0313] In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 96% sequence identity to 10, 15, 20, 25, 30, 35, 40, 45, 50, 51, 52, or 53 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozy me domain independently include a nucleotide sequence having 96% sequence identity to 10 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1 -22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include a nucleotide sequence having 96% sequence identity to 15 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 96% sequence identity to 20 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 96% sequence identity to 25 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 96% sequence identity to 30 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 96% sequence identity' to 35 contiguous nucleotides of the sequence of any one of SEQ ID NOs :1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 96% sequence identity to 40 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 96% sequence identity to 45 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 96% sequence identity to 50 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 96% sequence identity to 51 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 96% sequence identity to 52 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 96% sequence identity to 53 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22.
[0314] In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 97% sequence identity to 10, 15, 20. 25, 30, 35, 40, 45, 50, 51, 52, or 53 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 97% sequence identity to 10 contiguous nucleotides of the sequence of any one of SEQ ID NOs :1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 97% sequence identity to 15 contiguous nucleotides of the sequence of any one of SEQ ID NOs :1-22. In embodiments, the first twister ribozyme domain and the second twister ribozy me domain independently include a nucleotide sequence having 97% sequence identity to 20 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 97% sequence identity to 25 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 97% sequence identity to 30 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include a nucleotide sequence having 97% sequence identity to 35 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 97% sequence identity to 40 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 97% sequence identity to 45 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 97% sequence identity to 50 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 97% sequence identity to 51 contiguous nucleotides of the sequence of any one of SEQ ID NOs:l-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 97% sequence identity to 52 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozy me domain independently include a nucleotide sequence having 97% sequence identity to 53 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22.
[0315] In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 98% sequence identity to 10, 15, 20, 25, 30, 35, 40, 45, 50, 51, 52, or 53 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 98% sequence identity to 10 contiguous nucleotides of the sequence of any one of SEQ ID NOs :1-22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include a nucleotide sequence having 98% sequence identity to 15 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 98% sequence identity to 20 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 98% sequence identity to 25 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozy me domain independently include a nucleotide sequence having 98% sequence identity to 30 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister nbozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 98% sequence identity' to 35 contiguous nucleotides of the sequence of any one of SEQ ID NOs :1-22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include a nucleotide sequence having 98% sequence identity to 40 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include a nucleotide sequence having 98% sequence identity to 45 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 98% sequence identity to 50 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include a nucleotide sequence having 98% sequence identity to 51 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 98% sequence identity to 52 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 98% sequence identity to 53 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22.
[0316] In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 99% sequence identity to 10, 15, 20, 25, 30, 35, 40, 45, 50, 51, 52, or 53 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozy me domain independently include a nucleotide sequence having 99% sequence identity' to 10 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include a nucleotide sequence having 99% sequence identity' to 15 contiguous nucleotides of the sequence of any one of SEQ ID NOs :1-22. In embodiments, the first tyvister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 99% sequence identity to 20 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first tyvister ribozyme domain and the second tyvister ribozyme domain independently include a nucleotide sequence having 99% sequence identity to 25 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1 -22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 99% sequence identity to 30 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 99% sequence identity to 35 contiguous nucleotides of the sequence of any one of SEQ ID NOs :1-22. In embodiments, the first twister ribozyme domain and the second twister ribozy me domain independently include a nucleotide sequence having 99% sequence identity to 40 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 99% sequence identity7 to 45 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 99% sequence identity to 50 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include a nucleotide sequence having 99% sequence identity to 51 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 99% sequence identity to 52 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 99% sequence identity to 53 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22.
[0317] In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 100% sequence identity to 10, 15, 20, 25, 30, 35, 40, 45, 50, 51, 52, or 53 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 100% sequence identity to 10 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozy me domain independently include a nucleotide sequence having 100% sequence identity to 15 contiguous nucleotides of the sequence of any one of SEQ ID NOs:l-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 100% sequence identity to 20 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 100% sequence identity' to 25 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 100% sequence identity to 30 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include a nucleotide sequence having 100% sequence identity to 35 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 100% sequence identity7 to 40 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 100% sequence identity to 45 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozy me domain independently include a nucleotide sequence having 100% sequence identity to 50 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 100% sequence identity to 51 contiguous nucleotides of the sequence of any one of SEQ ID NOs :1-22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include a nucleotide sequence having 100% sequence identity to 52 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 100% sequence identity7 to 53 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22.
[0318] In embodiments, the first twister ribozyme domain and the second twister ribozyme domain include the nucleotide sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain include the nucleotide sequence of SEQ ID NO: 1. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain include the nucleotide sequence of SEQ ID N0:2. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain include the nucleotide sequence of SEQ ID N0:3. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain include the nucleotide sequence of SEQ ID N0:4. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain include the nucleotide sequence of SEQ ID NO:5. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain include the nucleotide sequence of SEQ ID N0:6. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain include the nucleotide sequence of SEQ ID NO:7. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain include the nucleotide sequence of SEQ ID NO: 8. In embodiments, the first twister ribozy me domain and the second twister ribozy me domain include the nucleotide sequence of SEQ ID NO:9. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain include the nucleotide sequence of SEQ ID NO: 10. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain include the nucleotide sequence of SEQ ID NO: 11. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain include the nucleotide sequence of SEQ ID NO: 12. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain include the nucleotide sequence of SEQ ID NO: 13. In embodiments, the first twister ribozyme domain and the second twister ribozy me domain include the nucleotide sequence of SEQ ID NO: 14. In embodiments, the first twister ribozy me domain and the second twister ribozy me domain include the nucleotide sequence of SEQ ID NO: 15. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain include the nucleotide sequence of SEQ ID NO: 16. In embodiments, the first twister ribozyme domain and the second twister ribozy me domain include the nucleotide sequence of SEQ ID NO: 17. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain include the nucleotide sequence of SEQ ID NO: 18. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain include the nucleotide sequence of SEQ ID NO: 19. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain include the nucleotide sequence of SEQ ID NO:20. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain include the nucleotide sequence of SEQ ID N0:21. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain include the nucleotide sequence of SEQ ID NO:22.
[0319] In embodiments, the first twister ribozyme domain or the second twister ribozy me domain includes the nucleotide sequence of any one of SEQ ID NOs:l-22. In embodiments, the first twister ribozyme domain or the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 1. In embodiments, the first twister ribozyme domain or the second twister ribozy me domain includes the nucleotide sequence of SEQ ID N0:2. In embodiments, the first twister ribozy me domain or the second twister ribozyme domain includes the nucleotide sequence of SEQ ID N0:3. In embodiments, the first twister ribozyme domain or the second twister ribozyme domain includes the nucleotide sequence of SEQ ID N0:4. In embodiments, the first twister ribozyme domain or the second twister ribozyme domain includes the nucleotide sequence of SEQ ID N0:5. In embodiments, the first twister ribozyme domain or the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO:6. In embodiments, the first twister ribozyme domain or the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO:7. In embodiments, the first twister ribozyme domain or the second twister ribozy me domain includes the nucleotide sequence of SEQ ID N0:8. In embodiments, the first twister ribozyme domain or the second twister ribozyme domain includes the nucleotide sequence of SEQ ID N0:9. In embodiments, the first twister ribozy me domain or the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NONO. In embodiments, the first twister ribozy me domain or the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 11. In embodiments, the first twister ribozy me domain or the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 12. In embodiments, the first twister ribozyme domain or the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 13. In embodiments, the first twister ribozyme domain or the second twister ribozy me domain includes the nucleotide sequence of SEQ ID NO: 14. In embodiments, the first twister ribozyme domain or the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 15. In embodiments, the first twister ribozy me domain or the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 16. In embodiments, the first twister ribozyme domain or the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 17. In embodiments, the first twister ribozyme domain or the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 18. In embodiments, the first twister ribozyme domain or the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 19. In embodiments, the first twister ribozyme domain or the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO:20. In embodiments, the first twister ribozyme domain or the second twister ribozy me domain includes the nucleotide sequence of SEQ ID N0:21. In embodiments, the first twister ribozy me domain or the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO:22.
[0320] In embodiments, the first twister ribozyme domain includes the nucleotide sequence of any one of SEQ ID NOs:1-22. In embodiments, the first twister ribozyme domain includes the nucleotide sequence of SEQ ID NO:1. In embodiments, the first twister ribozyme domain includes the nucleotide sequence of SEQ ID N0:2. In embodiments, the first twister ribozyme domain includes the nucleotide sequence of SEQ ID NO:3. In embodiments, the first twister ribozyme domain includes the nucleotide sequence of SEQ ID N0:4. In embodiments, the first twister ribozy me domain includes the nucleotide sequence of SEQ ID N0:5. In embodiments, the first twister ribozyme domain includes the nucleotide sequence of SEQ ID N0:6. In embodiments, the first twister ribozyme domain includes the nucleotide sequence of SEQ ID NO:7. In embodiments, the first twister ribozyme domain includes the nucleotide sequence of SEQ ID N0:8. In embodiments, the first twister ribozyme domain includes the nucleotide sequence of SEQ ID NO:9. In embodiments, the first twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 10. In embodiments, the first twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 11. In embodiments, the first twister ribozy me domain includes the nucleotide sequence of SEQ ID NO: 12. In embodiments, the first twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 13. In embodiments, the first twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 14. Tn embodiments, the first twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 15. In embodiments, the first twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 16. In embodiments, the first twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 17. In embodiments, the first twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 18. In embodiments, the first twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 19. In embodiments, the first twister ribozyme domain includes the nucleotide sequence of SEQ ID NO:20. In embodiments, the first twister ribozyme domain includes the nucleotide sequence of SEQ ID N0:21. In embodiments, the first twister ribozyme domain includes the nucleotide sequence of SEQ ID NO:22.
[0321] In embodiments, the first twister ribozyme domain is the nucleotide sequence of any one of SEQ ID NOs:l-22. In embodiments, the first twister ribozyme domain is the nucleotide sequence of SEQ ID NO: 1. In embodiments, the first twister ribozyme domain is the nucleotide sequence of SEQ ID N0:2. In embodiments, the first twister ribozyme domain is the nucleotide sequence of SEQ ID N0:3. In embodiments, the first twister ribozyme domain is the nucleotide sequence of SEQ ID N0:4. In embodiments, the first twister ribozyme domain is the nucleotide sequence of SEQ ID N0:5. In embodiments, the first twister ribozyme domain is the nucleotide sequence of SEQ ID N0:6. In embodiments, the first twister ribozyme domain is the nucleotide sequence of SEQ ID N0:7. In embodiments, the first twister ribozyme domain is the nucleotide sequence of SEQ ID NO:8. In embodiments, the first twister ribozy me domain is the nucleotide sequence of SEQ ID NO:9. In embodiments, the first twister ribozyme domain is the nucleotide sequence of SEQ ID NO: 10. In embodiments, the first twister ribozyme domain is the nucleotide sequence of SEQ ID NO:11. In embodiments, the first twister ribozyme domain is the nucleotide sequence of SEQ ID NO: 12. In embodiments, the first twister ribozy me domain is the nucleotide sequence of SEQ ID NO: 13. In embodiments, the first twister ribozyme domain is the nucleotide sequence of SEQ ID NO: 14. In embodiments, the first twister ribozyme domain is the nucleotide sequence of SEQ ID NO: 15. In embodiments, the first twister ribozyme domain is the nucleotide sequence of SEQ ID NO: 16. In embodiments, the first twister ribozy me domain is the nucleotide sequence of SEQ ID NO: 17. In embodiments, the first twister ribozyme domain is the nucleotide sequence of SEQ ID NO: 18. Tn embodiments, the first twister ribozyme domain is the nucleotide sequence of SEQ ID NO: 19. In embodiments, the first twister ribozyme domain is the nucleotide sequence of SEQ ID NO:20. In embodiments, the first twister ribozyme domain is the nucleotide sequence of SEQ ID N0:21. In embodiments, the first twister ribozyme domain is the nucleotide sequence of SEQ ID NO:22.
[0322] In embodiments, the second twister ribozy me domain includes the nucleotide sequence of any one of SEQ ID NOs: 1-22. In embodiments, the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 1. In embodiments, the second twister ribozyme domain includes the nucleotide sequence of SEQ ID N0:2. In embodiments, the second twister ribozy me domain includes the nucleotide sequence of SEQ ID N0:3. In embodiments, the second twister ribozyme domain includes the nucleotide sequence of SEQ ID N0:4. In embodiments, the second twister ribozyme domain includes the nucleotide sequence of SEQ ID N0:5. In embodiments, the second twister ribozyme domain includes the nucleotide sequence of SEQ ID N0:6. In embodiments, the second twister ribozy me domain includes the nucleotide sequence of SEQ ID N0:7. In embodiments, the second twister ribozy me domain includes the nucleotide sequence of SEQ ID NO: 8. In embodiments, the second twister ribozyme domain includes the nucleotide sequence of SEQ ID N0:9. In embodiments, the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 10. In embodiments, the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 11. In embodiments, the second twister ribozy me domain includes the nucleotide sequence of SEQ ID NO: 12. In embodiments, the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 13. In embodiments, the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 14. In embodiments, the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 15. In embodiments, the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 16. In embodiments, the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 17. In embodiments, the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 18. In embodiments, the second twister ribozy me domain includes the nucleotide sequence of SEQ ID NO: 19. In embodiments, the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO:20. In embodiments, the second twister ribozy me domain includes the nucleotide sequence of SEQ ID N0:21. In embodiments, the second twister ribozy me domain includes the nucleotide sequence of SEQ ID NO:22.
[0323] In embodiments, the second twister ribozyme domain is the nucleotide sequence of any one of SEQ ID NOs:l-22. In embodiments, the second twister ribozyme domain is the nucleotide sequence of SEQ ID NO:1. In embodiments, the second twister ribozyme domain is the nucleotide sequence of SEQ ID NO:2. In embodiments, the second twister ribozy me domain is the nucleotide sequence of SEQ ID N0:3. In embodiments, the second twister ribozyme domain is the nucleotide sequence of SEQ ID N0:4. In embodiments, the second twister ribozyme domain is the nucleotide sequence of SEQ ID N0:5. In embodiments, the second twister ribozy me domain is the nucleotide sequence of SEQ ID NO:6. In embodiments, the second twister ribozyme domain is the nucleotide sequence of SEQ ID N0:7. In embodiments, the second twister ribozyme domain is the nucleotide sequence of SEQ ID NO:8. In embodiments, the second twister ribozy me domain is the nucleotide sequence of SEQ ID NO:9. In embodiments, the second twister ribozyme domain is the nucleotide sequence of SEQ ID NO: 10. In embodiments, the second twister ribozyme domain is the nucleotide sequence of SEQ ID NO: 11. In embodiments, the second twister ribozyme domain is the nucleotide sequence of SEQ ID NO: 12. In embodiments, the second twister ribozyme domain is the nucleotide sequence of SEQ ID NO: 13. In embodiments, the second twister ribozy me domain is the nucleotide sequence of SEQ ID NO: 14. In embodiments, the second twister ribozyme domain is the nucleotide sequence of SEQ ID NO: 15. In embodiments, the second twister ribozyme domain is the nucleotide sequence of SEQ ID NO: 16. In embodiments, the second twister ribozyme domain is the nucleotide sequence of SEQ ID NO: 17. In embodiments, the second twister ribozyme domain is the nucleotide sequence of SEQ ID NO: 18. In embodiments, the second twister ribozyme domain is the nucleotide sequence of SEQ ID NO: 19. In embodiments, the second twister ribozyme domain is the nucleotide sequence of SEQ ID NO:20. In embodiments, the second twister ribozyme domain is the nucleotide sequence of SEQ ID NO:21. In embodiments, the second twister ribozy me domain is the nucleotide sequence of SEQ ID NO:22.
[0324] In embodiments, the method further includes allowing the first member of the split ligation stem and the second member of the split ligation stem to hybridize, thereby forming a reconstituted ligation stem.
[0325] In embodiments, the first member of the split ligation stem includes a 5' hydroxyl. In embodiments, the second member of the split ligation stem includes a 2'-3' cyclic phosphate. In embodiments, the first member of the split ligation stem includes a 5' hydroxyl and the second member of the split ligation stem includes a 2-3' cyclic phosphate.
[0326] In embodiments, the first member of the split ligation stem and the second member of the split ligation stem are capable of hybridizing and reconstituting the ligation stem. In embodiments, the reconsituted ligation stem brings the 5' end and 3' end of the linear RNA compound into close approximation. In embodiments, the reconstituted ligation stem allows for a ligase to ligate the 5' end and 3' end of a linear RNA compound, thereby forming a circularized RNA. In embodiments, the ligase enzyme ligates the 5' hydroxyl of the first member of the split ligation stem to the 2'-3' cyclic phosphate of the second member of the split ligation stem, thereby forming a circularized RNA. In embodiments, the ligase enzyme is an RNA 2',3'-cyclic phosphate and 5'-OH (RtcB) ligase. In embodiments, the first ribozy me domain includes the first member of the split ligation stem. In embodiments, the second ribozyme domain includes the second member of the split ligation stem. In embodiments, the first ribozyme domain does not the first member of the split ligation stem. In embodiments, the second ribozyme domain does not include the second member of the split ligation stem.
[0327] In embodiments, the first ribozyme domain includes the nucleotide sequence of SEQ ID NO: 1 and the second ribozyme domain includes the nucleotide sequence of SEQ ID NO:2. In embodiments, the first ribozyme domain is the nucleotide sequence of SEQ ID NO: 1 and the second ribozyme domain is the nucleotide sequence of SEQ ID NO:2.
[0328] In embodiments, the first member of the split ligation stem includes the nucleotide sequence of SEQ ID NO:23 and the second member of the split ligation stem includes the nucleotide sequence of SEQ ID NO:24.
[0329] In another aspect is provided a method of forming a circularized ribonucleic acid (RNA) including: incubating a linear RNA compound in vitro for between about 16 hours and about 24 hours under conditions conducive to group II intron cleavage, thereby forming a circularized RNA, wherein the RNA compound is a nucleic acid including from 5' to 3': a first member of a split group II intron domain, an internal ribosome entry site (IRES) domain, a 5' untranslated region (UTR), a protein-encoding nucleic acid sequence at least 750 nucleotides in length, a 3' UTR, a poly-adenine (poly A) domain, and a second member of the split group II intron domain, wherein the first member of the split group II intron includes domains V and VI and the second member of the split group II intron includes domains 1. II and III.
[0330] In embodiments, the linear RNA compound is incubated in vitro for between about 17 hours and about 24 hours under conditions conducive to group II intron cleavage. In embodiments, the linear RNA compound is incubated in vitro for between about 18 hours and about 24 hours under conditions conducive to group II intron cleavage In embodiments, the linear RNA compound is incubated in vitro for between about 19 hours and about 24 hours under conditions conducive to group II intron cleavage In embodiments, the linear RNA compound is incubated in vitro for between about 20 hours and about 24 hours under conditions conducive to group II intron cleavage In embodiments, the linear RNA compound is incubated in vitro for between about 21 hours and about 24 hours under conditions conducive to group II intron cleavage In embodiments, the linear RNA compound is incubated in vitro for between about 22 hours and about 24 hours under conditions conducive to group II intron cleavage In embodiments, the linear RNA compound is incubated in vitro for between about 23 hours and about 24 hours under conditions conducive to group II intron cleavage.
[0331] In embodiments, the linear RNA compound is incubated in vitro for at least 16 hours under conditions conducive to group II intron cleavage. In embodiments, the linear RNA compound is incubated in vitro for at least 17 hours under conditions conducive to group II intron cleavage. In embodiments, the linear RNA compound is incubated in vitro for at least 18 hours under conditions conducive to group II intron cleavage. In embodiments, the linear RNA compound is incubated in vitro for at least 19 hours under conditions conducive 149 to group IT intron cleavage. In embodiments, the linear RNA compound is incubated in vitro for at least 20 hours under conditions conducive to group II intron cleavage. In embodiments, the linear RNA compound is incubated in vitro for at least 21 hours under conditions conducive to group II intron cleavage. In embodiments, the linear RNA compound is incubated in vitro for at least 22 hours under conditions conducive to group II intron cleavage. In embodiments, the linear RNA compound is incubated in vitro for at least 23 hours under conditions conducive to group II intron cleavage. In embodiments, the linear RNA compound is incubated in vitro for at least 24 hours under conditions conducive to group II intron cleavage.
[0332] In embodiments, the conditions conducive to group II intron cleavage are well known in the art.
[0333] In embodiments, the group II intron domain is a Clostridium tetani group II intron domain, a Histoplasma capsulatum group II intron domain, a Coccidioides immitis group II intron domain, a Blastomyces dermatitidis group II intron domain, a Coccidioidomycosis posadasii group II intron domain, a Pylaeiella littoralis group II intron domain, a Saccharomyces cerevisiae group II intron domain, a Lactococcus lactis group II intron domain, aAnthoceros angustus group II intron domain, a Geobacillus stearothermophilus group II intron domain, a Bacillus megaterium group II intron domain, a Pseudomonas alcaligenes group II intron domain, a Chaetothyriales bantiana group II intron domain, or a Chaetothyricdes carrioni group II intron domain.
[0334] In embodiments, the group II intron domain includes the nucleotide sequence of any one of SEQ ID NOs:105-116. In embodiments, the group II intron domain includes the nucleotide sequence of SEQ ID NO: 105. In embodiments, the group II intron domain includes the nucleotide sequence of SEQ ID NO: 106. In embodiments, the group II intron domain includes the nucleotide sequence of SEQ ID NO: 107. In embodiments, the group II intron domain includes the nucleotide sequence of SEQ ID NO: 108. In embodiments, the group II intron domain includes the nucleotide sequence of SEQ ID NO: 109. In embodiments, the group II intron domain includes the nucleotide sequence of SEQ ID NO:110. In embodiments, the group II intron domain includes the nucleotide sequence of SEQ ID NO: 111. In embodiments, the group II intron domain includes the nucleotide sequence of SEQ ID NO: 112. In embodiments, the group II intron domain includes the nucleotide sequence of SEQ ID NO: 113. In embodiments, the group II intron domain includes the nucleotide sequence of SEQ ID NO: 114. In embodiments, the group II intron domain includes the nucleotide sequence of SEQ ID NO: 115. In embodiments, the group II intron domain includes the nucleotide sequence of SEQ ID NO:116.
[0335] In embodiments, the group II intron domain is the nucleotide sequence of any one of SEQ ID NOs: 105-116. In embodiments, the group II intron domain is the nucleotide sequence of SEQ ID NO: 105. In embodiments, the group II intron domain is the nucleotide sequence of SEQ ID NO: 106. In embodiments, the group II intron domain is the nucleotide sequence of SEQ ID NO: 107. In embodiments, the group II intron domain is the nucleotide sequence of SEQ ID NO: 108. In embodiments, the group II intron domain is the nucleotide sequence of SEQ ID NO: 109. In embodiments, the group II intron domain is the nucleotide sequence of SEQ ID NO: 110. In embodiments, the group II intron domain is the nucleotide sequence of SEQ ID NO: 111. In embodiments, the group II intron domain is the nucleotide sequence of SEQ ID NO: 112. In embodiments, the group II intron domain is the nucleotide sequence of SEQ ID NO: 113. In embodiments, the group II intron domain is the nucleotide sequence of SEQ ID NO: 114. In embodiments, the group II intron domain is the nucleotide sequence of SEQ ID NO: 115. In embodiments, the group II intron domain is the nucleotide sequence of SEQ ID NO: 116.
[0336] In embodiments, the group II intron domain is a Clostridium tetani group II intron domain. In embodiments, the Clostridium tetani group II intron domain includes the nucleotide sequence of any one of SEQ ID NOs: 105-107. In embodiments, the Clostridium tetani group II intron domain includes the nucleotide sequence of SEQ ID NO: 105. In embodiments, the Clostridium tetani group II intron domain includes the nucleotide sequence of SEQ ID NO: 106. In embodiments, the Clostridium tetani group II intron domain includes the nucleotide sequence of SEQ ID NO: 107. In embodiments, the Clostridium tetani group II intron domain is the nucleotide sequence of any one of SEQ ID NOs: 105-107. In embodiments, the Clostridium tetani group II intron domain is the nucleotide sequence of SEQ ID NO: 105. In embodiments, the Clostridium tetani group II intron domain is the nucleotide sequence of SEQ ID NO: 106. In embodiments, the Clostridium tetani group II intron domain is the nucleotide sequence of SEQ ID NO: 107.
[0337] In embodiments, the group II intron domain is a Histoplasma capsulatum group II intron domain. In embodiments, the Histoplasma capsulatum group II intron domain is a Histoplasma capsulatum H.c.LSU.Il group II intron domain. In embodiments, the Histoplasma capsulatum group II intron domain includes the nucleotide sequence of SEQ ID NO: 108. In embodiments, the Histoplasma capsulatum group II intron domain is the nucleotide sequence of SEQ ID NO: 108.
[0338] In embodiments, the group II intron domain is a Coccidioides immitis group II intron domain. In embodiments, the Coccidioides immitis group II intron domain is a Coccidioides immitis C.i.LSU.I3 group II intron domain or a Coccidioides immitis C.i.SSU.Il group II intron domain, or a Coccidioides immitis C.i.SSU.I2 group II intron domain. In embodiments, the Coccidioides immitis group II intron domain is a Coccidioides immitis C.i.LSU.I3 group II intron domain. In embodiments, the Coccidioides immitis C.i.LSU.13 group II intron domain includes the nucleotide sequence of SEQ ID NO: 109. In embodiments, the Coccidioides immitis C.i.LSU.I3 group II intron domain is the nucleotide sequence of SEQ ID NO: 109. In embodiments, the Coccidioides immitis group II intron domain is a Coccidioides immitis C.i.SSU.Il group II intron domain. In embodiments, the Coccidioides immitis C.i.SSU.Il group II intron domain includes the nucleotide sequence of SEQ ID NO: 110. In embodiments, the Coccidioides immitis C.i.SSU.Il group II intron domain is the nucleotide sequence of SEQ ID NO: 110. In embodiments, the Coccidioides immitis group II intron domain is a Coccidioides immitis C.i.SSU.12 group II intron domain. In embodiments, the Coccidioides immitis C.i.SSU.12 group II intron domain includes the nucleotide sequence of SEQ ID NO: 112. In embodiments, the Coccidioides immitis C.i.SSU.12 group II intron domain is the nucleotide sequence of SEQ ID NO: 112.
[0339] In embodiments, the group II intron domain is a Blastomyces dermatitidis group II intron domain. In embodiments, the Blastomyces dermatitidis group II intron domain is a Blastomyces dermatitidis B.d.LSU.I2 group II intron domain. In embodiments, the Blastomyces dermatitidis group II intron domain includes the nucleotide sequence of SEQ ID NO:111. In embodiments, the Blastomyces dermatitidis group II intron domain is the nucleotide sequence of SEQ ID NO:111.
[0340] In embodiments, the group II intron domain is a Coccidioidomycosis posadasii group II intron domain. In embodiments, the Coccidioidomycosis posadasii group II intron domain is a Coccidioidomycosis posadasii C.p.LSU.I3 group II intron domain. In embodiments, the Coccidioidomycosis posadasii group II intron domain includes the nucleotide sequence of SEQ ID NO: 113. In embodiments, the Coccidioidomycosisposadasii group II intron domain is the nucleotide sequence of SEQ ID NO: 113.
[0341] In embodiments, the group II intron domain is a Pylaeiella littoralis group II intron domain. In embodiments, the Pylaeiella littoralis group II intron domain is a Pylaeiella littoralis P.li.LSU.I2 group II intron domain or a Pylaeiella littoralis P.li.LSU.12 i|jU group II intron domain. In embodiments, the Pylaeiella littoralis group II intron domain is a Pylaeiella littoralis P.li.LSU.12 group II intron domain. In embodiments, the Pylaeiella littoralis P.li.LSU.12 group II intron domain is a modified Pylaeiella littoralis P.li.LSU.12 group II intron domain. In embodiments, the Pylaeiella littoralis P.li.LSU.12 group II intron domain includes the nucleotide sequence of SEQ ID NO: 114. In embodiments, the Pylaeiella littoralis P.li.LSU.12 group II intron domain is the nucleotide sequence of SEQ ID NO: 114. In embodiments, the Pylaeiella littoralis group II intron domain is a Pylaeiella littoralis P.li.LSU.12 ifrU group II intron domain. In embodiments, the Pylaeiella littoralis P.li.LSU.12 ijiU group II intron domain is a modified Pylaeiella littoralis P.li.LSU.12 xpU group II intron domain. In embodiments, the Pylaeiella littoralis P.li.LSU.12 rpu group II intron domain includes the nucleotide sequence of SEQ ID NO: 115 . In embodiments, the Pylaeiella littoralis P.li.LSU.12 xpU group II intron domain is the nucleotide sequence of SEQ ID NO:115.
[0342] In embodiments, the group II intron domain is a Saccharomyces cerevisiae group II intron domain. In embodiments, the Saccharomyces cerevisiae group II intron domain is a Saccharomyces cerevisiae bll group II intron domain. In embodiments, the Saccharomyces cerevisiae group II intron domain includes the nucleotide sequence of SEQ ID NO: 116. In embodiments, the Saccharomyces cerevisiae group II intron domain is the nucleotide sequence of SEQ ID NO: 116.
[0343] In embodiments, the group II intron domain is a Lactococcus lactis group II intron domain, aAnthoceros angustus group II intron domain. In embodiments, the group II intron domain is a Geobacillus stearothermophilus group II intron domain. In embodiments, the group II intron domain is a Bacillus megaterium group II intron domain. In embodiments, the group II intron domain is a Pseudomonas alcaligenes group II intron domain. In embodiments, the group II intron domain is a Chaetothyriales bantiana group II intron domain. In embodiments, the group II intron domain is or a Chaetothyriales carrioni group II intron domain.
[0344] In embodiments, the first member of the split group II intron domain includes the nucleotide sequence of SEQ ID NO: 105 and the second member of the split group II intron domain includes the nucleotide sequence of SEQ ID NO: 106. In embodiments, the first member of the split group II intron domain is the nucleotide sequence of SEQ ID NO: 105 and the second member of the split group II intron domain is the nucleotide sequence of SEQ ID NO: 106.
[0345] In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 800 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 850 to about 5000 nucleotides. In embodiments, the proteinencoding nucleic acid sequence includes about 900 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 1000 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 1100 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 1200 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 1300 to about 5000 nucleotides. In embodiments, the proteinencoding nucleic acid sequence includes about 1400 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 1500 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 1600 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 1700 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 1800 to about 5000 nucleotides. In embodiments, the proteinencoding nucleic acid sequence includes about 1900 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 2000 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 2100 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 2200 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 2300 to about 5000 nucleotides. In embodiments, the proteinencoding nucleic acid sequence includes about 2400 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 2500 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 2600 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 2700 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 2800 to about 5000 nucleotides. In embodiments, the proteinencoding nucleic acid sequence includes about 2900 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 3000 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 3100 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 3200 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 3300 to about 5000 nucleotides. In embodiments, the proteinencoding nucleic acid sequence includes about 3400 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 3500 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 3600 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 3700 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 3800 to about 5000 nucleotides. In embodiments, the proteinencoding nucleic acid sequence includes about 3900 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 4000 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 4100 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence 155 includes about 4200 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 4300 to about 5000 nucleotides. In embodiments, the proteinencoding nucleic acid sequence includes about 4400 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 4500 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 4600 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 4700 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 4800 to about 5000 nucleotides. In embodiments, the proteinencoding nucleic acid sequence includes about 4900 to about 5000 nucleotides.
[0346] In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 4900 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 4800 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 4700 nucleotides. In embodiments, the proteinencoding nucleic acid sequence includes about 750 to about 4600 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 4500 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 4400 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 4300 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 4200 nucleotides. In embodiments, the proteinencoding nucleic acid sequence includes about 750 to about 4100 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 4000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 3900 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 3800 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 3700 nucleotides. In embodiments, the proteinencoding nucleic acid sequence includes about 750 to about 3600 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 3500 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 3400 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 3300 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 3200 nucleotides. In embodiments, the proteinencoding nucleic acid sequence includes about 750 to about 3100 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 3000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 2900 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 2800 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 2700 nucleotides. In embodiments, the proteinencoding nucleic acid sequence includes about 750 to about 2600 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 2500 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 2400 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 2300 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 2200 nucleotides. In embodiments, the proteinencoding nucleic acid sequence includes about 750 to about 2100 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 2000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 1900 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 1800 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 1700 nucleotides. In embodiments, the proteinencoding nucleic acid sequence includes about 750 to about 1600 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 1500 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 1400 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 1300 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 1200 nucleotides. In embodiments, the proteinencoding nucleic acid sequence includes about 750 to about 1100 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 1000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 900 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 850 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 800 nucleotides.
[0347] In embodiments, the protein-encoding nucleic acid sequence includes at least 800 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 850 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 900 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 950 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 1000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 1100 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 1200 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 1300 nucleotides. In embodiments, the proteinencoding nucleic acid sequence includes at least 1400 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 1500 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 1600 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 1700 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 1800 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 1900 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 2000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 2100 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 2200 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 2300 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 2400 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 2500 nucleotides. In embodiments, the proteinencoding nucleic acid sequence includes at least 2600 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 2700 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 2800 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 2900 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 3000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 3100 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 3200 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 3300 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 3400 nucleotides. Tn embodiments, the protein-encoding nucleic acid sequence includes at least 3500 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 3600 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 3700 nucleotides. In embodiments, the proteinencoding nucleic acid sequence includes at least 3800 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 3900 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 4000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 4100 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 4200 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 4300 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 4400 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 4500 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 4600 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 4700 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 4800 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 4900 nucleotides. In embodiments, the proteinencoding nucleic acid sequence includes at least 5000 nucleotides.
[0348] In embodiments, the first member of the split group II intron does not include domains I, II, and II of the group II intron domain. In embodiments, the first member of the split group II intron only includes domains V and VI of the group II intron domain. In embodiments, the second member of the split group II intron does not include domains V and VI of the group II intron domain. In embodiments, the second member of the group II intron only includes domains I, II, and III of the group II intron domain.
[0349] In embodiments, the protein-encoding nucleic acid sequence encodes an epigenetic effector domain, a repressor domain, or an activator domain. In embodiments, the proteinencoding nucleic acid sequence encodes an epigenetic effector domain. In embodiments, the protein-encoding nucleic acid sequence encodes a repressor domain. In embodiments, the protein-encoding nucleic acid sequence encodes an activator domain.
[0350] In embodiments, the protein-encoding nucleic acid sequence encodes a deoxyribonucleic acid (DNA) methyltransferase domain, a CpG methyltransferase (M.SssI) domain, a Sin3 interacting repressor domain (SID4X), a protamine 2 (PRM2) domain, a protamine 1 (PRM1) domain, a VP64 domain, a Kruppel associated box (KRAB) domain, a coupled histone tail for autoinhibition release of methyltransferase (CHARM) domain, or a poly dactyl zinc finger protein domain.
[0351] In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of any one of SEQ ID NOs:25-54. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:25. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:26. In embodiments, the proteinencoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:27. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:28. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:29. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:30. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:31. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:32. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:33. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:34. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:35. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:36. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:37. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:38. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:39. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:40. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:41. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:42. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:43. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:44. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:45. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:46. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:47. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:48. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:49. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:50. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID N0:51. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:52. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:53. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:54.
[0352] In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of any one of SEQ ID NOs:25-54. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:25. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:26. In embodiments, the proteinencoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID 161 NO:27. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:28. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:29. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO: 30. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:31. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:32. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:33. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:34. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:35. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:36. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:37. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:38. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:39. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NONO. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:41. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:42. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:43. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:44. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:45. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:46. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:47. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino 162 acid sequence of SEQ ID NO:48. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:49. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:50. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:51. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:52. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:53. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:54.
[0353] In embodiments, the protein-encoding nucleic acid sequence encodes a deoxyribonucleic acid (DNA) methyltransferase domain. In embodiments, the DNA methyltransferase domain is a Dnmt3A-3L domain. In embodiments, the Dnmt3A-3L domain includes the amino acid sequence of SEQ ID NO:32. In embodiments, the Dnmt3A-3L domain is the amino acid sequence of SEQ ID NO:32.
[0354] In embodiments, the protein-encoding nucleic acid sequence encodes a a CpG methyltransferase (M.SssI) domain. In embodiments, the M.SssI domain includes the amino acid sequence of SEQ ID NO:26. In embodiments, the M.SssI domain is the amino acid sequence of SEQ ID NO:26.
[0355] In embodiments, the protein-encoding nucleic acid sequence encodes a Sin3 interacting repressor domain (SID4X. In embodiments, the SID4X domain includes the amino acid sequence of SEQ ID NO:27. In embodiments, the SID4X domain is the amino acid sequence of SEQ ID NO :27.
[0356] In embodiments, the protein-encoding nucleic acid sequence encodes a protamine 2 (PRM2) domain. In embodiments, the PRM2 domain includes the amino acid sequence of SEQ ID NO:28. In embodiments, the PRM2 domain is the amino acid sequence of SEQ ID NO:28.
[0357] In embodiments, the protein-encoding nucleic acid sequence encodes a protamine 1 (PRM1) domain. In embodiments, the PRM1 domain includes the amino acid sequence of 163 SEQ ID NO:29. In embodiments, the PRM1 domain is the amino acid sequence of SEQ ID NO:29.
[0358] In embodiments, the protein-encoding nucleic acid sequence encodes a VP64 domain. In embodiments, the VP64 domain includes the amino acid sequence of SEQ ID NO:30. In embodiments, the VP64 domain is the amino acid sequence of SEQ ID NO:30.
[0359] In embodiments, the protein-encoding nucleic acid sequence encodes a Krtippel associated box (KRAB) domain. In embodiments, the KRAB domain includes the amino acid sequence of SEQ ID NO:31. In embodiments, the KRAB domain is the amino acid sequence of SEQ ID NO:31.
[0360] In embodiments, the protein-encoding nucleic acid sequence encodes a coupled histone tail for autoinhibition release of methyltransferase (CHARM) domain. . In embodiments, the CHARM domain includes the amino acid sequence of SEQ ID NO:25. In embodiments, the CHARM domain is the amino acid sequence of SEQ ID NO:25.
[0361] In embodiments, the protein-encoding nucleic acid sequence encodes a poly dactyl zinc finger protein domain. In embodiments, the polydactyl zinc finger protein domain includes a first alpha helix domain including the amino acid sequence of SEQ ID NO:55 a second alpha helix domain including the amino acid sequence of SEQ ID NO:56, a third alpha helix domain including the amino acid sequence of SEQ ID NO:57, a fourth alpha helix domain including the amino acid sequence of SEQ ID NO: 58, a fifth alpha helix domain including the amino acid sequence of SEQ ID NO:59. and a sixth alpha helix domain including the amino acid sequence of SEQ ID NO: 60. In embodiments, the poly dactyl zinc finger protein domain includes the amino acid sequence of SEQ ID NO:61. In embodiments, the poly dactyl zinc finger protein domain is the amino acid sequence of SEQ ID NO:61.
[0362] In embodiments, the IRES domain is a cricket paralysis virus (CrPV) IRES domain, an insulin-like growth factor 2 (IGF2) IRES domain, a hepatitis C virus H77 IRES domain, a fibroblast growth factor 1 (FGF1) IRES domain, a bovine viral diarrhea virus (BVDV) 1 IRES domain, a human rhinovirus A89 IRES domain, a LIM domain and actin binding protein 1 (LIMA1) IRES domain, a human adenovirus 2 IRES domain, amontana Myotis leukoencephalitis virus (MMLV) IRES domain, a RAN binding protein 3 (RANBP3) IRES 164 domain, apestivirus giraffe 1 IRES domain, a TG-interacting factor 1 (TGIF1) TRES domain, a human poliovirus 1 Mahoney IRES domain, a Foot-and-Mouth disease virus type 0 IRES domain, an encephalomyocarditis virus (ECMV) IRES domain, an encephalomyocarditis virus 7A IRES domain, an encephalomyocarditis virus 6A IRES domain, an enterovirus 71 IRES domain, a Coxsackievirus B3 (CB3) IRES domain, a pegivirus A IRES domain, an equine rhinitis A virus (ERAV) IRES domain, a GB virus C (GBV-HGV) IRES domain, a human betaherpesvirus 5 IRES domain, a Senecavirus A (SVA) IRES domain, an equine rhinitis B virus 1 (ERBV-1) IRES domain, a Triticum mosaic virus (TriMV) IRES domain, a hepatovirus A (HAV) IRES domain, a hepatitis GB virus B (HGBV-B) IRES domain, a Giardia lamblia virus (GLV) IRES domain, a Cyrphonectria hypovirus 1 IRES domain, or an equine hepacivrus JPN3 / JAPAN / 2013 IRES domain.
[0363] In embodiments, the IRES domain includes the nucleotide sequence of any one of SEQ ID NOs:63-94. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:63. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:64. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:65. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:66. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:67. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:68. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:69. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:70. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:71. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:72. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:73. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:74. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:75. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:76. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:77. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:78. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:79. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:80. In embodiments, the TRES domain includes the nucleotide sequence of SEQ ID N0:81. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:82. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:83. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:84. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:85. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:86. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:87. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:88. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:89. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID N0:91. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:92. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:93. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:94.
[0364] In embodiments, the IRES domain is the nucleotide sequence of any one of SEQ ID NOs:63-94. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:63. In embodiments, the IRES domain includes the nucleotide sequence of any one of SEQ ID NO:64. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:65. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:66. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:67. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:68. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:69. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:70. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:71. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:72. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:73. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:74. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:75. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:76. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:77. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:78. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:79. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:80. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID N0:81. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:82. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:83. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:84. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:85. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:86. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:87. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:88. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:89. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID N0:91. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:92. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:93. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:94.
[0365] In embodiments, the IRES domain is a cricket paralysis virus (CrPV) IRES domain. In embodiments, the cricket paralysis virus (CrPV) IRES domain includes the nucleotide sequence of SEQ ID NO:63. In embodiments, the cricket paralysis virus (CrPV) IRES domain is the nucleotide sequence of SEQ ID NO:63.
[0366] In embodiments, the IRES domain is an insulin-like grow th factor 2 (IGF2) IRES domain. In embodiments, the insulin-like growth factor 2 (IGF2) IRES domain is a human IGF2 IRES domain. In embodiments, the insulin-like growth factor 2 (IGF2) IRES domain includes the nucleotide sequence of SEQ ID NO:64. In embodiments, the insulin-like growth factor 2 (IGF2) IRES domain is the nucleotide sequence of SEQ ID NO:64.
[0367] In embodiments, the IRES domain is a hepatitis C virus H77 IRES domain. In embodiments, the hepatitis C virus H77 IRES domain includes the nucleotide sequence of SEQ ID NO:66. In embodiments, the hepatitis C virus H77 IRES domain is the nucleotide sequence of SEQ ID NO:66.
[0368] In embodiments, the IRES domain is a fibroblast growth factor 1 (FGF1) IRES domain. In embodiments, the fibroblast grow th factor 1 (FGF1) IRES domain is a human FGF1 IRES domain. In embodiments, the fibroblast growth factor 1 (FGF1) IRES domain includes the nucleotide sequence of SEQ ID NO:67. In embodiments, the fibroblast growth factor 1 (FGF1) IRES domain is the nucleotide sequence of SEQ ID NO:67.
[0369] In embodiments, the IRES domain is a bovine viral diarrhea vims (BVDV) 1 IRES domain. In embodiments, the bovine viral diarrhea vims (BVDV) 1 IRES domain includes the nucleotide sequence of SEQ ID NO:68. In embodiments, the bovine viral diarrhea virus (BVDV) 1 IRES domain is the nucleotide sequence of SEQ ID NO:68.
[0370] In embodiments, the IRES domain is a human rhinovims A89 IRES domain. In embodiments, the human rhinovirus A89 IRES domain includes the nucleotide sequence of SEQ ID NO:69. In embodiments, the human rhinovirus A89 IRES domain is the nucleotide sequence of SEQ ID NO:69.
[0371] In embodiments, the IRES domain is a LIM domain and actin binding protein 1 (LIMA1) IRES domain. In embodiments, the LIM domain and actin binding protein 1 (LIMA1) IRES domain is a. Pan paniscus LIMA1 IRES domain. In embodiments, the LIM domain and actin binding protein 1 (LIMA1) IRES domain includes the nucleotide sequence of SEQ ID NO:70. In embodiments, the LIM domain and actin binding protein 1 (LIMA1) IRES domain is the nucleotide sequence of SEQ ID NO:70.
[0372] In embodiments, the IRES domain is a human adenovirus 2 IRES domain. In embodiments, the human adenovirus 2 IRES domain includes the nucleotide sequence of SEQ ID NO:71. In embodiments, the human adenovirus 2 IRES domain is the nucleotide sequence of SEQ ID NO:71.
[0373] In embodiments, the IRES domain is a Montana Myotis leukoencephalitis virus (MMLV) IRES domain. In embodiments, the Montana Myotis leukoencephalitis virus (MMLV) IRES domain includes the nucleotide sequence of SEQ ID NO:72. In embodiments, the Montana Myotis leukoencephalitis virus (MMLV) IRES domain is the nucleotide sequence of SEQ ID NO:72.
[0374] In embodiments, the IRES domain is a RAN binding protein 3 (RANBP3) IRES domain. In embodiments, the RAN binding protein 3 (RANBP3) IRES domain is a human RANBP3 IRES domain. In embodiments, the RAN binding protein 3 (RANBP3) IRES domain includes the nucleotide sequence of SEQ ID NO:73. In embodiments, the RAN binding protein 3 (RANBP3) IRES domain is the nucleotide sequence of SEQ ID NO:73.
[0375] In embodiments, the IRES domain is a pestivirus giraffe 1 IRES domain. In embodiments, the pestivirus giraffe 1 IRES domain includes the nucleotide sequence of SEQ ID NO:74. In embodiments, the pestivirus giraffe 1 IRES domain is the nucleotide sequence of SEQ ID NO: 74.
[0376] In embodiments, the IRES domain is a TG-interacting factor 1 (TGIF1) IRES domain. In embodiments, the TG-interacting factor 1 (TGIF1) IRES domain is a human TGIF1 IRES domain. In embodiments, the TG-interacting factor 1 (TGIF1) IRES domain includes the nucleotide sequence of SEQ ID NO:75. In embodiments, the TG-interacting factor 1 (TGIF1) IRES domain is the nucleotide sequence of SEQ ID NO:75.
[0377] In embodiments, the IRES domain is a human poliovirus 1 Mahoney IRES domain. In embodiments, the human poliovirus 1 Mahoney IRES domain includes the nucleotide sequence of SEQ ID NO:76. In embodiments, the human poliovirus 1 Mahoney IRES domain is the nucleotide sequence of SEQ ID NO:76.
[0378] In embodiments, the IRES domain is a Foot-and-Mouth disease virus type O IRES domain. In embodiments, the Foot-and-Mouth disease virus type O IRES domain includes the nucleotide sequence of SEQ ID NO:77. In embodiments, the Foot-and-Mouth disease virus type O IRES domain is the nucleotide sequence of SEQ ID NO: 77
[0379] In embodiments, the IRES domain is an encephalomyocarditis virus (ECMV) IRES domain. In embodiments, the encephalomyocarditis virus (ECMV) IRES domain includes the nucleotide sequence of SEQ ID NO:94. In embodiments, the encephalomyocarditis virus (ECMV) IRES domain is the nucleotide sequence of SEQ ID NO:94.
[0380] In embodiments, the IRES domain is an encephalomyocarditis virus 7A IRES domain. In embodiments, the encephalomyocarditis virus 7A IRES domain includes the nucleotide sequence of SEQ ID NO:78. In embodiments, the encephalomyocarditis virus 7A IRES domain is the nucleotide sequence of SEQ ID NO:78.
[0381] In embodiments, the IRES domain is an encephalomyocarditis virus 6A IRES domain. In embodiments, the encephalomyocarditis virus 6A IRES domain includes the nucleotide sequence of SEQ ID NO:79. In embodiments, the encephalomyocarditis virus 6A IRES domain is the nucleotide sequence of SEQ ID NO:79.
[0382] In embodiments, the IRES domain is an enterovirus 71 IRES domain. In embodiments, the enterovirus 71 IRES domain includes the nucleotide sequence of SEQ ID NO:80. In embodiments, the enterovirus 71 IRES domain is the nucleotide sequence of SEQ ID NO:80.
[0383] In embodiments, the IRES domain is a Coxsackievirus B3 (CB3) IRES domain. In embodiments, the Coxsackievirus B3 (CB3) IRES domain includes the nucleotide sequence of SEQ ID NO:81. In embodiments, the Coxsackievirus B3 (CB3) IRES domain is the nucleotide sequence of SEQ ID NO:81.
[0384] In embodiments, the IRES domain is a pegivirus A IRES domain. In embodiments, the pegivirus A IRES domain includes the nucleotide sequence of SEQ ID NO:82. In embodiments, the pegivirus A IRES domain is the nucleotide sequence of SEQ ID NO:82.
[0385] In embodiments, the IRES domain is an equine rhinitis A virus (ERAV) IRES domain. In embodiments, the equine rhinitis A virus (ERAV) IRES domain includes the nucleotide sequence of SEQ ID NO:83. In embodiments, the equine rhinitis A virus (ERAV) IRES domain is the nucleotide sequence of SEQ ID NO:83.
[0386] In embodiments, the IRES domain is a GB virus C (GBV-HGV) IRES domain. In embodiments, the GB virus C (GBV-HGV) IRES domain includes the nucleotide sequence of SEQ ID NO:84. In embodiments, the GB virus C (GBV-HGV) IRES domain is the nucleotide sequence of SEQ ID NO:84.
[0387] In embodiments, the IRES domain is a human betaherpesvirus 5 IRES domain. In embodiments, the human betaherpesvirus 5 IRES domain includes the nucleotide sequence of SEQ ID NO:85. In embodiments, the human betaherpesvirus 5 IRES domain is the nucleotide sequence of SEQ ID NO: 85
[0388] In embodiments, the IRES domain is a Senecavirus A (SVA) IRES domain. In embodiments, the Senecavirus A (SVA) IRES domain includes the nucleotide sequence of SEQ ID NO:86. . In embodiments, the Senecavirus A (SVA) IRES domain is the nucleotide sequence of SEQ ID NO: 86.
[0389] In embodiments, the IRES domain is an equine rhinitis B virus 1 (ERBV-1) IRES domain. In embodiments, the equine rhinitis B virus 1 (ERBV-1) IRES domain includes the nucleotide sequence of SEQ ID NO:87. In embodiments, the equine rhinitis B virus 1 (ERBV-1) IRES domain is the nucleotide sequence of SEQ ID NO:87.
[0390] In embodiments, the IRES domain is a Triticum mosaic virus (TriMV) IRES domain. In embodiments, the Triticum mosaic virus (TriMV) IRES domain includes the nucleotide sequence of SEQ ID NO:88. In embodiments, the Triticum mosaic virus (TriMV) IRES domain is the nucleotide sequence of SEQ ID NO:88.
[0391] In embodiments, the IRES domain is a hepatovirus A (HAV) IRES domain. In embodiments, the hepatovirus A (HAV) IRES domain includes the nucleotide sequence of SEQ ID NO:65 or SEQ ID NO:89. In embodiments, the hepatovirus A (HAV) IRES domain includes the nucleotide sequence of SEQ ID NO:65. In embodiments, the hepatovirus A (HAV) IRES domain includes the nucleotide sequence of SEQ ID NO: 89. In embodiments, the hepatovirus A (HAV) IRES domain is the nucleotide sequence of SEQ ID NO:65 or SEQ ID NO:89. In embodiments, the hepatovirus A (HAV) IRES domain includes the nucleotide sequence of SEQ ID NO:65. In embodiments, the hepatovirus A (HAV) IRES domain is the nucleotide sequence of SEQ ID NO:89.
[0392] In embodiments, the IRES domain is a hepatitis GB virus B (HGBV-B) IRES domain. In embodiments, the hepatitis GB virus B (HGBV-B) IRES domain includes the nucleotide sequence of SEQ ID NO:90. In embodiments, the hepatitis GB virus B (HGBV-B) IRES domain is the nucleotide sequence of SEQ ID NO:90.
[0393] In embodiments, the IRES domain is a Giardia lamblia virus (GLV) IRES domain. In embodiments, the Giardia lamblia vims (GLV) IRES domain includes the nucleotide sequence of SEQ ID NO:91. In embodiments, the Giardia lamblia vims (GLV) IRES domain is the nucleotide sequence of SEQ ID NO:91.
[0394] In embodiments, the IRES domain is a Cyrphonectria hypovirus 1 IRES domain. In embodiments, the Cyrphonectria hypovims 1 IRES domain includes the nucleotide sequence of SEQ ID NO:92. In embodiments, the Cyrphonectria hypovirus 1 IRES domain is the nucleotide sequence of SEQ ID NO:92.
[0395] In embodiments, the IRES domain is an equine hepacivms JPN3 / JAPAN / 2013 IRES domain. In embodiments, the equine hepacivms JPN3 / JAPAN / 2013 IRES domain includes the nucleotide sequence of SEQ ID NO:93. In embodiments, the equine hepacivms JPN3 / JAPAN / 2013 IRES domain is the nucleotide sequence of SEQ ID NO:93.
[0396] In embodiments, the 3' UTR is a poly-adenine (pA) 3' UTR, a mtRNRl-AES 3' UTR, a mtRNRl-LSPl 3' UTR, an AES-mtRNRl 3' UTR, an AES-hBg 3' UTR, a 2hBg 3' UTR, a FCGRT-hBg 3' UTR, or an HBA1 3' UTR.
[0397] In embodiments, the 3' UTR includes the nucleotide sequence of any one of SEQ ID NOs:95-103. In embodiments, the 3' UTR includes the nucleotide sequence of SEQ ID NO:95. In embodiments, the 3' UTR includes the nucleotide sequence of SEQ ID NO:96. In embodiments, the 3' UTR includes the nucleotide sequence of SEQ ID NO:97. In embodiments, the 3' UTR includes the nucleotide sequence of SEQ ID NO:98. In embodiments, the 3' UTR includes the nucleotide sequence of SEQ ID NO:99. In embodiments, the 3' UTR includes the nucleotide sequence of SEQ ID NO: 100. In embodiments, the 3' UTR includes the nucleotide sequence of SEQ ID NO: 101. In embodiments, the 3' UTR includes the nucleotide sequence of SEQ ID NO: 102. In embodiments, the 3' UTR includes the nucleotide sequence of SEQ ID NO: 103.
[0398] In embodiments, the 3' UTR is the nucleotide sequence of any one of SEQ ID NOs:95-103. In embodiments, the 3' UTR is the nucleotide sequence of SEQ ID NO:95. In embodiments, the 3' UTR is the nucleotide sequence of SEQ ID NO:96. In embodiments, the 3' UTR is the nucleotide sequence of SEQ ID NO:97. In embodiments, the 3' UTR is the 172 nucleotide sequence of SEQ ID NO:98. In embodiments, the 3' UTR is the nucleotide sequence of SEQ ID NO:99. In embodiments, the 3' UTR is the nucleotide sequence of SEQ ID NO: 100. In embodiments, the 3' UTR is the nucleotide sequence of SEQ ID NO: 101. In embodiments, the 3' UTR is the nucleotide sequence of SEQ ID NO: 102. In embodiments, the 3' UTR is the nucleotide sequence of SEQ ID NO: 103.
[0399] In embodiments, the 3' UTR is a poly-adenine (pA) 3' UTR. In embodiments, the 3' UTR is a poly-adenine (pA) 3' UTR is a poly(A) domain.
[0400] In embodiments, the poly-adenine (pA) 3' UTR includes from about 30 to about 200 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes about 40 to about 200 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes about 50 to about 200 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes about 60 to about 200 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes about 70 to about 200 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes about 80 to about 200 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes about 90 to about 200 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes about 100 to about 200 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes about 110 to about 200 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes about 120 to about 200 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes about 130 to about 200 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes about 140 to about 200 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes about 150 to about 200 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes about 160 to about 200 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes about 165 to about 200 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes about 170 to about 200 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes about 180 to about 200 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes about 190 to about 200 nucleotides.
[0401] In embodiments, the poly-adenine (pA) 3' UTR includes from about 30 to about 190 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes from about 30 to about 180 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes from about 30 to about 170 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes from about 30 to about 165 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes from about 30 to about 160 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes from about 30 to about 150 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes from about 30 to about 140 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes from about 30 to about 130 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes from about 30 to about 120 nucleotides. In embodiments, the polyadenine (pA) 3' UTR includes from about 30 to about 110 nucleotides. In embodiments, the poly-ademne (pA) 3' UTR includes from about 30 to about 100 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes from about 30 to about 90 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes from about 30 to about 80 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes from about 30 to about 80 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes from about 30 to about 80 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes from about 30 to about 80 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes from about 30 to about 80 nucleotides.
[0402] In embodiments, the poly-adenine (pA) 3' UTR includes at least 30 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes at least 40 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes at least 50 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes at least 60 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes at least 70 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes at least 80 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes at least 90 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes at least 100 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes at least 110 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes at least 120 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes at least 130 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes at least 140 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes at least 150 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes at least 160 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes at least 165 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes at least 170 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes at least 180 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes at least 190 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes at least 200 nucleotides.
[0403] In embodiments, the poly-adenine (pA) 3' UTR includes 165 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR is 165 nucleotides in length. In embodiments, the poly-adenine (pA) 3' UTR includes the nucleotide sequence of SEQ ID NO:95. In embodiments, the poly-adenine (pA) 3' UTR is the nucleotide sequence of SEQ ID NO:95.
[0404] In embodiments, the 3' UTR includes a wildtype Woodchuck Hepatitis Virus Posttranscriptional Regulatory Element (WPRE) 3' UTR, a modified WSPE 3' UTR, an mtRNRl 3' UTR, an AES 3' UTR, an LSP1 3' UTR, an hBg 3' UTR, a FCGRT 3' UTR, or a HBA1'3' UTR.
[0405] In embodiments, the 3' UTR includes a wildtype (WT) Woodchuck Hepatitis Virus Posttranscriptional Regulatory Element (WPRE) 3' UTR. In embodiments, the WT WPRE 3' UTR includes the nucleotide sequence of SEQ ID NO:96. In embodiments, the WT WPRE 3' UTR is the nucleotide sequence of SEQ ID NO:96.
[0406] In embodiments, the 3' UTR includes a modified WSPE 3' UTR. In embodiments, the modified WPRE 3' UTR includes the nucleotide sequence of SEQ ID NO:97. In embodiments, the modified WPRE 3' UTR is the nucleotide sequence of SEQ ID NO:97.
[0407] In embodiments, the 3' UTR includes an mtRNRl 3' UTR. In embodiments, the mtRNRl 3' UTR includes the nucleotide sequence of SEQ ID NO:98. In embodiments, the mtRNRl 3' UTR is the nucleotide sequence of SEQ ID NO:98.
[0408] In embodiments, the 3' UTR includes an AES 3' UTR. In embodiments, the AES 3' UTR includes the nucleotide sequence of SEQ ID NO:99. In embodiments, the AES 3' UTR is the nucleotide sequence of SEQ ID NO:99.
[0409] In embodiments, the 3' UTR includes an LSP1 3' UTR. In embodiments, the LSP1 3' UTR includes the nucleotide sequence of SEQ ID NO: 100. In embodiments, the LSP1 3' UTR is the nucleotide sequence of SEQ ID NO: 100.
[0410] In embodiments, the 3' UTR includes an hBg 3' UTR. In embodiments, the hBg 3' UTR includes the nucleotide sequence of SEQ ID NO: 101. In embodiments, the hBg 3' UTR is the nucleotide sequence of SEQ ID NO: 101.
[0411] In embodiments, the 3' UTR includes a FCGRT 3' UTR. In embodiments, the FCGRT 3' UTR includes the nucleotide sequence of SEQ ID NO: 102. In embodiments, the FCGRT 3' UTR is the nucleotide sequence of SEQ ID NO: 102.
[0412] In embodiments, the 3' UTR is a mtRNRl-AES 3' UTR. In embodiments, the mtRNRl-AES 3' UTR includes the nucleotide sequence of SEQ ID NO:98. In embodiments, the mtRNRl-AES 3' UTR includes the nucleotide sequence of SEQ ID NO:99. In embodiments, the mtRNRl-AES 3' UTR includes the nucleotide sequences of SEQ ID NO:98 and SEQ ID NO:99. In embodiments, the mtRNRl-AES 3' UTR includes from 5' to 3' the nucleotide sequences of SEQ ID NO:98 and SEQ ID NO:99.
[0413] In embodiments, the 3' UTR is a mtRNRl-LSPl 3' UTR. In embodiments, the mtRNRl-LSPl 3' UTR includes the nucleotide sequence of SEQ ID NO:98. In embodiments, the mtRNRl-LSPl 3' UTR includes the nucleotide sequence of SEQ ID NO: 100. In embodiments, the mtRNRl-LSPl 3' UTR includes the nucleotide sequences of SEQ ID NO:98 and SEQ ID NO: 100. In embodiments, the mtRNRl-LSPl 3' UTR includes from 5' to 3' the nucleotide sequences of SEQ ID NO:98 and SEQ ID NO:100..
[0414] In embodiments, the 3' UTR is an AES-mtRNRl 3' UTR. In embodiments, the AES-mtRNRl 3' UTR includes the nucleotide sequence of SEQ ID NO:99. In embodiments, the AES-mtRNRl 3' UTR includes the nucleotide sequence of SEQ ID NO:98. In embodiments, the AES-mtRNRl 3' UTR includes the nucleotide sequences of SEQ ID NO:99 and SEQ ID NO:98. In embodiments, the AES-mtRNRl 3' UTR includes from 5' to 3' the nucleotide sequences of SEQ ID NO:99 and SEQ ID NO:98.
[0415] In embodiments, the 3' UTR is an AES-hBg 3' UTR. In embodiments, the AES-hBg 3' UTR includes the nucleotide sequence of SEQ ID NO:99. In embodiments, the AES-hBg 3' UTR includes the nucleotide sequence of SEQ ID NO: 101. In embodiments, the AES-hBg 3' UTR includes the nucleotide sequences of SEQ ID NO:99 and SEQ ID NO: 101. In embodiments, the AES-hBg 3' UTR includes from 5' to 3' the nucleotide sequences of SEQ ID NO:99 and SEQ ID NO:101.
[0416] In embodiments, the 3' UTR is a 2hBg 3' UTR. In embodiments, the 2hBg 3' UTR includes two hBg 3' UTRs. In embodiments, the hBg 3' UTR includes the nucleotide sequence of SEQ ID NO: 101. In embodiments, the hBg 3' UTR is the nucleotide sequence of SEQ ID NO: 101.
[0417] In embodiments, the 3' UTR is a FCGRT-hBg 3' UTR. In embodiments, the FCGRT-hBg 3' UTR includes the nucleotide sequence of SEQ ID NO: 102. In embodiments, the FCGRT-hBg 3' UTR includes the nucleotide sequence of SEQ ID NO: 101. In embodiments, the FCGRT-hBg 3' UTR includes the nucleotide sequences of SEQ ID NO: 102 and SEQ ID NO: 101. In embodiments, the FCGRT-hBg 3' UTR includes from 5' to 3' the nucleotide sequences of SEQ ID NO: 102 and SEQ ID NO: 101.
[0418] In embodiments, the 3' UTR is an HBA1 3' UTR. In embodiments, the HBA1 3' UTR includes the nucleotide sequence of SEQ ID NO: 103. In embodiments, the HBA1 3' UTR is the nucleotide sequence of SEQ ID NO: 103.
[0419] In embodiments, the poly-adenine (polyA) domain includes from about 30 to about 200 nucleotides. In embodiments, the poly-adenine (polyA) domain includes about 40 to about 200 nucleotides. In embodiments, the poly-adenine (polyA) domain includes about 50 to about 200 nucleotides. In embodiments, the poly-adenine (polyA) domain includes about 60 to about 200 nucleotides. In embodiments, the poly-adenine (polyA) domain includes about 70 to about 200 nucleotides. In embodiments, the poly-adenine (polyA) domain includes about 80 to about 200 nucleotides. In embodiments, the poly-adenine (polyA) domain includes about 90 to about 200 nucleotides. In embodiments, the poly-adenine (polyA) domain includes about 100 to about 200 nucleotides. In embodiments, the polyadenine (polyA) domain includes about 110 to about 200 nucleotides. In embodiments, the poly-adenine (polyA) domain includes about 120 to about 200 nucleotides. Tn embodiments, the poly-adenine (polyA) domain includes about 130 to about 200 nucleotides. In embodiments, the poly-adenine (polyA) domain includes about 140 to about 200 nucleotides. In embodiments, the poly-adenine (polyA) domain includes about 150 to about 200 nucleotides. In embodiments, the poly-adenine (polyA) domain includes about 160 to about 200 nucleotides. In embodiments, the poly-adenine (polyA) domain includes about 165 to about 200 nucleotides. In embodiments, the poly-adenine (polyA) domain includes about 170 to about 200 nucleotides. In embodiments, the poly-adenine (polyA) domain includes about 180 to about 200 nucleotides. In embodiments, the poly-adenine (polyA) domain includes about 190 to about 200 nucleotides.
[0420] In embodiments, the poly-adenine (polyA) domain includes from about 30 to about 190 nucleotides. In embodiments, the poly-adenine (polyA) domain includes from about 30 to about 180 nucleotides. In embodiments, the poly-adenine (polyA) domain includes from about 30 to about 170 nucleotides. In embodiments, the poly-ademne (polyA) domain includes from about 30 to about 165 nucleotides. In embodiments, the poly-adenine (polyA) domain includes from about 30 to about 160 nucleotides. In embodiments, the poly-adenine (polyA) domain includes from about 30 to about 150 nucleotides. In embodiments, the polyadenine (polyA) domain includes from about 30 to about 140 nucleotides. In embodiments, the poly-adenine (polyA) domain includes from about 30 to about 130 nucleotides. In embodiments, the poly-adenine (polyA) domain includes from about 30 to about 120 nucleotides. In embodiments, the poly-adenine (polyA) domain includes from about 30 to about 110 nucleotides. In embodiments, the poly-adenine (polyA) domain includes from about 30 to about 100 nucleotides. In embodiments, the poly-adenine (polyA) domain includes from about 30 to about 90 nucleotides. In embodiments, the poly-adenine (polyA) domain includes from about 30 to about 80 nucleotides. In embodiments, the poly-adenine (polyA) domain includes from about 30 to about 80 nucleotides. In embodiments, the polyadenine (polyA) domain includes from about 30 to about 80 nucleotides. In embodiments, the poly-adenine (polyA) domain includes from about 30 to about 80 nucleotides. In embodiments, the poly-adenine (polyA) domain includes from about 30 to about 80 nucleotides.
[0421] In embodiments, the poly-adenine (polyA) domain includes at least 30 nucleotides. In embodiments, the poly-adenine (polyA) domain includes at least 40 nucleotides. In embodiments, the poly-adenine (polyA) domain includes at least 50 nucleotides. In embodiments, the poly-adenine (polyA) domain includes at least 60 nucleotides. In embodiments, the poly-adenine (polyA) domain includes at least 70 nucleotides. In embodiments, the poly-adenine (polyA) domain includes at least 80 nucleotides. In embodiments, the poly-adenine (polyA) domain includes at least 90 nucleotides. In embodiments, the poly-adenine (polyA) domain includes at least 100 nucleotides. In embodiments, the poly-adenine (polyA) domain includes at least 110 nucleotides. In embodiments, the poly-adenine (polyA) domain includes at least 120 nucleotides. In embodiments, the poly-adenine (polyA) domain includes at least 130 nucleotides. In embodiments, the poly-adenine (polyA) domain includes at least 140 nucleotides. In embodiments, the poly-adenine (polyA) domain includes at least 150 nucleotides. In embodiments, the poly-adenine (polyA) domain includes at least 160 nucleotides. In embodiments, the poly-adenine (polyA) domain includes at least 165 nucleotides. In embodiments, the poly-adenine (polyA) domain includes at least 170 nucleotides. In embodiments, the poly-adenine (polyA) domain includes at least 180 nucleotides. In embodiments, the poly-adenine (polyA) domain includes at least 190 nucleotides. In embodiments, the poly-adenine (polyA) domain includes at least 200 nucleotides.
[0422] In embodiments, the poly-adenine (polyA) domain includes 165 nucleotides. In embodiments, the poly-adenine (polyA) domain is 165 nucleotides in length. In embodiments, the poly-adenine (polyA) domain includes the nucleotide sequence of SEQ ID NO:95. In embodiments, the poly-adenine (poly A) domain is the nucleotide sequence of SEQ ID NO:95.
[0423] In embodiments, the linear RNA compound further includes a transcription termination domain. In embodiments, the transcription termination domain includes the nucleotide sequence of SEQ ID NO: 104. In embodiments, the transcription termination domain is the nucleotide sequence of SEQ ID NO: 104.
[0424] In embodiments, the linear RNA compound further includes a promoter. In embodiments, the promoter is a T7 promoter. In embodiments, the T7 promoter includes the nucleotide sequence of SEQ ID NO: 117. In embodiments, the T7 promoter is the nucleotide sequence of SEQ ID NO: 117.
[0425] In embodiments, the method of producing circularized RNA does not include RNaseR enrichment or high performance liquid chromatography (HPLC) purification. In embodiments, the method of producing circularized RNA does not include RNaseR enrichment. In embodiments, the method of producing circularized RNA does not include high performance liquid chromatography (HPLC) purification.
[0426] In embodiments, the circularized RNA is purified and isolated using known methods in the art. In embodiments, the purification and / or isolation of circularized RNA does not include RNaseR enrichment or high performance liquid chromatography (HPLC) purification. In embodiments, the purification and / or isolation of circularized RNA does not include RNaseR enrichment. In embodiments, the purification and / or isolation of circularized RNA does not include high performance liquid chromatography (HPLC) purification.
[0427] In embodiments, the nucleotide sequence of any one of SEQ ID NOs: 1 -24 or 63142 is an RNA sequence or a DNA sequence. In embodiments, the nucleotide sequence of any one of SEQ ID NOs: 1-24 or 63-142 is an RNA sequence. In embodiments, the thymine nucleotide residues of any one of SEQ ID NOs: 1-24 or 63-142 are uracil nucleotide residues. In embodiments, the nucleotide sequence of any one of SEQ ID NOs: 1 -24 or 63-142 is a DNA sequence. EXAMPLES Example 1
[0428] RNAs are a powerful therapeutic class, but their inherent transience limits activity. Circularization can improve their persistence, but simple and scalable approaches to achieve this are lacking. To facilitate the pursuit of circular RNAs (cRNAs), we developed two methods: “outside developed” cRNAs (ocRNAs) via in vitro circularization using group II introns, and in situ or “inside developed” cRNAs (icRNAs) via in cell circularization using the ubiquitously expressed RtcB protein. We also developed simple purification protocols for these enabling high yields (40-75%) while maintaining low immune responses. The resulting simplicity and scalability of production facilitated a range of applications from stem cell engineering to robust genome and epigenome targeting via zinc finger proteins. We anticipate our versatile and simple to implement circular RNA toolset will have broad utility in research and biotechnology applications. Example 2: Robust genome and cell engineering via in vitro and in situ circularized RNAs
[0429] Abstract
[0430] Circularization can improve RNA persistence, yet simple and scalable approaches to achieve this are lacking. Here we report two methods that facilitate the pursuit of circular RNAs (cRNAs): cRNAs developed via in vitro circularization using group II introns, and cRNAs developed via in-cell circularization by the ubiquitously expressed RtcB protein. We also report simple purification protocols that enable high cRNA yields (40-75%) while maintaining low immune responses. These methods and protocols facilitate a broad range of applications in stem cell engineering as well as robust genome and epigenome targeting via zinc finger proteins and CRISPR- Cas9. Notably, cRNAs bearing the encephalomyocarditis internal ribosome entry enabled robust expression and persistence compared with linear capped RNAs in cardiomyocytes and neurons, which highlights the utility of cRNAs in these non-dividing cells. We also describe genome targeting via deimmunized Cas9 delivered as cRNA and a long-range multiplexed protein engineering methodology for the combinatorial screening of deimmunized protein variants that enables compatibility between persistence of expression and immunogenicity in cRNA-delivered proteins. The cRNA toolset will aid research and the development of therapeutics.
[0431] Introduction
[0432] RNAs have emerged as a powerful therapeutic class. However, their typically short half-life impacts their activity both as an interacting moiety (such as short interfering RNAs) and a template (such as messenger RNAs). Towards this, RNA stability has been modulated using a host of approaches, including engineering untranslated regions (UTRs), modulating secondary structures, incorporating cap analogues, modifying nucleosides and optimizing 181 codons1 More recently, circularization strategies that remove free ends necessary for exonuclease-mediated degradation thereby rendering RNAs resistant to most mechanisms of turnover have emerged as a particularly promising methodology6 l5.
[0433] Simple and scalable approaches to achieve efficient production and purification of circular RNAs (cRNAs) are however lacking, thus limiting their broader application. Towards this, we developed two methods: ‘outside developed’ cRNA (ocRNA) via in vitro circularization using group II introns, and in situ or ‘inside developed’ cRNA (icRNA) via incell circularization using the ubiquitously expressed RtcB protein. As cRNA production commonly involves RNaseR enrichment and high performance liquid chromatography (HPLC) purification to reduce immune responses prompted by double-stranded RNA produced during in vitro transcription (IVT), resulting RNA yields are low (typically less than 1% of the input RNA6). We thus optimized protocols that incorporated alternatives and our resultant methods use an HPLC-free purification process while maintaining low immunogenicity, persistence and robust expression. Importantly, they enabled both high yields and cost-effectiveness, which we leveraged to demonstrate a range of applications from stem cell engineering to robust genome and epigenome targeting via zinc finger (ZF) proteins and CRISPR-Cas9 systems.
[0434] Common to all these applications enabled via cRNA delivery is the critical consideration of their immune system interactions. Although for some applications, such as vaccines, robust immune responses to a delivered transgene are desirable, for other applications, such as genome and epigenome targeting, immune responses can instead inhibit therapeutic effects1617. Inducing immune responses through transgene expression enabled by RNA delivery has been extensively researched in vaccine development and proven through the success of COVID vaccines based on this technology16 21. However, despite substantial engineering efforts, deimmunization remains a tougher problem to crack22. Thus, to facilitate compatibility between persistence of expression and immunogenicity, especially when delivering non-human payloads via cRNAs. we also concurrently developed a long-range multiplexed (LORAX) protein engineering methodology based on high-throughput screening of combinatorially deimmunized protein variants. We applied it to identify a Cas9 variant with seven key human leukocyte antigen (HLA) restricted epitopes simultaneously immunosilenced after a single round of screening, and showed that cRNA-mediated delivery of the same enabled genome targeting.
[0435] Results
[0436] Engineering ocRNAs and icRNAs
[0437] To engineer ocRNAs, we adopted the group II intron from Clostridium tetani23 25, where domains I, II and III were split from domains V and VI at domain IV, to preserve the catalytic and structurally relevant properties for intron splicing while permitting circularization of our RNA construct. We permuted the domains of the intron to have domains V and VI at the 5' end of our construct and domains I to III at the 3' end of our construct. Since the C. tetani group II intron is seen to have alternative splicing with four groups of DV / DVI domains (A, B, C and D), we kept a minor leader sequence to domain DV / DVI group A as it was likely to house the crucial exon binding site (EBS) required for the domain I intron sequence to bind for selective splicing. Importantly, the C. tetani intron is ORF-less, suggesting that splicing action is not dependent on an intron-encoded protein and thus compatible with IVT applications. Finally, to improve long-range interactions, we added flanking twister ribozymes that rapidly self-cleave during IVT (removing the immunogenic 5' phosphate) and enable hybridization of downstream complementary' ligation stems to one another. The above circularization mechanics are adjacent to an internal ribosome entry site (IRES) (25, 26) coupled to an mRNA of interest and a 3' UTR (FIG. 1 A). We analyzed the junction via Sanger sequencing of reverse transcription PCR (RT-PCR) product from IVT and found a small 26 nt splicing scar split between proposed EBS of each permuted domain group, confirming circularization (FIG. IB). In further characterization, we observed three species on tapestation that correspond with the lengths of each expected component: the intron arms, circularized product and precursor (FIG. IC). Evaluating IVT circularization efficiency of the group II intron, we found yields to be consistently around 70% (FIG. ID). In downstream dsRNA purification, we also noted that the ocRNA construct retained an approximate 60% yield (FIG. IE).
[0438] To engineer icRNAs. we took a simplified approach that did not require engineering beyond flanking the IRES and payload with twister ribozymes (FIG. IF). Specifically, to engineer icRNAs, we generated in vitro transcribed linear RNAs that bear a twister ribozyme flanked IRES26,27 coupled to an mRNA of interest and a 3' UTR. Once transcribed, the flanking twister ribozymes rapidly self-cleave, enabling hybridization of the complementary ligation stems to one another. Upon delivery into cells, these linear RNAs are then circularized in situ via ligation of the proximal 5' and 3' ends by the ubiquitous cellular RNA ligase RtcB. Thus, contrary to the ocRNA, this construct completely relies on the endogenous RtcB action to circularize. As a first check of circularization, we confirmed that the ligation junction was mapped correctly to its predicted location via Sanger sequencing of the RT-PCR product from total RNA (FIG. 1G). To assess in situ circularization efficiency. HEK293T cells were transfected with icRNAs encoding for green fluorescent protein (GFP) or with the same icRNAs pre-circularized in vitro using RTCB ligation in a test tube followed by RNaseR treatment to enrich circularized RNA species. RT-PCR was performed using outward-facing primers and normalized to GAPDH expression. Across independent icRNA treatments, we observed efficient (-20%) in situ circularization levels for icRNAs compared with pre-circularized RNA (FIG. 1H). However, since RNA degrades with time in situ, we wanted to observe the proportions of icRNA as best we could with cell proliferation. With RNA sequencing (RNAseq) data across the same three time points of 6 h, 24 h and 48 h, we observed a net increase of cRNA proportion, consistent with circularization increasing over time and depletion of linear RNA through endonucleolytic attack (FIG. II). Notably, the largest change in proportion occurs within the first 24 h, probably owing to the short half-life of linear RNA. As another side-by-side comparison, we also engineered linear in situ circularization defective RNAs (icdRNAs) by utilizing catalytically inactive mutants of the twister ribozymes. Specifically, HEK293Ts were transfected with circular GFP icRNA or linear icdRNA and RNA was isolated at 6h, 1 day, 2 days and 3 days after transfection. We observed similar amounts of GFP RNA at 6h (FIG. 7A, left), confirming that approximately equal quantities of icRNA and icdRNA were delivered to cells. However, GFP RNA with functional circularization was much higher at days 1, 2 and 3 than icdRNA, indicating improved RNA persistence via circularization. This improved RNA persistence also correlated with increased GFP translation at 3 days FIG. 7A, middle). To confirm that icRNAs were covalently circularized in cells upon delivery in vitro, we performed RT-PCR by designing outward-facing primers that selectively amplified only the circularized RNA molecules. Indeed, we only observed a PCR product for icRNAs, confirming successful circularization (FIG. 7A, right).
[0439] Engineering improved translation from cRNAs
[0440] To improve protein translation from cRNAs. we next screened a panel of 19 naturally occurring and synthetic IRES sequences2832. We found the 6A form of the encephalomyocarditis virus IRES (EMCV; FIG. 2A), the enterovirus 71 IRES (FIG. 2A, #18) and the Coxsackievirus B3 IRES (FIG. 2A, #19)6 to be the best at enabling cap-independent protein translation. To further improve protein translation, we also screened a panel of 3' UTRs in the context of the EMCV 6A IRES. Addition of a Woodchuck Hepatitis Virus Posttranscriptional Regulatory Element (WPRE) and a poly(A) stretch improved relative protein translation by over fivefold, but additional 3' UTRs5 did not improve translation efficiency (FIG. 2B). This is consistent with similar improvements in linear RNA, as poly(A) binding proteins (PABPs) are known to interact with the translation complex from the open 3' end of the RNA structure in a sandwich conformation33,34. Therefore, we hypothesized that a poly(A) tail in this conformation assists in recruiting PABPs in cRNAs. For all subsequent studies, we thus used cRNAs based on the EMCV IRES coupled to a modified WPRE and 165 bp poly (A) stretch (FIG. 2B, #15). Here a crucial ‘Y-stem’ secondary structure noted for recruiting the core translation complex eIF4G was found to have high stacking probability with a short sequence of the WPRE via RNAfold, and thus the sequence was removed from the WPRE, recovering an approximate twofold improvement in expression35,36. Notably, these designs showed robust activity' when benchmarked with contemporary' optimized cRNA constructs15 (FIG. 7B).
[0441] Developing facile purification protocols for cRNAs
[0442] Conventional purification strategies rely on HPLC or enzy me-based techniques, which can be complex to implement and / or costly and / or yield limiting and variable results (as seen by incomplete RNaseR reactions). Therefore, we sought to utilize simpler and more economical methods. Specifically, for icRNAs, we included urea (0.8 M) as a denaturing agent during the IVT process to reduce the highly immunogenic and spurious dsRNA production inherent to transcription via T7 polymerases37. For ocRNA constructs, optimal circularization yields were obtained via overnight IVT reactions but in the absence of urea (as crucial secondary structures for splicing are abrogated upon addition of denaturing agents) (FIG. 7C). For ocRNA purification, we thus instead relied on cellulose chromatography38 that can selectively remove dsRNA above 30 bp in length (which is above that of the crucial IRES secondary structures within cRNAs). Finally, resulting RNA was treated with phosphatase. Next, we evaluated the immune response of thus purified icRNAs and ocRNAs versus linear RNA benchmarks in A549 cells. As reported in previous literature, linear unmodified RNA demonstrated a large immune response, but linear RNA with N1-methyl pseudoundine-5'-triphosphate (mlT) modification demonstrated negligible responses from the retinoic acidinducible gene I (RIG-I), interleukin-6 (IL6) and interferon beta (IFNB) immune markers914 (FIG. 2C, top). Notably, icRNA and ocRNA constructs demonstrated mild immune responses, with the ocRNA (gel) demonstrating almost undetectable immune responses (similar to the linear capped and modified RNAs). Accordingly, all the circular constructs also demonstrated robust cell viability (FIG. 2D). While HPLC purification will be the method of choice of clinical translation, the cRNA purification approaches above provide a viable alternative for most in vitro and in vivo applications.
[0443] Functional characterization of cRNAs
[0444] We next sought to validate cRNA persistence across relevant in vitro and in vivo settings. We first confirmed that cRNAs were indeed improved in persistence compared with their linear counterparts in HEK293Ts (FIG. 2E). However, consistent with previous observations comparing IRES-based cRNAs versus 5' capped linear RNA14,13, we observed that the latter (FIG. 2B) enabled higher instantaneous protein translation 24 h post delivery (2-3-fold). Hypothesizing that IRES translation may be cell type specific39, we thus screened a variety of cells to determine the optimal cellular contexts to deploy our cRNAs (FIG. 3A). Specifically, since the IRES that we chose was the EMCV IRES, we rationalized that a test in neurons and cardiomyocytes could demonstrate higher efficiencies as it represents a closer cellular environment to what is natural for the parent cardiovirus type. To that end. we transfected icRNAs containing the EMCV, modified WPRE, and a 165 poly(A) stretch or a linear capped RNA with m IT nucleotide substitution and optimized UTRs into stem cell-derived neurons (FIG. 3B, left). Interestingly, we found that the icRNAs provide extended expression and increased translation after a short delay (FIG. 3B, right). Next, we transfected the same panel with the additional conditions of ocRNAs and defective icdRNAs into stem cell-derived cardiomyocytes (FIG. 3C, left). Interestingly, our observations confirmed that both EMCV-driven circularized RNA constructs exhibited higher or equivalent expression levels on day 1 compared with linear RNA. Notably, they again rapidly surpassed linear RNA in expression by day 2 and far past their own day 1 expression (FIG. 3C, middle, bottom). Moreover, linear and circularization defective icdRNAs demonstrated rapid and similar decreases over the course of a week. To ensure that trends were reproducible, the experiment was repeated and image analysis performed in triplicate at a lower exposure to optimize representation of GFP intensity, resulting in a nearly identical profile (FIG. 7D). On day 30, cells were collected, and qPCR analysis of GFP RNA was performed (FIG. 3C, right). Results indicated a considerably lower RNA level of linear RNAs compared with icRNAs and ocRNAs further confirming persistence of circular species.
[0445] Overall, these results suggest a potential cellular context dependence on the activity of IRESs and / or a preference for non-dividing cells, which compared with dividing cells do not as rapidly dilute RNA levels post transfection. On that note, we hypothesized that the differentiation of stem cells into non-dividing cells might similarly better leverage the heightened expression and apparent protein accumulation provided by persistent icRNA. Towards this, we were indeed able to robustly differentiate H1 human pluripotent stem cells (hPSCs) to neurons within 1 week via a single transfection of circular NEURODI RNAs (FIG. 3D).
[0446] Finally, we sought to extend these results in vivo. To ensure that icRNAs were also capable of circularization in vivo, we generated lipid nanoparticles (LNPs) containing either circular icRNA or icdRNA and retro-orbitally injected them into mice. LNPs for each condition w ere successfully generated of similar size (FIG. 8A, left). Livers w ere isolated 3 and 7 days after injection and quantification and visualization by RT-qPCR confirmed circularization in vivo persisting for at least 7 days (FIG. 8A, middle, right). To benchmark against true linear RNA, we synthesized LNPs40,41 bearing either icRNA or linear RNAs (FIG. 8B, left). RNA / LNPs were retro-orbitally injected into C57BL / 6 mice and liver transcriptomes assayed 7 days after injection. While linear RNA was largely degraded by day 187 7, RT-qPCR showed that icRNAs persisted (FIG. 8B, middle). Furthermore, RT-PCR also again confirmed successful in situ circularization of icRNAs in vivo (FIG. 8B, right).
[0447] Genome and epigenome engineering viaZF proteins delivered as cRNAs
[0448] Spurred by the above results, we hypothesized that this increased persistence of cRNAs, in addition to enabling applications entailing sustained transgene expression, could also facilitate efficient genome and especially epigenome targeting. Towards this, we first explored icRNA utility in the context of ZF proteins, as being a solely protein-based genome engineering toolset we anticipated that ZFs would be particularly suited for this mode of delivery. Indeed, we observed more efficient genome editing via zinc finger nuclease (ZFN) icRNAs compared with corresponding icdRNAs targeting the GFP and CCR5 genes42,43 (FIG. 4A). Building on this observation, we next investigated delivery of ZF-KRAB proteins to programmably repress genomic targets. We focused on PCSK9, a gene that encodes an enzyme that regulates low-density lipoprotein receptor degradation. Loss-of-function mutations in PCSK9 are associated with reduced risk of cardiovascular disease w ith no documented adverse side effects44 46. Antibodies, antisense oligonucleotides and CRISPRs have all been utilized to target PCSK947 53. and here we explored if a transient pulse of ZF epigenome regulators could enable repression of PCSK9. First, we screened in HeLa cells a panel of 67 ZF-KRAB proteins delivered as icRNAs (FIG. 4B, top). These ZFs tiled across 0-1,200 bp of the transcription start site (FIG. 4B, middle) of hPCSK9, and notably, 16 of these enabled robust gene repression (90-50%) (FIG. 4B, bottom). Next, to achieve inheritable repression, we fused a 3A3L DNA methylator with the ZF-KRAB modules’4'55. Notably, we observed sustained repression of PCSK9 over a 12-day period (FIG. 4C). Taken together, these results establish that a transient pulse of ZF epigenome regulators delivered as cRNAs can enable robust gene repression, enabling a genomically scarless and safe approach for modulating therapeutic gene expression.
[0449] Genome and epigenome engineering via deimmunized CRISPR-Cas delivered as cRNAs
[0450] Building on these results with ZF proteins, we next explored if cRNA persistence could similarly enhance the activity of CRISPR-Cas9 systems. However, unlike for ZFs, which are built on a human protein chassis, this feature of persistence may aggravate immune responses in therapeutic settings for CRISPR systems as those are derived from prokaryotes, including some residing in the human gut microbiome56 6". Thus, to enable compatibility between persistence of expression and immunogenicity, we sought first to develop a methodology to screen progressively deimmunized Cas9 proteins by combinatorially mutating particularly immunogenic epitopes61.
[0451] While variant library7 screening has proven to be an effective approach to protein engineering62 69. applying it to deimmunization faces three important technical challenges. One, the need to mutate multiple sites simultaneously across the full length of the protein; two, reading out the associated combinatorial mutations scattered across large (>1 kb) regions of the protein via typical short-read sequencing platforms; and three, engineering fully degenerate combinatorial libraries, which can very7 quickly balloon to unmanageable numbers of variants22,70. To overcome these challenges, we developed several methodological innovations, which, taken together, comprise a LORAX protein engineering method capable of screening millions of combinatorial variants simultaneously with mutations spread across the full length of arbitrarily large proteins (FIG. 5).
[0452] Towards library design, to narrow down the vast mutational space associated with combinatorial libraries, we utilize an approach guided by evolution and natural variation71,72. As deimmunizing protein engineering seeks to alter the amino acid sequence of a protein without disrupting functionality, it is extremely useful to narrow down mutations to those less likely to result in non-functional variants. To identify these mutants, we generated large alignments of Cas9 orthologues from publicly available data to identify low-frequency single nucleotide polymorphisms (SNPs) that have been observed in natural environments. Such variants are likely to have a limited effect on protein function, as highly deleterious alleles would tend to be quickly selected out of natural populations (if Cas9 activity is under purifying selection) and therefore not appear in sequencing data73. To further subset these candidate mutations, we evaluated for immunogenicity in silico using the netMHC epitope prediction software74,75, to determine to what degree the candidate mutations are likely to result in the deimmunization of the most immunogenic epitopes in which they appear. This is a critical step as many mutations may have little effect on overall immunogenicity76 7S. Screening for decreased peptide-major histocompatibility complex (MHC) class I binding filters out amino acid substitutions, which are likely immune neutral, substantially increasing the likelihood of functional hits with enough epitope variation to evade immune induction78-79.
[0453] Next, to enable readout, we applied long-read nanopore sequencing to measure the results of the screens of our combinatorial libraries. This circumvents the limit of short target regions and obviates the need for barcodes altogether by single-molecule sequencing of the entire target gene, enabling library design strategies that can explore any region of the protein in combination with any other region without any complicated cloning procedures required to facilitate barcoding80. So far, the adoption of nanopore sequencing has been limited by its high error rate, around 95% accuracy per DNA base81, compared with established short-read techniques, which are multiple orders of magnitude more accurate. To address this challenge, we designed our libraries such that each variant that we engineered would have multiple nucleotide changes for each single target amino acid change, effectively increasing the sensitivity of nanopore-based readouts with increasing numbers of nucleotide changes per library member. The large majority of amino acid substitutions are amenable to a library design paradigm in which each substitution is encoded by two, rather than one, nucleotide changes, owing to the degeneracy of the genetic code and the highly permissive third ‘wobble’ position of codons.
[0454] The scale of engineering that would be required to generate an effectively deimmunized Cas9 is not fully understood, as combinatorial deimmunization efforts at the scale of proteins thousands of amino acids long have not yet been possible. Therefore, to roughly estimate these parameters, we developed an immunogenicity scoring metric that takes into account all epitopes across a protein and the known diversity of MHC variants in a species weighted by population frequency to generate a single combined score representing the average immunogenicity of a full-length protein as a function of each of its immunogenic epitopes82. Formally, this score is calculated as Sty S";Wj (1 - log(kij x j^) Ix=y where Ix is the immunogenicity score of protein x, i is the epitopes, / is the HLA alleles, / is the allele specific standardization coefficient, w} is the HLA allele weights, kj is the predicted binding affinity of epitope / to allele / , and y is the protein specific scaling factor. We then predicted the overall effect of mutating the top epitopes in several Cas9 orthologues (FIG. 9A). As might be expected, this analysis suggests that single-epitope strategies are woefully inadequate to deimmunize a whole protein for multiple HLA ty pes, and also that there are diminishing returns as more and more epitopes are deimmunized. Our analysis suggests that it may require on the order of tens of deimmunized epitopes to make a notable impact on overall, population-wide protein immunogenicity. The scale of engineering demanded by these immunological facts has previously been intractable, but by applying LORAX, we conjectured that one could now make substantial steps, several mutations at a time, through the mutational landscape of the Cas9 protein.
[0455] Specifically, applying the procedure above, we designed a library of Cas9 variants based on the SpCas9 backbone containing 23 different mutations across 18 immunogenic epitopes (FIG. 5). Combining these in all possible combinations yields a library of 1,492,992 unique elements. With this design, we then constructed the library in a stepwise process. First, the full-length gene was broken up into short blocks of no more than 1,000 bp, which overlap by 30 bp on each end. Each block is designed such that it contains no more than four target epitopes to mutagenize. With few epitopes per block and few variant mutations per epitope, it becomes feasible to chemically synthesize each combination of mutations for each block. Each of these combinations was then synthesized and mixed at equal ratios to make a degenerate block mix. This was repeated for each of the blocks necessary to complete the full-length protein sequence via fusion PCR.
[0456] To identify functional variants still capable of editing DNA, we next designed and carried out a positive selection screen targeting the hypoxanthine phosphoribosyltransferase 1 (HPRT1) gene83. In the context of the screen, HPRT1 converts 6-thioguanine (6-TG), an analogue of the DNA base guanine, into 6-TG nucleotides that are cytotoxic to cells via incorporation into the DNA during S-phase84. Thus, only cells containing functional Cas9 variants capable of disrupting the HPRT1 gene can survive in 6-TG-containing cell culture media. To first identify the optimal 6-TG concentration, HeLa cells were transduced with lentivirus particles containing wild-type (WT) Cas9 and either an HPRT1-targeting guide RNA (gRNA) or a non-targeting guide. After selection with puromycin, cells were treated with 6-TG concentrations ranging from 0 pg ml'1 to 14 pg ml'1 for 1 week. Cells were stained with crystal violet at the end of the experiment and imaged. About 6 pg ml1 was selected as all cells containing anon-targeting guide had died while cells containing the HPRT1 guide remained viable (FIG. 9B).
[0457] To perform the screen, replicate populations of HeLa cells were transduced with lentiviral particles containing the variant SpCas9 library along with the HPRT1-targeting gRNA at 0.3 multiplicity of infection (MOI) and at greater than 75-fold coverage of the library elements. Cells were selected using puromycin after 2 days and 6-TG was added once cells reached 75% confluency. After 2 weeks, genomic DNA was extracted from remaining cells and full-length Cas9 amplicons were nanopore sequenced on the Oxford Nanopore (ONT) MinlON platform.
[0458] MinlON sequencing confirmed that the majority of the pre-screened library consists of Cas9 sequences with many mutations, with most falling into a broad peak between 6 and 14 mutations per sequence, each of which knocking out a key immunogenic epitope (FIG. 9C). Interestingly, the post-screening library was significantly shifted in the mutation density distribution, suggesting that the majority of the library’ with large (>4) numbers of mutations resulted in non-functional proteins that were unable to survive the screen. Meanwhile, WT, single and double mutants were generally enriched as these proteins proved more likely to retain functionality and pass through the screen (FIG. 9C). In addition, the two independent replicates of the screen showed strong correlation (R2 = 0.925) providing further evidence of robustness (FIG. 5). We also analyzed the change in overall frequency of mutations in the pre- and post-screen libraries to see if a pattern of mutation effects could be inferred. Although the WT allele was enriched at every site in the post-screen sequences, nearly every site retained a reasonable fraction of mutated alleles, suggesting that the mutations, at least individually, are fairly well tolerated and do not disrupt Cas9 functionality (FIG. 9D).
[0459] To select hits for downstream validation and analysis, we devised a method for differentiating high-support hits likely to be real from noise-driven false-positive hits. To do this, we hypothesized that the fitness landscape of the screen mutants is likely to be smooth; that is, variants that contain similar mutations are more likely to have similar fitness in terms of editing efficiency compared with randomly selected pairs85. We confirmed this by computing a predicted screen score for each variant based on a weighted regression of its nearest neighbors in the screen. This metric correlates well with the actual screen scores and approaches the screen scores even more closely as read coverage increases. This provides good evidence that the fitness landscape is indeed somewhat smooth (FIG. 10A). Next, we reasoned that because the fitness landscape is smooth, real hits should reside in broad fitness peaks, which include many neighbors that also show high screen scores, whereas hits that are less supported by near neighbors are more likely to be spurious as they represent non-smooth fitness peaks. Formalizing this logic, we performed a network analysis to differentiate noise-driven hits from bona fide hits by looking at the degree of connectivity with other hits (FIG. 6A).
[0460] Applying these analyses to the screen output led us to select and construct 20 variants (V1-20) for validation and characterization. We applied two independent methods to quantify editing of the deimmunized Cas9 variants. First, we performed a gene-rescue experiment using low-frequency homology-directed repair (HDR) to repair a genetically encoded broken GFP gene86 (FIG. 6B). Second, we quantified NHEJ-mediated editing by genomic DNA extraction and Illumina next-generation sequencing (NGS) using the CRISPResso2 package87 (FIG. 10B). Variants highly connected to neighbors were capable of editing, whereas those not connected were non-functional, validating the network-based approach that we used to select hits as enriching for truly functional sequences. Among the screen hits was the L616G mutation first identified in ref. 61 as a functional Cas9 variant with a critical immunodominant epitope deimmunized (VI). This concordance with previous work provided further confidence in our screening method. Interestingly, we discovered another deimmunizing mutation within the same epitope, L623Q (V2), which similarly retains Cas9 functionality, but appears to be more epistatically permissive, as many of our multi-mutation hits combine this mutation with other deimmunized epitopes. From these multi-mutation hits, we chose V4, which demonstrated high editing capability while still bearing simultaneous mutations across seven distinct epitopes, as well as family members V3, a variant bearing two mutations, and V5, a variant bearing the seven changes from V4 plus one additional mutation.
[0461] To confirm that mutation of these epitopes indeed elicited deimmunization, we assessed T cell response to WT and variant peptides by measuring IFNy secretion in the ELISpot assay17,61. We chose to use peripheral blood mononuclear cells (PBMCs) from three separate donors that carried the HLA-A*0201 allele as peptides were presented to cells using the TAP-deficient cell line T2 (HLA-A*0201 positive)88. Correspondingly, we synthesized peptides for epitopes 2, 7, 8, 9, 12, 15 and 16 as our predictions suggested that these epitopes would induce a reduction in immune response for the HLA-A*0201 allele (FIG. 11 A). Importantly, since SpCas9v4 carries four of these mutations, this assay would also provide confirmation of deimmunization for this variant. We found that mutant peptides for all epitopes tested indeed resulted in fewer spot-forming colonies for all three donors compared with WT peptides (FIG. 6C and FIG. 1 IB), thereby confirming our predictions. To assess whether the full-length protein is deimmunized, we next generated mRNA encoding for SpCas9WT and SpCas9V4 and electroporated this into the PBMC populations. As PBMCs have a mixture of both antigen-presenting cells (APCs) and T cells89 90, we are able to introduce the RNA to the APCs and measure T cell response via the ELISpot assay. Excitingly, we observed significantly fewer spot forming colonies for full-length SpCas9v4 compared with SpCas9WT and similarly when compared with SpCas9 L616G (FIG. 6D and FIGS. 11C-11D).
[0462] On the basis of this, we then further evaluated the efficacy of these mutants side by side with WT SpCas9 across a panel of genes and cell types, and assessed V4 activity across both targeted genome editing and epigenome regulation experiments91 (FIGS. 12A-12C). Together, these results confirmed that leveraging our unique combinatorial library design and screening strategy, we were able to produce Cas9 variants with multiple top immunogenic epitopes simultaneously mutated while still retaining genome targeting functionality. Spurred by this, we next evaluated delivery of SpCas9WT and SpCas9v4 and CRISPRoff versions of the same as icRNAs. CRISPRoff represents one of the newest additions to the CRISPR toolbox with the exciting capability to permanently silence gene expression upon transient expression55.
[0463] We conjectured that Cas9 and CRISPRoff would represent exciting applications of cRNAs for hit-and-run genome and epigenome targeting, as the prolonged persistence could enable robust targeting, while the use of partially deimmunized Cas9 proteins would enable greater safety in therapeutic contexts. Towards this. icRNAs for WT SpCas9 or SpCas9v4. along with single guide RNAs (sgRNAs) targeting the AAVS1 locus, were transfected into HEK293Ts. Excitingly, we observed AAVS1 genome editing and approximately 50% relative circularization rates for icRNAs as quantified by RNAseq. In addition, icRNAs and ocRNAs for WT dSpCas9 or dSpCas9v4 CRISPRoff along with sgRNAs targeting the B2M gene were transfected into HEK293Ts54, and we confirmed robust B2M gene repression via both cRNA formats (FIGS. 6E-6F). Importantly, the in vitro circularized ocRNAs were found to have a circularization efficiency of approximately 40% and 20% for SpCas9 and CRISPRoff inserts, respectively, as quantified via tapestation analyses. Lastly, to assess the specificity of SpCas9v4 targeting, we performed RNAseq on WT and SpCas9v4 CRISPRoff samples with and without the B2M guide. As expected, B2M was heavily downregulated in SpCas9WT and SpCas9v4 samples containing the sgRNA compared with samples with no guide (FIG. 12D). Importantly, all differentially expressed genes (DEGs) for V4 were also DEGs for WT, suggesting that SpCas9v4 and SpCas9WT are comparably specific in this assay (FIG. 12D).
[0464] Discussion
[0465] To facilitate the pursuit of cRNAs, we developed two methods: ‘outside developed’ cRNAs (ocRNAs) via in vitro circularization using group II introns (which, to the best of our knowledge, is the first demonstration of gene of interest expression from cRNAs prepared by a group II intron), and in situ or ‘inside developed’ cRNAs (icRNAs) via in-cell circularization using the ubiquitously expressed RtcB protein92,93. Furthermore, we developed HPLC-free purification protocols for these enabling high yields while maintaining low immune responses. The resulting simplicity and scalability of production facilitated a range of applications from stem cell engineering to robust genome and epigenome targeting via ZF proteins and CRISPR-Cas9 systems. In particular, icRNAs and ocRNAs bearing the EMCV IRES demonstrated robust expression (compared with linear capped and modified RNAs) in cardiomyocytes and neurons, and also prolonged RNA persistence highlighting substantial promise and utility in non-dividing cells. In addition, our circularization strategies allowed for efficient generation and delivery of large constructs, such as Cas9 and CRISPRoff, which would be otherwise cumbersome to deploy via lentiviruses and adeno-associated viruses owing to packaging limits55.
[0466] Concurrently, to enable compatibility between persistence of expression and immunogenicity, we also developed the LORAX protein engineering method that can be applied iteratively to tackle particularly challenging multiplexed protein engineering tasks by exploring huge swaths of combinatorial mutation space unapproachable using previous techniques. We demonstrated the power of this technique by creating a Cas9 variant with seven simultaneously deimmunized epitopes, which still retains functionality in a single round of screening. This opens up the application of gene editing to long-persistence therapeutic modalities such as AAV or icRNA delivery. Furthermore, whereas this methodology is particularly suited to the unique challenges of protein deimmunization, it is also applicable to any potential protein engineering goal, so long as there exists an appropriate screening procedure to select for the desired functionality.
[0467] While ocRNAs and icRNAs are a versatile system with broad applications, we anticipate multiple avenues for further engineering their efficacy: one, the upper limits of payload circularization for both icRNAs and ocRNAs require further exploring; two. despite comparable cell viability results, the full extent of immunogenicity presented by icRNAs and ocRNAs has yet to be explored in depth, and may serve to provide useful insights into future targeted improvements994 96: three, we observed persistence of cRNAs in non-dividing cells, but approaches to improve their persistence in dividing cells (where they are otherwise rapidly diluted in every cell division) could further broaden their utility; and four, as typically the 5' cap was observed to be more efficient than IRESs in recruiting the translation machinery, comprehensive screens of tissue specificity of IRESs97 101 will be crucial to deploy corresponding cRNAs in their optimal setting.
[0468] Similarly, the versatility of the LORAX method comes with a set of limitations and trade-offs that must be managed to leverage its utility. Naturally, library design is of critical importance. Here we have leveraged several features such as Cas9 evolutionary diversity, MHC-binding predictions, HLA allele frequencies and calculated immunogenicity scores to generate a useful library of variants to test. Other approaches may bring in more sources of information from places such as protein structure102, coevolutionary epistatic constraints10', amino acid signaling motifs104 or T / B cell receptor binding repertoires105, among other possibilities. Another critical factor is careful selection of hits downstream of screening, especially given the sparse coverage owing to nanopore sequencing. Here we have developed a network-based method for differentiating spurious from bona fide hits leveraging known aspects of protein epistasis and fitness landscapes. Similar customizations and tweaks relevant to the specific biolog}’ of a given problem may yield substantial returns in applying LORAX or other large-scale combinatorial screening methods to various protein engineering challenges.
[0469] Looking ahead, in addition to its core utility in applications entailing transgene delivery, we anticipate that cRNAs will be particularly useful in scenarios where a longer duration pulse of protein production is required. These include, for instance, epigenome engineering and cellular reprogramming, as well as transient healing and rejuvenation applications. Taken together, we anticipate that the highly simple and scalable ocRNA and icRNA methodologies could have broad utility in basic science and therapeutic applications.
[0470] Methods
[0471] Cell culture. HEK293T, A549. A375 and HeLa cells were cultured in DMEM supplemented with 10% FBS and 1% antibiotic-antimycotic (Thermo Fisher). K562 cells were cultured in RPMI supplemented with 10% FBS and 1% antibiotic-antimycotic (Thermo Fisher).
[0472] Hl human embryonic stem cells (hESCs) were maintained under feeder-free conditions in mTeSRl medium (StemCell Technologies). Before passaging, tissue-culture plates were coated with growth factor-reduced Matrigel (Coming) diluted in DMEM / F-12 medium (Thermo Fisher Scientific). Cells were dissociated and passaged using the dissociation reagent Versene (Thermo Fisher Scientific) and passaged at a 1:4 ratio.
[0473] All cells were cultured in an incubator at 37 °C and 5% CO2.
[0474] DNA transfections were performed by seeding HEK293T cells in 12-well plates at 25% confluency and adding 1 pg of each DNA construct and 4 pl of Lipofectamine 2000 (Thermo Fisher). RNA transfections were performed by adding a given pmol of each RNA construct and Lipofectamine MessengerMax (Thermo Fisher) (see Table 1 for specific details). Electroporations were performed in K562 cells using the SF Cell Line 4D-Nucleofector X Kit S (Lonza) per the manufacturer’s protocol. Table 11 RNA Transfection Conditions. Cell Line Number of Cells Seeded Lipofectamine MessengerMAX (uL / well) Well Format Transfected mRNA (pmol / well) Figure Hek293T 200k 3.5 12 well 3.75 FIG. 11 Hek293T 200k 3.5 12 well 1.5 FIGS. 1H, 2A, 2B, 2E, 4A, 7A Hek293T 125k 3.5 12 well 0.65 FIGS. 6E, 6F K562 - - 12 well 0.65 FIG. 6E Cardiomyocytes 300k 3.5 12 well 1.5 FIGS. 3B, 3A HeLa 150k 3.5 12 well 2 FIGS. 4B, 4C Hek293T 100k 1 24 well 0.6 FIG. 3A HeLa 100k 1 24 well 0.6 FIG. 3A A549 100k 1 24 well 0.6 FIGS. 2C, 7B, 3A A375 100k 1 24 well 0.6 FIG. 3A hPSCs 50k 1 24 well 0.6 FIGS. 3D, 3A Neurons 50k 1 24 well 1.2 FIG. 3B A549 7.5k 0.25 96 well 0.15 FIG. 2D
[0475] IVT, DNA templates for generating desired RNA products were created by PCR amplification from plasmids or gBlock gene fragments (IDT) and purified using a PCR purification kit (Qiagen). Plasmids were then generated with these templates containing a T7 promoter followed by 5' ribozyme sequence, a 5' ligation sequence, an IRES sequence linked 198 to the product of interest, a 3' UTR sequence, a 3' ligation sequence, a 3' ribozyme sequence and lastly a poly-T stretch to terminate transcription. Linearized plasmids were used as templates and RNA products were then produced using the HiScribe T7 Quick High Yield RNA Synthesis Kit (NEB E2040) per the manufacturer's protocol. In icRNA conditions, urea was added to a final concentration of 0.8 M in a 20 pl reaction through a freshly prepared 6 M urea solution. Linear mRNA was produced using the HiScribe T7 mRNA Kit with CleanCap Reagent AG (NEB E2080) unless otherwise mentioned. All UTP was replaced with N1-methylpseudouridine5'-triphosphate (Trilink Biotechnologies, N-1081) for mlT conditions.
[0476] Purification protocols. IVT RNA reactions were cleaned with the Monarch RNA Cleanup Kit (500 pg) (T2050) according to manufacturer's instructions. For icRNA and ocRNA, CIP treatment is preferred to ensure the removal of any remaining triphosphates from IVT, including cleaved twister ribozyme product. For ocRNA, cellulose chromatography was performed twice according to ref. 38 per 100 pg of RNA. Resulting RNA was again cleaned with the Monarch RNA Cleanup Kit (500 pg) (T2050) according to manufacturer’s instructions. For gel-extracted samples, RNA was further separated on precast 2% E-Gel EX agarose gels and then gel extraction performed with the Monarch RNA Cleanup Kit (50 pg) (T2040) according to manufacturer’s instructions. For RNaseR (Lucigen RNR07250)-treated samples, up to 89 pl of water was added to 20 pg of RNA. then heated to 70 °C for 3 min and briefly put on ice. RNaseR reaction buffer and 20 U of RNaseR were added and heated at 37 °C for 15 min, with an additional 10 U of RNaseR halfway. In the case of ref. 15, cRNA, no additional RNaseR was added and the reaction was heated at 37 °C for an hour. After RNaseR digestion, reactions were cleaned with the Monarch RNA Cleanup Kit (50 pg) (T2040) according to manufacturer’s instructions.
[0477] In vitro immune and cell viability experiments. To assess immune response, A549s were transfected with uncapped and unmodified linear RNA, linear ml'P RNA, icRNA, ocRNA (column) and ocRNA (gel), processed via previously mentioned methods additionally outlined in the associated figure schematic, and then RNA isolated at 6 h, 24 h and 48 h. RT-qPCR was performed to quantify IFNB, RIG-I and IL6 mRNA levels relative to GAPDH. To assess cell viability’, A549s were seeded and transfected with the same panel of RNA previously mentioned. A cell counting kit-8 (CCK-8) assay was performed before transfection on day 0, and then again on days 1 and 2, with media replacement before each measurement. For each well, 5 pl of CCK-8 reagent (Dojindo, CK04) was used per 100 pl of media and incubated at cell culture conditions for 1 h. Absorbance was measured at 450 nm on a plate reader.
[0478] In vitro persistence experiments. To assess persistence of circular icRNA, HEK293T cells were transfected with circular icRNA GFP or linear icdRNA and RNA was isolated 6 h, 1 day, 2 days and 3 days after transfection. RT-qPCR was performed to assess the amount of GFP RNA and RT-PCR was performed to confirm cRNA persistence in cells receiving icRNA.
[0479] Persistence of circular icRNA containing EMCV IRES. WPRE and a 50 adenosine poly(A) stretch compared with commercially sourced RNA (Trilink Biotechnologies, L-7601) with a 5' cap and a poly(A) tail was similarly performed, with additional time points at days 4 and 5. RNA was isolated from cells and RT-qPCR was performed to assess the amount of GFP RNA.
[0480] For cardiomyocyte experiments. Hl hESCs were differentiated into cardiomyocytes using established protocols106,107. In brief, stem cells were dissociated using Accutase and seeded into 12-well Matrigel-coated plates. Cells were maintained in mTeSRl (StemCell Technologies) for 3-4 days until cells reached about 95% confluence.
[0481] Media was changed to RPMI containing B27 supplement and 10 pM CHIR99021. After 24 h, media was changed to RPMI containing B27 supplement without insulin. Two days later, media was changed such that half of the cultured media was mixed with fresh RPMI containing B27 supplement without insulin and 5 pM IWP2. After 2 days, media was changed to RPMI containing B27 supplement without insulin. Media was then changed to RPMI containing B27 supplement every 2 days. After 6 days, media was replaced with cardiac metabolic enrichment media108. After 2 days, media was again changed to RPMI containing B27 supplement every 2 days. Cardiomyocytes were transfected with icRNA, ocRNA and icdRNA containing EMCV IRES, modified WPRE, and a 165 adenosine poly (A) stretch or linear m IT mRNA 2 weeks after CHIR99021 induction. Five images were taken for each biological replicate of each condition at each time point and GFP intensity was quantified for representative images using FIJI (NIH).
[0482] For neuron experiments, Hl hPSC clonal lines overexpressing NEURODI were seeded and neural differentiation media consisting of 1:1 DMEM / F12-Neurobasal media (Thermo Fisher Scientific) + lOOx Glutamax, 10 ng ml-1 BDNF (Peprotech). 10 ng ml-1 NT3 (Peprotech), 0.75 pg ml-1 puromycin, 1 pg ml-1 doxycycline (Sigma-Aldrich), 50x B27 supplement (Thermo Fisher Scientific) and lOOx N2 supplement (Thermo Fisher Scientific) was added. Media was changed every 2 days. Three images were taken for each biological replicate of each condition at each time point and GFP intensity was quantified using FIJI (NIH).
[0483] In vitro differentiation experiments. Hl hPSCs were seeded on day -1 in a Matrigel-coated 24-well plate in mTeSR. On day 0, mTeSR media was replaced and cells were transfected with RNA. On day 3. 2 pM Ara-C was added to mTeSR media to eliminate dividing hPSCs. mTeSR media was replaced every day until RNA collection or staining.
[0484] Immunostaining. To prepare for staining, medium was removed and cells were washed with 1 ml of PBS and 500 pl of 4% paraformaldehyde solution subsequently added for fixation at room temperature for 1 h, shielded from light. Post-fixation, cells were washed once with 500 pl of PBS. Afterwards, 500 pl of blocking buffer (3% FBS, 1% BSA, 0.5% Triton-X and 0.5% Tween) was added and incubated for 1 h at room temperature. Following blocking, primary antibody (diluted 1:1,000 in blocking buffer) was added and incubated overnight at 4 °C, shielded from light. After overnight incubation, cells w ere washed three times for 5 min each before the application of secondary antibody (1:2,000 dilution in blocking buffer) and incubated for 1 h at room temperature, shielded from light. Following secondary staining, cells w ere washed three times for 5 min each with PBS. Antibodies used include anti-tubulin [> 3 (TUBB3) (BioLegend, 801201) and anti-mouse IgG (H + L) Alexa Fluor Plus 647 (Invitrogen, A32728).
[0485] Flow cytometry experiments. To assess in vitro protein translation efficiencies, equimolar amounts of icRNA or linear mRNA were transfected into HEK293Ts. A549s. HeLas, hPSCs or A375s and GFP intensity was quantified 24 h later. GFP intensity, defined as the mean intensity of the cell population, was quantified after transfection using a BD LSRFortessa cell analyzer.
[0486] Quantifying circular efficiency. To assess circular efficiency, icRNA containing EMCV IRES, WPRE and a 165 adenosine poly (A) stretch was generated. RNA was then either frozen or pre-circularized using the RTCB ligase (NEB M0458S) per manufacturer’s instructions. To remove any linear RNA, pre-circularized RNA was treated with RNaseR (Lucigen RNR07250) per manufacturer’s instructions. icRNA or pre-circularized icRNA was then transfected into HEK293Ts and RNA was isolated from cells at 6 h, 24 h and 48 h. RT-PCR was performed and the intensity of the circular band for icRNA compared with precircularized icRNA was defined as the circular efficiency. All circular intensity' values were normalized to respective GAPDH band intensity.
[0487] To assess circularization efficiency of ocRNA, samples frozen overnight were run on an Agilent Tapestation to qualify ocRNA profiles. Then, gel analysis was performed on tapestation results to quantify circularization efficiency on ocRNA samples with band intensity normalization by nucleotide.
[0488] RNAseq. RNAseq was performed on HEK293T cells 6 h, 24 h and 48 h after transfection. Three biological replicates were sequenced. Total RNA was isolated from cells via an RNeasy kit (Qiagen) with on-column DNase I treatment. For RNA capture, f pg per sample of total RNA was used with NEBNext Poly(A) mRNA Magnetic Isolation Module (E7490S). Subsequently, an NEBNext Ultra RNA Library Prep Kit (E7530S) was used to generate Illumina-compatible RNAseq libraries. Sequencing was performed on an Illumina NovaSeq 6000, with paired end 100 bp reads. Reads were aligned using BWA-MEM2 to custom genomes derived from the GRCh38 Human Reference. Custom references were created using Cell Ranger to include GFP and Cas9 (720 bp) and ligation region (198 bp) for alignment. Featurecounts was used to quantity’ gene counts on BAM files output from BWA-MEM2. RPKM normalization was then performed and normalized counts evaluated to determine % circularization of transfected RNA.
[0489] LNP formulations. (6Z,9Z.28Z,3 lZ)-Heptatriaconta-6,9,28,31 -tetr aen-19-yl-4-(dimethylamino)butanoate (DLin-MC3-DMA) was purchased from BioFine International. L2-Distearoyl-.s7?-glycero-3-phosphocholine (DSPC) and 1,2-dimyristoyl-rac-glycero-3-methoxypoly ethylene glycol-2000 (DMG-PEG-2000) were purchased from Avanti Polar Lipids. Cholesterol was purchased from Sigma-Aldrich. mRNA LNPs were formulated with DLin-MC3-DMA:cholesterol:DSPC:DMG-PEGat amole ratio of 50:38.5:10:1.5 andaN / P ratio of 5.4. To prepare LNPs, lipids in ethanol and mRNA in 25 mM acetate buffer (pH 4.0) were combined at a flow rate of 1:3 in a PDMS staggered herringbone mixer109,110. The dimensions of the mixer channels were 200 by 100 pm, with herringbone structures 30 pm high and 50 pm wide. Immediately after formulation, 3 volumes of PBS was added and LNPs were purified in 100 kDa molecular weight cut-off (MWCO) centrifugal filters by exchanging the volume three times. Final formulations were passed through a 0.2 pm filter. LNPs were stored at 4 °C for up to 4 days before use. LNP hydrodynamic diameter and poly dispersity index were measured by dynamic light scattering (Malvern NanoZS Zetasizer). The mRNA content and percent encapsulation were measured with a Quant-iT RiboGreen RNA Assay (Invitrogen) with and without Triton X-100 according to the manufacturer’s protocol.
[0490] Animal experiments. All animal procedures were performed in accordance with protocols approved by the Institutional Animal Care and Use Committee of the University7 of California, San Diego. All mice were acquired from Jackson Labs.
[0491] To confirm circularization of icRNA constructs in vivo, 10 pg of circular GFP icRNA or linear GFP icdRNA LNPs was injected retro-orbitally into C57BL / 6J mice. After 3 and 7 days, livers were isolated and placed in RNAlater (Sigma-Aldrich). RNA was later isolated using QIAzol Lysis Reagent and purified using an RNeasy mini kit (Qiagen) according to the manufacturer’s protocol. The amount of circularized RNA was assessed by RT-qPCR.
[0492] To assess icRNA persistence in vivo, equal concentration of icRNA containing EMCV, WPRE, and a 165 adenosine poly(A) stretch (15 pg LNPs for EMCV) or linear RNA was injected retro-orbitally into C57BL / 6J mice. On day 7. livers were isolated and RNA was extracted. RT-qPCR was performed to assess mRNA expression among the conditions and RT-PCR was performed to ensure circularization for icRNA conditions.
[0493] Cas9 alignment and mutation selection. Naturally occurring variation in Cas9 sequence space was explored by aligning BLAST hits of the SpCas9 amino acid sequence. This set was then pruned by removing truncated, duplicated or engineered sequences, and those sequences whose origin could not be determined. At specified immunogenic epitopes and key anchor residues, top alternative amino acids were obtained using frequency in the alignment weighted by overall sequence identity to the WT SpCas9 sequence, such that commonly occurring amino acid substitutions appearing in sequences highly similar to the WT were prioritized for further analysis and potential inclusion in the LORAX library.
[0494] HL A frequency estimation and binding predictions. HLA-binding predictions were carried out using netMHC4.1 or netMHCpan3.1. Global HLA allele frequencies were estimated from data at allelefrequencies, net as follows. Data were divided into 11 geographical regions. Allele frequencies for each of those regions were estimated from all available data from populations therein. These regional frequencies were then averaged weighted by global population contribution. Alleles with greater than 0.001% frequency in the global population, or those with greater than 0.01% in any region, were included for further analysis and predictions.
[0495] Immunogenicity7 scores. The vector of predicted nM affinities output by netMHC was first normalized across alleles to account for the fact that some alleles have higher affinity across all peptides and to allow for the relatively equivalent contribution of all alleles. These values were then transformed using the 1 -log(affinity) transformation also borrowed from netMHC such that lower nM affinities will result in larger resulting values. These transformed, normalized affinities are then weighted by population allele frequency and summed across all alleles and epitopes. Finally, the scores are standardized across proteins to facilitate comparison.
[0496] Identification of HPRT1 guide. The lentiCRISPR-v2 plasmid (Addgene #52961) was first digested with Esp3I and a guide targeting the HPRT1 gene was cloned in via Gibson assembly. After lentivirus production, HeLa cells were seeded at 25% confluency in 96-well plates and transduced with virus (lentiCRISPR-v2 with or without HPRT1 guide) and 8 pg ml-1 polybrene (Millipore). Virus was removed the next day and 2.5 pg ml'1 puromycin was added to remove cells that did not receive virus 2 days later. After 2 days of puromycin selection, 0-14 pg ml’1 6-TG was added. After 5 days, cells were stained with crystal violet, solubilized using 1% sodium dodecyl sulfate, and absorbance was measured at 595 nm on a plate reader. Owing to the lack of cells in the negative control, 6 pg ml1 was chosen.
[0497] Generation of variant Cas9 library. Cas9 variant sequences were generated by separating the full-length gene sequence into small sections, where each section contained WT or variant Cas9 sequences. Degenerate pools of these gBlocks were PCR amplified and annealed together, yielding a final library size of 1,492,992 elements. Specifically, the Cas9 blocks used as input to the fusion PCR were synthesized as linear DNA. The nucleotide numbers that define the limits of each block are given in Table 2. Note the 30 bp overlaps to enable fusion. Table 2 | Cas9 block start and end positions. Block Start End 1 0 680 2 650 1030 3 1000 1630 4 1600 1990 5 1960 2400 6 2370 3000 7 2970 3790 8 3760 4140
[0498] The upper row of pie charts (FIG. 9) corresponds to the composition of each individually synthesized block, with each section corresponding to a singular block sequence. In the lower row, each pie corresponds to a single mutation site (note that one block may contain up to three mutation sites depending on the block). Once the fusion PCR was performed, the full-length library elements were purified by size selection using AmpureXP magnetic beads formulated to a very' low 0.4x concentration to enable selection of only high-molecular-weight DNA greater than ~3 kb. This insert was cloned into lentiCRISPRv2 (containing the HPRT1 guide) using Gibson assembly, and the plasmid product purified by 30 min dialysis before electroporation. Electroporation was done using NEB Stbl3 cells made electrocompetent as follows.
[0499] A 20 ml overnight liquid culture was inoculated from a single clone picked from an antibiotic-free plate streaked with Stbl3s. This culture was used to inoculate two 11 flasks, which were grown to a target ODeoo of 0.4. These cultures were centrifuged, washed with pure, cold water, and concentrated 10x to 2z 100 ml. This was repeated twice more to yield a final volume of 2 x 1 ml electrocompetent cells. These cells were electroporated in 200 pl aliquots, each with 100 ng of purified assembled library DNA. Dilutions of each of these electroporations were plated to estimate transformation efficiency before pooling into the final library. Plasmid DNA was isolated using the Qiagen Plasmid Maxi Kit and this DNA was then used to create lentivirus containing the variant Cas9 library.
[0500] Cas9 screen. HeLa cells were seeded in 15 15 cm plates at a density7 of 10 million cells per plate and transduced with virus containing the variant Cas9 library7 and 8 pg ml’1 polybrene the next day at an MOI of 0.3. Media was changed the next day and 2.5 pg ml’1 puromycin was added to remove cells that did not receive virus 2 days later. Once cells reached 90% confluency, 6 pg ml’1 6-TG was added to media. Media was changed every other day for 10 days to allow for selection of cells containing functional Cas9 variants. After 10 days, cells were lifted from the plates and DNA was isolated using the DNeasy Blood and Tissue Kit per the manufacturer’s protocol.
[0501] Nanopore sequencing. Pre-screen analysis of the Cas9 variant library elements was performed by amplifying the sequence from the plasmid. About 1 pg of the variant Cas9 sequences was used for library7 preparation using the Ligation Sequencing Kit (Oxford Nanopore Technologies, SQK-LSK109) per manufacturer’s instructions. DNA was then loaded into aMinlON flow cell (Oxford Nanopore Technologies. R9.4.1). Post-screen analysis of library elements was performed by amplifying the Cas9 sequences from 75 pg of genomic DNA. About 1 pg of the variant Cas9 sequences was similarly prepared using the Ligation Sequencing Kit and sequenced on a MinlON flow cell.
[0502] Base calling and genotyping. Raw7 reads coming off the MinlON flow cell were base-called using Guppy 3.6.0 and aligned to an SpCas9 reference sequence containing non-informative NNN bases at library mutation positions, so as not to bias calling towards WT or mutant library members, using Minimap2's map-ont presets. Reads covering the full length of the Cas9 gene with high mapping quality were genotyped at each individual mutation site and tabulated to the corresponding library member. Reads with ambiguous sites were excluded from further analysis.
[0503] Cluster analysis. Network analysis was performed by first thresholding genotypes to include only those identified as hits from the screen. These were genotypes appearing in the pre-screen plasmid library, both post-screen replicates, and having a fold change enrichment larger than the WT sequence (4.5-fold enrichment). These hits were used to create a graph with nodes corresponding to genotypes and node sizes corresponding to fold change enrichment. Edges were placed between nodes at most 4 mutations distant from each other, and edge weights were defined by 1 / J, where d is the distance between genotypes. Network analysis was done using Python bindings of igraph. Plots were generated using the Fruchterman-Reingold force-directed layout algorithm.
[0504] HDR validation. Lentivirus was produced from a plasmid containing a GFP sequence with a stop codon and 68 bp AAVS1 fragment. HEK293T cells were treated with 8 pg ml’1 polybrene and lentivirus. After puromycin selection to create a stable line, cells were transfected with plasmids containing variant Cas9 sequences, a guide targeting the AAVS locus and a GFP repair donor plasmid. After 3 days. FACS was performed and percent GFP-positive cells were quantified.
[0505] Genome engineering experiments. To validate variant Cas9 functional cutting, variant Cas9 and guides were transfected into HEK293T cells. After 2 days, genomic DNA was isolated. Genomic DNA was also isolated after 2 days from K562 cells after electroporation. To assess activity of CCR5 ZFNs delivered as icRNAs, HEK293Ts were transfected with circular icRNA or linear icdRNA and genomic DNA was isolated after 3 days. Assessment of GFP ZFN was performed by transfecting HEK293Ts stably expressing a broken GFP with circular icRNA or linear icdRNA and isolating genomic DNA after 3 days. To assess activity of Cas9 delivered as icRNAs, HEK293Ts and K562 were transfected or nucleofected with Cas9 WT or Cas9 v4 along with a gRNA (synthesized via Synthego) and genomic DNA was isolated after 3 days.
[0506] ZF experiments were performed by transfecting HEK293T cells with 0.5 pg of left and right arms of each ZF as either icRNA or icdRNA. After 3 days, genomic DNA was isolated.
[0507] Epigenome engineering experiments. ZF-KRAB icRNA experiments were performed by transfecting HeLa cells with icRNA encoding various in-house designed ZF sequences targeting the hPSCK9 gene. ZFs consisted of six variable DNA contacting regions inserted into the following backbone. MAPKKKRKVGIHGVPAAMAERPFQCRICMRNFS(ZF1)HIRTHTGEKPFACDICGRKF A(ZF2)HTKIHTGSQKPFQCR1CMRNFS(ZF3)H1RTHTGEKPFACDICGRKFA(ZF4)HTK IHTGSQKPFQCRICMRNFS(ZF5)HIRTHTGEKPFACDICGRKFA(ZF6)HTKIHLRQKDA ARGS
[0508] RNA was isolated 2 days later and repression of hPCSK9 was assessed by RT-qPCR. Similarly, 3A3L-ZF-KRAB ocRNA experiments were performed by transfecting HeLa cells with ocRNA encoding 3A3L-ZF-KRAB (specifically ZF10) targeting the hPSCK9 gene. Cells were passaged at 25% every 2 days or at 80-90% confluency, and the remaining 75% of cells at each passage collected for RNA isolation and repression of hPCSK9 was assessed by RT-qPCR. dCas9-VPR experiments were performed by transfecting HEK293T cells with dCas9wt-VPR or dCas9v4-VPR with or without a gRNA targeting the ASCL1 gene. Likewise, KRAB-dCas9 experiments were performed bytransfecting cells with KRAB-dCas9wt or KRAB-dCas9v4 with or without a gRNA targeting the CXCR4 gene. CRISPRoff experiments were performed by transfecting HEK293T cells with icRNA CRISPRoffwt or CRISPRoffv4 with or without a gRNA targeting the B2M gene (Synthego). RNA was isolated 3 days later and repression or activation of genes was assessed by RT-qPCR.
[0509] Quantification of editing using NGS. After extraction of genomic DNA, PCR was performed to amplify the target site. Amplicons were then indexed using the NEBNext Multiplex Oligos for the Illumina kit (NEB). Amplicons were then pooled and sequenced using a Miseq Nano with paired end 150 bp reads. Editing efficiency was quantified using CRISPResso2.
[0510] Cas9 specificity. RNA isolated from the CRISPRoff experiment was used to assess specificity. RNAseq libraries were generated from 300 ng of RNA using the NEBNext Poly(A) mRNA magnetic isolation module and NEBNext Ultra II Directional RNA Library Prep kit for Illumina and sequenced on the Illumina NovaSeq 6000 with paired end 100 bp reads. Fastq files were mapped to the reference human genome hg38 using STAR aligner. Differential gene expression was analyzed using the Bioconductor package DESeq2 with the cut-off of log2(fold change) greater than 0.5 or less than -0.5 and aF value less than 10A To identify’ DEGs. CRISPRoff WT and V4 samples containing the B2M guide were compared with samples not receiving the guide.
[0511] ELISpot assay. TAP-deficient T2 cells were a generous gift from Stephen Schoenberger lab. PBMCs were purchased from StemCell Technologies. All donors contained the HLA-A*0201 allele. Both cell lines were maintained in RPMI1640 media supplemented with 10% FBS, 1% penicillin-streptomycin, 10 mM HEPES and 1 mM sodium pyruvate. On the first day, PBMCs were thawed and rested overnight at a densify of 106 cells per ml. T2 cells were pulsed with peptides at 10 pg ml'1 overnight. Peptides were produced from Genscript’s Custom Peptide Synthesis service at crude purify. Lastly, 96-well plates (Immobilon-P, Millipore) were coated with 10 pg ml'1 anti-IFNy monoclonal antibody (1-D1K. Mabtech) overnight at 4 C. The next day, T2 cells were washed two times and 50,000 T2 cells and 100,000 PBMCs were added to each well. Four replicates were used per condition. After 22 h, cells were removed from the plate and 2 pg ml'1 biotinylated anti-IFNy secondary antibody (7-B6-1, Mabtech) was added for 2h. Plates were washed and 1:1,000 streptavidin-ALP (3310-10-1000, Mabtech) was added for 45 min. Plates were washed and colour was developed by adding BCIP / NBT-plus substrate (3650-10, Mabtech) for 10 min. Plates were thoroughly washed with water and dried at room temperature, and spots were automatically counted using an ELISpot plate reader.
[0512] To assess the immunogenicity of the full-length Cas9 WT and variant protein, in vitro transcribed RNA encoding for WT, L616G or V4 was electroporated into PBMCs as previously described89,90. As PBMCs contain both APCs and T cells, it is possible to electroporate RNA directly into these APCs and assess T cell response via the ELISpot. Electroporation was performed using the P3 Primary Cell 4D-Nucleofector X Kit (Lonza V4XP). In brief, PBMCs were first thawed and rested overnight at a density of 106 cells per ml. The next day, 1 x 106 PBMCs were resuspended in 20 pl of Lonza P3 nucleofector solution and mixed with 1 pg RNA. In some cases, electroporation was performed using the Lonza P3 Primary Cell 4D-Nucleofector X Kit L. with 3 * 106 PBMCs resuspended in 100 pl of Lonza P3 nucleofector solution and mixed with 6 pg RNA. After electroporation, 2 x 105 cells were added to each well of an ELISpot plate already coated with anti-IFNy monoclonal antibody as described above. After 28 h, cells were removed from the plate and the ELISpot assay and analysis was performed as described.
[0513] Lentivirus production. HEK293FT cells were seeded in 1 15 cm plate and transfected with 36 pl Lipofectamine 2000, 3 pg pMD2.G (Addgene #12259), 12 pg pCMV delta R8.2 (Addgene #12263) and 9 pg of the lentiCRISPR-v2 plasmid. The supernatant containing viral particles was collected after 48 h and 72 h, filtered with 0.45 pm Steriflip filters (Millipore), concentrated to a final volume of 1 ml using an Amicon Ultra-15 centrifugal filter unit with a 100,000 nominal molecular weight limit (NMWL) cut-off (Millipore) and frozen at -80 C.
[0514] RT-qPCR. cDNA was synthesized from RNA using the Protoscript II First Strand cDNA Synthesis Kit (NEB). qPCR was performed using a CFX Connect Real Time PCR Detection System (Bio-Rad). All samples were run in triplicates and results were normalized against GAPDH expression. Primers for qPCR are listed in Informal Sequence Listing (e.g., SEQ ID NOs: 143-176).
[0515] References
[0516] 1. Kariko, K., Muramatsu, FL, Ludwig, J. & Weissman, D. Generating the optimal mRNA for therapy: HPLC purification eliminates immune activation and improves translation of nucleoside-modified, protein-encoding mRNA. Nucleic Acids Res. 39, el 42 (2011).
[0517] 2. Presnyak, V. et al. Codon optimality is a major determinant of mRNA stability. Cell 160, 1111-1124 (2015).
[0518] 3. Kuhn, A. N. et al. Phosphorothioate cap analogs increase stability and translational efficiency of RNA vaccines in immature dendritic cells and induce superior immune responses in vivo. Gene Ther. 17, 961-971 (2010).
[0519] 4. Holtkamp, S. et al. Modification of antigen-encoding RNA increases stability, translational efficacy, and T-cell stimulatory capacity of dendritic cells. Blood 108, 40094017 (2006).
[0520] 5. Orlandini von Niessen. A. G. et al. Improving mRNA-based therapeutic gene delivery by expression-augmenting 3' UTRs identified by cellular library screening. Mol. Ther. 27, 824-836 (2019).
[0521] 6. Wesselhoeft, R. A., Kowalski, P. S. & Anderson, D. G. Engineering circular RNA for potent and stable translation in eukaryotic cells. Nat. Commun. 9, 2629 (2018).
[0522] 7. Petkovic, S. & Muller, S. RNA circularization strategies in vivo and in vitro. Nucleic Acids Res. 43. 2454-2465 (2015).
[0523] 8. Muller, S. & Appel, B. In vitro circularization of RNA. RNA Biol. 14, 10181027 (2017).
[0524] 9. Wesselhoeft, R. A. et al. RNA circularization diminishes immunogenicity' and can extend translation duration in vivo. Mol. Cell 14, 508-520.e4 (2019).
[0525] 10. Abe, N. et al. Rolling circle translation of circular RNA in living human cells. Sci. Rep. 5. 16435 (2015).
[0526] 11. Fan, X. et al. Pervasive translation of circular RNAs driven by short IRES-like elements. Nat. Commun. 13, 3751 (2022).
[0527] 12. Hansen, T. B. et al. Natural RNA circles function as efficient microRNA sponges. Nature 495, 384-388 (2013).
[0528] 13. Jeck, W. R. & Sharpless, N. E. Detecting and characterizing circular RNAs. Nai. Biotechnol. 32, 453-461 (2014).
[0529] 14. Kameda, S., Ohno, H. & Saito. H. Synthetic circular RNA switches and circuits that control protein expression in mammalian cells. Nucleic Acids Res. doi.org / 10.1093 / nar / gkacl252 (2023).
[0530] 15. Chen, R. et al. Engineering circular RNA for enhanced protein production. Nat. Biotechnol. doi.org / 10.1038 / s41587-02201393-0 (2022).
[0531] 16. Li, A. et al. AAV-CRISPR gene editing is negated by pre-existing immunity to Cas9. Mol. Ther. 28, 1432-1441 (2020).
[0532] 17. Charlesworth, C. T. et al. Identification of preexisting adaptive immunity to Cas9 proteins in humans. Nat. Med. 25, 249-254 (2019).
[0533] 18. Chaudhary, N., Weissman, D. & Whitehead, K. A. mRNA vaccines for infectious diseases: principles, delivery7 and clinical translation. Nat. Rev. Drug Discov. 20, 817-838(2021).
[0534] 19. Corbett, K. S. et al. SARS-CoV-2 mRNA vaccine design enabled by prototype pathogen preparedness. Nature 586, 567-571 (2020).
[0535] 20. Saunders, K. O. et al. Neutralizing antibody vaccine for pandemic and pre-emergent coronaviruses. Nature 594, 553-559 (2021).
[0536] 21. Thomas, S. J. et al. Efficacy and safety’ of the BNT162b2 mRNA COVID-19 vaccine in participants with a history’ of cancer: subgroup analysis of a global phase 3 randomized clinical trial. Vaccine doi.org / 10.1016 / j.vaccine.2021.12.046 (2021).
[0537] 22. Zinsli, L. V., Stierlin, N., Loessner, M. J. & Schmelcher, M. Deimmunization of protein therapeutics—recent advances in experimental and computational epitope prediction and deletion. Comput. Struct. Biotechnol. J. 19, 315-329 (2021).
[0538] 23. McNeil, B. A., Simon, D. M. & Zimmerly, S. Alternative splicing of a group II intron in a surface layer protein gene in Clostridium tetani. Nucleic Acids Res. 42, 19591969(2013).
[0539] 24. Pyle, A. M. Group II intron self-splicing. Annu. Rev. Biophys. 45, 183-205 (2016).
[0540] 25. Zimmerly, S. & Semper, C. Evolution of group II introns. Mob. DNA 6, 7 (2015).
[0541] 26. Chen, C. Y. & Sarnow, P. Initiation of protein synthesis by the eukaryotic translational apparatus on circular RNAs. Science 268, 415-417 (1995).
[0542] 27. Jang, S. K. et al. A segment of the 5' nontranslated region of encephalomyocarditis virus RNA directs internal entry of ribosomes during in vitro translation. J. Virol. 62, 2636-2643 (1988).
[0543] 28. Aitken, C. E. & Lorsch, J. R. A mechanistic ovendew of translation initiation in eukary otes. Nat. Struct. Mol. Biol. 19, 568-576 (2012).
[0544] 29. Alkemar, G. & Nygard, O. Secondary7 structure of two regions in expansion segments ES3 and ES6 with the potential of forming a tertiary7 interaction in eukaryotic 40S ribosomal subunits. RNA 10. 403-411 (2004).
[0545] 30. Bhat, P. et al. The beta hairpin structure within ribosomal protein S5 mediates interplay between domains II and IV and regulates HCV IRES function. Nucleic Acids Res. 43, 2888-2901 (2015).
[0546] 31. Chen, J. et al. Pervasive functional translation of noncanonical human open reading frames. Science 367. 1140-1146 (2020).
[0547] 32. Hershey, J. W. B., Sonenberg, N. & Mathews. M. B. Principles of translational control: an ovendew. Cold Spring Harb. Per sped. Biol. 4, aO 11528 (2012).
[0548] 33. Bradrick, S. S., Dobrikova, E. Y., Kaiser, C., Shveygert, M. & Gromeier, M. Poly(A)-binding protein is differentially required for translation mediated by viral internal ribosome entry sites. RNA 13, 1582-1593 (2007).
[0549] 34. Machida, K. et al. Dynamic interaction of poly(A)-binding protein with the ribosome. Sci. Rep. 8, 17435 (2018).
[0550] 35. Mailliot, J. & Martin, F. Viral internal ribosomal entry’ sites: four classes for one goal. Wiley Inlerdiscip. Rev. 9, el458 (2018).
[0551] 36. Imai, S., Kumar, P., Hellen, C. U. T., D'Souza, V. M. & Wagner, G. An accurately preorganized IRES RNA structure enables eIF4G capture for initiation of viral translation. Nat. Struct. Mol. Biol. 23, 859-864 (2016).
[0552] 37. Piao, X. et al. Double-stranded RNA reduction by chaotropic agents during in vitro transcription of messenger RNA. Mol. Ther. Nucleic Acids 29, 618-624 (2022).
[0553] 38. Baiersdorfer, M. et al. A facile method for the removal of dsRNA contaminant from in vitro-transcribed mRNA. Mol. Ther. Nucleic Acids 15, 26-35 (2019).
[0554] 39. Plank, T.-D. M., Whitehurst. J. T. & Kieft. J. S. Cell type specificity and structural determinants of IRES activity from the 5' leaders of different HIV-1 transcripts. Nucleic Acids Res. 41, 6698-6714 (2013).
[0555] 40. Jayaraman, M. et al. Maximizing the potency of siRNA lipid nanoparticles for hepatic gene silencing in vivo. Angew. Chem. Int. Ed. Engl. 51. 8529-8533 (2012).
[0556] 41. Sabnis, S. et al. A novel amino lipid series for mRNA deliver}’: improved endosomal escape and sustained pharmacology and safety in non-human primates. Mol. Ther. 26, 1509-1519 (2018).
[0557] 42. Lombardo, A. et al. Gene editing in human stem cells using zinc finger nucleases and integrase-defective lentiviral vector delivery. Nat. Biotechnol. 25, 1298-1306 (2007).
[0558] 43. Zou, J. et al. Gene targeting of a disease-related gene in human induced pluripotent stem and embryonic stem cells. Cell Stem Cell 5, 97-110 (2009).
[0559] 44. Abifadel, M. et al. Mutations in PCSK9 cause autosomal dominant hypercholesterolemia. Nat. Genet. 34, 154-156 (2003).
[0560] 45. Maxwell, K. N. & Breslow, J. L. Adenoviral-mediated expression of Pcsk9 in mice results in a low-density lipoprotein receptor knockout phenoty pe. Proc. Natl Acad. Sci. USA 101, 7100-7105 (2004).
[0561] 46. Cohen, J. C., Boerwinkle, E._ Mosley, T. H. Jr & Hobbs, H. H. Sequence variations in PCSK9, low LDL, and protection against coronary heart disease. N. Engl. J. Med. 354, 1264-1272 (2006).
[0562] 47. Thakore, P. I. et al. RNA-guided transcriptional silencing in vivo with S. aureus CRISPR-Cas9 repressors. Nat. Commun. 9, 1674 (2018).
[0563] 48. Ran, F. A. et al. In vivo genome editing using Staphylococcus aureus Cas9. Nature 520, 186-191 (2015).
[0564] 49. He, N.-Y. et al. Lowering serum lipids via PCSK9-targeting drugs: current advances and future perspectives. Acta Pharmacol. Sin. 38, 301-311 (2017).
[0565] 50. Ridker, P. M. et al. Cardiovascular efficacy and safety of bococizumab in high-risk patients. N. Engl. J. Med. 376, 1527-1539 (2017).
[0566] 51. Sabatine, M. S. et al. Evolocumab and clinical outcomes in patients with cardiovascular disease. N. Engl. J. Med. 376, 1713-1722 (2017).
[0567] 52. Fitzgerald, K. et al. A highly durable RNAi therapeutic inhibitor of PCSK9. N. Engl. J. Med. 376, 41-51 (2017).
[0568] 53. Ding, Q. et al. Permanent alteration of PCSK9 with in vivo CRISPR-Cas9 genome editing. Circ. Res. 115, 488-492 (2014).
[0569] 54. Amabile, A. et al. Inheritable silencing of endogenous genes by hit-and-run targeted epigenetic editing. Cell 167, 219-232.el4 (2016).
[0570] 55. Nunez, J. K. et al. Genome-wide programmable transcriptional memory by CRISPR-based epigenome editing. Cell 184, 2503- 2519.el7 (2021).
[0571] 56. Moreno, A. M. et al. Author correction: immune-orthogonal orthologues of AAV capsids and of Cas9 circumvent the immune response to the administration of gene therapy. Nat. Biomed. Eng. 3, 842 (2019).
[0572] 57. Chew, W. L. et al. A multifunctional AAV-CRISPR-Cas9 and its host response. Nat. Methods 13, 868-874 (2016).
[0573] 58. Jawa, V. et al. T-cell dependent immunogenicity of protein therapeutics pre-clinical assessment and mitigation-updated consensus and review 2020. Front. Immunol. 11, 1301 (2020).
[0574] 59. Moghadam, F. et al. Sy nthetic immunomodulation with a CRISPR superrepressor in vivo. Nat. Cell Biol. 22. 1143-1154 (2020).
[0575] 60. Hakim, C. H. et al. Cas9-specific immune responses compromise local and systemic AAV CRISPR therapy in multiple dystrophic canine models. Nat. Commun. 12, 6769 (2021).
[0576] 61. Ferdosi, S. R. et al. Multifunctional CRISPR-Cas9 with engineered immunosilenced human T cell epitopes. Nat. Commun. 10, 1842 (2019).
[0577] 62. Allen, B. D., Nisthal, A. & Mayo, S. L. Experimental library’ screening demonstrates the successful application of computational protein design to large structural ensembles. Proc. Natl Acad. Sci. USA 107, 19838-19843 (2010).
[0578] 63. Sun, M. G. F., Seo, M.-H., Nim, S., Corbi-Verge, C. & Kim, P. M. Protein engineering by highly parallel screening of computationally designed variants. Sci. Adv. 2, el600692 (2016).
[0579] 64. Cao, J. et al. High-throughput 5' UTR engineering for enhanced protein production in non-viral gene therapies. Nat. Commun. 12, 4138 (2021).
[0580] 65. Hu, J. H. et al. Evolved Cas9 variants with broad PAM compatibility and high DNA specificity. Nature 556, 57-63 (2018).
[0581] 66. Walton, R. T., Christie, K. A., Whittaker, M. N. & Kleinstiver, B. P. Unconstrained genome targeting with near-PAMless engineered CRISPR-Cas9 variants. Science 368, 290-296 (2020).
[0582] 67. Kleinstiver, B. P. et al. Engineered CRISPR-Cas9 nucleases with altered PAM specificities. Nature 523, 481-485 (2015).
[0583] 68. Kleinstiver, B. P. et al. High-fidelity CRISPR-Cas9 nucleases with no detectable genome-wide off-target effects. Nature 529, 490-495 (2016).
[0584] 69. Charles, E. J. et al. Engineering improved Cas 13 effectors for targeted post-transcriptional regulation of gene expression. Preprint at bioRxiv doi.org / 10.1101 / 2021.05.26.445687 (2021).
[0585] 70. Griswold. K. E. & Bailey-Kellogg. C. Design and engineering of deimmunized biotherapeutics. Curr. Opin. Struct. Biol. 39, 79-88 (2016).
[0586] 71. Doud, M. B., Lee, J. M. & Bloom, J. D. How single mutations affect viral escape from broad and narrow antibodies to Hl influenza hemagglutinin. Nat. Commun. 9, 1386(2018).
[0587] 72. Gasiunas. G. et al. A catalogue of biochemically diverse CRISPR-Cas9 orthologs. Nat. Commun. 11, 5512 (2020).
[0588] 73. Takeuchi, N., Wolf, Y. I., Makarova, K. S. & Koonin, E. V. Nature and intensity of selection pressure on CRISPR-associated genes. J. Bacterial. 194, 1216-1225 (2012).
[0589] 74. Andreatta, M. & Nielsen. M. Gapped sequence alignment using artificial neural networks: application to the MHC class I system. Bioinformatics 32, 511-517 (2016).
[0590] 75. Nielsen, M. et al. Reliable prediction of T-cell epitopes using neural networks with novel sequence representations. Protein Sci. 12, 1007-1017 (2003).
[0591] 76. Osipovitch, D. C. et al. Design and analysis of immune-evading enzymes for ADEPT therapy. Protein Eng. Des. Sei. 25, 613-623 (2012).
[0592] 77. Choi, Y., Verma, D., Griswold, K. E. & Bailey-Kellogg, C. in Computational Protein Design (ed. Samish, I.) 375-398 (Springer New York, 2017).
[0593] 78. King, C. et al. Removing T-cell epitopes with computational protein design. Proc. Natl Acad. Sci. USA 111, 8577-8582 (2014).
[0594] 79. Mazor, R. et al. Elimination of murine and human T-cell epitopes in recombinant immunotoxin eliminates neutralizing and anti-drug antibodies in vivo. Cell. Mol. Immunol. 14, 432-442 (2017).
[0595] 80. Wang, Y., Zhao, Y., Bellas, A., Wang, Y. & Au, K. F. Nanopore sequencing technology, bioinformatics and applications. Nat. Biotechnol. 39, 1348-1365 (2021).
[0596] 81. Rang, F. J., Kloosterman, W. P. & de Ridder, J. From squiggle to basepair: computational approaches for improving nanopore sequencing read accuracy. Genome Biol. 19, 90 (2018).
[0597] 82. Schubert, B. et al. Population-specific design of de-immunized protein biotherapeutics. PLoS Comput. Biol. 14, el005983 (2018).
[0598] 83. Liao, S., Tammaro, M. & Yan, H. Enriching CRISPR-Cas9 targeted cells by co-targeting the HPRT gene. Nucleic Acids Res. 43. el34 (2015).
[0599] 84. Yang. F. et al. HPRT1 activity loss is associated with resistance to thiopurine in ALL. Oncotarget 9, 2268-2278 (2018).
[0600] 85. Meini, M.-R , Tomatis, P. E., Weinreich, D. M. & Vila, A. J. Quantitative description of a protein fitness landscape based on molecular features. Mol. Biol. Evol. 32, 1774-1787 (2015).
[0601] 86. Mali, P. et al. RNA-guided human genome engineering via Cas9. Science 339, 823-826 (2013).
[0602] 87. Clement, K. et al. CRISPResso2 provides accurate and rapid genome editing sequence analysis. Nat. Biotechnol. 37, 224-226 (2019).
[0603] 88. Ninkovic, T. et al. Identification of G-glycosylated decapeptides within the MUC1 repeat domain as potential MHC class I (A2) binding epitopes. Mol. Immunol. 41, 131-140 (2009).
[0604] 89. Etschel. J. K. et al. HIV-1 mRNA electroporation of PBMC: a simple and efficient method to monitor T-cell responses against autologous HIV-1 in HIV-1-infected patients. J. Immunol. Methods 380, 40-55 (2012).
[0605] 90. Van Camp, K. et al. Efficient mRNA electroporation of peripheral blood mononuclear cells to detect memory' T cell responses for immunomonitoring purposes. J. Immunol. Methods 354, 1-10 (2010).
[0606] 91. Moreno, A. M. et al. In situ gene therapy via AAV-CRISPR- Cas9-mediated targeted gene regulation. Mol. Ther. 28, 1931 (2020).
[0607] 92. Litke, J. L. & Jaffrey, S. R. Highly efficient expression of circular RNA aptamers in cells using autocatalytic transcripts. Nat. Biotechnol. 37, 667-675 (2019).
[0608] 93. Katrekar, D. et al. Efficient in vitro and in vivo RNA editing via recruitment of endogenous ADARs using circular guide RNAs. Nat. Biotechnol. 40, 938-945 (2022).
[0609] 94. Chen. Y. G. et al. Sensing self and foreign circular RNAs by intron identity. Mol. Cell 67. 228-238.e5 (2017).
[0610] 95. Chen, Y. G. et al. N6-Methyladenosine modification controls circular RNA immunity'. Mol. Cell 16, 96-109.e9 (2019).
[0611] 96. Abe, B. T. et al. Circular RNA migration in agarose gel electrophoresis. Mol. Cell 82, 1768-1777 (2022).
[0612] 97. Chen. C.-K. et al. Structured elements drive extensive circular RNA translation. Mol. Cell 81. 4300-4318.el3 (2021).
[0613] 98. Yang, Y. et al. Extensive translation of circular RNAs driven by N6-methyladenosine. Cell Res. 27, 626-641 (2017).
[0614] 99. Meyer, K. D. et al. 5' UTR m6A promotes cap-independent translation. Cell 163, 999-1010(2015).
[0615] 100. Weingarten-Gabbay, S. et al. Comparative genetics. Systematic discovery' of cap-independent translation sequences in human and viral genomes. Science 351, aad4939 (2016).
[0616] 101. Sample, P. J. et al. Human 5' UTR design and variant effect prediction from a massively parallel translation assay. Nat. BiotechnoL 37, 803-809 (2019).
[0617] 102. Stiffler, M. A. et al. Protein structure from experimental evolution. Cell Syst. 10, 15-24.e5 (2020).
[0618] 103. Green, A. G. et al. Large-scale discovery of protein interactions at residue resolution using co-evolution calculated from genomic sequences. Nat. Commun. 12, 1396 (2021).
[0619] 104. Saylor, K, Gillam, F., Lohneis, T. & Zhang, C. Designs of antigen structure and composition for improved protein-based vaccine efficacy. Front. Immunol. 11, 283 (2020).
[0620] 105. Joglekar, A. V. et al. T cell antigen discovery via signaling and antigenpresenting bifunctional receptors. Nat. Methods 16. 191-198 (2019).
[0621] 106. Lian, X. et al. Directed cardiomyocyte differentiation from human pluripotent stem cells by modulating Wnt / p-catenin signaling under fully defined conditions. Nat. Protoc. 8, 162-175 (2013).
[0622] 107. Kumar, A. et al. Mechanical activation of noncoding-RNA-mediated regulation of disease-associated phenotypes in human cardiomyocytes. Nat. Biomed. Eng. 3. 137-146 (2019).
[0623] 108. Tohyama, S. et al. Distinct metabolic flow enables large-scale purification of mouse and human pluripotent stem cell-derived cardiomyocytes. Cell Stem Cell 12, 127-137 (2013).
[0624] 109. Chen, D. et al. Rapid discovery of potent siRNA-containing lipid nanoparticles enabled by controlled microfluidic formulation. JAm. Chem. Soc. 134, 69486951 (2012).
[0625] 1 10. Belliveau, N. M. et al. Microfluidic synthesis of highly potent limit-size lipid nanoparticles for in vivo delivery of siRNA. Mol. Ther. Nucleic Acids 1, e37 (2012). INFORMAL SEQUENCE LISTING SEQ ID NO: Peptide ID Amino Acid Sequence 1. 5' twister ribozyme sequence GC CAT CAGT C GC C GGT C C CAAGC C C GGATAAAATGGGAGGGGGC GGG AAACCGCCT 2 . 3' twister ribozyme sequence AAC AC T GC C AAT GC C GGT C C C AAGC C C GGAT AAAAGT GGAGGGT AC A GTCCACGC 3 . Sma-1-402 (652_NW_0030 37967.1 / 24595 1-245882) (twister ribozyme domain) CCAACCCTTCTCCTTTACCCGGGCTTGGGACCGGCAG T GAC T C TAGAAGAGC TACAGGCGGAGT GAGC GGc c 4 . env-112 (67_GYQO9XB0 2I4TUQ / 289231) (twister ribozyme domain) AAGAC AC T T C CC AT T T C T C AACC GGGC T T GAGAC C GG C TAC GTAT GACT GGGT TAT C T T c c 5. Slyc-2-1 (655_KX45376 5.1 / 4383-4447) (twister ribozyme domain) GGGTACACTCTGCTTT TACAAGGACT T GT GAC C T TAC C GT AT AAATACAGAGT AGC T T TC AC Ac c 6. Osin-1-1 (811_VCDQ010 11222.1 / 38513911) (twister ribozyme domain) CCAACCCGCCTCCTATTTTCATCCGGGCTTGAGACCG GC AT AGGC GGAGT T ATAAAAT TGAc C 7 . env-2 7 0 (103_SRS01991 0_WUGC_scaff old_41849 / 107 9-1133) (twister ribozyme domain) AATAGAC TTTGCATTTTT CAATTAGGCT T GT CAC TAA CTTTGGCTGCATTATATTcc 8 . env-9 4 (908_AglaG_G DN60OX02IW UIH / 74-133) (twister ribozyme domain) TTTTCCCTC GOAT T T TAC CAGAT T T GGGAC T GGCAGT AAC T T T AC T GGC AT GC T TAAAAGc c 9 . env-935 (1012_SRS013 687_Baylor_sc affold_18085 / 1567-1643) (twister ribozyme domain) GCCTCGCTCTGCATTTTCACGGGCTGGCGACCGTTCC C C C T GGGGGAGC T GCAT TAAGGC C GGGGGC C GAAGC C CCC 10 . Sma-1-66 (609_NW_00302 6403.1 / 11121172) (twister ribozyme domain) CCAACCCTCCTCCTTTACCCGGGCTTCGGACCGGCAG AAGC GC TACAGGC GGAGT TAC TGGc c 11. Dre-1-3 (1460_NW_001 878356.2 / 15906611590725) (twister ribozyme domain) TTAAACCCCTCCTCATTTACCCGGGCTCGGGACCGGC AC AAAAAAAC CTGTGGCT GAGT T C T GAAt c t c c 12 . Osa-1-4 (1319_NC_008 405.1 / 1767469 -1767418) (twister ribozyme domain) TGCCCCCTCCACTTTTATCCGGGCTTGGGACCGGCAC T GGCAGT GT TAGGCAc c 13 . Eana-1-1 (848_DWOK011 80159.1 / 293389) (twister ribozyme domain) TGAGATGCCCTCCCCATTTCCGGGCTTGGGACCGACA T C AC AGC GGC TAT T GC T AC C GAAGC AGT AAT C C C C GC T GT GAAGGC T GGGT T AC AT C T CAc c 14 . Osa-1-8 (1105_NC_008 404.1 / 2639065 5-26390734) (twister ribozyme domain) TATCCCCTCCTCCTTTTTCCGGGCTTGGGACCGGCTA T GT CAT T CAGATAAAC C TAAGAT GACATAGGC GGAGT TACCTAcc 15. Osa-1-3 (1099_NC_008 404.1 / 2636321 6-26363295) (twister ribozyme TATCCCCTCCTCCTTTTCCGGGCTTGGGACCGGCTAT GT CAT TAAAATCACAC C T CAAGTGACATAGGC GGAGT TACCTAcc domain) 16. env-13 (299_RUMENN ODE_2846306 _1 / 51290-51386) (twister ribozyme domain) AGCTCCCTCTACTTCGATTAGTTGTCACTCTAACCTT TAC GGAGT T GGGAC C GT GTAGATAGC C GCAC C T T GC T T GC T TAT C TAc c T GTAT TAAGCT c c 17 . Dre-1-4 (1461_NW_001 878356.2 / 1639 268-1639342) (twister ribozyme domain) T AGAC CCTCCTCATT TAT C C GGGC T T GGGAC C GACAC AAGGTGGTGTATAAAAATCCCTGTGGCCGAGTTCTGA ACC 18 . Aage-1-1 (60_DWGI0103 9204.1 / 398453) (twister ribozyme domain) CAATACCCTCCTCCTTTTGCATCCGGGCTTGGGACCG GCAAT GGC GGAGT TAAT T T c c 19. Spol-1-1 (406_JANEYH0 10000009.1 / 67451126745173) (twister ribozyme domain) CCTCCCCTCTTCCTTAATCCGGGCTTGGGACCGGCAT T GGAAGC CAATGGC GGAGT TAGAGGc C 20 . Ttru-1-1 (ABRN0238751 7.1 / 545-474) (twister ribozyme domain) CTTTCTCTCTGTGTTTACACGGGCTTGAGACCGTCTT GCAC C TAC GT TGTAGGTACAAGGC CACAT TAAAAAc c 21. Hsap-1-1(NC_000017. 11 / 4417665744176721) (twister ribozyme domain) AAGGT C C T GAGGAAAAT C C T GAAGGC T T GGGAC C C C T TCAGCTTGCTGTGACACTCCCTTGCCTTc 22 . Hsap-1-2 (NC_000002.1 2 / 121281196121281138) (twister ribozyme domain) GCGTCTCTTCAT GCAGC T T CACGGC T GGGGAC GT GGC AGCACACAACACATAT T GCAGGc c 23 . 5' ligation area ( 5' ligation stem sequence) AAC C AT GC C GAC T GAT GGC AG 24 . 3' ligation C T GC CAT C AGTC GGC GT GGAC TGTAG area ( 3' ligation stem sequence) 25. CHARM domain MART KQTARKST GGKAPRKQLAT KAARKSAPGGAS S GAGS S SGGS AA GS GS GGAS S GAAS S SAGSAAGSGS S PVEIYKTVSAWKRQ PIRVL S L F GNIDKELKSLGFLESGSHSEGGTLKYVEDVTNWRRDVEKWGPFDLL YGSTQPLGGSCDRCPGWYMFQFHRILQYARPRQESQPFFWIFMDNLL L T E D DQMT T VRF L QT E AVT L RDVRS RVL QNAVRVWS NIPGLKSKHEA LSPKEEESLQGHVRTRSKLAAPKVDPLVKNCLLPLREYFKYFSQSSL PLGGPSSGAPPPSGGSPAGSPTSTEEGTSESATPESGPGTSTEPSEG SAPGSPAGSPTSTEEGTSTEPSEGSAPGTSTEPSEKRTADGSEFESP KKKRKVGVPAA 26. M. Sssl domain MSKVENKTKKLRVFEAFAGIGAQRKALEKVRKDE YE IVGLAEWYVPA IVMYQAIHNNFHTKLEYKSVSREEMIDYLENKTLSWNSKNPVSNGYW KRKKDDELKIIYNAIKLSEKEGNIFDIRDLYKRTLKNIDLLTYSFPC QDLSQQGIQKGMKRGSGTRSGLLWEIERALDSTEKNDLPKYLLMENV GALLHKKNEEELNQWKQKLESLGYQNS IEVLNAADFGS SQARRRVFM ISTLNEFVELPKGDKKPKSIKKVLNKIVSEKDILNNLLKYNLTEFKK TKSNINKASLIGYSKFNSEGYVYDPEFTGPTLTASGANSRIKIKDGS NIRKMNSDETFLYMGFDSQDGKRVNEIEFLTENQKIFVCGNSISVEV LEAIIDKIGG 27 . Sid4x domain MN IQML L EAADYL E RRE RE AE HG YASML PGS GMNIQML L E AAD YL E R REREAEHGYASMLPGSGMNIQMLLEAADYLERREREAEHGYASMLPG S GMNIQML L EAADYL E RRE REAE HGYAS ML P S R 28 . PRM2 domain ML C DAS S GAC RS VF C AF C RS DAAAC GS GAAS FES SVC AY S E TGHGS S RRPRRDPGSALAPLGRGVLALAHATAACTPRASVRSGSSRGIGP 29 . PRM1 domain MARYRCCRSQSRSRYYRQRQRSRRRRRRSCQTRRRAMRCCRPRYRPR CRRH 30. VP64 domain GRADALDDFDLDMLGSDALDDFDLDMLGSDALDDFDLDMLGSDALDD FDLDML 31. KRAB domain RTLVTFKDVFVDFTREEWKLLDTAQQIVYRNVMLENYKNLVSLGYQL TKPDVILRLEKGEEPWLVE 32 . 3A3L domain MNHDQEFDPPKVYPPVPAEKRKPIRVLSLFDGIATGLLVLKDLGIQV DRYIASEVCEDS ITVGMVRHQGKIMYVGDVRSVTQKHIQEWGPFDLV IGGS PCNDLSIVNPARKGLYEGTGRLFFEFYRLLHDARPKEGDDRPF FWL F ENWAMGVS D KRDIS RF LE S N PVMIDAKEVSAAHRARYFWGNL PGMNRPLASTVNDKLELQECLEHGRIAKFSKVRTITTRSNSIKQGKD QHFPVFMNEKEDILWCTEMERVFGFPVHYTDVSNMSRLARQRLLGRS WSVPVIRHLFAPLKEYFACVSSGNSNANSRGPSFSSGLVPLSLRGSH MGPMEIYKTVSAWKRQPVRVLSLFRNIDKVLKSLGFLESGSGSGGGT LKYVEDVTNWRRDVEKWGPFDLVYGSTQPLGSSCDRCPGWYMFQFH RILQYALPRQESQRPFFWIFMDNLLLTEDDQETTTRFLQTEAVTLQD VRGRDYQNAMRVWSNIPGLKSKHAPLTPKEEEYLQAQVRSRSKLDAP KVDLLVKNCLLPLREYFKYFSQNSLPL 33 . Cas9 sequence (X represents position of variance based on wildtype sequence) MDKKYSIGLDIGTNSVGWAVITDEYKVXSKKFKVLGNTDRHSIKKNL IGAL L F D S GE TAEAT RL KRTARRRYT RRKNRIC YL QEIF S NEMAKVD DSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKK LVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQL VQT YNQL FEENPINAS GVDAKAIL SARL S KS RRL ENLIAQL PGEKKN GXFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQI GDQXADLFLAAKNLSDAILLSDILRVNTEITKAPLXASMIKRYDEHH QDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGAXQEEFYKFI KPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAIL RRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSE ETITPWNFEEWDKGASAQSFIERMTNXDKNLPNEKVLPKHSLXYEY FTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQL KEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEE NEDIXEDIVLTXTLFEDREMIEERXKTYAHLFDDKVMKQLKRRRYTG WGRL SRKLINGIRDKQS GKTILDFLKSDGFANRNFMQLIHDDS LTXK E DIQKAQVS GQGD S L HE HIANXAGS PAI KKGIL QT VKWD E L VKVMG RHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEH PVENTQLQNEKLYLYYXQNGRDMYVDQELDINRLSDYDVDHIVPQSF L KDD SIDNKVLT RS DKNRGKS DNVPS E E WKKMKNYWRQL LNAKLIT QRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMN TKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAHDA YLNAWGTALIKKYPKLESEFVYGDYKVXDVRKMIAKSEQEIGKATA KYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFA TVRKVL SMPQVNIVKKT EVQT GGFS KE SIL PKRNS DKLIARKKDWD P KKYGGFDSPTVAYSVLWAKVEKGKSKKLKSVKELLGITIMERSSFE KNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQK GNELALPSKYVNFLYLASHYEKXKGSPEDNEQKQLFVEQHKHYLDEI IEQXSEFSKRVIXADANLDKVLSAXNKHRDKPIREQAENIIHLFTLT NLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLS QLGGD 34 . Cas9V4 MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNL IGAL L F D S GE TAEAT RL KRTARRRYT RRKNRIC YL QEIF S NEMAKVD DSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKK LVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQL VQT YNQL FEENPINAS GVDAKAIL SARL S KS RRL ENLIAQL PGEKKN GLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQI GDQQADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHH QDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFI KPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAIL RRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSE ETITPWNFEEWDKGASAQSFIERMTNFDKNLPNEKVLPKHSLTYEY FTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQL KEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEE NEDILEDIVLTQTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTG WGRL SRKLINGIRDKQS GKTILDFLKS DGFANRNFMQLIHDDS LTFK EDIQKAQVSGQGDSLHEHIANGAGSPAIKKGILQTVKWDELVKVMG RHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEH PVENTQLQNEKLYLYYDQNGRDMYVDQELDINRLSDYDVDHIVPQSF L KDD SIDNKVLT RS DKNRGKS DNVPS E EWKKMKNYWRQL LNAKLIT QRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMN TKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAHDA YLNAWGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATA KYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFA TVRKVL SMPQVNIVKKT EVQT GGFS KE SIL PKRNS DKLIARKKDWD P KKYGGFDSPTVAYSVLWAKVEKGKSKKLKSVKELLGITIMERSSFE KNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQK GNELALPSKYVNFLYLASHYEKGKGSPEDNEQKQLFVEQHKHYLDEI IEQISEFSKRVIAADANLDKVLSAYNKHRDKPIREQAENIIHLFTLT NLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLS QLGGD 35. Cas9Vl MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNL IGAL L F D S GE TAEAT RL KRTARRRYT RRKNRIC YL QEIF S NEMAKVD DSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKK LVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQL VQT YNQL FEENPINAS GVDAKAIL SARL S KS RRL ENLIAQL PGEKKN GLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQI GDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHH QDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFI KPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAIL RRQEDFYPFLKDNREKTEKTLTFRIPYYVGPLARGNSRFAWMTRKSE ETITPWNFEEWDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEY FTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQL KEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEE NEDIGEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTG WGRL SRKLINGIRDKQS GKTILDFLKS DGFANRNFMQLIHDDS LTFK EDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKWDELVKVMG RHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEH PVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSF L KDD SIDNKVLT RS DKNRGKS DNVPS E E WKKMKNYWRQL LNAKLIT QRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMN TKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAHDA YLNAWGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATA KYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFA T VRKVL SMPQVNIVKKT EVQT GGFS KE S IL PKRNS DKLIARKKDWD P KKYGGFDSPTVAYSVLWAKVEKGKSKKLKSVKELLGITIMERSSFE KNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQK GNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEI IEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLT NLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLS QLGGD 36. Cas9V2 MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNL IGAL L F D S GE TAEAT RL KRTARRRYT RRKNRIC YL QEIF S NEMAKVD DSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKK LVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQL VQT YNQL FEENPINAS GVDAKAIL SARL S KS RRL ENLIAQL PGEKKN GLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQI GDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHH QDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFI KPTLEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAIL RRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSE ETITPWNFEEWDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEY FTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQL KEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEE NEDILEDIVLTQTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTG WGRL SRKLINGIRDKQS GKTILDFLKS DGFANRNFMQLIHDDS LTFK EDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKWDELVKVMG RHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEH PVENTQLQNEKLYLYYLQNGRDMYVDQELDTNRLSDYDVDHIVPQSF L KDD SIDNKVLT RS DKNRGKS DNVPS E EWKKMKNYWRQL LNAKLIT QRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMN TKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAHDA YLNAWGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATA KYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFA TVRKVL SMPQVNIVKKT EVQT GGFS KE SIL PKRNS DKLIARKKDWD P KKYGGFDSPTVAYSVLWAKVEKGKSKKLKSVKELLGITIMERSSFE KNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQK GNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEI IEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLT NLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLS QLGGD 37 . Cas9V3 MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNL IGAL L F D S GE TAEAT RL KRTARRRYT RRKNRIC YL QEIF S NEMAKVD DSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKK LVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQL VQT YNQL FEENPINAS GVDAKAIL SARL S KS RRL ENLIAQL PGEKKN GLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQI GDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHH QDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFI KPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAIL RRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSE ETITPWNFEEWDKGASAQSFIERMTNFDKNLPNEKVLPKHSLTYEY FTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQL KEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEE NEDILEDIVLTQTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTG WGRL SRKLINGIRDKQS GKTILDFLKS DGFANRNFMQLIHDDS LTFK EDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKWDELVKVMG RHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEH PVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSF L KDD SIDNKVLT RS DKNRGKS DNVPS E E WKKMKNYWRQL LNAKLIT QRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMN TKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAHDA YLNAWGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATA KYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFA T VRKVL SMPQVN IVKKT EVQT GGFS KE S IL PKRNS DKL IARKKDWD P KKYGGFDSPTVAYSVLWAKVEKGKSKKLKSVKELLGITIMERSSFE KNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQK GNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEI IEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLT NLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLS QLGGD 38 . Cas 9V5 MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNL IGAL L F D S GE TAEAT RL KRTARRRYT RRKNRIC YL QEIF S NEMAKVD DSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKK LVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQL VQT YNQL FEENPINAS GVDAKAIL SARL S KS RRL ENLIAQL PGEKKN GLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQI GDQQADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHH QDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGACQEEFYKFI KPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAIL RRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSE ETITPWNFEEWDKGASAQSFIERMTNFDKNLPNEKVLPKHSLTYEY FTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQL KEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKITKDKDFLDNEE NEDILEDIVLTQTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTG WGRL SRKLINGIRDKQS GKTILDFLKSDGFANRNFMQLIHDDS LTFK EDIQKAQVSGQGDSLHEHIANGAGSPAIKKGILQTVKWDELVKVMG RHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEH PVENTQLQNEKLYLYYDQNGRDMYVDQELDINRLSDYDVDHIVPQSF L KDD SIDNKVLT RS DKNRGKS DNVPS E EWKKMKNYWRQL LNAKLIT QRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMN TKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAHDA YLNAWGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATA KYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFA TVRKVL SMPQVNIVKKT EVQT GGFS KE SIL PKRNS DKLIARKKDWD P KKYGGFDSPTVAYSVLWAKVEKGKSKKLKSVKELLGITIMERSSFE KNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQK GNELALPSKYVNFLYLASHYEKGKGSPEDNEQKQLFVEQHKHYLDEI IEQISEFSKRVIAADANLDKVLSAYNKHRDKPIREQAENIIHLFTLT NLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLS QLGGD 39. Cas9V6 MDKKYSIGLDIGTNSVGWAVITDEYKVLSKKFKVLGNTDRHSIKKNL IGAL L F D S GE TAEAT RL KRTARRRYT RRKNRIC YL QEIF S NEMAKVD DSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKK LVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQL VQT YNQL FEENPINAS GVDAKATL SARL S KS RRL ENLIAQL PGEKKN GLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQI GDQQADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHH QDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGACQEEFYKFI KPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAIL RRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSE ETITPWNFEEWDKGASAQSFIERMTNFDKNLPNEKVLPKHSLTYEY FTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQL KEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEE NEDILEDIVLTQTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTG WGRL SRKLINGIRDKQS GKTILDFLKS DGFANRNFMQLIHDDS LTFK EDIQKAQVSGQGDSLHEHIANGAGSPAIKKGILQTVKWDELVKVMG RHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEH PVENTQLQNEKLYLYYDQNGRDMYVDQELDINRLSDYDVDHIVPQSF L KDD SIDNKVLT RS DKNRGKS DNVPS E EWKKMKNYWRQL LNAKLIT QRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMN TKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAHDA YLNAWGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATA KYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFA T VRKVL SMPQVN IVKKT EVQT GGFS KE S IL PKRNS DKL IARKKDWD P KKYGGFDSPTVAYSVLWAKVEKGKSKKLKSVKELLGITIMERSSFE KNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQK GNELALPSKYVNFLYLASHYEKGKGSPEDNEQKQLFVEQHKHYLDEI IEQISEFSKRVIAADANLDKVLSAYNKHRDKPIREQAENIIHLFTLT NLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLS QLGGD 40 . Cas9V7 MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNL IGAL L F D S GE TAEAT RL KRTARRRYT RRKNRIC YL QEIF S NEMAKVD DSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKK LVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQL VQT YNQL FEENPINAS GVDAKAIL S ARL S KS RRL ENLIAQL PGEKKN GLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQI GDQQADLFLAAKNLSDAILLSDILRVNTEITKAPLHASMIKRYDEHH QDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGACQEEFYKFI KPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAIL RRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSE ETITPWNFEEWDKGASAQSFIERMTNFDKNLPNEKVLPKHSLTYEY FTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQL KEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEE NEDILEDIVLTQTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTG WGRL SRKLINGIRDKQS GKTILDFLKSDGFANRNFMQLIHDDS LTFK EDIQKAQVSGQGDSLHEHIANGAGSPAIKKGILQTVKWDELVKVMG RHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEH PVENTQLQNEKLYLYYDQNGRDMYVDQELDINRLSDYDVDHIVPQSF L KDD SIDNKVLT RS DKNRGKS DNVPS E EWKKMKNYWRQL LNAKLIT QRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMN TKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAHDA YLNAWGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATA KYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFA TVRKVL SMPQVNIVKKT EVQT GGFS KE SIL PKRNS DKLIARKKDWD P KKYGGFDSPTVAYSVLWAKVEKGKSKKLKSVKELLGITIMERSSFE KNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQK GNELALPSKYVNFLYLASHYEKGKGSPEDNEQKQLFVEQHKHYLDEI IEQISEFSKRVIAADANLDKVLSAYNKHRDKPIREQAENIIHLFTLT NLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLS QLGGD 41. Cas9V8 MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNL IGAL L F D S GE TAEAT RL KRTARRRYT RRKNRIC YL QEIF S NEMAKVD DSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKK LVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQL VQT YNQL FEENPINAS GVDAKAIL SARL S KS RRL ENLIAQL PGEKKN GCFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQI GDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHH QDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGACQEEFYKFI KPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAIL RRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSE ETITPWNFEEWDKGASAQSFIERMTNTDKNLPNEKVLPKHSLLYEY FTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQL KEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEE NEDIGEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTG WGRL SRKLINGIRDKQS GKTILDFLKS DGFANRNFMQLIHDDS LTFK EDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKWDELVKVMG RHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEH PVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSF L KDD SIDNKVLT RS DKNRGKS DNVPS E EWKKMKNYWRQL LNAKLIT QRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMN TKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAHDA YLNAWGTALIKKYPKLESEFVYGDYKVGDVRKMIAKSEQEIGKATA KYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFA T VRKVL SMPQVN IVKKT EVQT GGFS KE S IL PKRNS DKL IARKKDWD P KKYGGFDSPTVAYSVLWAKVEKGKSKKLKSVKELLGITIMERSSFE KNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQK GNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEI IEQISEFSKRVILADANLDKVLSAQNKHRDKPIREQAENIIHLFTLT NLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLS QLGGD 42 . Cas9V9 MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNL IGAL L F D S GE TAEAT RL KRTARRRYT RRKNRIC YL QEIF S NEMAKVD DSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKK LVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQL VQTYNQL FEENPINAS GVDAKAIL SARL S KS RRL ENLIAQL PGEKKN GCFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQI GDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHH QDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGACQEEFYKFI KPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAIL RRQEDFYPFLKDNREKTEKTLTFRIPYYVGPLARGNSRFAWMTRKSE ETITPWNFEEWDKGASAQSFIERMTNTDKNLPNEKVLPKHSLGYEY FTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQL KEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEE NEDIGEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTG WGRL SRKLINGIRDKQS GKTILDFLKS DGFANRNFMQLIHDDS LTFK EDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKWDELVKVMG RHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEH PVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSF L KDD SIDNKVLT RS DKNRGKS DNVPS E EWKKMKNYWRQL LNAKLIT QRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMN TKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAHDA YLNAWGTALIKKYPKLESEFVYGDYKVGDVRKMIAKSEQEIGKATA KYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFA T VRKVL SMPQVNIVKKT EVQT GGFS KE S IL PKRNS DKLIARKKDWD P KKYGGFDSPTVAYSVLWAKVEKGKSKKLKSVKELLGITIMERSSFE KNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQK GNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEI IEQQSEFSKRVILADANLDKVLSAQNKHRDKPIREQAENIIHLFTLT NLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLS QLGGD 43 . Cas9V10 MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNL IGAL L F D S GE TAEAT RL KRTARRRYT RRKNRIC YL QEIF S NEMAKVD DSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKK LVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQL VQT YNQL FEENPINAS GVDAKAIL SARL S KS RRL ENLIAQL PGEKKN GLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQI GDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLHASMIKRYDEHH QDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFI KPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAIL RRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSE ETITPWNFEEWDKGASAQSFIERMTNTDKNLPNEKVLPKHSLGYEY FTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQL KEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEE NEDILEDIVLTQTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTG WGRL SRKLINGIRDKQS GKTILDFLKS DGFANRNFMQLIHDDS LTFK EDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKWDELVKVMG RHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEH PVENTQLQNEKLYLYYLQNGRDMYVDQELDTNRLSDYDVDHIVPQSF L KDD SIDNKVLT RS DKNRGKS DNVPS E E WKKMKNYWRQL LNAKLIT QRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMN TKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAHDA YLNAWGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATA KYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFA TVRKVL SMPQVNIVKKT EVQT GGFS KE SIL PKRNS DKLIARKKDWD P KKYGGFDSPTVAYSVLWAKVEKGKSKKLKSVKELLGITIMERSSFE KNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQK GNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEI IEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLT NLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLS QLGGD 44 . Cas9Vll MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNL IGAL L F D S GE TAEAT RL KRTARRRYT RRKNRIC YL QEIF S NEMAKVD DSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKK LVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQL VQT YNQL FEENPINAS GVDAKAIL SARL S KS RRL ENLIAQL PGEKKN GLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQI GDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHH QDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFI KPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAIL RRQEDFYPFLKDNREKTEKTLTFRIPYYVGPLARGNSRFAWMTRKSE ETITPWNFEEWDKGASAQSFIERMTNTDKNLPNEKVLPKHSLLYEY FTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQL KEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEE NEDILEDIVLTQTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTG WGRL SRKLINGIRDKQS GKTILDFLKS DGFANRNFMQLIHDDS LTFK EDIQKAQVSGQGDSLHEHIANGAGSPAIKKGILQTVKWDELVKVMG RHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEH PVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSF L KDD SIDNKVLT RS DKNRGKS DNVPS E EWKKMKNYWRQL LNAKLIT QRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMN TKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAHDA YLNAWGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATA KYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFA T VRKVL SMPQVN IVKKT EVQT GGFS KE S IL PKRNS DKL IARKKDWD P KKYGGFDSPTVA...
Claims
1. A method of forming a circularized ribonucleic acid (RNA) in a cell, themethod comprising transfecting a cell with a circularizable linear RNA compound that is capable of circularizing within the cell, thereby forming a circularized RNA,wherein the linear RNA compound is a nucleic acid comprising from 5' to 3': a first member of a split ligation stem, an internal ribosome entry site (IRES) domain, a 5' untranslated region (UTR), a protein-encoding nucleic acid sequence at least 750 nucleotides in length, a 3' UTR, a poly-adenine (poly A) domain, and a second member of the split ligation stem.
2. The method of claim 1, further comprising allowing the cell to translate the protein-encoding nucleic acid sequence, thereby forming a protein within the cell.
3. The method of claim 1 or 2, wherein the method comprises cleaving a ribozyme-cleavable linear RNA compound to form the circularizable linear RNA compound,wherein the ribozyme-cleavable linear RNA compound comprises from 5' to 3': a first ribozyme domain, an internal ribosome entry site (IRES) domain, a 5' untranslated region (UTR). a protein-encoding nucleic acid sequence at least 750 nucleotides in length, a 3' UTR, a poly-adenine (polyA) domain, and a second ribozyme domain,wherein the cleaving comprises allowing the first ribozy me domain to cleave the ribozyme-cleavable linear RNA compound in the 3' direction, thereby forming a 5' end comprising the first member of the split ligation stem; and the second ribozyme domain to cleave the ribozyme-cleavable linear RNA compound in the 5' direction, thereby forming a 3' end comprising the second member of the split ligation stem.
4. The method of claim 3, wherein the first ribozyme domain is a first twister ribozyme domain and the second ribozyme domain is a second twister ribozyme domain.
5. The method of claim 4, wherein the first twister ribozy me domain and the second twister ribozy me domain independently are a Sma-1-402 twister ribozyme domain, an env-112 twister ribozyme domain, a Slyc-2-1 twister ribozyme domain, an Osin-1-1 twister ribozyme domain, an env-270 twister ribozyme domain, an env-94 twister ribozy me domain, an env-935 twister ribozyme domain, a Sma-1-66 twister ribozyme domain, a Dre-1-3 twisterribozyme domain, an Osa-1-4 twister ribozyme domain, an Eana-1-1 twister ribozy me domain, an Osa-1-8 twister ribozyme domain, an Osa-1-3 twister ribozy me domain, an env-13 twister ribozy me domain, a Dre-1-4 twister ribozy me domain, an Aage-1-1 twister ribozyme domain, a Spol-1 -1 twister ribozyme domain, a Ttru-1-1 twister ribozyme domain, an Hsap-1-1 twister ribozyme domain, or an Hsap-1-2 twister ribozyme domain.
6. The method of any one of claims 3-5, wherein the first twister ribozy me domain and the second twister ribozyme domain independently comprise the nucleotide sequence of any one of SEQ ID NOs: 1-22.
7. The method of any one of claims 1-6, further comprising allowing thefirst member of the split ligation stem and the second member of the split ligation stem to react, thereby forming a ligation stem.
8. The method of any one of claims 1-7, wherein the first member of the split ligation stem comprises a 5' hydroxyl and the second member of the split ligation stem comprises a 2'-3' cyclic phosphate.
9. The method of claim 8, wherein the 5' hydroxyl group and the 2'-3' cyclic phosphate of the ligation stem are ligated together by a RtcB ligase, thereby forming a circularized RNA.
10. The method of any one of claims 3-9, wherein the first ribozyme domain comprises the nucleotide sequence of SEQ ID NO: 1 and the second ribozyme domain comprises the nucleotide sequence of SEQ ID NO:2.
11. The method of any one of claims 1-10, wherein the first member of the split ligation stem comprises the nucleotide sequence of SEQ ID NO:23 and the second member of the split ligation stem comprises the nucleotide sequence of SEQ ID NO:24.
12. The method of any one of claims 1-11, wherein the protein-encoding nucleic acid sequence encodes an epigenetic effector domain, a repressor domain, or an activator domain.
13. The method of any one of claims 1-12, wherein the protein-encoding nucleic acid sequence encodes a deoxyribonucleic acid (DNA) methyltransferase domain, a CpG 309methyltransferase (M.SssI) domain, a Sin3 interacting repressor domain (SID4X), a protamine 2 (PRM2) domain, a protamine 1 (PRM1) domain, a VP64 domain, a Kruppel associated box (KRAB) domain, a coupled histone tail for autoinhibition release of methyltransferase (CHARM) domain, or a poly dactyl zinc finger protein domain.
14. The method of any one of claims 1-13, wherein the protein-encoding nucleic acid sequence encodes a protein comprising the amino acid sequence of any one of SEQ ID NOs:25-54.
15. The method of claim 13, wherein the DNA methyltransferase domain is a is a Dnmt3A-3L domain. .
16. The method of claim 15, wherein the Dnmt3A-3L domain comprises the amino acid sequence of SEQ ID NO: 32.
17. The method of claim 13, wherein the poly dactyl zinc finger protein domain comprises a first alpha helix domain comprising the amino acid sequence of SEQ ID NO:55 a second alpha helix domain comprising the amino acid sequence of SEQ ID NO:56, a third alpha helix domain comprising the amino acid sequence of SEQ ID NO: 57, a fourth alpha helix domain comprising the amino acid sequence of SEQ ID NO:
58. a fifth alpha helix domain comprising the amino acid sequence of SEQ ID NO:59, and a sixth alpha helix domain comprising the amino acid sequence of SEQ ID NO:60.
18. The method of claim 13 or 17, wherein the poly dactyl zinc finger protein domain comprises the amino acid sequence of SEQ ID N0:61.
19. The method of any one of claims 1-16, wherein the IRES domain is a cricket paralysis virus (CrPV) IRES domain, an insulin-like growth factor 2 (IGF2) IRES domain, a hepatitis C virus H77 IRES domain, a fibroblast growth factor 1 (FGF1) IRES domain, a bovine viral diarrhea virus (BVDV) 1 IRES domain, a human rhinovirus A89 IRES domain, a LIM domain and actin binding protein 1 (LIMA1) IRES domain, a human adenovirus 2 IRES domain, a montana Myotis leukoencephalitis virus (MMLV) IRES domain, a RAN binding protein 3 (RANBP3) IRES domain, a pestivirus giraffe 1 IRES domain, a TG-interacting factor 1 (TGIF1) IRES domain, a human poliovirus 1 Mahoney IRES domain, a Foot-and-Mouth disease virus type 0 IRES domain, an encephalomyocarditis virus (ECMV)IRES domain, an encephalomyocarditis virus 7A IRES domain, an encephalomyocarditis virus 6A IRES domain, an enterovirus 71 IRES domain, a Coxsackievirus B3 (CB3) IRES domain, a pegivirus A IRES domain, an equine rhinitis A virus (ERAV) IRES domain, a GB virus C (GBV-HGV) IRES domain, a human betaherpesvirus 5 IRES domain, a Senecavirus A (SVA) IRES domain, an equine rhinitis B virus 1 (ERBV-1) IRES domain, a Triticum mosaic virus (TriMV) IRES domain, ahepatovirus A (HAV) IRES domain, a hepatitis GB virus B (HGBV-B) IRES domain, a Giardia lamblia virus (GLV) IRES domain, a Cyrphonectria hy povirus 1 IRES domain, or an equine hepacivrus JPN3 / JAPAN / 2013 IRES domain.
20. The method of any one of claims 1-19, wherein the IRES domain comprises the nucleotide sequence of any one of SEQ ID NOs :63-94.
21. The method of any one of claims 1-20, wherein the 3' UTR is a polyadenine (pA) 3' UTR, a mtRNRl-AES 3' UTR, a mtRNRl-LSPl 3' UTR, an AES-mtRNRl 3' UTR, an AES-hBg 3' UTR, a 2hBg 3' UTR, a FCGRT-hBg 3' UTR, or an HBA1 3' UTR.
22. The method of any one of claims 1-21, wherein the 3’ UTR comprises the nucleotide sequence of any one of SEQ ID NOs:95-103.
23. The method of any one of claims 1-22, wherein the linear RNA compound further comprises a transcription termination domain.
24. The method of claim 23, wherein the transcription termination domain comprises the nucleotide sequence of SEQ ID NO: 104.
25. A linear RNA compound comprising from 5’ to 3': a first ribozyme domain, an internal ribosome entry site (IRES) domain, a 5' untranslated region (UTR), a protein-encoding nucleic acid sequence at least 750 nucleotides in length, a 3' UTR, a polyadenine (polyA) domain, and a second ribozyme domain.
26. The linear RNA compound of claim 25, wherein the first ribozyme domain is a first twister ribozyme domain and the second ribozyme domain is a second twister ribozyme domain.
27. The linear RNA compound of claim 26, wherein the first twister ribozyme domain and the second twister ribozyme domain independently are a Sma-1-402 twisterribozyme domain, an env-112 twister ribozy me domain, a Slyc-2-1 twister ribozyme domain, an Osin-1-1 twister ribozyme domain, an env-270 twister ribozyme domain, an env-94 twister ribozy me domain, an env-935 twister ribozyme domain, a Sma-1-66 twister ribozyme domain, a Dre-1-3 twister ribozyme domain, an Osa-1-4 twister ribozyme domain, an Eana-1-1 twister ribozyme domain, an Osa-1-8 twister ribozyme domain, an Osa-1-3 twister ribozyme domain, an env-13 twister ribozyme domain, a Dre-1-4 twister ribozy me domain, an Aage-1-1 twister ribozy me domain, a Spol-1-1 twister ribozyme domain, aTtru-1-1 twister ribozyme domain, an Hsap-1-1 twister ribozyme domain, or an Hsap-1-2 twister ribozyme domain.
28. The linear RNA compound of any one of claims 26 or 27, wherein thefirst twister ribozyme domain and the second twister ribozyme domain independently comprise the nucleotide sequence of any one of SEQ ID NOs: 1-22.
29. The linear RNA compound of any one of claims 25-28, wherein the first ribozy me domain comprises the nucleotide sequence of SEQ ID NO: 1 and the second ribozyme domain comprises the nucleotide sequence of SEQ ID N0:2.
30. The linear RNA compound of any one of claims 25-29, wherein the protein-encoding nucleic acid sequence encodes an epigenetic effector domain, a repressor domain, or an activator domain.
31. The linear RNA compound of any one of claims 25-30, wherein theprotein-encoding nucleic acid sequence encodes a deoxyribonucleic acid (DNA) methyltransferase domain, a CpG methyltransferase (M.SssI) domain, a Sin3 interacting repressor domain (SID4X), a protamine 2 (PRM2) domain, a protamine 1 (PRM1) domain, a VP64 domain, a Kriippel associated box (KRAB) domain, a coupled histone tail for autoinhibition release of methyltransferase (CHARM) domain, or a poly dactyl zinc finger protein domain.
32. The linear RNA compound of any one of claims 25-31, wherein the protein-encoding nucleic acid sequence encodes a protein comprising the amino acid sequence of any one of SEQ ID NOs:25-54.
33. The linear RNA compound of claim 31, wherein the DNAmethyltransferase domain is a is a Dnmt3A-3L domain.
34. The linear RNA compound of claim 33, wherein the Dnmt3A-3L domain comprises the amino acid sequence of SEQ ID NO:32.
35. The linear RNA compound of claim 31, wherein the poly dactyl zinc finger protein domain comprises a first alpha helix domain comprising the amino acid sequence of SEQ ID NO:55 a second alpha helix domain comprising the amino acid sequence of SEQ ID NO:
56. a third alpha helix domain comprising the amino acid sequence of SEQ ID NO:
57. a fourth alpha helix domain comprising the amino acid sequence of SEQ ID NO:58, a fifth alpha helix domain comprising the amino acid sequence of SEQ ID NO:59, and a sixth alpha helix domain comprising the amino acid sequence of SEQ ID NO:60.
36. The linear RNA compound of claim 31 or 35, wherein the poly dactyl zinc finger protein domain comprises the amino acid sequence of SEQ ID NO: 61.
37. The linear RNA compound of any one of claims 25-36, wherein the IRES domain is a cricket paralysis virus (CrPV) IRES domain, an insulin-like growth factor 2 (IGF2) IRES domain, a hepatitis C virus H77 IRES domain, a fibroblast grow th factor 1 (FGF1) IRES domain, a bovine viral diarrhea virus (BVDV) 1 IRES domain, a human rhinovirus A89 IRES domain, a LIM domain and actin binding protein 1 (LIMA1) IRES domain, a human adenovirus 2 IRES domain, a montana Myotis leukoencephalitis virus (MMLV) IRES domain, a RAN binding protein 3 (RANBP3) IRES domain, a pestivirus giraffe 1 IRES domain, a TG-interacting factor 1 (TGIF1) IRES domain, a human poliovirus 1 Mahoney IRES domain, a Foot-and-Mouth disease virus type 0 IRES domain, an encephalomyocarditis virus (ECMV) IRES domain, an encephalomyocarditis virus 7A IRES domain, an encephalomyocarditis virus 6A IRES domain, an enterovirus 71 IRES domain, a Coxsackievirus B3 (CB3) IRES domain, a pegivirus A IRES domain, an equine rhinitis A virus (ERAV) IRES domain, a GB virus C (GBV-HGV) IRES domain, a human betaherpesvirus 5 IRES domain, a Senecavirus A (SVA) IRES domain, an equine rhinitis B virus 1 (ERBV-1) IRES domain, a Triticum mosaic virus (TriMV) IRES domain, ahepatovirus A (HAV) IRES domain, a hepatitis GB virus B (HGBV-B) IRES domain, a Giardia lamblia virus (GLV) IRES domain, a Cyrphonectria hypovirus 1 IRES domain, or an equine hepacivrus JPN3 / JAPAN / 2013 IRES domain.
38. The linear RNA compound of any one of claims 25-37, wherein the IRES domain comprises the nucleotide sequence of any one of SEQ ID NOs:63-94.
39. The linear RNA compound of any one of claims 25-38, wherein the 3' UTR is a poly-adenine (pA) 3' UTR, a mtRNRl-AES 3' UTR, a mtRNRl-LSPl 3' UTR, an AES-mtRNRl 3' UTR, an AES-hBg 3' UTR, a 2hBg 3' UTR, a FCGRT-hBg 3' UTR, or an HBA1 3' UTR.
40. The linear RNA compound of any one of claims 25-39, wherein the 3' UTR comprises the nucleotide sequence of any one of SEQ ID NOs:95-103.
41. The linear RNA compound of any one of claims 25-40, wherein the linear RNA compound further comprises a transcription termination domain.
42. The linear RNA compound of claim 41, wherein the transcription termination domain comprises the nucleotide sequence of SEQ ID NO: 104.
43. A method of forming a circularized ribonucleic acid (RNA) comprising : incubating a linear RNA compound in vitro for between about 16 hours and about 24 hours under conditions conducive to group II intron cleavage, thereby forming a circularized RNA,wherein the RNA compound is a nucleic acid comprising from 5' to 3': a first member of a split group II intron domain, an internal ribosome entry site (IRES) domain, a 5' untranslated region (UTR), a protein-encoding nucleic acid sequence at least 750 nucleotides in length, a 3' UTR, a poly-adenine (poly A) domain, and a second member of the split group II intron domain,wherein the first member of the split group II intron comprises domains V and VI and the second member of the split group II intron comprises domains I. Il and 111.
44. The method of claim 43, wherein the group II intron domain is aClostridium tetani group II intron domain. aHistoplasma capsulatum group II intron domain, a Coccidioides immitis group II intron domain, a Blastomyces dermatitidis group 11 intron domain, a Coccidioidomycosis posadasii group II intron domain, a Pylaeiella littoralis group II intron domain, a Saccharomyces cerevisiae group II intron domain, a Lactococcus lactis group IIintron domain, aAnthoceros angustus group II intron domain, a Geobacillus stearothermophilus group II intron domain, a Bacillus megaterium group II intron domain, a Pseudomonas alcaligenes group II intron domain, a Chaetothyriales bantiana group II intron domain, or a Chaetothyriales carrioni group II intron domain.
45. The method of claim 43 or 44, wherein the group II intron domain comprises the nucleotide sequence of SEQ ID NO: 105-116.
46. The method of any one of claims 43-45, wherein the first member of the split group II intron domain comprises the nucleotide sequence of SEQ ID NO: 105 and the second member of the split group II intron domain comprises the nucleotide sequence of SEQ ID NO: 106.
47. The method of any one of claims 4346, wherein the protein-encoding nucleic acid sequence encodes an epigenetic effector domain, a repressor domain, or an activator domain.
48. The method of any one of claims 43-47, wherein the protein-encoding nucleic acid sequence encodes a deoxyribonucleic acid (DNA) methyltransferase domain, a CpG methyltransferase (M.SssI) domain, a Sin3 interacting repressor domain (SID4X), a protamine 2 (PRM2) domain, a protamine 1 (PRM1) domain, a VP64 domain, a Kriippel associated box (KRAB) domain, a coupled histone tail for autoinhibition release of methyltransferase (CHARM) domain, or a poly dactyl zinc finger protein domain.
49. The method of any one of claims 43-45. wherein the protein-encoding nucleic acid sequence encodes a protein comprising the amino acid sequence of any one of SEQ ID NOs:25-54.
50. The method of claim 48 or 49, wherein the DNA methyltransferase domain is a is a Dnmt3A-3L domain.
51. The method of claim 50, wherein the Dnmt3A-3L domain comprises the amino acid sequence of SEQ ID NO: 32.
52. The method of claim 48, wherein the polydactyl zinc finger protein domain comprises a first alpha helix domain comprising the amino acid sequence of SEQ IDNO:55 a second alpha helix domain comprising the amino acid sequence of SEQ ID NO:56, a third alpha helix domain comprising the amino acid sequence of SEQ ID NO: 57, a fourth alpha helix domain comprising the amino acid sequence of SEQ ID NO:
58. a fifth alpha helix domain comprising the amino acid sequence of SEQ ID NO:
59. and a sixth alpha helix domain comprising the amino acid sequence of SEQ ID NO:60.
53. The method of claim 48 or 52, wherein the polydactyl zinc finger protein domain comprises the amino acid sequence of SEQ ID N0:61.
54. The method of any one of claims 43-53, wherein the IRES domain is a cricket paralysis virus (CrPV) IRES domain, an insulin-like growth factor 2 (IGF2) IRES domain, a hepatitis C virus H77 IRES domain, a fibroblast growth factor 1 (FGF1) IRES domain, a bovine viral diarrhea virus (BVDV) 1 IRES domain, a human rhinovirus A89 IRES domain, a LIM domain and actin binding protein 1 (LIMA1) IRES domain, a human adenovirus 2 IRES domain, a montana Myotis leukoencephalitis virus (MMLV) IRES domain, a RAN binding protein 3 (RANBP3) IRES domain, a pestivirus giraffe 1 IRES domain, a TG-interacting factor 1 (TG1F1) IRES domain, a human poliovirus 1 Mahoney IRES domain, a Foot-and-Mouth disease virus type 0 IRES domain, an encephalomyocarditis virus (ECMV) IRES domain, an encephalomyocarditis virus 7A IRES domain, an encephalomyocarditis vims 6A IRES domain, an enterovirus 71 IRES domain, a Coxsackievirus B3 (CB3) IRES domain, a pegivirus A IRES domain, an equine rhinitis A vims (ERAV) IRES domain, a GB virus C (GBV-HGV) IRES domain, a human betaherpesvirus 5 IRES domain, a Senecavims A (SVA) IRES domain, an equine rhinitis B vims 1 (ERBV-1) IRES domain, a Triticum mosaic virus (TriMV) IRES domain, a hepatovims A (HAV) IRES domain, a hepatitis GB virus B (HGBV-B) IRES domain, a Giardia lamblia vims (GLV) IRES domain, a Cyrphonectria hypovims 1 IRES domain, or an equine hepacivrus JPN3 / JAPAN / 2013 IRES domain.
55. The method of any one of claims 43-54, wherein the IRES domain comprises the nucleotide sequence of any one of SEQ ID NOs :63-94.
56. The method of any one of claims 43-55, wherein the 3' UTR is a polyadenine (pA) 3' UTR, a mtRNRl-AES 3' UTR, a mtRNRl-LSPl 3' UTR, an AES-mtRNRl 3' UTR, an AES-hBg 3' UTR, a 2hBg 3' UTR, a FCGRT-hBg 3' UTR, or an HBA1 3' UTR.
57. The method of any one of claims 43-56, wherein the 3' UTR comprises the nucleotide sequence of any one of SEQ ID NOs:95-103.
58. The method of any one of claims 43-57, wherein the linear RNA compound further comprises a transcription termination domain.
59. The method of claim 58, wherein the transcription termination domain comprises the nucleotide sequence of SEQ ID NO: 104.
60. A linear RNA compound comprising from 5' to 3': a first member of a split group II intron domain, an internal ribosome entry site (IRES) domain, a 5' untranslated region (UTR), a protein-encoding nucleic acid sequence at least 750 nucleotides in length, a 3' UTR, a poly-adenine (poly A) domain, and a second member of the split group II intron domain.
61. The linear RNA compound of claim 60, wherein the group II intron domain is a Clostridium tetani group II intron domain, a Histoplasma capsulatum group II intron domain, a Coccidioides immitis group II intron domain, a Blastomyces dermatitidis group II intron domain, a Coccidioidomycosis posadasii group II intron domain, a Pylaeiella littoralis group II intron domain, a Saccharomyces cerevisiae group II intron domain, a Lactococcus lactis group II intron domain, aAnthoceros angustus group II intron domain, a Geobacillus stearothermophilus group II intron domain, a Bacillus megaterium group II intron domain, a Pseudomonas alcaligenes group II intron domain, a Chaetothyriales bandana group II intron domain, or a Chaetothyriales carrioni group II intron domain.
62. The linear RNA compound of claim 60 or 61, wherein the group II intron domain comprises the nucleotide sequence of any one of SEQ IDNOs:105-116.
63. The linear RNA compound of any one of claims 60-62, wherein the first member of the split group II intron domain comprises the nucleotide sequence of SEQ ID NO: 105 and the second member of the split group II intron domain comprises the nucleotide sequence of SEQ ID NO: 106.
64. The linear RNA compound of claim 60 or 61, wherein the proteinencoding nucleic acid sequence encodes an epigenetic effector domain, a repressor domain, or an activator domain.
65. The linear RNA compound of any one of claims 60-64, wherein the protein-encoding nucleic acid sequence encodes a deoxyribonucleic acid (DNA) methyltransferase domain, a CpG methyltransferase (M.SssI) domain, a Sin3 interacting repressor domain (SID4X), a protamine 2 (PRM2) domain, a protamine 1 (PRM1) domain, a VP64 domain, a Kriippel associated box (KRAB) domain, a coupled histone tail for autoinhibition release of methyltransferase (CHARM) domain, or a polydactyl zinc finger protein domain.
66. The method of any one of claims 60-65, wherein the protein-encoding nucleic acid sequence encodes a protein comprising the amino acid sequence of any one of SEQ ID NOs:25-54.
67. The linear RNA compound of claim 65 or 66, wherein the DNA methyltransferase domain is a is a Dnmt3A-3L domain.
68. The method of claim 67, wherein the Dnmt3A-3L domain comprises the amino acid sequence of SEQ ID NO: 32.
69. The linear RNA compound of claim 65, wherein the poly dactyl zinc finger protein domain comprises a first alpha helix domain comprising the amino acid sequence of SEQ ID NO:55 a second alpha helix domain comprising the amino acid sequence of SEQ ID NO:56, a third alphahelix domain comprising the amino acid sequence of SEQ ID NO:57, a fourth alpha helix domain comprising the amino acid sequence of SEQ ID NO:58, a fifth alpha helix domain comprising the amino acid sequence of SEQ ID NO:
59. and a sixth alpha helix domain comprising the amino acid sequence of SEQ ID NO:60.
70. The linear RNA compound of claim 65 or 69, wherein the poly dactyl zinc finger protein domain comprises the amino acid sequence of SEQ ID NO:61.
71. The linear RNA compound of any one of claims 60-70, wherein the IRES domain is a cricket paralysis virus (CrPV) IRES domain, an insulin-like growth factor 2 (IGF2) IRES domain, a hepatitis C virus H77 IRES domain, a fibroblast growth factor 1 (FGF1) IRES domain, a bovine viral diarrhea virus (BVDV) 1 IRES domain, a human rhinovirus A89 IRES domain, a LIM domain and actin binding protein 1 (LIMA1) IRES domain, a human adenovirus2 IRES domain, a montana Myotis leukoencephalitis virus (MMLV) IRES domain, a RAN binding protein 3 (RANBP3) IRES domain, a pestivirus giraffe 1 IRES domain, a TG-interacting factor 1 (TGIF1) IRES domain, a human poliovirus 1 Mahoney IRES domain, a Foot-and-Mouth disease virus type O IRES domain, an encephalomyocarditis virus (ECMV) IRES domain, an encephalomyocarditis virus 7A IRES domain, an encephalomyocarditis virus 6A IRES domain, an enterovirus 71 IRES domain, a Coxsackievirus B3 (CB3) IRES domain, a pegivirus A IRES domain, an equine rhinitis A virus (ERAV) IRES domain, a GB virus C (GBV-HGV) IRES domain, a human betaherpesvirus 5 IRES domain, a Senecavirus A (SVA) IRES domain, an equine rhinitis B virus 1 (ERBV-1) IRES domain, a Triticum mosaic virus (TriMV) IRES domain, ahepatovirus A (HAV) IRES domain, a hepatitis GB virus B (HGBV-B) IRES domain, a Giardia lamblia virus (GLV) IRES domain, a Cyrphonectria hypovirus 1 IRES domain, or an equine hepacivrus JPN3 / JAPAN / 2013 IRES domain.
72. The linear RNA compound of any one of claims 60-71, wherein the IRES domain comprises the nucleotide sequence of any one of SEQ ID NOs:63-94.
73. The linear RNA compound of any one of claims 60-72, wherein the 3' UTR is a poly-adenine (pA) 3' UTR, a mtRNRl-AES 3' UTR, a mtRNRl-LSPl 3’ UTR, an AES-mtRNRl 3’ UTR, an AES-hBg 3’ UTR, a 2hBg 3’ UTR, a FCGRT-hBg 3' UTR, or an HBA1 3’ UTR.
74. The linear RNA compound of any one of claims 60-73, wherein the 3’ UTR comprises the nucleotide sequence of any one of SEQ ID NOs:95-103.
75. The linear RNA compound of any one of claims 60-74, wherein the linear RNA compound further comprises a transcription termination domain.
76. The linear RNA compound of claim 75, wherein the transcription termination domain comprises the nucleotide sequence of SEQ ID NO: 104.
77. A DNA endonuclease enzy me comprising the amino acid sequence of any one of SEQ ID NO:33-53.
78. A poly dactyl zinc finger protein comprising a first alpha helix domaincomprising the amino acid sequence of SEQ ID NO:55 a second alpha helix domain comprisingthe amino acid sequence of SEQ ID NO:56, a third alpha helix domain comprising the amino acid sequence of SEQ ID NO: 57, a fourth alpha helix domain comprising the amino acid sequence of SEQ ID NO:58, a fifth alpha helix domain comprising the amino acid sequence of SEQ ID NO:59, and a sixth alpha helix domain comprising the amino acid sequence of SEQ ID NO:60..79.. The poly dactyl zinc finger protein of claim 78, wherein the poly dactyl zinc finger protein comprises from N-terminus to C-terminus: the first alpha helix domain, the second alpha helix domain, the third alpha helix domain, the fourth alpha helix domain, the fifth alpha helix domain, and the sixth alpha helix domain.
80. The polydactyl zinc finger protein of claim 78 or 79. wherein the poly dactyl zinc finger protein comprises an amino acid sequence having at least 80% sequence identity' to the sequence of SEQ ID NO:61.
81. The poly dactyl zinc finger protein of any one of claims 78-80. wherein the poly dactyl zinc finger protein comprises the amino acid sequence of SEQ ID NO:61.