Repressor fusion protein system
A fusion protein system with a modified DNA-binding protein and repressor domain addresses gene repression inefficiencies, achieving stable and targeted gene silencing with reduced off-target effects for therapeutic and research applications.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SCRIBE THERAPEUTICS INC
- Filing Date
- 2024-03-28
- Publication Date
- 2026-05-26
AI Technical Summary
Existing gene repression methods, such as RNAi and CRISPR-based systems, suffer from off-target effects and inefficiencies, limiting their effectiveness in therapeutic and research applications.
A system comprising a modified DNA-binding protein fused with a repressor domain and optionally a guide RNA, designed for targeted gene silencing and epigenetic modification, delivered via vectors or formulations like lipid nanoparticles, to achieve heritable gene repression.
The system effectively represses target gene transcription with reduced off-target effects, providing a stable and efficient means for disease treatment and research by inducing heritable epigenetic modifications.
Smart Images

Figure 2026516568000001_ABST
Abstract
Description
[Technical Field]
[0001] Cross-references to related applications This application claims priority to and benefits of U.S. Provisional Patent Application No. 63 / 492,845, filed on 29 March 2023, and U.S. Provisional Patent Application No. 63 / 505,906, filed on 2 June 2023, the contents of which, by reference, are incorporated herein by reference in their entirety.
[0002] Reference to electronic sequence listings The contents of the electronic sequence listing (SCRB_054_02WO_SeqList_ST26.xml; size: 141,358,486 bytes; and creation date: March 26, 2024) are incorporated herein by reference in their entirety. [Background technology]
[0003] The methods by which the expression of target genes in cells is variable. In mammalian systems, cells use a system of chromatin regulators (CRs), as well as associated histone and DNA modifications, to regulate gene expression and establish long-term epigenetic memory. This system is crucial in development, aging, and disease and can provide essential capabilities for incorporating regulation into synthetic biology. In experimental systems, methods such as RNA interference (RNAi) have been useful for targeted gene knockdown and have been widely used for large-scale library screening. RNAi, however, has some limitations. In particular, RNAi-based knockdown is plagued by off-target effects, along with incomplete knockdown of the target (Jackson AL, et al. Expression profiling reveals off-target gene regulation by RNAi. Nat Biotechnol. 21:635 (2003)); Sigoillot FD, et al. A bioinformatics method identifies prominent off-targeted transcripts in RNAi screens. Nat Methods. 19:9(4):363 (2012)). Modified DNA-binding proteins linked to transcriptional repressor domains, such as zinc finger proteins or transcription activator-like effectors (TALEs), can mediate selective gene repression, but are limited by the fact that each desired target gene requires the generation of a new protein.
[0004] The emergence of DNA editing systems and their programmable nature has facilitated their use as multipurpose technologies for genome manipulation and engineering. CRISPR proteins, in particular, are well-suited for such operations. For example, certain Class 2 CRISPR / Cas systems have a compact size, offer ease of delivery, and have relatively short protein-coding nucleotide sequences, which are advantages for their incorporation into viral vectors for cell delivery. However, in certain disease manifestations, gene silencing, or transcriptional repression, is preferred over gene editing. The ability to catalytically inactivate CRISPR nucleases, such as Cas9 and CasX, has been demonstrated (International Publication No. 2020247882A1 and U.S. Patent No. 20200087641A1, incorporated herein by reference), thereby making these systems an attractive platform for the creation of fusion proteins with repressor domains capable of gene silencing. While certain repressor systems have been described, there remains a need for additional gene repressor systems that are optimized and / or offer improvements over earlier generation gene repressor systems, such as Cas9-based ones for use in various therapeutic, diagnostic, and research applications.
[0005] Provided herein are systems and methods for addressing this need, including epigenetic modifications, as well as delivery vectors and formulations for targeting and repressing genes in cells. [Overview of the Initiative]
[0006] Aspects of this disclosure are directed to systems and methods for regulating the expression of target nucleic acids in cells.
[0007] This disclosure provides a system comprising or encoding a fusion protein comprising a modified DNA-binding protein and a ligated repressor domain protein in a defined configuration, used in some cases with a guide ribonucleic acid (gRNA) in transcriptional repression and / or epigenetic modification of a target nucleic acid sequence. The components of the system can be modified for formulations for passive entry into target cells and are useful in various methods for gene silencing or gene transcriptional repression in disease, where the repression of the gene product is useful for reversing the root cause of the disease or for relieving the signs or symptoms of the disease, and such methods are also provided. The systems of this disclosure can result in heritable epigenetic modification or silencing of the gene targeted by the system. This disclosure provides a method for repressing gene transcription in a population of cells, the method comprising introducing into the cells of the population a repressor fusion protein comprising or encoding a modified DNA-binding protein and a ligated repressor domain protein, and in some cases gRNA, wherein gene transcription is repressed by the repressor fusion protein. The disclosure also provides vectors and particle formulations (e.g., lipid nanoparticles, or LNPs, and synthetic nanoparticles) that encode or encapsulate system components for delivery to cells for silencing or transcriptional repression of target nucleic acids in cells.
[0008] In another aspect, the present invention provides a composition comprising a system or a vector encoding a system for use in the manufacture of a pharmaceutical product for the treatment of a disease in a subject requiring it.
[0009] Further features and advantages of certain embodiments of this disclosure will become more readily apparent in the following description of the embodiments and their drawings, as well as from the claims.
[0010] Inclusion by reference All publications, patents, and patent applications referenced herein are incorporated by reference to the same extent as each individual publication, patent, or patent application would be specifically and individually indicated as being incorporated by reference. The contents of WO2020 / 247882, WO2020 / 247883, WO2021 / 113772, WO2022 / 120095, WO2022 / 125843, WO2022 / 261150, WO2022 / 261149, WO2023 / 049872, WO2023 / 235818, and WO2023 / 049742, WO2023 / 240162, which disclose CasX variants and gRNA variants, as well as methods for delivering them, are incorporated by reference in their entirety herein. [Brief explanation of the drawing]
[0011] Novel features of this disclosure are described in detail in the appended claims. A better understanding of the features and merits of this disclosure will be obtained from the following detailed description illustrating exemplary embodiments in which the principles of this disclosure are utilized, and from reference to the appended drawings.
[0012] [Figure 1A] Figure 1A is a time-course plot showing the percentage of mouse Hepa1-6 cells negatively stained for intracellular PCSK9 at 4, 12, 18, 24, 41, and 53 days post-delivery, as described in Example 1. Hepa1-6 cells were treated with LTRP5-ZIM3 or LTRP5-ADD-ZIM3 mRNA paired with PCSK9-targeting gRNA accompanied by spacer 27.88. Untargeted (NT) spacers were used as experimental controls.
[0013] [Figure 1B]Figure 1B is a time-course plot showing the percentage of mouse Hepa1-6 cells negatively stained for intracellular PCSK9 at 4, 12, 18, 24, 41, and 53 days post-delivery, as described in Example 1. Hepa1-6 cells were treated with LTRP5-ZIM3 or LTRP5-ADD-ZIM3 mRNA paired with PCSK9-targeting gRNA accompanied by spacer 27.94. Untargeted (NT) spacers were used as experimental controls.
[0014] [Figure 2A] Figure 2A is a bar graph showing the quantification of standardized secreted PCSK9 levels on day 4 after transfection in HepG2 cells lipofected with mRNA encoding CasX 676, dCasX fused to ZIM3-KRAB(dXR1), or LTRP5-ADD-ZIM3, paired with the indicated targeted gRNAs, as described in Example 2. Secreted PCSK9 levels were standardized relative to the total cell number. Naive, untreated cells served as experimental controls.
[0015] [Figure 2B] Figure 2B is a bar graph showing the quantification of standardized secreted PCSK9 levels on day 4 after transfection in Huh7 cells lipofected with mRNA encoding CasX 676, dXR1, or LTRP5-ADD-ZIM3, paired with the indicated targeted gRNA, as described in Example 2. Secreted PCSK9 levels were standardized relative to the total cell number. Naive, untreated cells served as experimental controls.
[0016] [Figure 2C]Figure 2C is a bar graph showing the quantification of secreted PCSK9 levels 4 days after transfection in Hep3B cells lipofected with mRNA encoding CasX 676, dXR1, or LTRP5-ADD-ZIM3 when paired with the indicated targeted gRNA, as described in Example 2. Secreted PCSK9 levels were normalized to the total cell number. Naïve, untreated cells served as the experimental control.
[0017] [Figure 3] Figure 3 is a bar graph showing the quantification of secreted PCSK9 levels at 4, 14, and 27 days after transfection in Huh7 cells lipofected with mRNA encoding CasX 676, dXR1, or LTRP5-ADD-ZIM3 when paired with the indicated targeted gRNA, as described in Example 2. Quantification of secreted PCSK9 levels is shown relative to the secreted levels detected in naïve, untreated cells at the 4-day time point.
[0018] [Figure 4] Figure 4 illustrates a schematic of the LTRP5 molecule without the ADD domain of DNMT3A, as described in Example 5. "D3A CD" and "D3L ID" denote the catalytic domain of DNMT3A and the interaction domain of DNMT3L, respectively. "L1", "L2", "L3A", and "L3B" are linkers. "NLS" is a nuclear localization signal. "RD1" denotes a repressor domain.
[0019] [Figure 5]Figure 5 is a bar graph illustrating the results of a time-course experiment comparing the levels of B2M suppression (expressed as the average percentage of HLA-negative cells) in HEK293T cells transfected with LTRP5 variant plasmids containing linker sets 1-11, as described in Example 5. Data for each time point (day 8, day 15, and day 45) are overlaid and presented as mean values with standard deviation, N = 3. A non-targeting (NT) spacer was included as an experimental control.
[0020] [Figure 6] Figure 6 is a bar graph illustrating the results of a time-course experiment comparing the levels of target 1 suppression (expressed as the percentage of total cells with knockdown of target 1) of LTRP5 variants containing linker sets 1-11, measured in HEK293T cells, as described in Example 5. Data for each time point (day 8, day 15, and day 45) are overlaid and presented as mean values with standard deviation.
[0021] [Figure 7] Figure 7 is a bar graph illustrating the results of a time-course experiment comparing the levels of target 2 suppression (expressed as the percentage of total cells with knockdown of target 2) by LTRP5 variants containing linker sets 1-11. The suppression levels were measured in HEK293T cells as described in Example 5. Data for each time point (day 8, day 15, and day 45) are overlaid and presented as mean values with standard deviation, N = 3. A non-targeting (NT) spacer was included as an experimental control.
[0022] [Figure 8]Figure 8 is a bar graph illustrating the results of a time-course experiment comparing the level of B2M suppression (expressed as the mean percentage of HLA-negative cells) in HEK293T cells transfected with LTRP5 variant plasmids containing linker sets 12-28, as described in Example 5. The data for each time point (day 7 and day 17) are superimposed and presented as the mean with standard deviation, N=3. A non-targeted (NT) spacer was included as an experimental control.
[0023] [Figure 9] Figure 9 is a bar graph illustrating the results of a time-course experiment comparing the level of suppression of target 1 (expressed as the percentage of total cells with target 1 knockdown) by LTRP5 variants containing linker sets 12-28, as measured in HEK293T cells, as described in Example 5. The data for each time point (day 7 and day 17) are superimposed and presented as the mean with standard deviation, N=3. Non-targeting (NT) spacers were included as experimental controls.
[0024] [Figure 10] Figure 10 is a bar graph illustrating the results of a time-course experiment comparing the level of suppression of target 2 (expressed as the percentage of total cells with target 2 knockdown) by LTRP5 variants containing linker sets 12-28, as measured in HEK293T cells, as described in Example 5. The data for each time point (days 7 and 17) are superimposed and presented as the mean with standard deviation, N=3. Non-targeting (NT) spacers were included as experimental controls.
[0025] [Figure 11]Figure 11 is a bar graph showing the percentage of mouse Hepa1-6 cells treated with either dXR1 or LTRP1-ZIM3 mRNA paired with the indicated PCSK9-targeting gRNA, which were negatively stained for intracellular PCSK9 on day 6, as described in Example 7. Spacer 6.7, which targets the human PCSK9 locus, served as a non-targeting control.
[0026] [Figure 12] Figure 12 is a time-course plot showing the percentage of mouse Hepa1-6 cells treated with dXR1 mRNA paired with the indicated PCSK9-targeting gRNA, which were negatively stained for intracellular PCSK9 at days 6, 13, and 25 after delivery, as described in Example 7. Spacer 6.7, which targets the human PCSK9 locus, served as a non-targeting control, and treatment with water served as a negative control.
[0027] [Figure 13] Figure 13 is a time-course plot showing the percentage of mouse Hepa1-6 cells treated with LTRP1-ZIM3 mRNA paired with the indicated PCSK9-targeting gRNA, which were negatively stained for intracellular PCSK9 at days 6, 13, and 25 after delivery, as described in Example 7. Spacer 6.7, which targets the human PCSK9 locus, served as a non-targeting control, and treatment with water served as a negative control.
[0028] [Figure 14] Figure 14 is a time-course plot showing the percentage of mouse Hepa1-6 cells treated with in-house in vitro transcription (IVT) produced LTRP1-ZIM3 vs. LTRP5-ZIM3 mRNA paired with the indicated PCSK9-targeting gRNA, which was negatively stained for intracellular PCSK9 at the indicated time points after delivery, as described in Example 7.
[0029] [Figure 15] Figure 15 is a time-course plot showing the percentage of mouse Hepa1-6 cells treated with LTRP1-ZIM3 paired with the indicated PCSK9-targeted gRNA paired with dCas9-ZNF10-DNMT3A / 3L mRNA, which was negatively stained for intracellular PCSK9 at the indicated time points after delivery, as described in Example 7.
[0030] [Figure 16] Figure 16 is a plot illustrating the percentage of HEK293T cells transfected with plasmids encoding the indicated CasX or LTRP:gRNA constructs, expressing B2M six days after treatment with the DNMT1 inhibitor 5-azadC at varying concentrations, as described in Example 7.
[0031] [Figure 17] Figure 17 is a plot comparing the quantification of B2M suppression in HEK293T cells transfected with plasmids encoding the indicated CasX or LTRP:gRNA constructs and cultured for 58 days, as described in Example 7, with the quantification of B2M reactivation upon treatment of transfected cells with 5-azadC.
[0032] [Figure 18] Figure 18 is a bar graph showing the quantification of secreted PCSK9 levels at 6, 18, and 36 days post-transfection in Huh7 cells lipofected with mRNA encoding CasX 676, dXR1, or LTRP5-ADD-ZIM3, paired with the indicated targeted gRNA, as described in Example 8. Secreted PCSK9 levels were normalized against total cell number. Naive, untreated cells served as experimental controls.
[0033] [Figure 19]Figure 19 illustrates schematic diagrams of various configurations of the LTRP molecule with the DNMT3A ADD domain. "D3A ADD," "D3A CD," and "D3L ID" represent the ADD domain of DNMT3A, the catalytic domain of DNMT3A, and the interaction domain of DNMT3L, respectively. "DBP" represents the DNA-binding protein. "L1," "L2," "L3A," "L3B," and "L4" are linkers. "NLS" is the nuclear localization signal. "RD1" represents the repressor domain, and "RD1a" and "RD1b" represent repressor domain variants.
[0034] [Figure 20A] Figure 20A is a schematic diagram illustrating versions 1–3 of chemical modifications made to gRNA scaffold variant 235, as described in Example 11. Structural motifs are highlighted. Standard ribonucleotides are shown as white circles, and 2'OMe-modified ribonucleotides are shown as black circles. Phosphothioate bonds are indicated with an asterisk (*) below or beside the bond. For the v2 profile, the addition of three 3'uracil (3'UUU) is annotated with a "U" in the corresponding circle.
[0035] [Figure 20B] Figure 20B is a schematic diagram illustrating versions 4–6 of chemical modifications prepared for gRNA scaffold variant 235, as described in Example 11. Structural motifs are highlighted. Standard ribonucleotides are shown as white circles, and 2'OMe-modified ribonucleotides are shown as black circles. Phosphothioate bonds are indicated with an asterisk (*) below or beside the bond.
[0036] [Figure 21]Figure 21 is a plot illustrating the quantification of B2M percent knockout in HepG2 cells co-transfected with 100 ng of CasX 491 mRNA and the indicated doses of terminally modified (v1) or unmodified (v0) B2M-targeted gRNA with spacer 7.37, as described in Example 11. The editing level was determined by flow cytometry as a population of cells with loss of HLA complex surface presentation due to successful editing at the B2M locus.
[0037] [Figure 22] Figure 22 is a schematic diagram illustrating versions 7–9 of chemical modifications prepared for gRNA scaffold variant 316, as described in Example 11. Structural motifs are highlighted. Standard ribonucleotides are shown as white circles, and 2'OMe-modified ribonucleotides are shown as black circles. Phosphothioate bonds are indicated with an asterisk (*) below or beside the bond.
[0038] [Figure 23A] Figure 23A is a schematic diagram of gRNA scaffold variant 174 (SEQ ID NO: 1744), as described in Example 11. Structural motifs are highlighted.
[0039] [Figure 23B] Figure 23B is a schematic diagram of gRNA scaffold variant 235 (SEQ ID NO: 1745) as described in Example 11. The highlighted structural motifs are the same as those in Figure 20A. The differences between gRNA variant 174 and variant 235 lie in the elongation stem motif and several single nucleotide changes (indicated by asterisks). Variant 316 retains the shorter elongation stem from variant 174, but retains the four substitutions found in scaffold 235.
[0040] [Figure 23C]Figure 23C is a schematic diagram of gRNA scaffold variant 316 (SEQ ID NO: 1746) as described in Example 11. The highlighted structural motif is the same as that in Figure 20A. Variant 316 retains the shorter elongated stem from gRNA variant 174 (Figure 23A), but retains the four substitutions found in scaffold 235 (Figure 23B).
[0041] [Figure 24] Figure 24 is a plot showing the correlation between the indel rate at the PCSK9 locus (depicted as edit fraction rate) (x axis), measured by NGS, and the secreted PCSK9 level (ng / mL) (y axis), detected by ELISA, in HepG2 cells lipofected with CasX 491 mRNA and PCSK9-targeted gRNA containing the indicated scaffold variant and spacer combination, as described in Example 11.
[0042] [Figure 25A] Figure 25A is a plot illustrating the results of an editing assay measured as the indel rate detected by NGS at the human B2M locus in HepG2 cells treated with the indicated doses of LNP formulated with CasX 491 mRNA and the indicated B2M-targeted gRNA, as described in Example 11.
[0043] [Figure 25B] Figure 25B is a plot illustrating the quantification of B2M knockout percentage in HepG2 cells treated with the indicated doses of LNP formulated with CasX 491 mRNA and the indicated B2M-targeted gRNA, as described in Example 11. The editing level was determined by flow cytometry as the population of cells that lacked surface presentation of the HLA complex due to successful editing at the B2M locus.
[0044] [Figure 26A]Figure 26A is a plot depicting the results of an editing assay measured as the indel rate detected by NGS at the mouse ROSA26 locus in Hepa1-6 cells treated with the indicated dose of LNP formulated with CasX 676 mRNA #2 and the indicated ROSA26-targeted gRNA with either a v1 or v5 modification profile, as described in Example 11.
[0045] [Figure 26B] Figure 26B is a plot illustrating the quantification of edit percentages, measured as indel rates detected by NGS at the ROSA26 locus, in mice treated with LNPs formulated with CasX 676 mRNA #2 and the shown chemically modified ROSA26-targeted gRNA, as described in Example 11.
[0046] [Figure 27] Figure 27 is a bar graph showing the results of an editing assay measured as the indel rate detected by NGS as the mouse PCSK9 locus in mice treated with LNP formulated with CasX 676 mRNA #1 and the indicated chemically modified PCSK9-targeted gRNA, as described in Example 11. Untreated mice served as experimental controls.
[0047] [Figure 28A] Figure 28A is a schematic diagram illustrating versions 1–3 of chemical modifications made to gRNA scaffold variant 316, as described in Example 11. Structural motifs are highlighted. Standard ribonucleotides are shown as white circles, and 2'OMe-modified ribonucleotides are shown as black circles. Phosphothioate bonds are indicated with an asterisk (*) below or beside the bond. For the v2 profile, the addition of three 3'uracils (3'UUU) is annotated with a "U" in the corresponding circle.
[0048] [Figure 28B] Figure 28B is a schematic diagram illustrating versions 4–6 of chemical modifications prepared for gRNA scaffold variant 316, as described in Example 11. Structural motifs are highlighted. Standard ribonucleotides are shown as white circles, and 2'OMe-modified ribonucleotides are shown as black circles. Phosphothioate bonds are indicated with an asterisk (*) below or beside the bond.
[0049] [Figure 29] Figure 29 is a violin plot with points representing the mean methylation percentage at individual CpG motifs. The median methylation is shown by a dashed line, with the superior and inferior quartiles shown by dotted lines. Proximal DNA methylation at the transcription start site (TSS) was measured from homogenized liver-extracted gDNA by amplicon enzyme methylation sequencing (EM-seq) from N=3 mice sacrificed at 7, 14, and 42 days post-treatment, as described in Example 12.
[0050] [Figure 30] Figure 30 is a violin plot with points representing the mean methylation percentage at individual CpGs. The median methylation is shown by a dashed line, accompanied by the superior and inferior quartiles shown by dotted lines. Proximal DNA methylation at the transcription start site (TSS) was measured from homogenized liver-extracted gDNA by amplicon enzyme methylation sequencing (EM-seq) from N=3 mice sacrificed 7 days after treatment, as described in Example 13. [Modes for carrying out the invention]
[0051] While exemplary embodiments are shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided only as examples. Numerous variations, alterations, and substitutions will be conjured upon hereby to those skilled in the art without departing from the claimed invention. It should be understood that various alternatives to the embodiments described herein may be employed when practicing the embodiments of this disclosure. The claims define the scope of the invention, and methods and structures within the scope of these claims and their equivalents are intended to be covered thereby.
[0052] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those generally understood by those skilled in the art in the field to which this invention pertains. Methods and materials similar to or equivalent to those described herein may be used in the practice or testing of this embodiment, but suitable methods and materials are described below. In case of any inconsistency, the patent specification containing the definitions shall prevail. Furthermore, the materials, methods, and examples are illustrative only and not intended to be limiting. Numerous variations, changes, and substitutions will be recalled herein by those skilled in the art without departing from the present invention.
[0053] definition As used herein and in the appended claims, the singular forms “a,” “an,” and “the” include plural referents unless the context clearly indicates otherwise. Thus, for example, a reference to “host cell” includes two or more such host cells; a reference to “manipulated CasX protein” includes one or more manipulated CasX proteins; a reference to “nucleic acid sequence” includes one or more nucleic acid sequences and the like.
[0054] As used herein, the term “about” is to be understood by those skilled in the art and may vary to some extent depending on the context in which it is used. In cases where the use of the term “about” is not obvious to those skilled in the art, considering the context in which it is used, “about” means a range of plus or minus 10% of the given term.
[0055] As will be understood by those skilled in the art, for any and all purposes, the entire scope disclosed herein also includes any and all possible sub-scopes and combinations thereof. Furthermore, as will be understood by those skilled in the art, the scope includes each individual component. Thus, for example, a base having 1 to 3 members refers to a base having 1, 2, or 3 members. Similarly, a base having 1 to 5 members refers to a base having 1, 2, 3, 4, or 5 members, and so on.
[0056] The term "those combinations" includes all possible combinations of the elements that the term refers to.
[0057] As used herein, the term “exemplary” refers to an example or illustration and is not intended to imply any priority or value.
[0058] The terms “polynucleotide” and “nucleic acid” are used interchangeably herein and refer to macromolecular forms of nucleotides of any length, which are either ribonucleotides or deoxyribonucleotides. Thus, the terms “polynucleotide” and “nucleic acid” encompass macromolecules including single-stranded DNA; double-stranded DNA; multi-stranded DNA; single-stranded RNA; double-stranded RNA; multi-stranded RNA; genomic DNA; cDNA; DNA-RNA hybrids; and purine and pyrimidine bases, or other natural, chemically or biochemically modified, unnatural, or derivatized nucleotide bases.
[0059] The terms "hybridable" or "complementary" are used interchangeably and mean that a nucleic acid (e.g., RNA, DNA) contains a sequence of nucleotides that allows it to non-covalently bind to another nucleic acid in a sequence-specific and antiparallel manner (i.e., the nucleic acid specifically binds to its complementary nucleic acid), i.e., to form Watson-Crick base pairs and / or G / U base pairs, i.e., to "anneal," or "hybridize," under suitable in vitro and / or in vivo conditions of temperature and solution ionic strength. It is understood that a polynucleotide sequence does not need to be 100% complementary to its target nucleic acid sequence in order to be specifically hybridizable, but may have at least about 70%, at least about 80%, at least about 90%, or at least about 95% sequence identity and still be able to hybridize to the target nucleic acid. Furthermore, a polynucleotide may hybridize across one or more segments such that intervening or adjacent segments are not included in the hybridization event (e.g., loop or hairpin structures, "bulges," "bubbles," and similar). Therefore, those skilled in the art will understand that while individual bases within a sequence may not be complementary to another sequence, the sequence as a whole is still considered complementary.
[0060] For the purposes of this disclosure, “gene” includes the DNA region encoding a gene product (e.g., protein, RNA), as well as all DNA regions that regulate the production of the gene product, regardless of whether such regulatory sequences are adjacent to the coding sequence and / or transcription sequence. Therefore, a gene may include, but is not limited to, accessory element sequences, including promoter sequences, terminators, translation regulatory sequences, such as ribosome binding sites and internal ribosome entry sites, enhancers, silencers, insulators, boundary elements, origins of replication, matrix attachment sites, and locus regulatory regions. The coding sequence codes for the gene product during transcription or transcription and translation; the coding sequences of this disclosure may include fragments but do not necessarily include a full-length open reading frame. A gene may include both the transcribed strand and a complementary strand containing an anticodon.
[0061] The term "downstream" refers to a nucleotide sequence located at the 3' position of a reference nucleotide sequence. In certain embodiments, the downstream nucleotide sequence relates to the sequence following the transcription start site. For example, the translation start codon of a gene is located downstream of the transcription start site.
[0062] The term "upstream" refers to a nucleotide sequence located 5' of the reference nucleotide sequence. In certain embodiments, the upstream nucleotide sequence refers to a sequence located 5' of the coding region or transcription start site. For example, most promoters are located upstream of the transcription start site.
[0063] The term “adjacent to” in relation to polynucleotide or amino acid sequences refers to sequences that are adjacent to or adjacent to each other in a polynucleotide or polypeptide. Those skilled in the art will understand that two sequences can be considered adjacent to each other and still contain a limited number of intervening sequences, e.g., one, two, three, four, five, six, seven, eight, nine, or ten nucleotides or amino acids.
[0064] The term “regulatory element” is used herein interchangeably with the term “regulatory sequence” and is intended to include promoters, enhancers, and other expression regulatory elements. It will be understood that the selection of a suitable regulatory element depends on whether the coding component (e.g., protein or RNA) or nucleic acid being expressed contains multiple components that require different polymerases or are not intended to be expressed as a fusion protein.
[0065] The term “accessory element” is used herein interchangeably with the term “accessory sequence” and includes coding sequences and non-coding sequences that enhance nucleic acid expression, transport, or the function of mRNA or protein, and is intended to include, among other things, polyadenylation signals (poly(A) signals), enhancer elements, introns, post-transcriptional regulatory elements (PTREs), nuclear localization signals (NLSs), deaminases, DNA glycosylase inhibitors, additional promoters, factors that stimulate CRISPR-mediated homology-directed repair (e.g., in cis or trans), self-cleavage sequences, and fusion domains, e.g., fusion domains fused to CRISPR proteins. It will be understood that the selection of a suitable accessory element or element depends on whether the coding component to be expressed (e.g., protein or RNA), or nucleic acid, contains multiple components that require different polymerases or are not intended to be expressed as a fusion protein.
[0066] The term “promoter” refers to a DNA sequence containing a transcription initiation site and additional sequences to facilitate polymerase binding and transcription. Exemplary eukaryotic promoters include elements, such as a TATA box and / or a B-recognition element (BRE), that assist or promote the transcription and expression of associated transcribed polynucleotide sequences and / or genes (or transgenes). Promoters can be produced synthetically, or they can be known, naturally occurring, or derived from other promoter sequences. Promoters can be located proximal or distal to the gene being transcribed. Promoters can also include chimeric promoters, which include a combination of two or more heterologous sequences to confer a particular characteristic. Promoters in this disclosure may include variants of promoter sequences that are similar in composition to, but not identical to, other promoter sequences known or provided herein. Promoters can be classified according to criteria relating to the pattern of expression of associated codes or transcribed sequences or genes operably linked to the promoter, for example, constitutively, developmentally, tissue-specifically, or inducibly. Promoters can also be classified according to their strength. When used in relation to promoters, "strength" refers to the transcription rate of a gene controlled by the promoter. A "strong" promoter means a high transcription rate, while a "weak" promoter means a relatively low transcription rate.
[0067] The promoters of this disclosure may be polymerase II (Pol II) promoters. Polymerase II transcribes all protein-coding genes and many non-coding genes. Typical Pol II promoters include a core promoter, which is a sequence of approximately 100 base pairs surrounding the transcription start site and serves as a binding platform for Pol II polymerase and associated basal transcription factors. The promoter may include one or more core promoter elements, such as a TATA box, BRE, initiator (INR), motif 10 element (MTE), downstream core promoter element (DPE), downstream core element (DCE), etc., although core promoters lacking these elements are known in the art. All Pol III promoters are assumed to be within the scope of this disclosure.
[0068] The promoters of this disclosure may be polymerase III (Pol III) promoters. Pol III transcribes DNA to synthesize small ribosomal RNAs, such as 5S rRNA, tRNA, and other small RNAs. Typical Pol III promoters use internal regulatory sequences (sequences within the transcription region of a gene) to support transcription, but upstream elements, such as TATA boxes, are also sometimes used. All Pol III promoters are assumed to be within the scope of this disclosure.
[0069] The term "enhancer" refers to a regulatory DNA sequence that, when bound by a specific protein called a transcription factor, modulates the expression of the associated gene. Enhancers can be located in the intron of a gene, or at the 5' or 3' of the gene's coding sequence. Enhancers can be located proximal to the gene (i.e., within tens or hundreds of base pairs (bp) of the promoter) or distal to the gene (i.e., thousands, tens of thousands, or even millions of bp away from the promoter). A single gene can be regulated by more than one enhancer, all of which are assumed to be within the scope of this disclosure.
[0070] As used herein, “post-transcriptional regulatory elements (PTREs),” such as hepatitis PTREs, refer to DNA sequences that, when transcribed, form tertiary structures capable of exhibiting post-transcriptional activity to enhance or promote the expression of related genes operably linked to them.
[0071] "Operatively coupled" refers to the juxtaposition of two or more components (e.g., sequence elements) in which the components are positioned, both components function normally, and at least one component may mediate the function exerted by at least one of the other components, such as a promoter and a code sequence. Those skilled in the art will understand that the two components do not need to be physically coupled to be operationally coupled.
[0072] In the context of this disclosure and with respect to genes, the terms “repress,” “repression,” “transcriptional repression,” “repression,” “inhibition of gene expression,” “downregulation,” and “silencing” are used interchangeably herein to refer to the inhibition or blockage of transcription of a gene or a portion thereof. Thus, transcriptional repression can result in a reduction in the production of a gene product. Examples of gene repression processes that reduce transcription include, but are not limited to, inhibiting the formation of a transcription initiation complex, reducing the rate of transcription initiation, reducing the rate of transcription elongation, reducing the forward momentum of transcription, and antagonizing transcriptional activation (e.g., by blocking the binding of transcription activators). Gene repression can constitute, for example, the prevention of activation and the inhibition of expression below existing levels. Transcriptional repression includes both reversible and irreversible inactivation of gene transcription, the latter of which may result from epigenetic modifications of a gene.
[0073] The terms “repressor” or “repressor domain” are interchangeable and refer to polypeptide factors that act as regulatory elements on DNA to inhibit, repress, or block DNA transcription, resulting in repression of gene expression. In the context of this disclosure, this refers to the ligation of a repressor domain to a DNA-binding protein that, when bound to a target nucleic acid, can prevent transcription from a promoter or otherwise inhibit gene expression. Without being constrained by theory, transcriptional repressors are thought to function by various mechanisms, including physically blocking RNA polymerase passage through steric hindrance, altering the post-translational modification state of polymerase, modifying the epigenetic state of nascent RNA, altering the epigenetic state of DNA through methylation, altering the epigenetic state of DNA through histone deacetylation, or regulating nucleosome remodeling, or preventing enhancer-promoter interactions, thereby leading to gene silencing or a reduction in the level of gene expression.
[0074] "Long-term repressor fusion protein" or "LTRP" is used herein interchangeably with "repressor fusion protein" and refers to a fusion protein comprising a DNA-binding protein (or DNA-binding domain of a protein) fused to one or more domains capable of repressing the transcription of a target nucleic acid sequence. Optionally, the repressor fusion protein of this disclosure may include additional elements, such as linkers between any of the domains of the fusion protein, nuclear localization signals, nuclear export signals, and additional protein domains that confer additional activity to the repressor fusion protein.
[0075] As used herein, “LTRP:gRNA system” is a system for transcriptional repression comprising a non-catalyzed CRISPR protein and a long-term repressor fusion protein comprising one or more linked repressor domains, as well as a guide nucleic acid (gRNA) that binds to the non-catalyzed CRISPR protein. For clarity, the system also comprises any coding DNA, RNA, or vector and similar that can be used to produce the system’s repressor fusion protein and gRNA components.
[0076] As used herein, a DNA-binding protein refers to a protein or protein domain capable of binding to DNA. Exemplary DNA-binding proteins include zinc finger (ZF) proteins, activator-like effectors (TALEs), and clustered regularly spaced short palindromic repeat (CRISPR) proteins. Those skilled in the art will understand that in multifunctional proteins capable of both binding to DNA and performing other activities, such as DNA cleavage, such as CRISPR proteins, the DNA-binding function can be separated from other functions of the protein to result in a DNA-binding protein without catalytic activity.
[0077] As used herein, “non-catalytic DNA-binding protein” refers to a protein that can bind to DNA but cannot cleave or break it. As used herein, “non-catalytic CRISPR protein” refers to a CRISPR protein that lacks endonuclease activity. Those skilled in the art will understand that a CRISPR protein may lack catalytic activity but can still perform additional protein functions, such as DNA binding. Similarly, “non-catalytic CasX” refers to a CasX protein that lacks endonuclease activity but can still perform additional protein functions, such as DNA binding.
[0078] "Recombinant," as used herein, means that a particular nucleic acid (DNA or RNA) is the product of various combinations of cloning, restriction, and / or ligation steps, resulting in a construct having a structurally coding or non-coding sequence that is distinguishable from endogenous nucleic acids found in the natural system. Generally, DNA sequences encoding structurally coding sequences can be assembled from cDNA fragments and short oligonucleotide linkers, or from a series of synthetic oligonucleotides, to provide synthetic nucleic acids that can be expressed from recombinant transcription units contained in cells or cell-free transcription and translation systems. Such sequences can typically be provided in the form of an open reading frame uninterrupted by internal non-coding sequences or introns present in eukaryotic genes. Genomic DNA containing the relevant sequences can also be used to form recombinant genes or transcription units. The non-coding DNA sequence may be located at 5' or 3' of the open reading frame, and such sequences may not interfere with the manipulation or expression of the coding region, but rather may act to regulate the production of desired products by various mechanisms (see "enhancers" and "promoters" above).
[0079] The term “recombinant polynucleotide” or “recombinant nucleic acid” refers to a substance that does not occur naturally, but is created through artificial combinations of two originally separate sequence segments, for example, through human intervention. This artificial combination is often achieved either by chemical synthesis or by artificial manipulation of isolated nucleic acid segments, for example, by genetic engineering techniques. Such combinations are typically performed to replace codons with redundant codons encoding the same or conserved amino acids, usually while introducing or removing sequence recognition sites. Alternatively, this is carried out to join nucleic acid segments of desired function together to produce a desired combination of function. This artificial combination is often achieved either by chemical synthesis or by artificial manipulation of isolated nucleic acid segments, for example, by genetic engineering techniques.
[0080] Similarly, the terms “recombinant polypeptide” or “recombinant protein” refer to polypeptides or proteins that do not occur naturally, but are created by the artificial combination of two originally separate amino acid sequence segments, for example, through human intervention. Thus, for example, a protein containing heterologous amino acid sequences is recombinant.
[0081] As used herein, “lipid nanoparticles” or “LNPs” refer to particles having at least one dimension on the order of nanometers (e.g., 1 to 1,000 nm) and containing one or more lipids (e.g., cationic lipids, non-cationic lipids, helper phospholipids, and PEG-modified lipids). Specific components of LNPs are described more fully below. Lipid nanoparticles may be included in formulations that can be used to deliver active or therapeutic agents, such as nucleic acids (e.g., mRNA), to a target site of interest (e.g., cells, tissues, organs, tumors, and similar). Lipid nanoparticles of this disclosure may include nucleic acids. Such lipid nanoparticles typically include neutral lipids, charged lipids, steroids, and high molecular weight conjugate lipids. Active or therapeutic agents, such as nucleic acids, may be encapsulated in an aqueous space surrounded by the lipid portion of the lipid nanoparticle, or some or all of the lipid portion of the lipid nanoparticle, thereby protecting them from enzymatic degradation or other undesirable effects induced by host organism or cellular mechanisms, such as harmful immune responses.
[0082] As used herein, the term “to bring into contact” means to establish a physical connection between two or more entities. For example, bringing a target nucleic acid into contact with a guide nucleic acid means that the target nucleic acid and the guide nucleic acid share a physical connection, for example, they can hybridize if their sequences share sequence similarity.
[0083] "Dissociation constant" or "K" dThe term "L" is used interchangeably and refers to the affinity between ligand "L" and protein "P," i.e., how tightly the ligand binds to a particular protein. It is represented by formula K. d It can be calculated using the formula =[L][P] / [LP], where [P], [L], and [LP] represent the molar concentrations of the protein, ligand, and complex, respectively.
[0084] A polynucleotide or polypeptide has a certain percentage of "sequence similarity" or "sequence identity" with another polynucleotide or polypeptide, meaning that when aligned, the percentage of bases or amino acids are the same and that they are in the same relative positions when comparing the two sequences. Sequence similarity (sometimes referred to as similarity percentage, identity percentage, or homology) can be determined in a number of different ways. To determine sequence similarity, sequences can be aligned using methods and computer programs known in the art, including BLAST, which is available on the World Wide Web at ncbi.nlm.nih.gov / BLAST. The complementarity rate between specific stretches of nucleic acid sequences within a nucleic acid can be determined using any convenient method. Exemplary methods include using the BLAST program (a basic local sorting search tool) and the PowerBLAST program (Altschul et al., J.Mol.Biol., 1990, 215, 403~410; Zhang and Madden, Genome Res., 1997, 7, 649~656), or the Gap program (Wisconsin Sequence Analysis Package, Version 8 for Unix, Genetics Computer Group, University Research Park, Madison Wis.), for example, by using the default settings that employ the Smith and Waterman algorithm (Adv.Appl.Math., 1981, 2, 482~489).
[0085] The terms “polypeptide” and “protein” are used interchangeably herein and refer to macromolecular forms of amino acids of any length, which may include coding and non-coding amino acids, chemically or biochemically modified or derivatized amino acids, and polypeptides having a modified peptide backbone. The term includes, but is not limited to, fusion proteins, including fusion proteins with heterologous amino acid sequences.
[0086] A "vector" or "expression vector" is a replicon, such as a plasmid, phage, virus, or cosmid, which may contain another DNA segment, i.e., an expression cassette, and can result in the replication or expression of another DNA segment in a cell.
[0087] The terms “naturally occurring,” “unmodified,” or “wild-type,” as used herein, applied to nucleic acids, polypeptides, cells, or organisms, refer to nucleic acids, polypeptides, cells, or organisms found in nature.
[0088] As used herein, “mutation” means an insertion, deletion, substitution, duplication, or inversion of one or more amino acids or nucleotides compared to the wild-type or reference amino acid sequence, or the wild-type or reference nucleotide sequence.
[0089] As used herein, the term “isolated” is used to describe polynucleotides, polypeptides, or cells that are in an environment different from the environment in which they naturally occur. Isolated recombinant host cells may be present in a mixed population of recombinant host cells.
[0090] "Host cell" as used herein refers to a cell from a multicellular organism (e.g., a cell line) cultured as a single-cell entity, whether eukaryotic or prokaryotic, which is used as a recipient for nucleic acid (e.g., an AAV vector) and includes offspring of the original cell that have been genetically modified by the nucleic acid. Single-cell offspring may not necessarily be completely identical to the original parent in morphology or in genomic or total DNA complement due to spontaneous, accidental, or intentional mutations. "Recombinant host cell" (also referred to as "genetically modified host cell") is a host cell into which a heterologous nucleic acid, such as an AAV vector, has been introduced.
[0091] The term "conservative amino acid substitution" refers to the interchangeability of amino acid residues with similar side chains in proteins. For example, amino acids with aliphatic side chains include glycine, alanine, valine, leucine, and isoleucine; amino acids with aliphatic-hydroxyl side chains include serine and threonine; amino acids with amide-containing side chains include asparagine and glutamine; amino acids with aromatic side chains include phenylalanine, tyrosine, and tryptophan; amino acids with basic side chains include lysine, arginine, and histidine; and amino acids with sulfur-containing side chains include cysteine and methionine. Exemplary conservative amino acid substitutions are valine-leucine-isoleucine, phenylalanine-tyrosine, lysine-arginine, alanine-valine, and asparagine-glutamine.
[0092] As used herein, “treatment” or “treating” are interchangeable and not limited herein, but refer to an approach to obtain a beneficial or desired outcome, including therapeutic and / or preventive benefits. Therapeutic benefits mean the eradication or remission of the underlying disorder or disease being treated. Therapeutic benefits can also be achieved with the eradication or remission of one or more symptoms, or improvement in one or more clinical parameters associated with the underlying disease, such that improvement is observed in the subject, even though the subject may still suffer from the underlying disorder.
[0093] The terms “therapeutic dose” and “therapeutic dosage” mean, as used herein, a certain amount of a drug or bioagent, alone or as part of a composition, that, when administered in a single or repeated dose to a subject, e.g., a human or an experimental animal, is capable of producing any detectable and beneficial effect on any symptom, aspect, measured parameter, or characteristic of a disease state or condition. Such an effect does not need to be absolute to be beneficial.
[0094] As used herein, “administer” means a method of giving a certain dose of a compound (e.g., a composition of this disclosure) or a composition (e.g., a pharmaceutical composition) to a subject.
[0095] The "subjects" are mammals. Mammals include, but are not limited to, domesticated animals, non-human primates, humans, dogs, rabbits, mice, rats, and other rodents.
[0096] All publications, patents, and patent applications referenced herein are incorporated by reference to the same extent as each individual publication, patent, or patent application would be specifically and individually indicated to be incorporated by reference.
[0097] I. General Methods The practice of this invention employs conventional techniques of immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA, unless otherwise indicated, and these are referenced in Molecular Cloning: A Laboratory Manual, 3rd Ed. (Sambrook et al., Harbor Laboratory Press 2001) and Short Protocols in Molecular Biology, 4 th This information can be found in standard textbooks such as Ed. (Ausubel et al. eds., John Wiley & Sons 1999), Protein Methods (Bollag et al., John Wiley & Sons 1996), Nonviral Vectors for Gene Therapy (Wagner et al. eds., Academic Press 1999), Viral Vectors (Kaplift & Loewy eds., Academic Press 1995), Immunology Methods Manual (I. Lefkovits ed., Academic Press 1997), and Cell and Tissue Culture: Laboratory Procedures in Biotechnology (Doyle & Griffiths, John Wiley & Sons 1998), but their disclosures are incorporated herein by reference.
[0098] Where a range of values is provided, it is understood that, unless the context otherwise explicitly determines, each intermediary value up to one-tenth of the lower limit unit is included between the upper and lower limits of that range, and between any other listed values or intermediary values within that listed range. The upper and lower limits of these smaller ranges may independently be included within smaller ranges and may also be included by any specifically excluded limits within the listed range. If a listed range includes one or both limits, it also includes ranges that exclude one or both of those included limits.
[0099] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which the present invention pertains. All publications referenced herein are incorporated herein by reference to disclose and describe methods and / or materials in relation to those cited in the publications.
[0100] For clarity, it will be understood that certain features of the Disclosure described in the context of separate embodiments may also be provided in combination in a single embodiment. In other cases, various features of the Disclosure, for brevity, are described in the context of a single embodiment and may be provided separately or in any suitable secondary combination. All combinations of embodiments relating to the Disclosure are specifically encompassed by the Disclosure and are intended to be disclosed herein in the same way that each and all combinations may be disclosed individually and explicitly. In addition, all secondary combinations of various embodiments and their elements are also specifically encompassed by the Disclosure and are disclosed herein in the same way that each and all such secondary combinations may be disclosed individually and explicitly.
[0101] II. Systems for Epigenetic Modification and Repression of Genes This disclosure provides a particularly configured system having utility in transcriptional repression and / or epigenetic modification of genes in cells. As used herein, “system” is used interchangeably with “composition.” In some cases, the system is designed to repress the transcription of genes in eukaryotic cells having mutations. In other cases, the system is designed to repress or silence the transcription of wild-type genes in eukaryotic cells, which nevertheless contribute to disease or condition. Generally, any portion of a gene can be targeted using the programmable systems and methods of this disclosure, which are described more fully herein.
[0102] This disclosure provides components of a system of long-term repressor fusion proteins comprising or encoding various configurations of DNA-binding proteins and linked repressor domains capable of binding to target nucleic acid sequences of genes targeted for transcriptional repression and / or epigenetic modification. Such fusion proteins are referred herein as long-term repressor fusion proteins (LTRPs) and enable long-term repression or silencing effects on targeted genes. This disclosure also provides nucleic acids encoding the system. Also provided herein are methods for gene repression and / or epigenetic modification, as well as methods for constructing systems including methods for treating diseases or disorders in which gene repression or silencing is desired, and methods for using the system.
[0103] In some embodiments, DNA-binding proteins include zinc finger (ZF) or TALE (transmission activator-like effector) proteins that bind to target nucleic acids but do not cleave them. The DNA-binding domain of TALE consists of a tandem array of customizable monomers 33-34 amino acids (aa) long, which can theoretically be assembled to recognize any gene sequence according to a recognition code in which one repeat binds to one base pair (see Jain, S., et al. TALEN outperforms Cas9 in editing heterochromatin target sites. Nat.Commun. 12:606 (2021)). The specificity of TALE for binding to DNA arises from two polymorphic amino acids, so-called repeat variable duodecimal residues (RVDs) located at positions 12 and 13 of the repeat unit. By rearranging the repeats, the DNA-binding specificity of TALE can be arbitrarily altered. Zinc finger proteins are transcription factors, where each finger recognizes 3-4 bases. By mixing and matching these finger modules, ZF proteins can be customized for the target sequence.
[0104] In some embodiments, the DNA-binding protein comprises a non-catalyzed class 1 or class 2 CRISPR protein. Non-catalyzed CRISPR proteins are also referred to in the art as "catalyzably inactive" CRISPR proteins. In some embodiments, the class 2, type II protein is a non-catalyzed Cas9. In other embodiments, the class 2 CRISPR protein is selected from the group consisting of type II, type V, or type VI proteins. In some embodiments, the class 2, type V protein is selected from the group consisting of Cas12a (Cpf1), Cas12b (C2c1), Cas12c (C2c3), Cas12d (CasY), Cas12e (CasX), Cas12f, Cas12g, Cas12h, Cas12i, Cas12j, Cas12k, Cas14, and / or CasΦ, each in which specific mutations make the protein catalytically inactive as described herein. The CRISPR-based system further includes a guide ribonucleic acid (gRNA) with a targeting sequence complementary to the target sequence of the gene to be bound and repressed by a complex of a fusion protein (CRISPR protein and linked repressor domain) and gRNA.
[0105] In some embodiments, the present invention provides a system comprising, or encoding, a long-term repressor fusion protein comprising a CasX nuclease protein without catalytic activity and a linked repressor domain, and a guide ribonucleic acid (gRNA) comprising a targeting sequence complementary to a target nucleic acid sequence of a gene targeted for transcriptional repression, silencing, or epigenetic modification. In some embodiments, the system comprises a long-term repressor fusion protein and the gRNA of the present disclosure as a gene repressor pair ("LTRP:gRNA system") capable of forming a ribonucleoprotein (RNP) complex and binding to the target nucleic acid. In some embodiments, the target nucleic acid is located in a eukaryotic cell. In other cases, the present disclosure provides a system of a long-term repressor fusion protein and a nucleic acid encoding the gRNA. In yet other cases, the present disclosure provides a system of gRNA and mRNA encoding a long-term repressor fusion protein for use in certain particle formulations (e.g., LNPs) as described herein.
[0106] Furthermore, provided herein are methods using the LTRP:gRNA system, including methods for producing long-term repressor fusion proteins and gRNAs, as well as methods for gene repression and / or epigenetic modification and therapeutic methods. The gRNA components and their characteristics, as well as delivery modalities and methods for using the system for gene transcriptional repression, epigenetic modification, or silencing, are described more fully below.
[0107] III. CRISPR proteins without catalytic activity for use in long-term repressor fusion protein systems In some embodiments, the DNA-binding protein for use in the systems of this disclosure is a non-catalyzed class 1 or class 2 CRISPR protein. In some embodiments, the class 2 CRISPR protein is a class 2, type II protein, such as non-catalyzed Cas9. In other embodiments, the non-catalyzed class 2 CRISPR protein is selected from the group consisting of type II, type V, or type VI proteins. In one embodiment, the class 2 CRISPR type V protein is selected from the group consisting of Cas12a (Cpf1), Cas12b (C2c1), Cas12c (C2c3), Cas12d (CasY), Cas12e (CasX), Cas12f, Cas12g, Cas12h, Cas12i, Cas12j, Cas12k, Cas14, and / or CasΦ, in each case, catalytic inactivation is achieved by specific mutations as described herein. In another embodiment, the class 2 CRISPR type V protein is a CasX (dCasX) protein that does not exhibit catalytic activity.
[0108] The term “CasX protein” as used herein refers to a family of proteins that, in addition to those that render CasX catalytically inactive (dCasX), possess one or more improved characteristics compared to the non-catalyzative reference CasX protein, which is described more fully below, all naturally occurring CasX proteins ("reference CasX") and engineered CasX proteins with multiple sequence modifications. The CasX proteins of this disclosure include the following domains: non-target chain binding (NTSB) domain, target chain loading (TSL) domain, helix I domain, helix II domain, oligonucleotide binding (OBD) domain, and RuvC domain, and in some cases, the domains may be further classified into subdomains, as listed in Table 1.
[0109] In the context of this disclosure, CasX for use in the system is catalytically inactive (dCasX), which is achieved by a mutation introduced at a selected site in the RuvC sequence, as described below.
[0110] a. Reference CasX protein This disclosure provides a naturally occurring CasX protein (referred to herein as the “reference CasX protein”) which was subsequently modified to produce the engineered dCasX of this disclosure. For example, the reference CasX protein can be isolated from naturally occurring prokaryotes, such as Deltaproteobacteria, Plantomycetes, or Candidatus Sungbacteria. The reference CasX protein (referred to herein as the reference CasX polypeptide) is a class 2, type V CRISPR / Cas endonuclease belonging to the CasX (referred to herein as Cas12e) family of proteins that interact with a guide RNA to form a ribonucleoprotein (RNP) complex.
[0111] In some cases, the reference CasX protein is isolated from or derived from Deltaproteobacter and contains the sequence of Sequence ID No. 1.
[0112] In some cases, the reference CasX protein is isolated from or derived from Plantomycetes and contains the sequence of SEQ ID NO: 2.
[0113] In some cases, the reference CasX protein is isolated from or derived from Candidatus Sungbacteria and contains the sequence of Sequence ID No. 3.
[0114] b. Class 1 or Class 2 CRISPR proteins that do not exhibit catalytic activity. In the long-term repressor protein system of this disclosure, non-catalyzed class 1 or class 2 CRISPR proteins are catalytically inactive in that they cannot cleave DNA, but they retain the ability to bind to target nucleic acids when complexed with guide RNA (gRNA). This disclosure provides non-catalyzed variants of class 1 or class 2 CRISPR proteins, wherein the non-catalyzed variants include multiple modifications in the selection domain. In some embodiments, this disclosure provides non-catalyzed CasX variants (hereinafter interchangeably referred to as “dCasX variant” or “dCasX variant protein”), wherein the non-catalyzed CasX variants include multiple modifications in the RuvC domain compared to the sequences of SEQ ID NOs. 1–3 (described above). In some embodiments, the non-catalyzed reference CasX protein includes substitutions at residues 672, 769, and / or 935, with reference to SEQ ID NO. 1. In one embodiment, the non-catalyzed reference CasX protein includes substitutions D672A, E769A, and / or D935A with reference to SEQ ID NO: 1. In other embodiments, the non-catalyzed reference CasX protein includes substitutions at amino acids 659, 756, and / or 922 with reference to SEQ ID NO: 2. In some embodiments, the non-catalyzed reference CasX protein includes substitutions D659A, E756A, and / or D922A with reference to SEQ ID NO: 2. An exemplary RuvC domain of a dCasX variant of this disclosure includes amino acids 661-824 and 935-986 of SEQ ID NO: 1, or amino acids 648-812 and 922-978 of SEQ ID NO: 2, and is accompanied by one or more amino acid modifications compared to the RuvC cleavage domain sequence, wherein the dCasX variant exhibits one or more improved features compared to the reference dCasX. In further embodiments, the CasX variant protein lacking catalytic activity comprises the deletion of all or part of the RuvC domain of the reference CasX protein.The same aforementioned substitutions or deletions may similarly be introduced into CasX variants known in the Art, and it will be understood that this results in dCasX variants (for example, see International Publication No. 2022120095A1 and U.S. Patent Publication No. 11,560,555, incorporated herein by reference, for exemplary sequences).
[0115] In some embodiments, a long-term repressor fusion protein containing a dCasX variant with a linked repressor domain exhibits at least one improved feature compared to a long-term repressor fusion protein containing a reference dCasX protein with an equivalent linked repressor domain. All dCasX variants that improve one or more functions or features of a long-term repressor fusion protein containing a dCasX variant protein compared to an equivalent long-term repressor fusion protein containing a reference dCasX protein are assumed to be within the scope of this disclosure. In some embodiments, the modification is a mutation in one or more amino acids of the reference dCasX, other than one that would render dCasX catalytically inactive. For example, a dCasX variant may include one or more amino acid substitutions, insertions, deletions, or swapped domains, or any combination thereof, compared to the reference dCasX protein sequence. Any amino acid may be substituted for any other amino acid in the substitutions described herein. Substitutions may be conserved substitutions (e.g., a basic amino acid is substituted for another basic amino acid). The substitutions may be non-conservative substitutions (e.g., a basic amino acid being substituted for an acidic amino acid, or vice versa). For example, proline in the reference dCasX protein may be substituted for any of arginine, histidine, lysine, aspartic acid, glutamic acid, serine, threonine, asparagine, glutamine, cysteine, glycine, alanine, isoleucine, leucine, methionine, phenylalanine, tryptophan, tyrosine, or valine to generate the dCasX variant protein of this disclosure.Exemplary improved features of dCasX variant embodiments include, but are not limited to, improved folding of the variant, increased binding affinity to target nucleic acids, improved ability to utilize a broader spectral PAM sequence in transcriptional repression and / or binding to target nucleic acids, improved unwinding of target DNA, increased target strand loading, increased binding of non-target strands of DNA, improved protein stability, increased ability to complex with gRNA, increased binding affinity to gRNA, improved protein:gRNA(RNP) complex stability, and, with linked repressor domains, increased repressor activity when complexed as RNP, improved repressor specificity for target nucleic acids, reduced off-target repression, and an increased percentage of eukaryotic genomes that can be efficiently repressed and / or epigenetically modified. In some embodiments, the improved features of the dCasX variant are improved by at least about 1.1 to about 100,000 times compared to the reference dCasX protein.In some embodiments, the improved features of the dCasX variant are improved by at least approximately 1.1 to 10,000 times compared to the reference dCasX protein, at least approximately 1.1 to 1,000 times, at least approximately 1.1 to 500 times, at least approximately 1.1 to 400 times, at least approximately 1.1 to 300 times, at least approximately 1.1 to 200 times, at least approximately 1.1 to 100 times, at least approximately 1.1 to 50 times, at least approximately 1.1 to 40 times, at least approximately 1.1 to 30 times, at least approximately 1.1 to 20 times, at least approximately 1.1 to 10 times, at least approximately 1.1 to 9 times, and at least The improvements are approximately 1.1 to 8 times, at least 1.1 to 7 times, at least 1.1 to 6 times, at least 1.1 to 5 times, at least 1.1 to 4 times, at least 1.1 to 3 times, at least 1.1 to 2 times, at least 1.1 to 1.5 times, at least 1.5 to 3 times, at least 1.5 to 4 times, at least 1.5 to 5 times, at least 1.5 to 10 times, at least 5 to 10 times, at least 10 to 20 times, at least 10 to 30 times, at least 10 to 50 times, or at least 10 to 100 times. In some embodiments, the improved features of the dCasX variant are at least 10 to 1000 times improved compared to the reference dCasX protein. Further disclosures regarding the improved features are described below in this specification.
[0116] In other embodiments, modifications are substitutions of one or more domains of a reference dCasX using one or more domains from different CasX proteins. In some embodiments, insertions include insertions of some or all of a domain from a different CasX protein. Mutations can be placed in any one or more domains of a dCasX variant and may include, for example, deletions of some or all of one or more domains, or one or more amino acid substitutions, deletions, or insertions in any domain. The domains of a dCasX protein include non-target chain binding (NTSB) domains, target chain loading (TSL) domains, helix I domains, helix II domains, oligonucleotide binding domains (OBDs), and RuvC DNA domains, which may further include subdomains, as described below.
[0117] In some embodiments, the dCasX variant protein contains amino acids between 800 and 1100, or between 900 and 1000.
[0118] The long-term repressor fusion protein comprising dCasX and a ligated repressor domain of this disclosure, when complexed with gRNA as an RNP, has an enhanced ability to utilize and bind to a PAM TC motif comprising a PAM sequence selected from TTC, ATC, GTC, or CTC, and to efficiently bind to the target nucleic acid, compared to an RNP of a reference dCasX protein and an equivalent ligated repressor domain and gRNA fusion protein in an equivalent assay system. As described above, the PAM sequence is positioned at least 1 nucleotide 5' relative to the non-target strand of the protospacer, which is identical to the targeting sequence of the gRNA.
[0119] In some embodiments, the RNP comprising a long-term repressor fusion protein containing a dCasX variant protein and gRNA with a linked repressor domain of the present disclosure is capable of binding to a double-stranded DNA target with an efficiency of at least 70%, at least 80%, at least 85%, at least 90%, or at least 95% at concentrations of 20 pM or less. In one embodiment, the RNP comprising a long-term repressor fusion protein containing a dCasX variant and gRNA variant with a linked repressor domain exhibits greater binding affinity to a target sequence in a target nucleic acid compared to an RNP comprising a reference dCasX protein and gRNA with a linked repressor domain in an equivalent assay system, wherein the PAM sequence of the target nucleic acid is TTC. In another embodiment, an RNP of a long-term repressor fusion protein containing a dCasX variant and a gRNA variant with a ligated repressor domain exhibits greater binding affinity for the target sequence in the target nucleic acid compared to an RNP containing a reference dCasX protein and a reference gRNA with a ligated repressor domain in an equivalent assay system, wherein the PAM sequence of the target nucleic acid is ATC. In another embodiment, an RNP of a long-term repressor fusion protein containing a dCasX variant and a gRNA variant with a ligated repressor domain exhibits greater binding affinity for the target sequence in the target nucleic acid compared to an RNP containing a reference dCasX protein and a reference gRNA with a ligated repressor domain in an equivalent assay system, wherein the PAM sequence of the target nucleic acid is CTC. In another embodiment, an RNP of a long-term repressor fusion protein containing a dCasX variant and a gRNA variant with a linked repressor domain exhibits greater binding affinity to a target sequence in a target nucleic acid compared to an RNP containing a reference dCasX protein and a reference gRNA with a linked repressor domain in an equivalent assay system, where the PAM sequence of the target nucleic acid is GTC.In other embodiments, an RNP of a long-term repressor fusion protein containing a dCasX variant with a ligated repressor domain and gRNA exhibits greater binding affinity for target sequences in target nucleic acids in an equivalent assay system compared with an RNP containing an equivalent repressor fusion protein containing a reference dCasX protein with a ligated repressor domain and gRNA, where the PAM sequence of the target nucleic acid is GTC, TTC, ATC, or CTC. In the embodiments described above, the increased binding affinity for one or more PAM sequences is at least 1.5 times greater for the PAM sequences compared with the binding affinity of any one RNP of any of the reference dCasX proteins (modified from SEQ ID NOs: 1-3) with a ligated repressor domain and gRNA as shown in Table 8.
[0120] c. dCasX variant protein with domains from multiple source proteins In certain embodiments, the present disclosure provides a chimeric dCasX variant protein for use in repressor fusion proteins.
[0121] As used herein, the term “chimeric dCasX” protein refers to both a dCasX protein containing at least two domains from different sources, and a dCasX protein containing at least one domain that is itself a chimeric protein. Thus, in some embodiments, a chimeric dCasX protein contains at least two domains isolated or derived from different sources, for example, from two different naturally occurring CasX proteins (e.g., from two different CasX reference proteins). In other embodiments, a chimeric dCasX protein contains at least one domain that is a chimeric domain, for example, in some embodiments, a portion of the domain contains substitutions from a different CasX protein (a reference CasX protein, or another CasX variant protein).
[0122] In some embodiments, at least one chimeric domain may be one of the NTSB, TSL, helix I, helix II, OBD, or RuvC domains described herein. In the case of segmented or discontinuous domains, such as helix I, RuvC, and OBD, a portion of the discontinuous domain may be replaced with a corresponding portion from any other source. In some embodiments, the helix I-II domain of the dCasX variant derived from SEQ ID NO: 2 is replaced with the corresponding helix I-II sequence from SEQ ID NO: 1 to yield a chimeric dCasX protein. In some embodiments, the helix I-II domain and NTSB domain of the dCasX variant derived from SEQ ID NO: 2 are replaced with the corresponding helix I-II sequence and NTSB sequence from SEQ ID NO: 1 to yield a chimeric dCasX protein.
[0123] Chimeric dCasX variant proteins may include the NTSB, TSL, helical II, helical I-II, helical II, OBD-I, and OBD-II domains from the CasX protein of SEQ ID NO: 2, and the RuvC-I and / or RuvC-II domains from the CasX protein of SEQ ID NO: 1, or vice versa, and mutations or other sequence modifications are introduced to create non-catalyzed variants with improved variant properties compared to the reference dCasX protein. As an example, the chimeric RuvC domain includes amino acids 660-823 from SEQ ID NO: 1 and amino acids 921-978 from SEQ ID NO: 2. As an alternative example, the chimeric RuvC domain includes amino acids 647-810 from SEQ ID NO: 2 and amino acids 934-986 from SEQ ID NO: 1. In certain embodiments, dCasX for use in long-term repressor fusion proteins comprises the NTSB domain and helix I-II domain from SEQ ID NO: 1, and the helix II domain from SEQ ID NO: 2, the latter being a chimeric domain, and it is understood that the dCasX variant has additional amino acid changes at a selected position (compared to the reference sequence), and the resulting chimeric dCasX protein has improved characteristics compared to the reference dCasX protein. Sequences in Table 2 that have the NTSB domain and helix I-II domain from SEQ ID NO: 1, and the helix II domain from SEQ ID NO: 2 include dCasX 491 (SEQ ID NO: 4), 515 (SEQ ID NO: 6), 516 (SEQ ID NO: 7), 518-520 (SEQ ID NO: 9-11), 522-527 (SEQ ID NO: 12-17), 532 (SEQ ID NO: 22), 593 (SEQ ID NO: 25), 676 (SEQ ID NO: 28 with L169K substitution in the NTSB domain), and 812 (SEQ ID NO: 29). Table 1 below provides the coordinates of the CasX domain in the reference CasX protein of Sequence ID No. 1 and Sequence ID No. 2. Those skilled in the art will understand that the domain boundaries shown in Table 1 below are approximate, and that protein fragments whose boundaries differ by only one, two, or three amino acids from those shown in the table below may have the same activity as the domains described below. [Table 1]
[0124] In some embodiments, the dCasX variant protein used in the long-term repressor fusion protein of this disclosure comprises a sequence selected from the group consisting of sequences SEQ ID NOs: 4-29 shown in Table 2, wherein the long-term repressor fusion protein containing dCasX retains the ability to form RNPs with gRNA. In other embodiments, the dCasX variant protein used in the repressor fusion protein of the present disclosure includes a sequence that is at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 86%, at least 87%, at least 88%, at least 89%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, and at least 99.5% identical to a sequence selected from the group consisting of sequences SEQ ID NOs: 4-29 shown in Table 2, wherein the long-term repressor fusion protein containing dCasX retains the ability to form RNPs with gRNA. In some embodiments, the dCasX variant protein used in the repressor fusion protein of the disclosure comprises a sequence selected from the group consisting of sequences SEQ ID NOs: 4 to 29, wherein the long-term repressor fusion protein containing dCasX retains the ability to form RNPs with gRNA. In a particular embodiment, the dCasX variant protein used in the long-term repressor fusion protein of the gene repressor system of the disclosure comprises the sequence of SEQ ID NO: 4 (dCasX 491). In another particular embodiment, the dCasX variant protein used in the long-term repressor fusion protein of the gene repressor system of the disclosure comprises the sequence of SEQ ID NO: 6 (dCasX 515). In yet another particular embodiment, the dCasX variant protein used in the long-term repressor fusion protein of the gene repressor system of the disclosure comprises the sequence of SEQ ID NO: 28 (dCasX 676).In another specific embodiment, the dCasX variant protein used in the long-term repressor fusion protein of the gene repressor system of this disclosure comprises the sequence of SEQ ID NO: 29 (dCasX 812). [Table 2-1] [Table 2-2] [Table 2-3] [Table 2-4] [Table 2-5] [Table 2-6] [Table 2-7] [Table 2-8] [Table 2-9]
[0125] Affinity for d.gRNA In some embodiments, a long-term repressor fusion protein containing dCasX with a linked repressor domain exhibits improved affinity for gRNA compared to a corresponding long-term repressor fusion protein containing a reference dCasX protein with a linked repressor domain, leading to the formation of a ribonucleoprotein complex (RNP). The increased affinity of the long-term repressor fusion protein for gRNA leads to, for example, lower K for RNP complex formation. dThis may result in the formation of a more stable ribonucleoprotein complex in some cases. In some embodiments, the K of the long-term repressor fusion protein for gRNA d This increases by at least approximately 1.1, at least approximately 1.2, at least approximately 1.3, at least approximately 1.4, at least approximately 1.5, at least approximately 1.6, at least approximately 1.7, at least approximately 1.8, at least approximately 1.9, at least approximately 2, at least approximately 3, at least approximately 4, at least approximately 5, at least approximately 6, at least approximately 7, at least approximately 8, at least approximately 9, at least approximately 10, at least approximately 15, at least approximately 20, at least approximately 25, at least approximately 30, at least approximately 35, at least approximately 40, at least approximately 45, at least approximately 50, at least approximately 60, at least approximately 70, at least approximately 80, at least approximately 90, or at least approximately 100 times compared to the reference dCasX protein and the ligated repressor domain. In some embodiments, the long-term repressor fusion protein, comprising a dCasX variant and a linked repressor domain, exhibits approximately 1.1 to 10-fold increased binding affinity to gRNA compared to the corresponding repressor fusion protein, which comprises a variant of the reference CasX protein of SEQ ID NO: 2 that does not exhibit catalytic activity.
[0126] In some embodiments, the increased affinity of the long-term repressor fusion protein for gRNA results in increased stability of the ribonucleoprotein complex when delivered to mammalian cells, including in vivo delivery to the target. This increased stability may affect the function and utility of the complex in the target cells and, when delivered to the target, may result in improved pharmacokinetic properties in the blood. In some embodiments, the increased affinity of the repressor fusion protein, and the resulting increased stability of the ribonucleoprotein complex, allows for lower doses of the long-term repressor fusion protein delivered to the target or cells, while still possessing the desired activity, e.g., gene repression and / or epigenetic modification in vivo or in vitro. The increased ability to form RNPs and maintain them in a stable form can be evaluated using in vitro assays known in the art.
[0127] In some embodiments, the higher affinity (tighter binding) of the long-term repressor fusion protein, which includes the dCasX variant protein and the linked repressor domain, to the gRNA allows for a greater amount of transcriptional repression and / or epigenetic modification events if both the long-term repressor fusion protein and the gRNA remain in the RNP complex. The increased transcriptional repression events can be evaluated using the assays described herein.
[0128] Methods for measuring the binding affinity of long-term repressor fusion proteins to gRNAs include in vitro methods using purified long-term repressor fusion proteins and gRNAs. Binding affinity for long-term repressor fusion proteins can be measured by fluorescence polarization if the gRNA or long-term repressor fusion protein is tagged with a fluorophore. Alternatively, or in addition, binding affinity can be measured by biolayer interferometry, electrophoretic mobility shift assay (EMSA), or filter binding. Additional standard techniques for quantifying the absolute binding affinity of the repressor fusion proteins of this disclosure to specific gRNAs include, but are not limited to, isothermal calorimetry (ITC) and surface plasmon resonance (SPR).
[0129] e. Improved specificity regarding target nucleic acid sequences In some embodiments, long-term repressor fusion proteins containing a dCasX variant protein with a ligated repressor domain exhibit improved specificity for target nucleic acid sequences complementary to the gRNA targeting sequence compared to a reference dCasX protein with a ligated repressor domain. As used herein, “specificity” is sometimes referred to as “target specificity” and refers to the degree to which the RNP complex binds to off-target sequences that are similar to, but not identical to, the target nucleic acid sequence. For example, a long-term repressor fusion protein RNP with a higher degree of specificity may exhibit reduced off-target methylation of the sequence compared to an RNP of a reference dCasX with a ligated repressor domain. The specificity of long-term repressor fusion proteins, and the reduction of potentially harmful off-target effects, may be important for achieving an acceptable therapeutic index for use in mammalian subjects. Without being constrained by theory, amino acid changes in the helical I and II domains can increase the specificity of dCasX for the target nucleic acid strand, and thereby increase the specificity of the long-term repressor fusion protein for the entire target nucleic acid. In some embodiments, amino acid changes that increase the specificity of the repressor fusion protein for the target nucleic acid may also result in a decreased affinity of the repressor fusion protein for DNA, however, the overall benefits and safety of the composition are enhanced.
[0130] f. Repressor fusion proteins involving heterologous proteins Furthermore, within the scope of this disclosure are repressor fusion proteins comprising heterologous proteins fused to long-term repressor fusion proteins for use in the systems of this disclosure. These include repressor fusion proteins comprising N-terminal and / or C-terminal fusions to heterologous proteins or their domains. In some embodiments, the long-term repressor fusion protein is fused to one or more proteins or their domains having different desired activities.
[0131] In some cases, heterologous polypeptides (fusion partners) for use with long-term repressor fusion proteins provide intracellular localization, i.e., the heterologous polypeptide contains intracellular localization sequences (e.g., nuclear localization signals (NLS) for targeting to the nucleus, sequences for keeping the fusion protein outside the nucleus, nuclear export sequences (NES), sequences for keeping the fusion protein retained in the cytoplasm, mitochondrial localization signals for targeting to mitochondria, chloroplast localization signals for targeting to chloroplasts, ER retention signals, and similar).
[0132] In some cases, the long-term repressor fusion protein contains (is fused to) a nuclear localization signal (NLS). In some cases, the long-term repressor fusion protein is fused to two or more, three or more, four or more, or five or more, six or more, seven or more, or eight or more NLSs. In some cases, one or more NLSs (two or more, three or more, four or more, or five or more NLSs) are located at or near the N-terminus and / or C-terminus of the repressor fusion protein (e.g., within 20 amino acids). In some cases, one or more NLSs (two or more, three or more, four or more, or five or more NLSs) are located at or near the N-terminus of the repressor fusion protein (e.g., within 20 amino acids). In some cases, one or more NLSs (two or more, three or more, four or more, or five or more NLSs) are located at or near the C-terminus of the repressor fusion protein (e.g., within 20 amino acids). In some cases, one or more NLSs (three or more, four or more, or five or more NLSs) are located at or near both the N-terminus and C-terminus of the repressor fusion protein (e.g., within 20 amino acids). In some cases, a single NLS is located at the N-terminus of the repressor fusion protein, and a single NLS is located at the C-terminus. Those skilled in the art will understand that an NLS at or near the N-terminus or C-terminus of a protein may be within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids of the N-terminus or C-terminus. In some embodiments, an NLS ligated to the N-terminus of a long-term repressor fusion protein is identical to an NLS ligated to the C-terminus. In other embodiments, an NLS ligated to the N-terminus of a long-term repressor fusion protein is different from an NLS ligated to the C-terminus. A typical structure of a repressor fusion protein associated with NLS is shown in Figure 19.In some embodiments, suitable NLSs for use with long-term repressor fusion proteins in the systems of the present disclosure include, or are identical to, sequences having at least about 85%, at least about 90%, or at least about 95% identity with sequences derived from nucleoplasmic NLS (e.g., nucleoplasmic bifurcation NLS with sequence KRPAATKKAGQAKKKK (SEQ ID NO: 31)); sequences derived from c-MYC NLS having amino acid sequences PAAKRVKLD (SEQ ID NO: 32) or RQRRNELKRSP (SEQ ID NO: 33). In some embodiments, the NLS and short peptide linker ligated to the N-terminus of the long-term repressor fusion protein is sequence PKKKRKVSR (SEQ ID NO: 34). In some embodiments, the NLS and short peptide linker ligated to the N-terminus of the long-term repressor fusion protein is sequence PKKKRKVSRVNGSGSGGG (SEQ ID NO: 21840). In some embodiments, the NLS and short peptide linker ligated to the C-terminus of the long-term repressor fusion protein is sequence TSPKKKRKV (SEQ ID NO: 21841). In some embodiments, the NLS ligated to or adjacent to the N-terminus of the long-term repressor fusion protein is selected from the group consisting of SEQ ID NOs: 34-67. In some embodiments, the NLS ligated to or adjacent to the N-terminus of the long-term repressor fusion protein is selected from the group consisting of SEQ ID NOs: 68-97. In some embodiments, an NLS suitable for use with the long-term repressor fusion protein in the system of this disclosure contains a sequence having at least about 80%, at least about 90%, or at least about 95% identity, or is identical to one or more sequences in Table 3 or Table 4. Those skilled in the art will understand that any of the NLS sequences shown in Tables 3 and 4 may be fused to or adjacent to either the N-terminus or C-terminus of the repressor fusion protein described herein. [Table 3] [Table 4]
[0133] In some embodiments, one or more NLSs are linked to a long-term repressor fusion protein, or to an adjacent NLS accompanied by a linker peptide. In some embodiments, the linker peptide is SR, GS, GP, TS, VGS, GGS, (G)n (SEQ ID NO: 98), (GS)n (SEQ ID NO: 99), (GSGGS)n (SEQ ID NO: 100), (GGSGGS)n (SEQ ID NO: 101), (GGGS)n (SEQ ID NO: 102), GGSG (SEQ ID NO: 103), GGSGG (SEQ ID NO: 104), GSGSG (SEQ ID NO: 105), GSGGG (SEQ ID NO: 106), GGGSG (SEQ ID NO: 107), GSSSG (SEQ ID NO: 108), GP (SEQ ID NO: 109), GGP, PPP, VPPP, PPAPPA (SEQ ID NO: 110), PPPG (SEQ ID NO: 111), PPPGPPP (SEQ ID NO: 112), PPP(GGGS)n (SEQ ID NO: 113), (GGGS)nPPP (SEQ ID NO: 114), AEAAAKEAAAKEAAAKA (SEQ ID NO: 115), The formula is selected from the group consisting of VPPPGGGSGGGSGGGS (sequence number 116), TGGGPGGGAAAGSGS (sequence number 117), GGGSGGGSGGGSPPP (sequence number 118), TPPKTKRKVEFE (sequence number 119), GGSGGGS (sequence number 120), GGSGGGG (sequence number 121), SSGNSNANSRGPSFSSGLVPLSLRGSH (sequence number 122), GGPSSGAPPPSGGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAPGTSTEPSE (sequence number 123), GGSGGG (sequence number 124), GSGS (sequence number 1988), GGSGSSG (sequence number 2130), and GGSGGGSA (sequence number 2131), where n is between 1 and 5.
[0134] Generally, NLS (or multiple NLSs) are strong enough to drive the accumulation of long-term repressor fusion proteins in the nuclei of eukaryotic cells. Detection of accumulation in the nucleus can be carried out by any suitable technique known in the art. For example, a detectable marker can be fused to the long-term repressor fusion protein so that its intracellular location can be visualized. The cell nucleus can also be isolated from the cell, and its contents can then be analyzed by any suitable process for detecting proteins, such as immunohistochemistry, Western blotting, or enzyme activity assays. Accumulation in the nucleus can also be determined indirectly.
[0135] IV. Long-term repressor domain fusion protein The present invention provides a system comprising a long-term repressor fusion protein (LTRP) containing a DNA-binding protein linked to multiple repressor domains in a designed configuration, wherein the system is capable of binding to a target nucleic acid of a gene and repressing its transcription by epigenetic modification of the target nucleic acid. Exemplary DNA-binding proteins for use in fusion proteins include zinc finger (ZF), TALE (transmission activator-like effector) proteins, and CRISPR proteins without catalytic activity.
[0136] In some embodiments, the present invention provides a system of repressor fusion proteins comprising a non-catalyzed CasX variant protein (dCasX) linked to multiple repressor domains, which, when complexed with a guide ribonucleic acid (gRNA) containing a targeting sequence complementary to the target nucleic acid sequence of a gene, can bind to the target nucleic acid to repress or silence transcription and / or affect the epigenetic modifications of the target nucleic acid. Examples of gene repression processes that reduce transcription include, but are not limited to, inhibiting the formation of the transcription initiation complex, reducing the transcription initiation rate, reducing the transcription elongation rate, reducing the forward momentum of transcription, and antagonizing transcriptional activation (e.g., by blocking the binding of transcription activators). Gene repression may constitute, for example, prevention of activation and inhibition of expression below existing levels. Transcriptional repression includes both reversible and irreversible inactivation of gene transcription, the latter of which may result from epigenetic modifications of the target nucleic acid.
[0137] Among repressor domains capable of repressing or silencing genes, the Kruppel-associated box (KRAB) repressor domain is among the most potent in the human genome system (Alerasool, N., et al. An efficient KRAB domain for CRISPRi applications. Nat. Methods 17:1093 (2020)). KRAB-like domains are present in approximately 400 human zinc finger protein-based transcription factors, and they can recruit additional repressor domains, such as Trim28 (also known as Kap1 or Tif1-beta), upon binding of linked dCasX to target nucleic acids, which in turn assemble protein complexes with chromatin regulators, such as CBX5 / HP1α and SETDB1, which induce repression of gene transcription, but in a limited, time-dependent manner, involving modifications of DNA-associated histones rather than DNA modifications. By reducing histone H3-acetylation and increasing H3-lysine 9 trimethylation at the cellular level, KRAB / KAP1 mediates reversible and long-range transcriptional repression through heterochromatin diffusion. Representative, non-limiting examples of KRAB domains include ZIM3 (SEQ ID NOs. 129 and 1892) and ZNF10 (SEQ ID NOs. 128 and 1891). This disclosure provides repressor domains from human sources, as well as repressor domains from non-human sources with a distinctly different sequence (referred to herein as "RD1") that, when incorporated into embodiments of long-range repressor fusion protein constructs, more fully described below, have been found to result in enhanced transcriptional repression compared to ZIM3 and ZNF10.
[0138] In some embodiments, this disclosure provides a system in which the modification of a gene conferred by the use of an LTRP:gRNA system is epigenetic, and therefore the silencing of the gene is inheritable by a mechanism other than DNA editing replication. As used herein, “epigenetic modification” means a modification of either DNA or a histone associated with DNA other than a change in the DNA sequence itself (e.g., substitution, deletion, or rearrangement), wherein the modification is either a direct modification by a component of the system or indirectly by the recruitment of one or more additional cellular components, but the DNA target nucleic acid sequence itself is not edited to alter its sequence. For example, while DNA methyltransferase 3A (DNMT3A) (or its catalytic domain) directly modifies DNA by methylating it, KRAB can recruit the KAP-1 / TIF1β corepressor complex, which acts as a potent transcriptional repressor, and further recruit factors associated with DNA methylation and repressive chromatin formation, such as heterochromatin protein 1 (HP1), histone deacetylase, and histone methyltransferase (Ying, Y., et al. The Kruppel-associated box repressor domain induces reversible and irreversible regulation of endogenous mouse genes by mediating different chromatin states. Nucleic Acids Res. 43(3):1549 (2015)). Furthermore, catalytically inactive DNMT3-like (DNMT3L) cofactors, together with endogenous DNMT1 in cells, help establish post-replicated genetic methylation patterns.The ATRX-DNMT3-DNMT3L domain (ADD) of DNMT3A is known to have two main functions: 1) allosterically regulates the catalytic activity of DNMT3A by acting as a methyltransferase autoinhibitory domain, and 2) specifically interacts with the lysine (K)4-unmethylated histone H3 tail (H3K4me0), leading to preferential methylation of DNA bound to the K4-unmethylated chromatin H3 tail (Zhang, Y., et al. Chromatin methylation activity of Dnmt3a and Dnmt3a / 3L is guided by interaction of the ADD domain with the histone H3 tail. Nucleic Acids Research 38:4246 (2010)). In some embodiments, the inclusion of the ADD domain enhances transcriptional repression of targeted genes when incorporated into the design of an LTRP, compared to an otherwise equivalent LTRP lacking the ADD domain. In other embodiments, the inclusion of the ADD domain enhances the specificity of transcriptional repression of the targeted gene when incorporated into the design of the LTRP, compared to an otherwise equivalent LTRP lacking the ADD domain. Supporting data for the foregoing are provided in the examples and in WO2023049742A2, which is incorporated herein by reference.
[0139] In some embodiments, the long-term repressor fusion protein comprises DNA-binding proteins linked to first, second, and third repressor domains, wherein each of the repressor domains is distinct. In some embodiments, the long-term repressor fusion protein comprises DNA-binding proteins linked to first, second, third, and fourth repressor domains, wherein each of the repressor domains is distinct. In any of the embodiments described above, the long-term fusion protein is capable of forming an RNP with a gRNA system that binds to a target nucleic acid.
[0140] In some embodiments, the DNA-binding protein is a TALE that binds to the target nucleic acid but cannot cleave it. In some embodiments, the DNA-binding protein is a zinc finger protein modified to bind to the target nucleic acid but not to cleave it. In some embodiments, the DNA-binding protein is a non-catalyzed CRISPR protein that binds to the target nucleic acid but cannot cleave it. In some embodiments, the long-term repressor fusion protein comprises a non-catalyzed CRISPR protein sequence, a first repressor domain (hereinafter referred to as "RD1"), a DNMT3A catalytic domain from a DNMT3A protein as a second domain (hereinafter referred to as "DNMT3A"), and a DNMT3L interaction domain from a DNMT3L protein as a third domain (hereinafter referred to as "DNMT3L"). In some embodiments, the long-term repressor fusion protein comprises a CRISPR protein sequence without catalytic activity, RD1, DNMT3A as a second domain, DNMT3L as a third domain, and an ATRX-DNMT3-DNMT3L domain (hereinafter referred to as "ADD") from the DNMT3A protein as a fourth domain. In some embodiments, the long-term repressor fusion protein further comprises the first and second NLSs described herein and one or more linker peptides. In some embodiments, the long-term repressor fusion protein is capable of forming an RNP with a gRNA that binds to the target nucleic acid. It has been found that the use of the aforementioned domains, when configured in a selective orientation toward the DNA-binding protein in the long-term repressor fusion protein, can result in significant epigenetic modification of the target nucleic acid when bound to a defined region of the gene to be silenced, and that the combination of repressor domains works synchronously and, depending on the configuration, produces an additive or synergistic effect on the transcriptional silencing of the targeted gene.
[0141] Representative amino acid sequences of components used in long-term repressor fusion protein constructs are provided herein. In some embodiments, the DNA-binding protein of the long-term repressor fusion protein is a dCasX containing a sequence selected from the group consisting of SEQ ID NOs: 4-29, or a sequence having at least about 70%, at least about 80%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In some embodiments, the DNA-binding protein of the long-term repressor fusion protein is a dCasX containing the sequence of SEQ ID NO: 4, or a sequence having at least about 70%, at least about 80%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In some embodiments, the DNA-binding protein of the long-term repressor fusion protein is a dCasX containing the sequence of SEQ ID NO: 4. In some embodiments, the DNA-binding protein of the long-term repressor fusion protein is a dCasX containing the sequence of SEQ ID NO: 5, or a sequence having at least about 70%, at least about 80%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In some embodiments, the DNA-binding protein of the long-term repressor fusion protein is a dCasX containing the sequence of SEQ ID NO: 28, or a sequence having at least about 70%, at least about 80%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto.In some embodiments, the DNA-binding protein of the long-term repressor fusion protein is a dCasX containing the sequence of SEQ ID NO: 29, or a sequence having at least about 70%, at least about 80%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto.
[0142] In some embodiments, the RD1 of the long-term repressor fusion protein includes a sequence selected from the group consisting of SEQ ID NOs: 1891 and 1892, or a sequence having at least about 70%, at least about 80%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In some embodiments, the RD1 of the long-term repressor fusion protein includes a sequence selected from the group consisting of SEQ ID NOs: 1891 and 1892. In another embodiment, RD1 of the long-term repressor fusion protein includes a sequence selected from the group consisting of SEQ ID NOs: 130-1726, or a sequence having at least about 70%, at least about 80%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In another embodiment, RD1 of the long-term repressor fusion protein includes a sequence selected from the group consisting of SEQ ID NOs: 130-1726. In another embodiment, RD1 of the long-term repressor fusion protein includes a sequence selected from the group consisting of SEQ ID NOs: 130-224, or a sequence having at least about 70%, at least about 80%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In another embodiment, RD1 of the long-term repressor fusion protein includes a sequence selected from the group consisting of SEQ ID NOs: 130-224. In another embodiment, the long-term repressor fusion protein RD1 includes a sequence selected from the group consisting of SEQ ID NOs. 130-138, or a sequence having at least about 70%, at least about 80%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto.In another embodiment, RD1 of the long-term repressor fusion protein includes a sequence selected from the group consisting of SEQ ID NOs. 130 to 138. In another embodiment, RD1 of the long-term repressor fusion protein includes the sequence of SEQ ID NO. 135, or a sequence that is at least about 70%, at least about 80%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical thereto. In another embodiment, RD1 of the long-term repressor fusion protein includes the sequence of SEQ ID NO. 131, or a sequence that is at least about 70%, at least about 80%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical thereto. In another embodiment, RD1 of the long-term repressor fusion protein includes the sequence of SEQ ID NO: 130, or a sequence having at least about 70%, at least about 80%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In another embodiment, RD1 of the long-term repressor fusion protein includes a sequence selected from the group consisting of SEQ ID NOs: 130, 131, and 135.
[0143] In some embodiments, the second repressor domain of the long-term repressor fusion protein is a DNMT3A containing the sequence of SEQ ID NO: 126, or a sequence having at least about 70%, at least about 80%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In some embodiments, the second repressor domain of the long-term repressor fusion protein contains the sequence of SEQ ID NO: 126.
[0144] In some embodiments, the third repressor domain of the long-term repressor fusion protein is DNMT3L containing the sequence of SEQ ID NO: 127, or a sequence having at least about 70%, at least about 80%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In some embodiments, the third repressor domain of the long-term repressor fusion protein contains the sequence of SEQ ID NO: 127.
[0145] In some embodiments, the fourth repressor domain of the long-term repressor fusion protein is an ADD containing the sequence of SEQ ID NO: 125, or a sequence having at least about 70%, at least about 80%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In some embodiments, the fourth repressor domain of the long-term repressor fusion protein contains the sequence of SEQ ID NO: 125. In some embodiments, the C-terminus of the ADD is ligated to the N-terminus of DNMT3A. In a surprising finding, the addition of the ADD to a construct containing the DNA-binding proteins RD1, DNMT3A, and DNMT3L has been found to significantly enhance or increase the long-term repression and / or epigenetic modification of the target nucleic acid, as well as the specificity of the repression, compared to a construct lacking the ADD. Exemplary data on the improved transcriptional repression and specificity of constructs containing ADD are presented in the examples and described in WO2023049742A2, which is incorporated herein by reference.
[0146] This disclosure provides a system comprising a long-term repressor fusion protein comprising first, second, third, and optionally fourth repressor domains operably linked to a DNA-binding protein, wherein the DNA-binding protein is a dCasX comprising the sequence of SEQ ID NO: 4, or a sequence having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto, and the second repressor domain comprising the sequence of SEQ ID NO: 126, or a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, The third repressor is a DNMT3A domain containing a sequence that is at least approximately 98% or at least approximately 99% identical to sequence number 127, or a DNMT3L domain containing a sequence that is at least approximately 70%, at least approximately 80%, at least approximately 85%, at least approximately 90%, at least approximately 91%, at least approximately 92%, at least approximately 93%, at least approximately 94%, at least approximately 95%, at least approximately 96%, at least approximately 97%, at least approximately 98%, or at least approximately 99% identical to it, and the fourth repressor is an ADD domain containing a sequence that is at least approximately 70%, at least approximately 80%, at least approximately 85%, at least approximately 90%, at least approximately 91%, at least approximately 92%, at least approximately 93%, at least approximately 94%, at least approximately 95%, at least approximately 96%, at least approximately 97%, at least approximately 98%, or at least approximately 99% identical to it.In some embodiments, the long-term repressor fusion protein comprises first, second, third, and optionally fourth repressor domains operably linked to a DNA-binding protein, wherein the DNA-binding protein is a dCasX containing a sequence selected from the group consisting of SEQ ID NOs: 4-29, the second repressor domain is a DNMT3A domain containing the sequence of SEQ ID NO: 126, the third repressor is a DNMT3L domain containing the sequence of SEQ ID NO: 127, and the fourth repressor is an ADD domain containing the sequence of SEQ ID NO: 125. In some embodiments, the long-term repressor fusion protein comprises one or more linker peptides selected from the group consisting of SEQ ID NOs: 98-124, 1823-1874, 1988, and 2130-2131 (Table 5). In some embodiments, the long-term repressor fusion protein comprises one or more NLS containing a sequence selected from the group consisting of SEQ ID NOs: 30, 32, 34, and 21841. In some embodiments, the first RD1 includes a sequence selected from the group consisting of SEQ ID NOs: 130 to 1726, or a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In some embodiments, the first RD1 includes a sequence selected from the group consisting of SEQ ID NOs: 130 to 1726. In other embodiments of the long-term repressor protein, the first RD1 includes a sequence selected from the group consisting of SEQ ID NOs: 130-224, or a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto.In other embodiments of the long-term repressor protein, the first RD1 includes a sequence selected from the group consisting of SEQ ID NOs. 130-138, or a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In other embodiments of the long-term repressor protein, the first RD1 includes a sequence selected from the group consisting of SEQ ID NOs. 130-138. In other embodiments of the long-term repressor protein, the first RD1 includes the sequence of SEQ ID NO: 135, or a sequence that is at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical thereto. In other embodiments of the long-term repressor protein, the first RD1 includes the sequence of SEQ ID NO: 135. In other embodiments of the long-term repressor protein, the first RD1 includes the sequence of SEQ ID NO: 131, or a sequence that is at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical thereto. In other embodiments of the long-term repressor protein, the first RD1 includes the sequence of SEQ ID NO: 131. In other embodiments of the long-term repressor protein, the first RD1 includes the sequence of SEQ ID NO: 130, or a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In other embodiments of the long-term repressor protein, the first RD1 includes the sequence of SEQ ID NO: 130.In other embodiments of the long-term repressor protein, the first RD1 comprises a sequence selected from the group consisting of SEQ ID NOs: 130, 131, and 135. In some embodiments, embodiments of the long-term repressor protein of this paragraph are capable of forming an RNP with the gRNA of the disclosure that binds to a target gene and represses or silences its expression.
[0147] In some embodiments, the long-term repressor fusion protein comprises DNMT3A, DNMT3L, a DNA-binding protein, and RD1 from the N-terminus to the C-terminus. In some embodiments, the long-term repressor fusion protein comprises ADD, DNMT3A, DNMT3L, a DNA-binding protein, and RD1 from the N-terminus to the C-terminus. In some embodiments, the DNA-binding protein may be a zinc finger, TALE, or a CRISPR protein without catalytic activity. In some embodiments, the long-term repressor fusion protein comprises an NLS at the N-terminus, C-terminus, or both. In some embodiments, the long-term repressor fusion protein comprises one or more linkers between DNMT3A and DNMT3L, between DNMT3L and the DNA-binding protein, and / or between the DNA-binding protein and RD1.
[0148] In some embodiments, the long-term repressor fusion protein comprises a DNA-binding protein, RD1, DNMT3A, and DNMT3L from the N-terminus to the C-terminus. In some embodiments, the long-term repressor fusion protein comprises a DNA-binding protein, RD1, ADD, DNMT3A, and DNMT3L from the N-terminus to the C-terminus. In some embodiments, the DNA-binding protein may be a zinc finger, TALE, or a CRISPR protein without catalytic activity. In some embodiments, the long-term repressor fusion protein comprises an NLS at the N-terminus. In some embodiments, the long-term repressor fusion protein comprises an NLS between RD1 and DNMT3A. In some embodiments, the long-term repressor fusion protein comprises an NLS at the N-terminus, the C-terminus, or both. In some embodiments, the long-term repressor fusion protein comprises one or more linkers between the N-terminal NLS and the DNA-binding protein, between the DNA-binding protein and RD1, between RD1 and DNMT3A, or optionally between ADD, and / or between DNMT3A and DNMT3A.
[0149] In some embodiments, the long-term repressor fusion protein comprises a DNA-binding protein, DNMT3A, DNMT3L, and RD1 from the N-terminus to the C-terminus. In some embodiments, the long-term repressor fusion protein comprises a DNA-binding protein, ADD, DNMT3A, DNMT3L, and RD1 from the N-terminus to the C-terminus. In some embodiments, the DNA-binding protein may be a zinc finger, TALE, or a CRISPR protein without catalytic activity. In some embodiments, the long-term repressor fusion protein comprises an NLS at the N-terminus, C-terminus, or both. In some embodiments, the long-term repressor fusion protein comprises one or more linkers between the N-terminal NLS and the DNA-binding protein, between the DNA-binding protein and DNMT3A, or optionally ADD, between DNT3A and DNMT3L, and / or between DNMT3L and RD1.
[0150] In some embodiments, the long-term repressor fusion protein comprises RD1, DNMT3A, DNMT3L, and a DNA-binding protein from the N-terminus to the C-terminus. In some embodiments, the long-term repressor fusion protein comprises RD1, ADD, DNMT3A, DNMT3L, and a DNA-binding protein from the N-terminus to the C-terminus. In some embodiments, the DNA-binding protein may be a zinc finger, TALE, or a CRISPR protein without catalytic activity. In some embodiments, the long-term repressor fusion protein comprises an NLS at the N-terminus, C-terminus, or both. In some embodiments, the long-term repressor fusion protein comprises one or more linkers between RD1 and DNMT3A, or optionally ADD, between DNMT3A and DNMT3L, between DNMT3L and the DNA-binding protein, and / or between the DNA-binding protein and the C-terminal NLS.
[0151] In some embodiments, the long-term repressor fusion protein comprises DNMT3A, DNMT3L, RD1, and a DNA-binding protein from the N-terminus to the C-terminus. In some embodiments, the long-term repressor fusion protein comprises ADD, DNMT3A, DNMT3L, RD1, and a DNA-binding protein from the N-terminus to the C-terminus. In some embodiments, the DNA-binding protein may be a zinc finger, TALE, or a CRISPR protein without catalytic activity. In some embodiments, the long-term repressor fusion protein comprises an NLS at the N-terminus, C-terminus, or both. In some embodiments, the long-term repressor fusion protein comprises one or more linkers between DNMT3A and DNMT3L, between DNMT3L and RD1, between RD1 and the DNA-binding protein, and / or between the DNA-binding protein and the C-terminal NLS.
[0152] In some embodiments, the long-term repressor fusion protein comprises DNMT3A, DNMT3L, RD1, a DNA-binding protein, and a second RD1 from the N-terminus to the C-terminus. In some embodiments, the long-term repressor fusion protein comprises ADD, DNMT3A, DNMT3L, RD1, a DNA-binding protein, and a second RD1 from the N-terminus to the C-terminus. In one embodiment described above, the second RD1 may be sequenced identical to the first RD1. In another embodiment described above, the second RD1 may be sequenced different from the first RD1. In some embodiments, the DNA-binding protein may be a zinc finger, TALE, or a CRISPR protein without catalytic activity. In some embodiments, the long-term repressor fusion protein comprises NLS at the N-terminus, C-terminus, or both. In some embodiments, the long-term repressor fusion protein includes one or more linkers between DNMT3A and DNMT3L, between DNMT3L and RD1, between RD1 and the DNA-binding protein, between the DNA-binding protein and the second RD1, and / or between the second RD1 and the C-terminal NLS.
[0153] In some cases, the long-term repressor fusion protein also comprises one or more NLSs. In some embodiments, the long-term repressor fusion protein comprises the following configurations from N-terminus to C-terminus: NLS-ADD-DNMT3A-DNMT3L-DNA-binding protein-RD1-NLS, NLS-DNA-binding protein-RD1-NLS-ADD-DNMT3A-DNMT3L, NLS-DNA-binding protein-ADD-DNMT3A-DNMT3L-RD1-NLS), NLS-RD1-ADD-DNMT3A-DNMT3L-DNA-binding protein-NLS, NLS-ADD-DNMT3A-DNMT3L-RD1-DNA-binding protein-NLS, or NLS-ADD-DNMT3A-DNMT3L-RD1-DNA-binding protein-RD1-NLS. In some embodiments, the DNA-binding protein may be a zinc finger, TALE, or a CRISPR protein without catalytic activity. In some embodiments, the long-term repressor fusion protein can form an RNP with the gRNA embodiment of this disclosure, and can bind to and repress or silence a gene target nucleic acid.
[0154] In some embodiments, the long-term repressor fusion protein, one or more linker peptides, can be inserted between any two adjacent domains of the long-term repressor fusion protein. In some embodiments, the long-term repressor fusion protein comprises the configuration NLS-ADD-DNMT3A-linker2-DNMT3L-linker1-linker3A-DNA-binding protein-linker3B-RD1-NLS (configuration 1) from the N-terminus to the C-terminus. In some embodiments, the long-term repressor fusion protein comprises the configuration NLS-linker3A-DNA-binding protein-linker3B-RD1-NLS-linker1-ADD-DNMT3A-linker2-DNMT3L (configuration 2) from the N-terminus to the C-terminus. In some embodiments, the long-term repressor fusion protein comprises the configuration NLS-linker3A-DNA-binding protein-linker1-ADD-DNMT3A-linker2-DNMT3L-linker3B-RD1-NLS (configuration 3) from the N-terminus to the C-terminus. In some embodiments, the long-term repressor fusion protein includes the configuration NLS-RD1-linker3A-ADD-DNMT3A-linker2-DNMT3L-linker1-DNA-binding protein-linker3B-NLS (configuration 4) from the N-terminus to the C-terminus. In some embodiments, the long-term repressor fusion protein includes the configuration NLS-ADD-DNMT3A-linker2-DNMT3L-linker3A-RD1-linker1-DNA-binding protein-linker3B-NLS (configuration 5) from the N-terminus to the C-terminus. In some embodiments, the DNA-binding protein in the above configurations may be a zinc finger, TALE, or a CRISPR protein without catalytic activity. Schematic diagram of the configurations shown in Figure 19. In some embodiments of the LTRP of configurations 1 to 5, the NLS may include sequences selected from the group consisting of sequence numbers 30 to 97 (Tables 3 and 4), and the linker sequences may include sequences independently selected from the group consisting of sequence numbers 98 to 124, 1823 to 1874, 1988, and 2130 to 2131 (representative linkers shown in Table 5).In some embodiments of LTRP, one or more NLSs may include the sequence of SEQ ID NO: 30, and one or more linker sequences may independently include sequences selected from the group consisting of SEQ ID NOs: 120 and 122-124. In some embodiments of LTRP, the DNA-binding protein may be a dCasX sequence selected from the group consisting of SEQ ID NOs: 4-29, or a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In some embodiments of the LTRP of configurations 1 to 5, the second repressor domain is a DNMT3A domain containing the sequence of sequence number 126, or a sequence variant having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In some embodiments of the LTRP, the third repressor is a DNMT3L domain containing the sequence of sequence number 127, or a sequence variant having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In some embodiments of LTRP, the fourth repressor is an ADD domain containing the sequence of SEQ ID NO: 125, or a sequence variant having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In some embodiments, the long-term repressor fusion protein is in one of configurations 1-5, as shown in Figure 19.In some embodiments, the LTRP can form an RNP with the gRNA of the Disclosure, and the RNP can bind to a target nucleic acid and repress or silence it.
[0155] In some embodiments, the disclosure provides a system comprising a long-term repressor fusion protein comprising a first repressor domain (RD1), a second repressor domain, a third repressor domain, and two copies of a fourth repressor domain, operably linked to a DNA-binding protein, such as a zinc finger, TALE, or a non-catalyzed CRISPR protein. In some embodiments, the DNA-binding protein may be a non-catalyzed CasX selected from the group consisting of SEQ ID NOs: 4-29, or a sequence having at least 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity. In some embodiments, the DNA-binding protein may be the non-catalyzed CasX sequence of SEQ ID NO: 4. In some embodiments, the sequences of the two copies of RD1 are identical. In other embodiments, the two RD1s have different sequences. In some embodiments, two copies of RD1 are positioned at the N-terminus of the DNA-binding protein. In some embodiments, two copies of RD1 are positioned at the C-terminus of the DNA-binding protein. In some embodiments, one copy of RD1 is positioned at the N-terminus of the DNA-binding protein, and one copy of RD1 is positioned at the C-terminus of the DNA-binding protein.
[0156] In some embodiments, the long-term repressor fusion protein contains two RD1s. In some embodiments, the long-term repressor fusion protein has the configuration NLS-ADD-DNMT3A-linker2-DNMT3L-linker3A-a RD1a-linker1-DNA binding protein-linker3B-RD1a-linker4-NLS (configuration 6a) from the N-terminus to the C-terminus, where the RD1a sequences are identical (see Figure 19 for a schematic diagram of the fusion protein). In some embodiments, the long-term repressor fusion protein has the configuration NLS-ADD-DNMT3A-linker2-DNMT3L-linker3A-a RD1a-linker1-DNA binding protein-linker3B-RD1b-linker4-NLS (configuration 6b) from the N-terminus to the C-terminus, where the RD1a and RD1b sequences are different (see Figure 19 for a schematic diagram of the fusion protein). In some embodiments, the DNA-binding protein may be a zinc finger, a TALE, or a CRISPR protein without catalytic activity. In some embodiments, embodiments of the long-term repressor protein of the paragraph may form an RNP with the gRNA of the Disclosure that binds to a target nucleic acid and represses or silences its expression.
[0157] In some embodiments, the disclosure provides a long-term repressor fusion protein of configuration 6a, in which two RD1 sequences are identical. In some embodiments of the long-term repressor fusion protein of configuration 6a, the two copies of RD1 are identical, and the DNA-binding protein is a dCasX sequence having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. The first and second copies of the first repressor domain (RD1a) include a sequence selected from the group consisting of SEQ ID NOs: 130-1726, or a sequence having at least approximately 70%, at least approximately 80%, at least approximately 85%, at least approximately 90%, at least approximately 91%, at least approximately 92%, at least approximately 93%, at least approximately 94%, or at least approximately 95% identity thereto, and the second repressor domain includes the sequence of SEQ ID NO: 126, or a sequence having at least approximately 70%, at least approximately 80%, at least DNMT3A containing a sequence that is approximately 85%, at least approximately 90%, at least approximately 91%, at least approximately 92%, at least approximately 93%, at least approximately 94%, at least approximately 95%, at least approximately 96%, at least approximately 97%, at least approximately 98%, or at least approximately 99%, and the third repressor is the sequence of sequence number 127, or at least approximately 70%, at least approximately 80%, at least approximately 85%, at least approximately 90%, at least approximately 91%, at least approximately 92%, at least approximately 93%, at least DNMT3L containing sequence variants having approximately 94%, at least approximately 95%, at least approximately 96%, at least approximately 97%, at least approximately 98%, or at least approximately 99% identity, the fourth repressor being the sequence of sequence number 125, or at least approximately 70%, at least approximately 80%, at least approximately 85%, at least approximately 90%, at least approximately 91%, at least approximately 92%, at least approximately 93%, at least approximately 94%, at least approximately 95%, at least approximately 96%, at least approximately 97%, at least approximately 98%,Alternatively, the ADD contains sequences having at least approximately 99% identity, and the NLS contains sequences independently selected from the group consisting of SEQ ID NOs: 30-97 (Tables 3 and 4), the L1 linker contains the sequence of SEQ ID NO: 123, the L2 linker contains the sequence of SEQ ID NO: 122, the L3A linker contains the sequence of SEQ ID NO: 124, the L3B linker contains the sequence of SEQ ID NO: 120, and the L4 linker contains the sequence of SEQ ID NO: 1988 or SEQ ID NO: 2130. In some embodiments of the long-term repressor fusion protein of configuration 6a, the linker sequences are independently selected from the group consisting of SEQ ID NOs: 98-124, 1823-1874, 1988, and 2130-2131 (exemplary linkers shown in Table 5). In some embodiments of the long-term repressor fusion protein of configuration 6a, two copies of RD1 are identical, and each RD1 contains a sequence independently selected from the group consisting of SEQ ID NOs: 130, 131, and 135, or a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, and at least about 95% identity thereto. In some embodiments, the DNA-binding protein may be the non-catalyzed CasX sequence of SEQ ID NO: 4. A schematic diagram of configuration 6a is shown in Figure 19. In some embodiments, the long-term repressor fusion protein of configuration 6a is capable of forming an RNP with the gRNA of this disclosure, and is capable of binding to and repressing or silencing a gene target nucleic acid.
[0158] In some embodiments of the long-term repressor fusion protein of configuration 6b, the two copies of RD1 are different, and the DNA-binding protein comprises a dCasX sequence having at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with it, and the first The first copy of the repressor domain (RD1a) contains a sequence selected from the group consisting of SEQ ID NOs. 130-1726, or a sequence that is at least approximately 70%, at least approximately 80%, at least approximately 85%, at least approximately 90%, at least approximately 91%, at least approximately 92%, at least approximately 93%, at least approximately 94%, or at least approximately 95% identical thereto, and the second copy of the first repressor domain (RD1b) contains a sequence selected from the group consisting of SEQ ID NOs. 130-1726, or a sequence that is at least approximately 70%, at least approximately 80%, at least Both contain sequences that are approximately 85%, at least approximately 90%, at least approximately 91%, at least approximately 92%, at least approximately 93%, at least approximately 94%, or at least approximately 95% identical, and the second repressor domain contains the sequence of sequence number 126, or at least approximately 70%, at least approximately 80%, at least approximately 85%, at least approximately 90%, at least approximately 91%, at least approximately 92%, at least approximately 93%, at least approximately 94%, at least approximately 95%, at least approximately 96%, at least approximately 97%, at least approximately 98%, or at least approximately 99% identical thereto. DNMT3A containing a sequence with identity, the third repressor is DNMT3L containing a sequence variant that is at least approximately 70%, at least approximately 80%, at least approximately 85%, at least approximately 90%, at least approximately 91%, at least approximately 92%, at least approximately 93%, at least approximately 94%, at least approximately 95%, at least approximately 96%, at least approximately 97%, at least approximately 98%, or at least approximately 99% identical thereto, the fourth repressor is the sequence of sequence number 125, or at least approximately 70% identical thereto,The ADD contains sequences having at least approximately 80%, at least approximately 85%, at least approximately 90%, at least approximately 91%, at least approximately 92%, at least approximately 93%, at least approximately 94%, at least approximately 95%, at least approximately 96%, at least approximately 97%, at least approximately 98%, or at least approximately 99% identity, the NLS contains sequences independently selected from the group consisting of sequence numbers 30-97 in Tables 3 and 4, the L1 linker contains the sequence of sequence number 123, the L2 linker contains the sequence of sequence number 122, the L3A linker contains the sequence of sequence number 124, the L3B linker contains the sequence of sequence number 120, and the L4 linker contains the sequence of sequence number 1988 or sequence number 2130. In some embodiments of the long-term repressor fusion protein of configuration 6b, the linker sequence is independently selected from the group consisting of SEQ ID NOs. 98-124, 1823-1874, 1988, and 2130-2131 (exemplary sequences shown in Table 5). In some embodiments of the long-term repressor fusion protein of configuration 6b, the two copies of RD1 are different, with RD1a comprising the sequence of SEQ ID NO 130, or a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, or at least about 95% identity thereto, and RD1b comprising the sequence of SEQ ID NO 131, or a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, or at least about 95% identity thereto. In some embodiments of the long-term repressor fusion protein of configuration 6b, the two copies of RD1 are different, with RD1a comprising the sequence of SEQ ID NO: 130, or a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, or at least about 95% identity thereto, and RD1b comprising the sequence of SEQ ID NO: 135, or a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%,Or it includes a sequence that is at least about 95% identical. In some embodiments of the long-term repressor fusion protein of configuration 6b, the two copies of RD1 are different, with RD1a including the sequence of SEQ ID NO: 131, or a sequence that is at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, or at least about 95% identical thereto, and RD1b including the sequence of SEQ ID NO: 130, or a sequence that is at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, or at least about 95% identical thereto. In some embodiments of the long-term repressor fusion protein of configuration 6b, the two copies of RD1 are different, with RD1a comprising the sequence of SEQ ID NO: 131, or a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, or at least about 95% identity thereto, and RD1b comprising the sequence of SEQ ID NO: 135, or a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, or at least about 95% identity thereto. In some embodiments of the long-term repressor fusion protein of configuration 6b, the two copies of RD1 are different, with RD1a comprising the sequence of SEQ ID NO: 135, or a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, or at least about 95% identity thereto, and RD1b comprising the sequence of SEQ ID NO: 130, or a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, or at least about 95% identity thereto. In some embodiments of the long-term repressor fusion protein of configuration 6b, the two copies of RD1 are different, with RD1a comprising the sequence of SEQ ID NO: 135,Or it comprises a sequence that is at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, or at least about 95% identical thereto, and RD1b comprises the sequence of SEQ ID NO: 131, or a sequence that is at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, or at least about 95% identical thereto. In some embodiments of the long-term repressor fusion protein of configuration 6b, therein the two copies of RD1 are different, and RD1a and Rd1b are independently selected from the group consisting of sequences SEQ ID NOs: 132-134 and 136-138. In some embodiments of the long-term repressor fusion protein of configuration 6b, the two copies of RD1 are different, and RD1a and Rd1b are independently selected from the group consisting of sequences 130, 131, and 135. A schematic diagram of configuration 6b is shown in Figure 19. In some embodiments of the long-term repressor fusion protein of configuration 6b, the DNA-binding protein may be the non-catalyzed CasX sequence of sequence number 4. In some embodiments, the long-term repressor fusion protein is capable of forming an RNP with the gRNA of this disclosure, and is capable of binding to and repressing or silencing a gene target nucleic acid. [Table 5]
[0159] In some embodiments, the long-term repressor fusion protein comprises dCasX and optionally ADD, and constitutes configuration 1. In some embodiments, the long-term repressor protein comprises a sequence selected from the group consisting of SEQ ID NOs. 21903 to 21922, or a sequence having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least 99% identity thereto. In some embodiments of the long-term repressor fusion protein comprising dCasX and optionally ADD, and constituted as configuration 1, the long-term repressor protein comprises a sequence selected from the group consisting of SEQ ID NOs. 21903 to 21922. In some embodiments of the long-term repressor fusion protein comprising dCasX and optionally ADD, and constituted as configuration 1, the long-term repressor protein comprises a sequence selected from the group consisting of SEQ ID NOs. 21905 and 21914. In some embodiments of the long-term repressor fusion protein comprising dCasX and optionally ADD, and configured as component 1, the long-term repressor protein comprises a sequence selected from the group consisting of SEQ ID NOs: 21906 and 21915. In some embodiments of the long-term repressor fusion protein comprising dCasX and optionally ADD, and configured as component 1, the long-term repressor protein comprises a sequence selected from the group consisting of SEQ ID NOs: 21907 and 21916.
[0160] This disclosure provides long-term repressor fusion proteins of configuration 1, configuration 2, configuration 3, configuration 4, configuration 5, configuration 6a and configuration 6b, as shown in Figure 19 and described above, wherein the long-term repressor fusion proteins are capable of complexing with a gRNA having a targeting sequence complementary to the target nucleic acid of a gene in a cell to form an RNP. Upon binding of the RNP to the target nucleic acid of the targeted gene in the cell, the nucleic acid of the gene is epigenetically modified and the transcription of the gene is repressed. In some embodiments, the transcription of the gene is repressed by at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least 99%. In some embodiments, gene transcription is repressed in at least about 1%, at least about 2%, at least about 3%, at least about 4%, at least about 5%, at least about 6%, at least about 7%, at least about 8%, at least about 9%, or at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, or at least about 60% or more of cells in a population of cells. Most preferably, gene repression results in complete inhibition of gene expression, making the gene product undetectable. However, those skilled in the art will understand that incomplete inhibition may still be useful and desirable for various applications. In some embodiments, the repression of gene transcription is maintained for at least about 8 hours, at least about 1 day, at least about 7 days, at least 2 weeks, at least about 3 weeks, at least about 1 month, or at least about 2 months when assayed in an in vitro assay, including a cell-based assay. In some embodiments, the repression of gene transcription persists in target cells for at least about 7 days, at least about 2 weeks, at least about 3 weeks, at least about 1 month, at least about 2 months, at least about 3 months, at least about 4 months, at least about 5 months, or at least about 6 months or longer.In some embodiments, the use of long-term repressor fusion protein constructs 1, 4, 5, 6a, and 6b, when used in the LTRP:gRNA system of the embodiments, results in off-target methylation or off-target activity of less than approximately 10%, less than approximately 9%, less than approximately 8%, less than approximately 7%, less than approximately 6%, less than approximately 5%, less than approximately 4%, less than approximately 3%, less than approximately 2%, less than approximately 1%, less than 0.5%, or less than 0.1% in cells. In some embodiments, transcriptional repression in cells treated with the LTRP:gRNA system of the embodiments is heritable and stable through one or more cell divisions. In some embodiments, transcriptional repression is stable through 1, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 cell divisions or more. In some embodiments, transcriptional repression is assayed in in vitro assays, including cell-based assays, and transcriptional repression is compared to untreated cells or cells treated with an equivalent system in which the gRNA contains a non-targeting spacer. In some embodiments, transcriptional repression is assayed in vivo in cells obtained from subjects to whom a long-term repressor fusion protein and a gRNA with a targeting sequence complementary to the target nucleic acid of the gene in the cell are administered as a protein and gRNA, or as nucleic acids (e.g., gRNA and mRNA encoding the long-term repressor fusion protein), where the subjects are selected from a group consisting of mice, rats, pigs, non-human primates, and humans.
[0161] V. mRNA compositions encoding long-term repressor fusion proteins In another aspect, the Disclosure relates to messenger RNA (mRNA) compositions comprising the sequences of individual components, as well as the full-length mRNA sequences of the long-term repressor domain fusion protein constructs of the Disclosure. When used to express long-term repressor fusion proteins, the mRNA compositions are useful in transcriptional repression and epigenetic modification of genes. In some cases, the mRNA compositions are designed for use in certain delivery formulations, such as nanoparticles, such as synthetic nanoparticles or lipid nanoparticles (LNPs). The Disclosure also provides mRNA sequences of mRNA used in the compositions and methods used to design formulations for delivering the compositions. In some cases, the mRNA encoding the long-term repressor fusion protein may be co-formulated in nanoparticles with a gRNA containing a targeting sequence complementary to the target nucleic acid sequence of the gene to be transcriptionally repressed or silenced, wherein upon delivery of the nanoparticles to target cells, the long-term repressor domain fusion protein can be expressed from the mRNA and complexed with the gRNA as an RNP capable of binding to the target nucleic acid. In other cases, the mRNA encoding the long-term repressor domain fusion protein and gRNA may be formulated in separate nanoparticles and delivered separately or as a mixture.
[0162] In some embodiments, mRNA compositions are modified to yield one or more improved characteristics compared to unmodified mRNA encoding the same long-term repressor protein, and thus can have a significant impact on the efficacy of mRNA-based delivery. Exemplary, but not limited to, improved characteristics of mRNAs described herein include, compared to unmodified mRNA, improved expression upon delivery to cells, reduced immunogenicity, increased stability, and enhanced manufacturability. In some cases, the modification of mRNA yields improved characteristics that are at least about 1.1 to about 100,000 times better than unmodified mRNA. In some embodiments, the improved features of modified mRNA are improved by at least approximately 1.1 to 10,000 times compared to unmodified mRNA, at least approximately 1.1 to 1,000 times, at least approximately 1.1 to 500 times, at least approximately 1.1 to 400 times, at least approximately 1.1 to 300 times, at least approximately 1.1 to 200 times, at least approximately 1.1 to 100 times, at least approximately 1.1 to 50 times, at least approximately 1.1 to 40 times, at least approximately 1.1 to 30 times, at least approximately 1.1 to 20 times, at least approximately 1.1 to 10 times, at least approximately 1.1 to 9 times, and at least approximately 1. The improvement is 1 to approximately 8 times, at least 1.1 to approximately 7 times, at least 1.1 to approximately 6 times, at least 1.1 to approximately 5 times, at least 1.1 to approximately 4 times, at least 1.1 to approximately 3 times, at least 1.1 to approximately 2 times, at least 1.1 to approximately 1.5 times, at least 1.5 to approximately 3 times, at least 1.5 to approximately 4 times, at least 1.5 to approximately 5 times, at least 1.5 to approximately 10 times, at least 5 to approximately 10 times, at least 10 to approximately 20 times, at least 10 to approximately 30 times, at least 10 to approximately 50 times, or at least 10 to approximately 100 times. In some embodiments, the improved features of modified mRNA are at least 10 to approximately 1000 times better than those of unmodified mRNA.
[0163] Optimization of coding sequences and untranslated regions (UTRs) can be useful when delivering mRNA encoding a target protein, as opposed to DNA templates that will be transcribed into mRNA. DNA templates are long-lived, can replicate, and can produce many RNA transcripts throughout their lifetime. For DNA templates, the efficiency of transcription and mRNA precursor processing is a major determinant of protein expression levels. In contrast, mRNA generally has a much shorter half-life, typically in the range of hours, because it is vulnerable to degradation in the cytoplasm and cannot produce many copies of itself. As such, mRNA stability and translation efficiency can be important determinants of protein expression levels for mRNA-based delivery, and the effectiveness of mRNA-based delivery can be improved by enhancing specific sequences of the UTR and coding sequence that indicate mRNA stability and translation efficiency.
[0164] a.5' cap In some embodiments of the mRNA encoding the LTRP of this disclosure, the mRNA includes a 5' cap ligated at 5' to the 5'UTR of the mRNA sequence of any of the embodiments described herein. In some embodiments, the 5' cap is a 7-methylguanylate cap. In some embodiments, the 5' cap includes m7G(5')ppp(5')mAG. In other embodiments, the 5' cap includes m7G(5')ppp(5'(A, G(5')ppp(5')A or G(5')ppp(5')G. Exemplary caps are known in the art and are described, for example, in International Publication No. 2017 / 053297, the contents of which are incorporated herein by reference.
[0165] b.5' Untranslated Region (UTR) The 5'UTR of an mRNA molecule can be a determinant of both mRNA stability and how efficiently it is translated into protein. Specifically, in conjunction with the 5' cap structure, the 5'UTR acts as a binding site and recruitment platform for pre-translational complexes and additional regulatory proteins that can positively or negatively affect translation. Structures within the 5'UTR can enhance translation by recruiting initiation factors or other protein or RNA factors, degrade translation by physically blocking ribosome binding and scanning, and contribute to mRNA stability by influencing both hydrolysis and nuclease digestion.
[0166] Exemplary 5'UTR sequences for use in mRNA according to this disclosure are provided in Table 6A. Table 6A lists the RNA sequence, the RNA sequence with N1-methylpseudridine substituted in place of uridine, and the DNA sequence of the 5'UTR. [Table 6]
[0167] In some embodiments, the 5'UTR includes the sequence of sequence number 21831, or a sequence that is at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, or at least about 95% identical thereto. In some embodiments, the 5'UTR includes the sequence of sequence number 21831. In some embodiments, the 5'UTR includes the sequence of sequence number 21831. In some embodiments, the 5'UTR includes the sequence of sequence number 21842, or a sequence that is at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, or at least about 95% identical thereto. In some embodiments, the 5'UTR includes the sequence of sequence number 21842. In some embodiments, the 5'UTR consists of the sequence of sequence number 21842.
[0168] c.3'UTR 3'UTR sequences can have a significant impact on mRNA stability and translation efficiency, as well as determining both intracellular localization and tissue-specific expression. Factors influencing these properties include microRNA binding sites, AU-rich elements that recruit arrays of RNA-binding proteins, Pumilio binding elements, and other binding sites for RNA-binding proteins. While many of these interactions with 3'UTR are known to negatively affect stability or expression, some can enhance translation. The effects of 3'UTR sequences can be highly cell-type specific due to differential expression of microRNAs and RNA-binding proteins, but this provides an opportunity to manipulate tissue-specific expression in therapeutic mRNA. In some embodiments, the 3'UTR for use in mRNA of this disclosure is the mouse 3'UTR. In some embodiments, the 3'UTR is the mouse HBA gene 3'UTR shown in Table 6B.
[0169] Exemplary 3'UTR sequences of this disclosure are provided in Table 6B. Table 6B lists RNA sequences and RNA sequences with N1-methylpseudridine substituted in place of uridine. [Table 7]
[0170] In some embodiments, the 3'UTR includes the sequence of sequence number 21844, or a sequence that is at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, or at least about 95% identical thereto. In some embodiments, the 3'UTR consists of the sequence of sequence number 21844. In some embodiments, the 3'UTR consists of the sequence of sequence number 21844. In some embodiments, the 3'UTR includes the sequence of sequence number 21845, or a sequence that is at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, or at least about 95% identical thereto. In some embodiments, the 3'UTR consists of the sequence of sequence number 21845. In some embodiments, the 3'UTR consists of the sequence of sequence number 21845.
[0171] d. Poly(A) array The inclusion of a 3' poly(A) tail in mRNA sequences can contribute to mRNA stability and translation efficiency. Generally, longer poly(A) tails are associated with increased mRNA stability, thereby enabling their translation and promoting higher protein expression.
[0172] In some embodiments, the mRNA of the Disclosure comprises a poly(A) sequence having at least about 40, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 100, at least about 110, at least about 120, at least about 130, at least about 140, at least about 150, at least about 160, at least about 170, at least about 180, at least about 185, or at least about 190 adenine nucleotides. In some embodiments, the poly(A) sequence of the mRNA of the Disclosure comprises 80 adenine nucleotides. In some embodiments, the poly(A) sequence comprises the nucleic acid sequence AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA (Sequence ID 1961). In some embodiments, the poly(al) sequence includes the nucleic acid sequence AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAGCAUGCAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA (Sequence ID 21851).
[0173] Sequence modification of e mRNA In some embodiments, the mRNA sequences of this disclosure have been modified by codon optimization of the sequences encoding long-term repressor proteins using one or more parameters to enhance expression in target cells. Non-limiting examples of such parameters include codon use in human host cells (e.g., utilizing codon adaptation indices (CAIs)), codon use tables derived from biologics intended for therapeutic use, mRNA stability indices, or GC content. Methods of codon optimization and codon use in various organisms are known in the art. See, for example, www.genscript.com / tools / codon-frequency-table. In some embodiments, what is provided herein is mRNA sequences for long-term repressor protein constructs codon-optimized for expression in human cells. In other cases, modified mRNAs according to this disclosure may be produced using various naturally occurring or modified nucleosides.
[0174] In some embodiments, mRNA is a natural nucleoside (e.g., adenosine, guanosine, cytidine, uridine); nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3-methyladenosine, 5-methylcytidine, C-5 propynylcytidine, C-5 propynyluridine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-pripynyluridine, C5-pripynylcytidine, C5-methylcytidine, 2-aminoadenosine, 7-deazaadenosine, 7 -Deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, O(6)-methylguanine, pseudouridine (e.g., N-1-methyl-pseudridine), 2-thiouridine, and 2-thiocytidine); chemically modified bases; biologically modified bases (e.g., methylated bases); intercalated bases; modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose); and / or modified phosphate groups (e.g., phosphorothioates and 5'-N-phosphoramidite linkages), or comprising them. In some embodiments, mRNA contains one or more non-standard nucleotide residues. Non-standard nucleotide residues may include, for example, 5-methylcytidine ("5mC"), N1-methyl-pseudridine ("ψU"), and / or 2-thiouridine ("2sU"). In certain embodiments, one, more, or all uridine residues of the mRNA of this disclosure are replaced with N1-methyl-pseudouridine. In some embodiments, all uridine residues of the mRNA of this disclosure are replaced with N1-methyl-pseudouridine. For example, for consideration of such residues and their incorporation into mRNA, see U.S. Patent Application No. 8,278,036 or International Publication No. 2011012316, which are incorporated herein by reference. In some embodiments, the modification of mRNA results in improved properties of at least about 1.1 to about 100,000 times improvement compared to unmodified mRNA.In some embodiments, the improved features of modified mRNA are improved by at least approximately 1.1 to 10,000 times compared to unmodified mRNA, at least approximately 1.1 to 1,000 times, at least approximately 1.1 to 500 times, at least approximately 1.1 to 400 times, at least approximately 1.1 to 300 times, at least approximately 1.1 to 200 times, at least approximately 1.1 to 100 times, at least approximately 1.1 to 50 times, at least approximately 1.1 to 40 times, at least approximately 1.1 to 30 times, at least approximately 1.1 to 20 times, at least approximately 1.1 to 10 times, at least approximately 1.1 to 9 times, and at least approximately 1. The improvement is 1 to approximately 8 times, at least 1.1 to approximately 7 times, at least 1.1 to approximately 6 times, at least 1.1 to approximately 5 times, at least 1.1 to approximately 4 times, at least 1.1 to approximately 3 times, at least 1.1 to approximately 2 times, at least 1.1 to approximately 1.5 times, at least 1.5 to approximately 3 times, at least 1.5 to approximately 4 times, at least 1.5 to approximately 5 times, at least 1.5 to approximately 10 times, at least 5 to approximately 10 times, at least 10 to approximately 20 times, at least 10 to approximately 30 times, at least 10 to approximately 50 times, or at least 10 to approximately 100 times. In some embodiments, the improved features of modified mRNA are at least 10 to approximately 1000 times better than those of unmodified mRNA.
[0175] f.LTRP mRNA component sequence This disclosure provides mRNAs containing sequences encoding components used in long-term repressor fusion proteins described herein. In some embodiments, the mRNAs contain sequences encoding DNA-binding proteins, including TALE, ZF, and non-catalyzed CRISPR proteins. In some embodiments, the mRNAs contain sequences of SEQ ID NO: 2211 encoding dCasX 515 (SEQ ID NO: 6), or sequences having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto. In some embodiments, the mRNA includes the sequence of SEQ ID NO: 2213 or SEQ ID NO: 2214 encoding dCasX 812 (SEQ ID NO: 29), or a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto. In some embodiments, the mRNA includes the sequence of SEQ ID NO: 1948 or SEQ ID NO: 2405 encoding dCasX 491 (SEQ ID NO: 4), or a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto. In some embodiments, the mRNA includes a DNA-binding protein encoded by a sequence essentially derived from the sequence of SEQ ID NO: 2405 encoding dCasX 491 (SEQ ID NO: 4). In some embodiments, the mRNA includes the sequence of SEQ ID NO: 2407 or SEQ ID NO: 2408 encoding dCasX 676 (SEQ ID NO: 28), or a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto. The sequence encoding dCasX is as described above and is incorporated into the mRNA sequence encoding the long-term repressor fusion protein.
[0176] In some embodiments, the mRNA sequence encoding dCasX 491 (SEQ ID NO: 4) has a pseudouridine nucleoside that substitutes one or more or all of the uridines in the sequence (SEQ ID NO: 2406). In some embodiments, the mRNA sequence encoding dCasX 515 (SEQ ID NO: 6) has a pseudouridine nucleoside that substitutes one or more or all of the uridines in the sequence. In some embodiments, the mRNA sequence encoding dCasX 676 (SEQ ID NO: 28) has a pseudouridine nucleoside that substitutes one or more or all of the uridines in the sequence. In some embodiments, the mRNA sequence encoding dCasX 812 (SEQ ID NO: 29) has a pseudouridine nucleoside that substitutes one or more or all of the uridines in the sequence.
[0177] In some embodiments, the mRNA sequence encoding dCasX is selected from the group consisting of SEQ ID NOs: 1948, 2211, 2213-2214, and 2405-2408. In some embodiments, the mRNA sequence encoding dCasX has a pseudouridine nucleoside substituting one or more uridines. In some embodiments, the mRNA sequence encoding dCasX has a pseudouridine nucleoside substituting all uridines.
[0178] This disclosure provides mRNA sequences encoding the RD1 domain. In some embodiments, the sequence encoding RD1 includes a sequence selected from the group consisting of SEQ ID NOs. 1946, 21846, and 18637-20233, or a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, or at least about 95% identity thereto. In other embodiments, the sequence encoding RD1 includes a sequence selected from the group consisting of SEQ ID NOs. 18637-18731, or a sequence having at least about 70%, at least about 80%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In another embodiment, the sequence encoding RD1 includes a sequence selected from the group consisting of sequence numbers 18637 to 18645, or a sequence having at least about 70%, at least about 80%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In another embodiment, the sequence encoding RD1 includes the sequence of sequence number 18642, or a sequence having at least about 70%, at least about 80%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In another embodiment, the sequence encoding RD1 includes the sequence of sequence number 18638, or a sequence having at least about 70%, at least about 80%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto.In another embodiment, the sequence encoding RD1 includes the sequence of sequence number 18637, or a sequence having at least about 70%, at least about 80%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In some embodiments, the sequence encoding RD1 includes sequence numbers 18637, 18638, or 18642.
[0179] In some embodiments, the mRNA includes a sequence encoding a second repressor domain. In some embodiments, the second repressor domain includes DNMT3A. In some embodiments, the sequence encoding DNMT3A includes the sequence of SEQ ID NO: 1923, or a sequence having at least about 70%, at least about 80%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In some embodiments, the sequence encoding DNMT3A includes the sequence of SEQ ID NO: 1923.
[0180] In some embodiments, the mRNA includes a sequence encoding a third repressor domain. In some embodiments, the third repressor domain includes DNMT3L. In some embodiments, the sequence encoding DNMT3L includes the sequence of Sequence ID No. 1945, or a sequence having at least about 70%, at least about 80%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In some embodiments, the sequence encoding DNMT3L includes Sequence ID No. 1945.
[0181] In some embodiments, the mRNA includes a sequence encoding a fourth repressor domain. In some embodiments, the fourth repressor domain includes ADD. In some embodiments, the sequence encoding ADD includes the sequence of SEQ ID NO: 1954, or a sequence having at least about 70%, at least about 80%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In some embodiments, the fourth repressor domain includes ADD. In some embodiments, the sequence encoding ADD includes the sequence of SEQ ID NO: 1954.
[0182] In some embodiments, the mRNA of this disclosure includes a Kosack sequence. In some embodiments, the Kosack sequence includes GCCACCAUGG (SEQ ID NO: 21848). In some embodiments, the mRNA of this disclosure includes a Kosack sequence and two bases upstream of the NLS (to produce methionine and alanine upstream of the NLS). In some embodiments, the mRNA includes the sequence GCCACCAUGGCC (SEQ ID NO: 21832) between the 5'UTR and the sequence encoding the NLS.
[0183] In another embodiment, the mRNA includes an NLS. In some embodiments, the sequence encoding the NLS includes a sequence selected from the group consisting of SEQ ID NOs. 21833 and 21849, or a sequence having at least about 70%, at least about 80%, or at least about 90% identity thereto. In another embodiment, the mRNA includes a sequence selected from the group consisting of SEQ ID NOs. 21833 and 21849. In another embodiment, for example, in an embodiment containing an NLS with a long-term repressor protein greater than 1, the mRNA includes two or more sequences independently selected from the group consisting of SEQ ID NOs. 21833 and 21849.
[0184] In some embodiments, the mRNA includes a sequence encoding a linker of the long-term repressor fusion protein. In some embodiments, the linker encoding sequence includes a sequence selected from the group consisting of SEQ ID NOs. 21857 to 21873, or a sequence having at least about 80%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, and at least about 95% identity with them. In some embodiments, for example, in embodiments where the long-term repressor fusion protein includes more than one linker, the linker encoding sequence is independently selected from the group consisting of SEQ ID NOs. 21857 to 21873, or a sequence having at least about 80%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, and at least about 95% identity with them.
[0185] In some embodiments, the mRNA includes a sequence encoding a long-term repressor fusion protein of the configuration provided herein. In some embodiments, the mRNA encodes a long-term repressor fusion protein of configuration 1, and the mRNA includes a sequence selected from the group consisting of SEQ ID NOs. 2521-7311, or a sequence having at least about 80%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, and at least about 99% identity thereto. In some embodiments, the mRNA includes a sequence selected from the group consisting of SEQ ID NOs. 2521-7311.
[0186] In some embodiments, the mRNA includes a sequence encoding the long-term repressor fusion protein of composition 1, and a sequence encoding RD1 including the sequence of SEQ ID NO: 130. In some embodiments, the mRNA includes the sequence of SEQ ID NO: 18637 or 20234, or a sequence having at least about 80%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In some embodiments, the mRNA includes the sequences of SEQ ID NO: 18637 and 20234. In some embodiments, the mRNA includes a sequence encoding the long-term repressor fusion protein of composition 1, and a sequence encoding RD1 including the sequence of SEQ ID NO: 131. In some embodiments, the mRNA includes the sequence of SEQ ID NO: 18638 or 20235, or a sequence that is at least about 80%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical thereto. In some embodiments, the mRNA includes the sequences of SEQ ID NO: 18638 and 20235. In some embodiments, the mRNA includes the sequence encoding the long-term repressor fusion protein of composition 1, and the sequence encoding RD1, including SEQ ID NO: 132. In some embodiments, the mRNA includes the sequence of SEQ ID NO: 18639 or 20236, or a sequence that is at least about 80%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical thereto. In some embodiments, the mRNA includes the sequence encoding the long-term repressor fusion protein of composition 1, and the sequence encoding RD1, which includes the sequence of SEQ ID NO: 133.In some embodiments, the mRNA includes the sequence of SEQ ID NO: 18640 or 20237, or a sequence that is at least about 80%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical thereto. In some embodiments, the mRNA includes the sequence of SEQ ID NO: 18640 or 20237. In some embodiments, the mRNA includes the sequence encoding the long-term repressor fusion protein of composition 1, and includes the sequence encoding RD1, which includes the sequence of SEQ ID NO: 134. In some embodiments, the mRNA includes the sequence of SEQ ID NO: 18641 or 20238, or a sequence that is at least about 80%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical thereto. In some embodiments, the mRNA includes the sequence of SEQ ID NO: 18641 or 20238. In some embodiments, the mRNA includes the sequence encoding the long-term repressor fusion protein of composition 1, and includes the sequence encoding RD1, which includes the sequence of SEQ ID NO: 135. In some embodiments, the mRNA includes the sequence of SEQ ID NO: 18642 or 20239, or a sequence that is at least about 80%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical thereto. In some embodiments, the mRNA includes the sequence of SEQ ID NO: 18642 or 20239. In some embodiments, the mRNA includes the sequence encoding the long-term repressor fusion protein of composition 1, and includes the sequence encoding RD1, which includes the sequence of SEQ ID NO: 136.In some embodiments, the mRNA includes the sequence of SEQ ID NO: 18643 or 20240, or a sequence that is at least about 80%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical thereto. In some embodiments, the mRNA includes the sequence of SEQ ID NO: 18643 or 20240. In some embodiments, the mRNA includes the sequence encoding the long-term repressor fusion protein of composition 1, and includes the sequence encoding RD1, which includes the sequence of SEQ ID NO: 137. In some embodiments, the mRNA includes the sequence of SEQ ID NO: 18644 or 20241, or a sequence that is at least about 80%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical thereto. In some embodiments, the mRNA includes the sequence of SEQ ID NO: 18644 or 20241. In some embodiments, the mRNA includes the sequence encoding the long-term repressor fusion protein of composition 1, and includes the sequence encoding RD1, which includes the sequence of SEQ ID NO: 138. In some embodiments, the mRNA contains the sequence of SEQ ID NO: 18645 or 20242, or a sequence that is at least about 80%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical thereto. In some embodiments, the mRNA contains the sequence of SEQ ID NO: 18645 or 20242.
[0187] This disclosure provides mRNA encoding a long-term repressor fusion protein of configuration 5. In some embodiments, the mRNA includes a sequence selected from the group consisting of SEQ ID NOs. 8909 to 12102, or a sequence having at least about 80%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In some embodiments, the mRNA includes a sequence selected from the group consisting of SEQ ID NOs. 8909 to 12102.
[0188] This disclosure provides mRNA encoding a long-term repressor fusion protein of configuration 6a. In some embodiments, the mRNA includes a sequence selected from the group consisting of SEQ ID NOs. 15297 to 16893, or a sequence having at least about 80%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In some embodiments, the mRNA includes a sequence selected from the group consisting of SEQ ID NOs. 15297 to 16893.
[0189] This disclosure provides mRNA encoding a long-term repressor fusion protein of configuration 6b. In some embodiments, the mRNA comprises a sequence selected from the group consisting of SEQ ID NOs. 18491 to 18563, or a sequence having at least about 80%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In some embodiments, the mRNA comprises a sequence selected from the group consisting of SEQ ID NOs. 18491 to 18563.
[0190] In some embodiments of the mRNA encoding the long-term repressor fusion protein of the LTRP:gRNA system of this disclosure, upon delivery of the system and expression in cells, the long-term repressor fusion protein is capable of complexing with the gRNA and binding to the target DNA of the cell's targeted gene, resulting in repression or silencing of gene transcription in the cell.
[0191] VI. Guide nucleic acids In another aspect, the disclosure relates to a specifically designed guide ribonucleic acid (gRNA) comprising a scaffold and a targeting sequence ligated to (and thus hybridizable with) a target nucleic acid sequence of a gene. The gRNAs described herein can be used in conjunction with long-term repressor proteins and systems comprising them to repress the transcription of target nucleic acids in eukaryotic cells. As used herein, the term “gRNA” encompasses naturally occurring molecules and gRNA variants and includes chimeric gRNA variants containing domains from different gRNAs. The gRNAs of the disclosure comprise a scaffold and a targeting sequence ligated to the 3' end of the scaffold and complementary to the target nucleic acid of the cell.
[0192] In some embodiments, a system comprising an mRNA encoding a long-term repressor fusion protein containing a dCasX protein, and one or more gRNAs, forms a ribonucleoprotein (RNP) complex containing the LTRP and gRNA upon expression of the long-term repressor fusion protein in a cell, which can target and bind to a specific site in the target nucleic acid sequence of the gene to be repressed or silenced in the cell. The gRNA provides target specificity to the complex by including a targeting sequence (or "spacer") having a nucleotide sequence complementary to the sequence of the target nucleic acid sequence, while the long-term repressor fusion protein of the system provides site-specific activity, such as binding and transcriptional repression of the target gene, which is induced to a target site in the target nucleic acid sequence (e.g., stabilized at the target site) upon its association with the gRNA.
[0193] Embodiments of gRNA for use in transcriptional repression and / or epigenetic modification of target nucleic acids, as well as formulations of mRNA and gRNA, are described below.
[0194] a. Reference gRNA and gRNA variants As used herein, “reference gRNA” refers to a CRISPR guide ribonucleic acid containing a wild-type sequence of a naturally occurring gRNA. In some embodiments, the gRNA scaffold of this disclosure may be subjected to one or more mutagenesis methods, such as those described in International Publication No. 2023235818A2, International Publication No. 2022120095A1, and International Publication No. 2020247882A1, which are incorporated by reference herein, and which may include deep mutation evolution (DME), deep mutation scanning (DMS), error-prone PCR, cassette mutagenesis, random mutagenesis, staggered extension PCR, gene shuffling, domain swapping, or chemical modification, to generate one or more gRNA variants with enhanced or altered properties relative to the modified gRNA scaffold. The activity of a gRNA scaffold derived from a gRNA variant can be used as a benchmark to measure improvements in the function or other characteristics of the gRNA scaffold by comparing it to the activity of the gRNA variant.
[0195] Table 7 provides reference gRNA tracr sequences and scaffold sequences. In some embodiments, the disclosure provides gRNA variants having a scaffold that includes a sequence having one or more nucleotide modifications to any one of the reference gRNA sequences SEQ ID NOs. 1731-1743 in Table 7. [Table 8]
[0196] b.gRNA domains and their functions The gRNAs of the systems of this disclosure comprise two segments: a targeting sequence and a protein-binding segment. The targeting segment of the gRNA comprises a nucleotide sequence (referred interchangeably as a spacer, target factor, or targeting sequence) that is complementary to (and therefore hybridizes with) a specific sequence (target site) within a target nucleic acid sequence (e.g., a strand of double-stranded target DNA, target ssRNA, target ssDNA, etc.), as fully described below. In the context of this disclosure, the targeting sequence of the gRNA can bind to target nucleic acid sequences, including coding sequences, complements of coding sequences, non-coding sequences, and accessory elements. The protein-binding segment of the gRNA (or “activator” or “protein-binding sequence”) interacts with (e.g., binds to) the CasX protein as a complex to form an RNP (as fully described below). As used herein, “scaffold” refers to all parts of the guide, with the exception of the targeting sequence, which consists of several regions as fully described below. The properties and characteristics of both wild-type and variant CasX gRNAs are described in International Publication No. 2020247882A1, U.S. Patent Publication No. 20220220508A1, International Publication No. 2022120095A1, and International Publication No. 2023235818A2, which are incorporated herein by reference.
[0197] In the case of a reference gRNA, the gRNA occurs naturally as a dual guide RNA (dgRNA), in which the target factor portion and the activator factor portion each have a double-stranded segment that hybridizes with each other to form a double-stranded duplex (dsRNA double-stranded for gRNA). The terms “target factor” or “target factor RNA” are used herein to refer to the crRNA-like molecule (crRNA: “CRISPR RNA”) of the CasX dual guide RNA (and therefore of the CasX single guide RNA when the “activator” and “target factor” are linked together, for example, by intervening nucleotides). The crRNA has a tracrRNA, followed by a 5' region that anneals with the nucleotides of the targeting sequence. In the case of gRNAs for use in the systems of this disclosure, the scaffold is designed to consist of a single molecule, with the activator and target factor moieties covalently linked to each other (rather than hybridizing to each other), and may be referred to as "single-molecule gRNA," "single-guide RNA," "single-molecule guide RNA," "single-molecule guide RNA," or "sgRNA." All gRNA variants of this disclosure for use in systems are single-molecule versions.
[0198] Collectively, the assembled gRNAs of this disclosure comprise different structured regions, or domains, namely, RNA triples, scaffolding stem-loops, elongation stem-loops, pseudoknots, and, in embodiments of this disclosure, a targeting sequence that is specific to the target nucleic acid and located at the 3' end of the gRNA. The RNA triples, scaffolding stem-loops, pseudoknots, and elongation stem-loops, together with an unstructured triple-loop that cross-links the triple-stranded portions together, are referred to as the “scaffold” of the gRNA. In some cases, the scaffolding stem further comprises a bubble. In other cases, the scaffold further comprises a triple-loop region. In yet another case, the scaffold further comprises a 5' unstructured region. In some embodiments, the gRNA scaffolds of this disclosure for use in an LTRP:gRNA system comprise a scaffolding stem-loop having the sequence CCAGCGACUAUGUCGUAGUGG (SEQ ID NO: 1822), or a sequence with at least one, two, three, four, or five mismatches.
[0199] Each of the structured domains contributes to the establishment of the guide's overall RNA folding and maintains the guide's functionality, particularly its ability to appropriately complex with the dCasX protein. For example, the guide scaffold stem interacts with the helical I domain of the dCasX protein, while residues within the triple helix, triple loop, and pseudoknot stems interact with the OBD of the dCasX protein. Collectively, these interactions confer the guide's ability to bind to dCasX and form a stable RNP, while spacers (or targeting sequences) direct and define the RNP's specificity for binding to specific sequences of DNA.
[0200] Site-specific binding of a target nucleic acid sequence (e.g., genomic DNA) by the dCasX protein may occur at one or more locations (e.g., the target nucleic acid sequence) determined by base pair complementarity between the gRNA targeting sequence and the target nucleic acid sequence. Thus, for example, the gRNA of this disclosure has sequence complementarity to a target nucleic acid adjacent to a TC protospacer adjacency motif (PAM) motif or PAM sequence, such as ATC, CTC, GTC, or TTC, and can therefore hybridize with it. Since the targeting sequence of the guide sequence hybridizes with the sequence of the target nucleic acid sequence, the targeting sequence may be modified by the user to hybridize with a particular target nucleic acid sequence, insofar as the position of the PAM sequence is taken into consideration. In some embodiments, for the design of the targeting sequence, the target nucleic acid includes a PAM sequence located at 5' of the targeting sequence and includes at least a single nucleotide separating the PAM from the first nucleotide of the target nucleic acid that is complementary to the first nucleotide of the targeting sequence. This feature distinguishes the systems described herein from the Cas9 system and results in the ability of the systems disclosed herein to modify different positions in the DNA sequence compared to the Cas9 system. In some embodiments, the PAM is located on the non-targeting strand of the target region, i.e., the strand complementary to the target nucleic acid. In some embodiments, the targeting sequence of the gRNA is complementary to a one-nucleotide target nucleic acid sequence from the ATC PAM sequence. In some embodiments, the targeting sequence of the gRNA is complementary to a one-nucleotide target nucleic acid sequence from the CTC PAM sequence. In some embodiments, the targeting sequence of the gRNA is complementary to a one-nucleotide target nucleic acid sequence from the GTC PAM sequence. In some embodiments, the targeting sequence of the gRNA is complementary to a one-nucleotide target nucleic acid sequence from the TTC PAM sequence. By selecting the targeting sequence of the gRNA, a defined region of the sequence surrounding a target nucleic acid sequence, or a specific position within the target nucleic acid, can be suppressed using the LTRP:gRNA system described herein.
[0201] In some embodiments, the targeting sequence of the gRNA has a sequence of 15 to 22 nucleotides. In some embodiments, the targeting sequence has a sequence of 15, 16, 17, 18, 19, 20, 21, and 22 nucleotides. In some embodiments, the targeting sequence consists of 22 nucleotides. In some embodiments, the targeting sequence consists of 21 nucleotides. In some embodiments, the targeting sequence consists of 20 nucleotides. In some embodiments, the targeting sequence consists of 19 nucleotides. In some embodiments, the targeting sequence consists of 18 nucleotides. In some embodiments, the targeting sequence consists of 17 nucleotides. In some embodiments, the targeting sequence consists of 16 nucleotides. In some embodiments, the targeting sequence consists of 15 nucleotides. By selecting the targeting sequence of the gRNA, a defined region of the target nucleic acid sequence can be repressively and / or epigenetically modified using the LTRP:gRNA system described herein.
[0202] The gene repressor systems of this disclosure may be designed to target any region of a gene or gene region whose transcriptional repression is desired, or any region proximal thereto. When an entire gene is to be repressed, the disclosure intends to design a guide containing a targeting sequence complementary to the transcription start site (TSS) or a sequence proximal thereto. TSS selection occurs at different locations within the promoter region, depending on the promoter sequence and the initiation substrate concentration. The core promoter serves as a binding platform for the transcription mechanism, including Pol II and its associated general transcription factors (GTFs) (Haberle, V. et al. Eukaryotic core promoters and the functional basis of transcription initiation (Nat Rev Mol Cell)). Biol. 19(10):621(2018)). Variability in TSS selection has been proposed to include DNA "scrunching" and "anti-scrunching," characterized by (i) forward and reverse movement of the RNA polymerase leading edge, but not the trailing edge, relative to the DNA, and (ii) expansion and contraction of the transcription bubble. In some embodiments, the target nucleic acid sequence bound by the RNP of the LTRP:gRNA system is within 1.5 kilobases (kb) of the transcription start site (TSS) in the gene. In some embodiments, the target nucleic acid sequence bound by the RNP of the system The target nucleic acid sequence is located within 20 base pairs (bp), 50 bp, 100 bp, 150 bp, 200 bp, 250 bp, 500 bp, 1 kb, or 1.5 kb upstream of the gene's TSS. In some embodiments, the target nucleic acid sequence bound by the system's RNP is located within 20 bp, 50 bp, 100 bp, 150 bp, 200 bp, 250 bp, 500 bp, 1 kb, or 1.5 kb downstream of the gene's TSS. In some embodiments, the target nucleic acid sequence bound by the system's RNP is located within 500 bp upstream to 500 bp downstream of the gene's TSS, or 300 bp upstream to 300 bp downstream, or 100 bp upstream to 100 bp downstream.In some embodiments, the target nucleic acid sequence bound by the system's RNP is within 20 bp, 50 bp, 100 bp, 150 bp, 200 bp, 250 bp, 500 bp, or 1 kb of the gene's enhancer. In some embodiments, the target nucleic acid sequence bound by the system's RNP is within 1 kb of the 3' region relative to the 5' untranslated region of the gene. In other embodiments, the target nucleic acid sequence bound by the system's RNP is within the open reading frame of the gene, including any introns. In some embodiments, the targeting sequence of the system's gRNA is complementary to an exon of the gene. In certain embodiments, the targeting sequence of the system's gRNA is complementary to exon 1 of the gene. In other embodiments, the targeting sequence of the system's gRNA is complementary to an intron of the gene. In other embodiments, the targeting sequence of the system's gRNA is complementary to an intron-exon junction of the gene. In other embodiments, the targeting sequence of the gRNA system of the Disclosure is complementary to the intron-exon junction of the gene. In other embodiments, the targeting sequence of the system of the Disclosure is complementary to the regulatory element of the gene. In other embodiments, the targeting sequence of the gRNA system of the Disclosure is complementary to the sequence of the intergenetic region of the gene. In other embodiments, the targeting sequence of the gRNA system of the Disclosure is specific to the junctions of the exons, introns, and / or regulatory elements of the gene. Where the targeting sequence is specific to a regulatory element, such regulatory elements include, but are not limited to, regions containing promoter regions, enhancer regions, intergenetic regions, 5' untranslated regions (5'UTR), 3' untranslated regions (3'UTR), conserved elements, and cis-regulatory elements. The promoter region is intended to contain nucleotides within 5 kb of the start of the coding sequence, or, in the case of a gene enhancer element or conserved element, may be thousands of bp, hundreds of thousands of bp, or even millions of bp away from the coding sequence of the gene. As described above, the target is intended to suppress the gene containing the target nucleic acid, so that the gene product is not expressed in the cell or is expressed at a lower level.In some embodiments, upon binding of the RNP of the system to the binding site of a target nucleic acid, the system can repress the transcription of the gene at the 5' position relative to the RNP binding site. In other embodiments, upon binding of the RNP of the system to the binding site of a target nucleic acid, the system can repress the transcription of the gene at the 3' position relative to the RNP binding site.
[0203] c.gRNA modification In another embodiment, the Disclosure relates to gRNA variants for use in the System of the Disclosure, which include modifications compared to the reference gRNA from which the gRNA is derived. In some embodiments, the gRNA variants for use in the System of the Disclosure include one or more nucleotide substitutions, insertions, deletions, or exchanges or replaced domains compared to the gRNA sequence of the Disclosure, which improve the characteristics compared to the reference gRNA. Exemplary regions and exchange regions or domains for modification include RNA triple helices, pseudoknots, scaffolding stem-loops, and elongation stem-loops. In some embodiments, the gRNA variants of the Disclosure include at least one exchange region from a different gRNA, resulting in a chimeric gRNA. A representative example of such a chimeric gRNA is guide 316 (sequence number 1746), in which the elongation stem loop of gRNA scaffold 235 (sequence number 1745) is replaced with the elongation stem loop of gRNA scaffold 174 (sequence number 1744), and the resulting 316 variant retains the ability to form RNPs with dCasX and long-term repressor fusion proteins and exhibits improved properties compared to parent 235 when evaluated in in vitro or in vivo assays under equivalent conditions.
[0204] All gRNAs that possess one or more improved functions, features, or add one or more novel functions, while retaining the functional property of being able to form a complex with a long-term repressor fusion protein and induce a ribonucleoprotein holoRNP complex to a target nucleic acid, when compared to the gRNA scaffold variant from which they are derived, are envisioned within the scope of this disclosure. In some embodiments, the gRNA has improved features selected from the group consisting of increased pseudoknot stem stability, increased triple-stranded region stability, increased scaffold stem stability, elongation stem stability, decreased off-target folding intermediates, increased binding affinity to repressor fusion proteins, and increased transcriptional repressive activity when complexed with a repressor fusion protein, or any combination thereof. In some of the aforementioned cases, the improved features are evaluated in an in vitro assay, including the assay of the Examples. In other cases, the improved features are evaluated in vivo.
[0205] Table 8 provides exemplary gRNA variant scaffold sequences of the present disclosure that can be used as gRNA scaffolds or for generating gRNAs for use in the LTRP:gRNA system of the present disclosure. In some embodiments, the gRNA variant scaffold for use in the system comprises one of the sequences of SEQ ID NOs. 1744–1746 listed in Table 8, or a sequence having at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity therewith, wherein the gRNA variant retains the ability to form RNPs with dCasX of the present disclosure. In other embodiments, the gRNA variant scaffold for use in the LTRP:gRNA system comprises one of the sequences of SEQ ID NOs. 1744–1746 therewith, wherein the gRNA variant retains the ability to form RNPs with the long-term repressor fusion protein of the present disclosure. In embodiments where the vector contains a DNA coding sequence for gRNA, it will be understood that thymine (T) bases can be substituted for uracil (U) bases in any of the gRNA sequence embodiments described herein. Similarly, any RNA sequence disclosed herein can be encoded by DNA in which uracil bases are substituted with thymine. In some embodiments, this disclosure provides chemically modified gRNA variants, as described below. In some embodiments, the gRNA is chemically modified and includes a scaffold containing the sequences of SEQ ID NOs. 1744–1746. [Table 9]
[0206] Additional gRNA variants considered for use in the systems of this disclosure are selected from the group consisting of SEQ ID NOs: 1747–1821. In some embodiments, the gRNA is chemically modified and comprises a scaffold containing the sequences of SEQ ID NOs: 1747–1821.
[0207] Guide scaffolds can be prepared by several methods, including recombinantly or by solid-phase RNA synthesis. However, while scaffold length can affect manufacturability when using solid-phase RNA synthesis, longer lengths have resulted in increased manufacturing costs, decreased purity and yield, and higher synthesis failure rates. For use in particulate formulations, such as lipid nanoparticle (LNP) formulations, solid-phase RNA synthesis of scaffolds is preferred to produce the quantities required for commercial development. Previous experiments had identified gRNA scaffold 235 (SEQ ID NO: 1745) as having enhanced properties compared to gRNA scaffold 174 (SEQ ID NO: 1744), but its increased length (in nucleotides) made its use for LNP formulations potentially problematic due to synthetic manufacturing constraints. Therefore, alternative sequences were sought. In some embodiments, this disclosure provides gRNA variant scaffolds with improved manufacturability compared to the gRNA scaffolds from which they are derived. In some embodiments, the disclosure provides gRNAs in which the gRNA scaffold and the ligated targeting sequence have sequences of less than approximately 115 nucleotides, less than approximately 110 nucleotides, or less than approximately 100 nucleotides. In certain embodiments, the 316 gRNA scaffold (SEQ ID NO: 1746) has a shorter sequence compared to the 235 scaffold from which it is derived. The 316 gRNA scaffold was designed so that the scaffold 235 sequence is modified by domain exchange, in which the elongation stem loop of scaffold 174 replaces the elongation stem loop of scaffold 235, resulting in a chimeric gRNA scaffold 316 having the sequence ACUGGCGCUUCUAUCUGAUUACUCUGAGCGCCAUCACCAGCGACUAUGUCGUAGUGUGGUAAAGCUCCCUCUUCGGAGGGAGCAUCAGAG (SEQ ID NO: 1746), which is 89 nucleotides compared to the 99 nucleotides of gRNA scaffold 235. The resulting 316 scaffolds had further advantages in that the elongated stem-loops did not contain CpG motifs; and possessed enhanced properties that conferred reduced potential to induce an immune response.In some embodiments, shorter sequence lengths of the 316 scaffolds confer higher fidelity in the ability to synthetically construct guides with accurate and complete sequences, as well as enhanced ability to be successfully incorporated into LNPs. In some embodiments, this disclosure provides chemically modified gRNA316 variants, as described below.
[0208] d. Chemically modified gRNA In some embodiments, the gRNA has one or more chemical modifications. In some embodiments, the chemical modification is the addition of a 2'O-methyl group to one or more nucleotides in the sequence. In some embodiments, one or more nucleotides on each end of the gRNA are modified by the addition of a 2'O-methyl group. In some embodiments, the chemical modification is the substitution of a phosphorothioate bond between two or more nucleotides in the sequence. In some embodiments, the chemical modification is the substitution of a phosphorothioate bond between two or more nucleotides on each end of the gRNA. In some embodiments, the gRNA includes the substitution of a phosphorothioate bond between two or more nucleotides located at the 5' end, 3' end, or 1, 2, 3, or 4 nucleotides from both ends of the gRNA. In some embodiments, the gRNA includes the addition of a 2'O-methyl group to one or more nucleotides in the gRNA. In some embodiments, one or more nucleotides located at the 5' end, 3' end, or 1, 2, 3, or 4 nucleotides from both ends of the gRNA are modified by the addition of a 2'O-methyl group. In some embodiments, the first one, two, or three nucleotides at the 5' end of the scaffold (i.e., A, C, and U in the case of gRNA 174, 235, and 316) are modified by the addition of a 2'O-methyl group, and each of the modified nucleotides is ligated to an adjacent nucleotide by a phosphorothioate bond. Similarly, the last one, two, or three nucleotides at the 3' end of the targeted sequence ligated to the 3' end of the scaffold are similarly modified. In some embodiments, the gRNAs containing one or more chemical modifications contain a sequence selected from the group consisting of sequences 2136-2144, 2146-2154, and 2156-2164, or a sequence having at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, and at least about 99% sequence identity thereto.In some embodiments, the chemically modified gRNAs include the scaffolds of SEQ ID NOs. 2136-2144, 2146-2154, and 2156-2164, i.e., the sequences of SEQ ID NOs. 2136-2144, 2146-2154, and 2156-2164 without the 20-nucleotide spacer at the 3' end, which is represented as an undefined nucleotide in the aforementioned sequences. Schematic diagrams of the structures of gRNA variants 174, 235, and 316 are shown in Figures 23A-23C, respectively, and schematic diagrams of the chemically modified gRNAs are shown in Figures 22, 28A, and 28B, respectively. In some embodiments, the chemically modified gRNAs exhibit improved stability compared to the unmodified gRNAs.
[0209] e. Complex formation using long-term repressor fusion proteins Upon delivery or expression of the system's components in target cells, gRNA variants can conjugate with long-term repressor fusion proteins as RNPs and bind to target nucleic acids targeted by the gRNA's targeting sequence. In some embodiments, gRNA variants have an improved ability to form RNP complexes with long-term repressor fusion proteins when compared to a reference gRNA or another gRNA variant from which it is induced. By improving ribonucleoprotein complex formation, the efficiency of assembly of functional RNPs can be improved in some embodiments. In some embodiments, more than 90%, more than 93%, more than 95%, more than 96%, more than 97%, more than 98%, or more than 99% of the RNPs containing the gRNA variant and its targeting sequence are capable of gene repression of target nucleic acids.
[0210] VII. Polynucleotides and Vectors This disclosure provides polynucleotides encoding long-term repressor fusion proteins and / or gRNAs. This disclosure provides polynucleotides encoding mRNAs encoding long-term repressor fusion proteins, for example, DNA polynucleotides encoding the corresponding mRNAs.
[0211] Long-term repressor fusion proteins or mRNA encoding the long-term repressor fusion proteins of this disclosure can be prepared by in vitro synthesis using conventional methods known in the art. Various commercial synthesis apparatuses are available, e.g., automated synthesis apparatuses by Applied Biosystems, Inc., Beckman, etc. By using a synthesis apparatus, naturally occurring amino acids or nucleotides (where applicable) can be substituted with non-natural amino acids or nucleotides. The specific sequence and mode of preparation are determined by convenience, cost-effectiveness, required purity, and similar factors. gRNA can also be produced synthetically; for example, by using T7 RNA polymerase systems known in the art.
[0212] Long-term repressor fusion proteins and / or gRNAs can also be prepared by recombinantly producing polynucleotide sequences encoding any of the long-term repressor fusion proteins or gRNAs of the embodiments described herein, by incorporating the encoding gene into an expression vector suitable for host cells using standard recombination techniques known in the art. For the production of the encoded long-term repressor fusion proteins and / or gRNAs of any of the embodiments described herein, the method comprises transforming suitable host cells with an expression vector containing the encoding polynucleotide, and culturing the host cells under conditions that enable or result in the expression or transcription of the long-term repressor fusion proteins or gRNAs resulting from any of the embodiments described herein, which are then recovered by the method described herein, or by standard purification methods known in the art, or as described in the examples. The polynucleotides and expression vectors of this disclosure are prepared using standard recombination techniques in molecular biology.
[0213] The long-term repressor fusion proteins and / or gRNAs of this disclosure may also be isolated and purified according to conventional methods of recombinant synthesis. Lysates may be prepared from the expression host, but the lysates may be purified using high-performance liquid chromatography (HPLC), exclusion chromatography, gel electrophoresis, affinity chromatography, or other purification techniques. For the majority, the compositions used will contain 50% by weight or more of the desired product, more commonly 75% by weight or more, preferably 95% by weight or more, and for therapeutic purposes, usually 99.5% by weight or more, with respect to the method of product preparation and the impurities associated with its purification. Typically, the percentages are based on total protein. Thus, in some cases, the long-term repressor fusion proteins or gRNAs of this disclosure will be at least 80% pure, at least 85% pure, at least 90% pure, at least 95% pure, at least 98% pure, or at least 99% pure (e.g., free from impurities or other macromolecules).
[0214] In addition, this disclosure provides vectors comprising a repressor fusion protein and a polynucleotide encoding the gRNA described herein. In some cases, the vector is used for the expression and recovery of CasX and gRNA components of the LTRP:gRNA system when the system is delivered as a repressor fusion protein and gRNA, or as an RNP. In other cases, the vector is used for the delivery of the encoding polynucleotide to target cells for transcriptional repression and / or epigenetic modification of the target nucleic acid, as described more fully below. In some embodiments, the sequences encoding the long-term repressor fusion protein and gRNA are templated on the same vector. In some embodiments, the sequences encoding the long-term repressor fusion protein and gRNA are templated on different vectors. Suitable vectors are described, for example, in International Publication No. 2023235818A2, International Publication No. 2022120095A1, and International Publication No. 2020247882A1, which are incorporated herein by reference. As described in International Publication Nos. 2023235818A2, 2022120095A1, and 2020247882A1, depending on the host / vector system used, any of a number of suitable transcriptional and translational regulatory elements, including constitutive and inducible promoters, transcriptional enhancer elements, and transcriptional terminators, may be used in the expression vector.
[0215] In some embodiments, the Disclosure provides polynucleotide sequences encoding long-term repressor fusion proteins of any embodiment described herein, comprising long-term repressor fusion proteins comprising sequences of SEQ ID NOs. 1883-1903, 1909-1912, or 1915-1924, or sequences having at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto. In some embodiments, the polynucleotide comprises sequences encoding long-term repressor fusion proteins comprising sequences of SEQ ID NOs. 1883-1903, 1909-1912, or 1915-1924. In some embodiments, the polynucleotide comprises sequences of mRNA encoding long-term repressor fusion proteins of any embodiment described herein for use in a particle system for delivery to cells. In some embodiments, mRNA sequences encoding long-term repressor fusion proteins for use in LNP particle formulations for delivery to cells. In certain embodiments, the disclosure provides gRNA and mRNA sequences encoding long-term repressor fusion proteins for use in LNP particle formulations for delivery to cells, which are described more fully below.
[0216] In some embodiments, the Disclosure provides isolated polynucleotide sequences encoding gRNA variants of any of the embodiments described herein. In some embodiments, the Disclosure provides polynucleotides encoding gRNAs comprising scaffold sequences including SEQ ID NOs. 1744-1746, or sequences having at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity therewith, wherein the expressed gRNA variant retains the ability to form RNPs with repressor fusion proteins. In the embodiments described above, the gRNA further comprises a targeting sequence complementary to the target nucleic acid of the gene to be repressed.
[0217] In some embodiments, this disclosure relates to methods for producing polynucleotide sequences encoding long-term repressor fusion proteins or gRNAs, including variants thereof, of any of the embodiments described herein, and methods for expressing proteins or RNAs transcribed by these polynucleotide sequences. Generally, the methods include producing polynucleotide sequences encoding long-term repressor fusion proteins or gRNAs of any of the embodiments described herein, and incorporating the encoding gene into an expression vector. In some embodiments, the vector is designed for transduction of cells for transcriptional repression and / or epigenetic modification of a target nucleic acid. Such vectors may include retroviral vectors, lentiviral vectors, adenovirus vectors, adeno-associated virus (AAV) vectors, herpes simplex virus (HSV) vectors, plasmids, minicircles, nanoplasmides, DNA vectors, and RNA vectors. In other embodiments, the expression vector is designed for the production of repressor fusion proteins and mRNAs or gRNAs encoding long-term repressor fusion proteins, either in a cell-free system or in host cells. For the production of the long-term repressor fusion protein or gRNA encoded in any of the embodiments described herein in host cells, the method comprises transforming suitable host cells with an expression vector comprising the coding polynucleotide, and culturing the transformed host cells under conditions that enable or allow the expression or transcription of the long-term repressor fusion protein or gRNA resulting from any of the embodiments described herein, thereby producing the long-term repressor fusion protein or gRNA, which is recovered by the method described herein (e.g., in the examples below) or by standard purification methods known in the art. The polynucleotides and expression vectors of this disclosure are prepared using standard recombination techniques in molecular biology.
[0218] In accordance with the present invention, nucleic acid sequences encoding long-term repressor fusion proteins or gRNAs of any embodiment described herein are used to generate recombinant nucleic acid molecules directed for expression in suitable host cells. Several cloning strategies are suitable for carrying out the present disclosure, many of which are used to generate constructs containing genes encoding the long-term repressor fusion proteins or gRNAs of the present disclosure, or their complements. In some embodiments, the cloning strategy is used to construct a gene encoding a construct containing nucleotides encoding the long-term repressor fusion protein or gRNA. In some embodiments, the gene is used, for example, as part of a vector to transform host cells for gene expression, e.g., expression of the long-term repressor fusion protein or gRNA.
[0219] One approach involves first preparing a construct that includes a DNA sequence encoding a long-term repressor fusion protein or gRNA. Exemplary methods for preparing such constructs are described in the examples. The construct is then used to create an expression vector suitable for transforming host cells, such as prokaryotic or eukaryotic host cells, for the expression and recovery of the protein construct, in the case of the repressor fusion protein or gRNA. If desired, the host cell is E. coli. In other embodiments, the cell is a eukaryotic cell. The eukaryotic host cells can be selected from baby hamster kidney fibroblast (BHK) cells, human embryonic kidney 293 (HEK293), human embryonic kidney 293T (HEK293T), NS0 cells, SP2 / 0 cells, YO myeloma cells, P3X63 mouse myeloma cells, PER cells, PER.C6 cells, hybridoma cells, NIH3T3 cells, CV-1 (monkey) (COS) derived from the SV40 gene material, HeLa, Chinese hamster ovary (CHO), yeast cells, or other eukaryotic cells known in the art that are suitable for the production of recombinant products. Exemplary methods for the preparation of expression vectors, transformation of host cells, and expression and recovery of long-term repressor fusion proteins or gRNAs are described in the examples.
[0220] Genes encoding long-term repressor fusion proteins or gRNA constructs can be prepared in one or more steps, either entirely synthetically or by synthesis in combination with enzymatic processes, such as restriction enzyme-mediated cloning, PCR, and duplication extension, including methods more fully described in the examples. The methods disclosed herein can be used, for example, to ligate sequences of polynucleotides encoding various components into a gene of a desired sequence. Genes encoding polypeptide compositions are assembled from oligonucleotides using standard gene synthesis techniques.
[0221] In some embodiments, the nucleotide sequence encoding the long-term repressor fusion protein is codon-optimized. This type of optimization may involve mutations in the coding nucleotide sequence to mimic the codon priority of the intended host organism or cell while encoding the same protein, as well as other parameters including the codon adaptation index (CAI), a codon use table derived from the biologic intended for use as a therapeutic agent, the mRNA stability index, or GC content. Thus, the codons can be altered, but the encoded protein remains immutable. For example, if the intended target cells of the long-term repressor fusion protein are human cells, a human codon-optimized long-term repressor fusion protein coding nucleotide sequence may be used. As another non-limiting example, if the intended host cells are mouse cells, then a mouse codon-optimized long-term repressor fusion protein coding nucleotide sequence may be generated. Gene design can be carried out using algorithms that optimize codon use and amino acid composition suitable for the host cells utilized in the production of the long-term repressor fusion protein or gRNA. In one method of the present disclosure, a library of polynucleotides encoding components of a construct is prepared and then assembled as described above. The resulting gene is then assembled and used to transform host cells and produce and recover long-term repressor fusion proteins or gRNA compositions for characterization or use in modification of target nucleic acids, as described herein.
[0222] In some embodiments, the nucleotide sequence encoding the long-term repressor fusion protein or gRNA is operably ligated to a regulatory element, such as a transcriptional regulatory element, such as a promoter. In some embodiments, the promoter is a constitutively active promoter. In some embodiments, the promoter is a regulatory promoter. In some embodiments, the promoter is an inductive promoter. In some embodiments, the promoter is a tissue-specific promoter. In some embodiments, the promoter is a cell-type-specific promoter. In some embodiments, the transcriptional regulatory element (e.g., promoter) is functional in a targeted cell type or targeted cell population. For example, in some embodiments, the transcriptional regulatory element may be functional in eukaryotic cells, such as hepatocytes or hepatic sinusoidal endothelial cells.
[0223] Non-limiting examples of Pol II promoters operably linked to polynucleotides encoding the long-term repressor fusion proteins of this disclosure include, but are not limited to, EF-1 alpha, EF-1 alpha core promoter, Jens Tornoe (JeT), promoter from cytomegalovirus (CMV), CMV early-stage promoter (CMVIE), CMV enhancer, herpes simplex virus (HSV) thymidine kinase, early and late simian virus 40 (SV40), SV40 enhancer, long terminal repeat (LTR) from retrovirus, mouse metallothionein-I, and adenovirus major late-stage promoter (Ad MLP), CMV promoter full-length promoter, minimal CMV promoter, chicken β-actin promoter (CBA), CBA hybrid (CBh), chicken β-actin promoter with cytomegalovirus enhancer (CB7), chicken β-actin promoter and rabbit β-globin splice receptor site fusion (CAG), Roussarcoma virus (RSV) promoter, HIV-Ltr promoter, hPGK promoter, HSVTK promoter, 7SK promoter, Mini-TK promoter, human synapsin I (SYN) promoter conferring neuron-specific expression, beta-actin promoter, supercore promoter 1 (SCP1), Mecp2 promoter for selective expression in neurons, minimal IL-2 promoter, Roussarcoma virus enhancer / promoter (single), splenic fossaforming virus long-term repeat (LTR) promoter, TBG promoter, promoter from human thyroxine-binding globulin gene (liver-specific), PGK promoter, human ubiquitin C promoter (UBC), UCOE promoter (promoter for HNRPA2B1-CBX3), synthetic CAG promoter, histone H2 promoter, histone H3 promoter This includes the U1a1 micronuclear RNA promoter (226nt), U1a1 micronuclear RNA promoter (226nt), U1b2 micronuclear RNA promoter (246nt), GUSB promoter, CBh promoter, rhodopsin (Rho) promoter, silencing-prone splenic fossilizing virus (SFFV) promoter, human H1 promoter (H1), POL1 promoter, TTR minimal enhancer / promoter, β-kinesin promoter, mouse mammary tumor virus long-terminal repeat (LTR) promoter, human eukaryotic initiation factor 4A (EIF4A1) promoter, ROSA26 promoter, glyceraldehyde 3-phosphate dehydrogenase (GAPDH) promoter, tRNA promoter, as well as the aforementioned cleavage versions and sequence variants. In certain embodiments, the Pol II promoter is EF-1 alpha, in which the promoter enhances transfection efficiency, CRISPR nuclease transgene transcription or expression, the percentage of expression-positive clones, and the copy number of the episomal vector in long-term culture.
[0224] Non-limiting examples of Pol III promoters operably ligated to polynucleotides encoding gRNA variants of this disclosure include, but are not limited to, U6, mini-U6, U6 cleavage promoters, 7SK, and H1 variants, BiH1 (bidirectional H1 promoter), BiU6, Bi7SK, BiH1 (bidirectional U6, 7SK, and H1 promoters), gorilla U6, rhesus monkey U6, human 7SK, human H1 promoter, and their cleavage versions and sequence variants. In the embodiments described above, the Pol III promoter enhances gRNA transcription. In certain embodiments, the Pol III promoter is U6, in which the promoter enhances gRNA expression. Experimental details and data for the use of such promoters are provided in the examples.
[0225] The selection of suitable vectors and promoters is well within the normal skill level of the art, because it concerns controlling expression. Expression vectors may also contain ribosome binding sites for translation initiation and transcription terminators. Expression vectors may also contain appropriate sequences for amplifying expression. Expression vectors may also contain nucleotide sequences encoding protein tags (e.g., 6xHis tags, hemagglutinin tags, fluorescent proteins, etc.) which can be fused to long-term repressor fusion proteins, thus resulting in chimeric proteins used for purification or detection.
[0226] The recombinant expression vectors of this disclosure may also contain elements that promote robust expression of the proteins and gRNAs of this disclosure. For example, the recombinant expression vector may contain one or more polyadenylation (Poly(A)), intron sequences, or post-transcriptional regulatory elements, such as the Woodchuck hepatitis post-transcriptional regulatory element (PTRE). Exemplary Poly(A) sequences include the hGH Poly(A) signal (short), the HSV TK Poly(A) signal, the synthetic polyadenylation signal, the SV40 Poly(A) signal, the β-globin Poly(A) signal, and the split Poly(A) sequence with an SphI restriction site between two 60A stretches (SEQ ID NO: 21851), and similar sequences. Those skilled in the art will be able to select appropriate elements to include in the recombinant expression vectors described herein.
[0227] Polynucleotides encoding long-term repressor fusion proteins or gRNA sequences can be individually cloned into expression vectors. The selection of suitable vectors and promoters is well within the realm of normal skill in the art, as it concerns, for example, controlling expression to repress gene expression and / or epigenetic modifications. Expression vectors may also contain ribosome binding sites and transcriptional terminators for translation initiation. Expression vectors may also contain appropriate sequences for amplification of expression.
[0228] Nucleic acid sequences are inserted into vectors by various procedures. Generally, DNA is inserted into a suitable restriction endonuclease site using techniques known in the art. Vector components generally include, but are not limited to, one or more signal sequences, origins of replication, one or more marker genes, enhancer elements, promoters, and transcription termination sequences. The construction of a suitable vector containing one or more of these components employs standard ligation techniques known to those skilled in the art. Such techniques are well known in the art and are well described in the scientific and patent literature. A variety of vectors are publicly available. Vectors may be in the form of plasmids, cosmids, viral particles, or phages that are conveniently provided for recombinant DNA procedures, and the choice of vector often depends on the host cell into which it is introduced. Thus, vectors may be autonomously replicating vectors, i.e., vectors that exist as extrachromosomal entities, and their replication is independent of chromosomal replication, such as plasmids. Alternatively, when introduced into a host cell, a vector may be integrated into the host cell genome and replicate together with the chromosome into which it is integrated. Once introduced into suitable host cells, the expression of the long-term repressor fusion protein or gRNA can be determined using any nucleic acid or protein assay known in the art. For example, the presence of transcribed mRNA of the long-term repressor fusion protein can be detected and / or quantified using a probe complementary to any region of the CasX polynucleotide by conventional hybridization assays (e.g., Northern blot analysis), amplification procedures (e.g., RT-PCR), SAGE (U.S. Patent No. 5,695,937), and array-based techniques (e.g., see U.S. Patents No. 5,405,783, 5,412,087, and 5,445,934).
[0229] In some embodiments, vectors are prepared for the transcription of long-term repressor fusion protein genes, and for the expression and recovery of the resulting coding mRNA. In some embodiments, mRNA is produced by PCR product or in vitro transcription (IVT) using a linearized plasmid DNA template and T7 RNA polymerase, where the plasmid contains a T7 promoter. When using a PCR product, the DNA sequence encoding the candidate mRNA is cloned into a plasmid containing the T7 promoter, the plasmid DNA template is linearized, and then used to carry out the IVT reaction for mRNA expression. Exemplary methods for producing such vectors, as well as mRNA production and recovery, are provided in the examples below.
[0230] VIII. LTRP: Delivery particles for gRNA systems In another embodiment, the Disclosure provides particle compositions for delivery to cells or targets of gene repressor systems, such as the LTRP:gRNA system described herein, for transcriptional repression or silencing of genes. Particles envisioned within the scope of the Disclosure include, but are not limited to, nanoparticles, such as synthetic nanoparticles, polymer nanoparticles, lipid nanoparticles, viral particles, and virus-like particles. The particles of the Disclosure may encapsulate a payload, such as a gRNA variant, in combination with mRNA encoding a repressor fusion protein of any of the embodiments described herein, as described herein. Alternatively, or in addition, the particles of the Disclosure may encapsulate a payload of a gRNA variant and a repressor fusion protein, for example, when associated as a ribonucleoprotein (RNP) complex. In some embodiments, the particles are synthetic nanoparticles encapsulating a payload of a gRNA variant and mRNA encoding a repressor fusion protein of any of the embodiments described herein. In some embodiments, the synthetic nanoparticles include biodegradable polymer nanoparticles (PNPs). In some embodiments, materials for the production of biodegradable polymer nanoparticles (PNPs) include polylactide, poly(lactic acid-co-glycolic acid) (PLGA), poly(ethyl cyanoacrylate), poly(butyl cyanoacrylate), poly(isobutyl cyanoacrylate), and poly(isohexyl cyanoacrylate), polyglutamic acid (PGA), poly(ε-caprolactone) (PCL), cyclodextrin, and natural polymers such as chitosan, albumin, gelatin, and alginates, which are the most commonly used polymers for PNP synthesis (Production and clinical development of nanoparticles for gene delivery. Molecular Therapy-Methods&Clinical Development 3:16023;doi:10.1038(2016)).In other embodiments, the particle is a lipid nanoparticle (LNP) encapsulating mRNA encoding a gRNA variant and any of the long-term repressor fusion proteins of the embodiments described herein, which is described more fully below. In other embodiments, the particles are lipid nanoparticles that separately encapsulate mRNA encoding a gRNA variant and any of the long-term repressor fusion proteins of the embodiments described herein in different particles, which are co-formulated as a mixture for administration and are described more fully below. In other embodiments, the particles are lipid nanoparticles that separately encapsulate mRNA encoding a gRNA variant and any of the long-term repressor fusion proteins of the embodiments, and the two types of particles are administered separately.
[0231] a. Lipid nanoparticle (LNP) The present disclosure provides lipid nanoparticles (LNPs) for delivery of the LTRP:gRNA systems described herein to cells or to a subject. In some embodiments, the LNPs of the invention are tissue-specific, have excellent biocompatibility, and can deliver the LTRP:gRNA systems with high efficiency and thus can be used for transcriptional repression of a targeted gene.
[0232] The present disclosure further provides LNP compositions and pharmaceutical compositions comprising a plurality of LNPs described herein.
[0233] In their native forms, nucleic acid polymers are generally unstable in biological fluids and cannot penetrate into the cytoplasm of target cells, thus requiring delivery systems. Lipid nanoparticles (LNPs) have proven useful for both protecting nucleic acids and delivering them to tissues and cells. Furthermore, the use of mRNA in LNPs encoding long-term repressor fusion proteins eliminates the possibility of unwanted genomic integration compared to DNA vectors. Additionally, mRNA efficiently transfects both mitotic and non-mitotic cells because it does not require entry into the nucleus and because it functions in the cytoplasmic compartment. LNPs as delivery platforms thus provide the added advantage that both mRNA encoding long-term repressor fusion proteins and gRNA can be co-formulated into a single LNP particle.
[0234] Thus, in various embodiments, the present disclosure encompasses lipid nanoparticles and compositions that can be used for a variety of purposes, including delivery in both in vitro and in vivo of encapsulated or associated (e.g., complexed) therapeutic agents, such as nucleic acids to cells. In certain embodiments, the present disclosure provides a method of treating or preventing a disease or disorder in a subject that needs it by contacting the subject with a lipid nanoparticle that encapsulates or is associated with a suitable therapeutic agent complexed through various physical, chemical, or electrostatic interactions among one or more of the lipid components used in the composition to make the LNP.
[0235] In some embodiments, the LNP comprises an LTRP:gRNA system described herein for the repression or silencing of a gene associated with a disease or disorder. In some embodiments, the disclosure provides an LNP in which a gRNA and an mRNA encoding a long-term repressor fusion protein are incorporated into a single LNP particle. In some embodiments, the LNP comprises an mRNA containing a sequence selected from the group consisting of SEQ ID NOs: 2409-18636, or a sequence having at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, and at least about 99% sequence identity thereto, and a gRNA containing a sequence selected from the group consisting of SEQ ID NOs: 1744-1746, 2136-2144, 2146-2154, and 2156-2164, accompanied by a ligated targeting sequence complementary to the sequence of the gene to be targeted for repression or silencing. In one embodiment, the LNP comprises an mRNA containing a sequence selected from the group consisting of SEQ ID NOs: 2411, 2421, 2467, and 2477, and a gRNA containing the sequence of SEQ ID NO: 2156 with a ligated targeting sequence complementary to the sequence of the gene targeted for repression or silencing. In other embodiments, the disclosure provides an LNP in which the gRNA and the mRNA encoding a long-term repressor fusion protein are incorporated into a population of separate LNPs, which can be formulated together in various ratios for administration.
[0236] The lipid nanoparticles and lipid nanoparticle compositions of this disclosure may be used to suppress the expression of a desired protein both in vitro and in vivo by contacting cells with lipid nanoparticles containing one or more ionizable lipids described herein, wherein the lipid nanoparticles encapsulate or associate with nucleic acids (e.g., messenger RNA encoding a repressor fusion protein) to be expressed to produce the desired protein. In some embodiments, the lipid nanoparticles and compositions may be used to suppress the expression of a target gene both in vitro and in vivo by contacting cells with lipid nanoparticles containing one or more novel cationic lipids described herein, wherein the lipid nanoparticles encapsulate or associate with one or more nucleic acids of the LTRP:gRNA system of this disclosure. The lipid nanoparticles and compositions of embodiments of this disclosure may also be used for the co-delivery of different nucleic acids (e.g., mRNA and plasmid DNA), either separately or in combination, and may be useful, for example, to provide an effect requiring the co-localization of different nucleic acids (e.g., mRNA encoding a suitable gene repressor or enzyme and gRNA for gene targeting).
[0237] In some embodiments, the LNPs and LNP compositions described herein include at least one cationic lipid, at least one conjugate lipid, at least one steroid or its derivative, at least one additional lipid, or any combination thereof. Alternatively, the lipid compositions of this disclosure may include ionizable lipids, such as ionizable cationic lipids, helper lipids (usually phospholipids), cholesterol, and polyethylene glycol lipid conjugates (PEG lipids), for example, to reduce the specific absorption of plasma proteins and to improve colloidal stability in the biological environment by forming a hydrate layer on nanoparticles. Such lipid compositions can be formulated in typical molar ratios of 50:10:37-39:1.5-2.5 or 20-50:8-65:25-40:1-2.5, with variations created to tune individual properties.
[0238] The LNPs and LNP compositions of this disclosure are configured to protect the encapsulated payload of the systems of this disclosure and to deliver it to tissues and cells, both in vitro and in vivo. Various embodiments of the LNPs and LNP compositions of this disclosure are described in further detail herein.
[0239] Cationic lipids In some embodiments, the LNPs and LNP compositions of this disclosure include at least one cationic lipid. The term “cationic lipid” refers to a lipid species having a net positive charge. In some embodiments, the cationic lipid is an ionizable cationic lipid having a net positive charge at a selected pH, e.g., physiological pH. In some embodiments, the ionizable cationic lipid has a pKa of less than 7, so that the LNPs and LNP compositions achieve efficient encapsulation of the payload at relatively low pH. In some embodiments, the cationic lipid has a pKa of 5–8, 5.5–7.5, 6–7, or 6.5–7. In some embodiments, the cationic lipid may be protonated at a pH below the pKa of the cationic lipid, and it may be substantially neutral at pH above the pKa. LNPs and LNP compositions are safely delivered in vivo to target organs (e.g., liver, lungs, heart, spleen, and tumors) and / or cells (hepatocytes, LSECs, cardiac cells, cancer cells, etc.), exhibit a positive charge after endocytosis, and release the encapsulated payload through electrostatic interactions with anionic proteins of the endosomal membrane.
[0240] Initial formulations of LNPs utilizing permanently cationic lipids resulted in LNPs with a positive surface charge that were proven toxic in vivo and were rapidly removed by phagocytic cells. By converting them to tertiary amines, particularly ionizable cationic lipids with pKa < 7, LNPs were obtained that achieve efficient encapsulation of nucleic acid macromolecules at low pH by electrostatically interacting with the negative charge of the mRNA phosphate backbone, which also results in a system that is largely neutral at physiological pH values, thus mitigating the problems associated with permanently charged cationic lipids.
[0241] As used herein, “ionizable lipid” means an amine-containing lipid that can be readily protonated, and which may be a lipid that changes its charge state depending on the ambient pH. Ionizable lipids may be protonated (positively charged) at pH below the pKa of cationic lipids, and may be substantially neutral at pH above the pKa. In one example, LNPs may include protonated ionizable lipids and / or ionizable lipids that exhibit neutralization. In some embodiments, LNPs have pKas of 5–8, 5.5–7.5, 6–7, or 6.5–7. The pKa of the LNP is important for the in vivo stability and release of the nucleic acid payload of the LNP in target cells or organs. In some embodiments, LNPs having the aforementioned pKa range are safely delivered in vivo to target organs (e.g., liver, lungs, heart, spleen, and tumors) and / or target cells (hepatocytes, LSECs, cardiac cells, cancer cells, etc.), exhibit a positive charge after endocytosis, and release the encapsulated payload through electrostatic interactions with anionic proteins of the endosomal membrane.
[0242] Ionizable lipids are ionizable compounds that generally have similar characteristics to lipids and can play a role in encapsulating nucleic acid payloads within LNPs with high efficiency through electrostatic interactions with nucleic acids (e.g., mRNA as disclosed herein).
[0243] Depending on the types of amine and tail groups contained in the ionizable lipid, (i) nucleic acid encapsulation efficiency, (ii) PDI (polydispersion index), and / or (iii) nucleic acid delivery efficiency of LNPs to tissues and / or cells constituting organs (e.g., hepatocytes or hepatic sinusoidal endothelial cells in the liver) may differ. In certain embodiments, the ionizable lipid is an ionizable cationic lipid, comprising about 46 mol% to about 66 mol% of the total lipids present in the particles.
[0244] LNPs containing ionizable lipids containing amines may have one or more of the following characteristics: (1) the ability to encapsulate nucleic acids with high efficiency; (2) uniform size of the prepared particles (or having a low PDI value); and / or (3) excellent nucleic acid delivery efficiency to organs such as the liver, lungs, heart, spleen, bone marrow, etc., as well as to tumors and / or cells constituting such organs (e.g., hepatocytes, LSECs, cardiac cells, cancer cells, etc.).
[0245] In certain embodiments, cationic lipid forms play a crucial role in both nucleic acid encapsulation via electrostatic interactions and intracellular release by disrupting the endosomal membrane. Nucleic acid payloads are encapsulated within LNPs by ionic interactions they form with positively charged cationic lipids. Non-limiting examples of cationic lipid components utilized in the LNPs of this disclosure are selected from DLin-MC3-DMA (heptatriaconta-6,9,28,31-tetraen-19-yl-4-(dimethylamino)butanoate), DLin-KC2-DMA (2,2-dilinoleyl-4-(2-dimethylaminoethyl)-[1,3]-dioxolane), and TNT (1,3,5-triazinan-2,4,6-trione) and TT (N1,N3,N5-tris(2-aminoethyl)benzene-1,3,5-tricarboxamide). Non-limiting examples of helper lipids used in the LNPs of this disclosure include DSPC (1,2-distearoyl-sn-glycero-3-phosphocholine), PQC (2-oleoyl-1-palmitoyl-sn-glycero-3-phosphocholine), and DOPE (1,2-dioleoyl-sn-glycero-3-phosphoethanolamine), 1,2-dioleoyl-sn-glycero-3-phospho(1'-rac-glycerol)DOPG, 1,2-dimiristoyl-sn-glycero-3-phosphoethanolamine (DMPE), 1,2-dilauroyl-sn-glycero-3-phosphocholine (DLPC), sphingolipids, and ceramides. Cholesterol and PEG-DMG((R)-2,3-bis(octadecyloxy)propyl-1-(methoxypolyethylene glycol 2000)carbamate), PEG-DSG(1,2-distearoyl-rac-glycero-3-methylpolyoxyethylene glycol 2000), or DSPE-PEG2k(1,2-distearoyl-sn-glycero-3-phosphoethanolamine-N-[amino(polyethylene glycol)-2000]) are components used in the LNPs of this disclosure for the stability, circulation, and size of the LNPs.
[0246] In some embodiments, the cationic lipid in the LNP of this disclosure comprises a tertiary amine. In some embodiments, the tertiary amine comprises an alkyl chain connected to the nitrogen of the tertiary amine with an ether linkage. In some embodiments, the alkyl chain comprises a C12-C30 alkyl chain having 0-3 double bonds. In some embodiments, the alkyl chain comprises a C16-C22 alkyl chain. In some embodiments, the alkyl chain comprises a C18 alkyl chain. Numerous cationic lipids and related analogues are described in U.S. Patent Publications 20060083780, 20060240554, 20110117125, 20190336608, 20190381180, and 20200121809, and U.S. Patents 5,208,036, 5,264,618, 5,279,833, and 5,283,18. These disclosures are described in publications 5, 5,753,613, 5,785,992, 9,738,593, 10,106,490, 10,166,298, 10,221,127, and 11,219,634, as well as in PCT Publication 96 / 10390, the disclosures of which are incorporated herein by reference in their entirety.
[0247] In some embodiments, the cationic lipid in the LNP of this disclosure may include, for example, one or more ionizable cationic lipids, where the ionizable cationic lipid is a dialkyl lipid. In other embodiments, the ionizable cationic lipid is a trialkyl lipid.
[0248] In some embodiments, the cationic lipids in the LNPs of this disclosure are 1,2-dilinoleyloxy-N,N-dimethylaminopropane (DLinDMA), 1,2-dilinolenyloxy-N,N-dimethylaminopropane (DLenDMA), 2,2-dilinoleyl-4-(2-dimethylaminoethyl)-[1,3]-dioxolane (DLin-K-C2-DMA), and 2,2-dilinoleyl-4-(3-dimethylaminopropyl)-[1,3]-dioxolane. Solan (DLin-K-C3-DMA), 2,2-dilinoleyl-4-(4-dimethylaminobutyl)-[1,3]-dioxolane (DLin-K-C4-DMA), 2,2-dilinoleyl-5-dimethylaminomethyl-[1,3]-dioxane (DLin-K6-DMA), 2,2-dilinoleyl-4-N-methylpepiazino (methylpepiazino)-[1,3]-dioxolane (DLin-K-MPZ), 2,2-dilinoleyl -4-dimethylaminomethyl-[1,3]-dioxolane (DLin-K-DMA), 1,2-dilinoleylcarbamoyloxy-3-dimethylaminopropane (DLin-C-DAP), 1,2-dilinoleyoxy-3-(dimethylamino)acetoxypropane (DLin-DAC), 1,2-dilinoleyoxy-3-morpholinopropane (DLin-MA), 1,2-dilinoleo Il-3-dimethylaminopropane (DLinDAP), 1,2-dilinoleylthio-3-dimethylaminopropane (DLin-S-DMA), 1-linoleyl-2-linoleyloxy-3-dimethylaminopropane (DLin-2-DMAP), 1,2-dilinoleyloxy-3-trimethylaminopropane chloride salt (DLin-TMA.Cl), 1,2-dilinoleyl-3-trimethylaminopropane chloride salt (DLin-TAP.Cl), 1,2-dilinoleyloxy-3-(N-methylpiperazino)propane (DLin-MPZ), 3-(N,N-dilinoleylamino)-1,2-propanediol (DLinAP), 3-(N,N-dioleylamino)-1,2-propanediol (DOAP), 1,2-dilinoleyloxo-3-(2-N,N-dimethylamino)ethoxypropane (DLin-EG-DMA), N,N-dioleyl-N,N-dimethylammonium chloride (DODAC), 1,2-dioleyloxy-N,N- Dimethylaminopropane (DODMA), 1,2-distearyloxy-N,N-dimethylaminopropane (DSDMA), N-(1-(2,3-dioleyloxy)propyl)-N,N,N-trimethylammonium chloride (DOTMA), N,N-distearyl-N,N-dimethylammonium bromide (DDAB), N-(1-(2,3-dioleyloxy)propyl)-N,N,N-trimethylammonium chloride (DOTAP), 3-(N-(N',N'-dimethylaminoethane)-carbamoyl)cholesterol (DC-Ch ol), N-(1,2-dimyrityloxyprop-3-yl)-N,N-dimethyl-N-hydroxyethylammonium bromide (DMRIE), 2,3-dioleyloxy-N-[2(spermine-carboxamide)ethyl]-N,N-dimethyl-1-propanaminonium trifluoroacetic acid (DOSPA), dioctadecylamide glycylspermine (DOGS), 3-dimethylamino-2-(cholest-5-en-3-beta-oxybutane-4-oxy)-1-(cis,cis-9,12-octadecadienoxy)propane (CLinDM) A) Selected from 2-[5'-(cholest-5-ene-3-beta-oxy)-3'-oxapentoxy)-3-dimethyl-1-(cis,cis-9',1-2'-octadecadienoxy)propane (CpLinDMA), N,N-dimethyl-3,4-dioleyloxybenzylamine (DMOBA), 1,2-N,N'-dioleylcarbamyl-3-dimethylaminopropane (DOcarbDAP), 1,2-N,N'-dilinoleylcarbamyl-3-dimethylaminopropane (DLincarbDAP), and any combination thereof.
[0249] In some embodiments, the cationic lipids in the LNPs of this disclosure are selected from heptatriaconta-6,9,28,31-tetraen-19-yl-4-(dimethylamino)butanoate (DLin-MC3-DMA), 2,2-dilinoleyl-4-(2-dimethylaminoethyl)-[1,3]-dioxolane (DLin-KC2-DMA), (1,3,5-triazinan-2,4,6-trione) (TNT), N1,N3,N5-tris(2-aminoethyl)benzene-1,3,5-tricarboxamide (TT), and any combination thereof.
[0250] In some embodiments, the N / P ratio (nitrogen from cationic / ionizable lipids and phosphate from nucleic acids) in the LNPs of the Disclosure is in the range of about 3:1 to 7:1, or about 4:1 to 6:1, or 3:1, or 4:1, or 5:1, or 6:1, or 7:1.
[0251] Conjugate lipids In some embodiments, the LNPs and LNP compositions of the Disclosure include at least one conjugated lipid. In some embodiments, the conjugated lipid may be selected from polyethylene glycol (PEG)-lipid conjugates, polyamide (ATTA)-lipid conjugates, cationic polymeric lipid conjugates (CPL), and any combination thereof. In some cases, the conjugated lipid can inhibit the aggregation of the LNPs of the Disclosure.
[0252] In some embodiments, the conjugate lipids of the LNPs of this disclosure include pegylated lipids. The terms “polyethylene glycol (PEG)-lipid conjugate,” “pegylated lipid,” “lipid-PEG conjugate,” “lipid-PEG,” “PEG-lipid,” “PEG-lipid,” or “lipid-PEG” are used interchangeably herein and refer to lipids attached to a polyethylene glycol (PEG) polymer, which is a hydrophilic polymer. Pegylated lipids contribute to the stability of LNPs and LNP compositions and reduce LNP aggregation.
[0253] PEG-lipids can form surface lipids, so the size of the LNP can be easily varied by varying the ratio of surface (PEG) lipids to core (ionizable cationic) lipids. In some embodiments, the PEG-lipids of the LNPs of the present disclosure can be varied up to about 1-5 mol% to modify particle properties such as size, stability, and circulation time.
[0254] Lipid-PEG conjugates contribute to the particle stability of the nanoparticles in serum within the LNP and play a role in preventing aggregation between the nanoparticles. In addition, lipid-PEG conjugates protect nucleic acids, such as the mRNA encoding the repressor fusion protein of the present disclosure, or the gRNA of the present disclosure, etc., from degrading enzymes during in vivo delivery of the nucleic acid, enhance the stability of the nucleic acid in vivo, and can increase the half-life of the delivered nucleic acid encapsulated in the nanoparticles. Examples of PEG-lipid conjugates include, but are not limited to, PEG-DAG conjugates, PEG-DAA conjugates, and mixtures thereof. In certain embodiments, the PEG-lipid conjugate is selected from the group consisting of PEG-diacylglycerol (PEG-DAG) conjugates, PEG-dialkyloxypropyl (PEG-DAA) conjugates, PEG-phospholipid conjugates, PEG-ceramide (PEG-Cer) conjugates, and mixtures thereof.
[0255] In some embodiments, the pegylated lipids of the LNPs of the present disclosure are selected from PEG-ceramide, PEG-diacylglycerol, PEG-dialkyloxypropyl, PEG-dialkoxypropyl carbamate, PEG-phosphatidylethanolamine, PEG-phospholipid, PEG-diacylglycerol succinate, and any combination of the foregoing.
[0256] In some embodiments, the pegylated lipid of the LNP of this disclosure is PEG-dialkyloxypropyl. In some embodiments, the pegylated lipid is selected from PEG-didecyloxypropyl (C10), PEG-dilauryloxypropyl (C12), PEG-dimyristyloxypropyl (C14), PEG-dipalmityloxypropyl (C16), PEG-distearyloxypropyl (C18), and any combination thereof.
[0257] In other embodiments, the lipid-PEG conjugate of the LNP of the present disclosure may be phospholipid-conjugated PEG, such as phosphatidylethanolamine (PEG-PE), PEG conjugated to ceramide (PEG-CER, ceramide-PEG conjugate, ceramide-PEG, PEG conjugated to cholesterol or its derivatives, PEG-c-DOMG, PEG-DMG, PEG-DLPE, PEG-DMPE, PEG-DPPC, PEG-DSPE (DSPE-PEG), and mixtures thereof, for example, C16-PEG2000 ceramide (N-palmitoyl-sphingosine-1-{succinyl[methoxy(polyethylene glycol)2000]}), DMG-PEG2000, 14:0 PEG2000 PE.
[0258] In some embodiments, the pegylated lipids of the LNPs of this disclosure are selected from 1-(monomethoxy-polyethylene glycol)-2,3-dimyristoylglycerol, 4-O-(2',3'-di(tetradecanoyloxy)propyl-1-O-(ω-methoxy(polyethoxy)ethyl)butanediate (PEG-S-DMG), ω-methoxy(polyethoxy)ethyl-N-(2,3-di(tetradecanoxy)propyl)carbamate, 2,3-di(tetradecanoxy)propyl-N-(ω-methoxy(polyethoxy)ethyl)carbamate, and any combination thereof.
[0259] In some embodiments, the pegylated lipids of the LNPs of this disclosure are selected from mPEG2000-1,2-di-O-alkyl-sn3-carbomoylglyceride (PEG-C-DOMG), 1-[8'-(1,2-dimiristoyl-3-propanoxy)-carboxamide-3',6'-dioxaoctanyl]carbamoyl-w-methyl-poly(ethylene glycol) (2KPEG-DMG), and any combination thereof.
[0260] In some embodiments, PEG is directly attached to the lipids of the pegylated lipid. In other embodiments, PEG is bound to the lipids of the pegylated lipid by a linker moiety selected from ester-free or ester-containing linker moieties. Non-limiting examples of ester-free linker moieties include amide (-C(O)NH-), amino (-NR-), carbonyl (-C(O)-), carbamate (-NHC(O)O-), urea (-NHC(O)NH-), disulfide (-SS-), ether (-O-), succinyl (-(O)CCH2CH2C(O)-), succinamidyl (-NHC(O)CH2CH2C(O)NH-), ethers, disulfides, and combinations thereof. For example, the linker may contain carbamate linker moieties and amide linker moieties. Non-limiting examples of ester-containing linker moieties include carbonates (-OC(O)O-), succinoyl, phosphate esters (-O-(O)POH-O-), sulfonic acid esters, and combinations thereof.
[0261] The PEG portion of the pegylated lipids of the LNPs disclosed herein may have an average molecular weight in the range of about 550 daltons to about 10,000 daltons. In certain embodiments, the PEG portion may have an average molecular weight of about 750 daltons to about 5,000 daltons, about 1,000 daltons to about 4,000 daltons, about 1,500 daltons to about 3,000 daltons, about 750 daltons to about 3,000 daltons, or about 1,750 daltons to about 2,000 daltons.
[0262] In some embodiments, the conjugated lipids (e.g., pegylated lipids) comprise about 1 mol% to about 60 mol%, about 2 mol% to about 50 mol%, about 5 mol% to about 40 mol%, or about 5 mol% to about 20 mol% of the total lipids present in the LNP and / or LNP composition. In certain embodiments, the conjugated lipids comprise about 0.5 mol% to about 3 mol% of the total lipids present in the particles.
[0263] In additional embodiments, the conjugate lipids of the LNPs of this disclosure (e.g., pegylated lipids) comprise at least about 1, 2, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, or 60 mol% of the total lipids present in the LNPs and / or LNP compositions, or any intermediate range as described above.
[0264] In the case of lipids in the lipid-PEG conjugate of the LNP of this disclosure, any lipid capable of binding to polyethylene glycol may be used, without limitation, and other elements of the LNP, such as phospholipids and / or cholesterol, may also be used. In some embodiments, the lipids in the lipid-PEG conjugate may be, but are not limited to, ceramide, dimyristoyl glycerol (DMG), succinoyl diacylglycerol (s-DAG), distearoyl phosphatidylcholine (DSPC), distearoyl phosphatidylethanolamine (DSPE), or cholesterol.
[0265] In the lipid-PEG conjugates of LNPs of this disclosure, PEG can be directly conjugated to the lipid or linked to the lipid via a linker moiety. Any suitable linker moiety can be used to link PEG to the lipid, including, for example, ester-free linker moieties and ester-containing linker moieties. Ester-free linker moieties include, but are not limited to, amides (-C(O)NH-), aminos (-NR-), carbonyls (-C(O)-), carbamates (-NHC(O)O-), ureas (-NHC(O)NH-), disulfides (-SS-), ethers (-O-), succinyls (-(O)CCH2CH2C(O)-), succinamidyls (-NHC(O)CH2CH2C(O)NH-), ethers, and disulfides, as well as combinations thereof (e.g., linkers containing both carbamate linker moieties and amide linker moieties). The ester-containing linker portion includes, but is not limited to, carbonates (-OC(O)O-), succinoyl, phosphate esters (-O-(O)POH-O-), sulfonic acid esters, and combinations thereof.
[0266] steroid In some embodiments, the LNPs and LNP compositions of this disclosure comprise at least one steroid or a derivative thereof. In some embodiments, the steroid comprises cholesterol. In some embodiments, the LNPs and LNP compositions comprise cholestanol, cholestanone, cholestenone, coprostanol, cholesteryl-2'-hydroxyethyl ether, cholesteryl-4'-hydroxybutyl ether, and cholesterol derivatives selected from any combination thereof.
[0267] In some embodiments, the steroid of the LNP (e.g., cholesterol) of the Disclosure comprises about 1 mol% to about 60 mol%, about 2 mol% to about 50 mol%, about 5 mol% to about 40 mol%, or about 5 mol% to about 20 mol% of the total lipids present in the LNP and / or LNP composition. In other embodiments, the steroid of the LNP (e.g., cholesterol) of the Disclosure comprises at least about 1, 2, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, or 60 mol%, or any intermediate range as described above, of the total lipids present in the LNP and / or LNP composition.
[0268] Additional lipids In some embodiments, the LNPs and LNP compositions of this disclosure include at least one additional lipid. In some embodiments, the additional lipid is a non-cationic lipid selected from anionic lipids, neutral lipids, or both. In some embodiments, the additional lipid includes at least one phospholipid. In some embodiments, the phospholipid is selected from anionic phospholipids, neutral phospholipids, or both. The phospholipids of the elements of the LNPs and LNP compositions can play a role in covering and protecting the core of the LNP formed by the interaction of cationic lipids and nucleic acids in the LNP, and can facilitate cell membrane translocation and endosomal escape during intracellular delivery of nucleic acids by binding to the phospholipid bilayer of target cells. Phospholipids that can facilitate the fusion of LNPs into cells may include, without limitation, any of the phospholipids selected from the group described below.
[0269] In some embodiments, LNPs and LNP compositions are, but are not limited to, dipalmitoyl-phosphatidylcholine (DPPC), distearoyl-phosphatidylcholine (DSPC), dioleoyl-phosphatidylethanolamine (DOPE), dioleoyl-phosphatidylcholine (DOPC), dioleoyl-phosphatidylglycerol (DOPG), palmitoyloleoyl-phosphatidylcholine (POPC), palmitoyloleoyl-phosphatidylethanolamine (POPE), palmitoyloleyol-phosphatidylglycerol (POPG), dipalmitoyl-phosphatidylethanolamine (DPPE), dipalmitoyl-phosphatidylglycerol (DPPG), and dimyristoyl-phosphatidylethanolamine (DM The LNP comprises PE, distearoyl-phosphatidylethanolamine (DSPE), monomethyl-phosphatidylethanolamine, dimethyl-phosphatidylethanolamine, dielidoyl-phosphatidylethanolamine (DEPE), stearoyloleoyl-phosphatidylethanolamine (SOPE), egg phosphatidylcholine (EPC), phosphatidylethanolamine (PE), 1,2-dioleoyl-sn-glycero-3-phosphoethanolamine, 1-palmitoyl-2-oleoyl-sn-glycero-3-phosphocholine (POPC), 1,2-dioleoyl-sn-glycero-3-[phospho-L-serine] (DOPS), 1,2-dioleoyl-sn-glycero-3-[phospho-L-serine], and at least one phospholipid selected from any combination of the above. In one example, an LNP containing DOPE may be effective in mRNA delivery.
[0270] In some embodiments, the additional lipids (e.g., phospholipids) of the LNPs of this disclosure comprise about 1 mol% to about 60 mol%, about 2 mol% to about 50 mol%, about 5 mol% to about 40 mol%, or about 5 mol% to about 20 mol% of the total lipids present in the LNPs and / or LNP compositions. In other embodiments, the additional lipids (e.g., phospholipids) of the LNPs of this disclosure comprise at least about 1, 2, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, or 60 mol%, or any intermediate range as described above, of the total lipids present in the LNPs and / or LNP compositions.
[0271] It will be understood that the total lipids present in LNPs and / or LNP compositions include a combination of cationic lipids or ionizable cationic lipids, conjugate lipids (e.g., pegylated lipids), steroids (e.g., cholesterol), and additional lipids (e.g., phospholipids).
[0272] LNPs and / or LNP compositions can be prepared by dissolving total lipids (or a portion thereof) in an organic solvent (e.g., ethanol), followed by mixing with a payload (e.g., nucleic acids of the system) dissolved in an acidic buffer (e.g., pH 4) via a micromixer. At this pH, the cationic lipids are positively charged and interact with negatively charged nucleic acid polymers. The resulting nanostructures, containing nucleic acids, are then converted to neutral LNPs when dialyzed against a neutral buffer, which can then be followed by the removal of the organic solvent (e.g., ethanol) and the exchange of the LNPs into a physiologically relevant buffer. The LNPs and / or LNP compositions thus formed have a distinct electron-density nanostructure core in which the cationic lipids are organized in inverse micelles around the encapsulated payload, in contrast to conventional bilayer liposome structures. In another embodiment, the LNPs may form blister-like structures having nucleic acids in aqueous pockets along a non-electron-density lipid core.
[0273] b. Lipid nanoparticle properties In some embodiments, the LNP and / or LNP composition comprises about 50 mol% to about 85 mol% of cationic or ionizable cationic lipids, about 0.5 mol% to about 10 mol% of conjugate lipids (e.g., pegylated lipids), about 0.5 mol% to about 10 mol% of steroids (e.g., cholesterol), and about 5 mol% to about 50 mol% of additional lipids (e.g., phospholipids). In some embodiments, the LNP and / or LNP composition comprises about 50 mol% to about 85 mol% of cationic or ionizable cationic lipids, about 0.5 mol% to about 5 mol% of conjugate lipids (e.g., pegylated lipids), about 0.5 mol% to about 5 mol% of steroids (e.g., cholesterol), and about 5 mol% to about 20 mol% of additional lipids (e.g., phospholipids).
[0274] In some embodiments, the LNPs and / or LNP compositions of the present disclosure comprise cationic lipids: additional lipids (e.g., phospholipids): steroids (e.g., cholesterol): conjugate lipids (e.g., pegylated lipids) in molar ratios of 20-50:10-30:30-60:0.5-5, 25-45:10-25:40-50:0.5-3, 25-45:10-20:40-55:0.5-3, or 25-45:10-20:40-55:1.0-1.5.
[0275] In some embodiments, the LNPs and / or LNP compositions of this disclosure have a total lipid:payload ratio (mass / mass) of about 1 to about 100. In some embodiments, the total lipid:payload ratio is about 1 to about 50, about 2 to about 25, about 3 to about 20, about 4 to about 15, or about 5 to about 10. In some embodiments, the total lipid:payload ratio is about 5 to about 15, for example, about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or any intermediate range as described above.
[0276] In certain embodiments, the LNPs of this disclosure contain a total lipid:nucleic acid mass ratio of about 5:1 to about 15:1. In some embodiments, the weight ratio of cationic lipids to nucleic acids contained in the LNP may be 1 to 20:1, 1 to 15:1, 1 to 10:1, 5 to 20:1, 5 to 15:1, 5 to 10:1, 7.5 to 20:1, 7.5 to 15:1, or 7.5 to 10:1.
[0277] In some embodiments, the LNPs of this disclosure may comprise 20 to 50 parts by weight of cationic lipids, 10 to 30 parts by weight of phospholipids, 20 to 60 parts by weight (or 20 to 60 parts by weight) of cholesterol, and 0.1 to 10 parts by weight (or 0.25 to 10 parts by weight, 0.5 to 5 parts by weight) of lipid-PEG conjugate. Alternatively, the LNPs may comprise, based on the total weight of nanoparticles, 20 to 50% by weight of cationic lipids, 10 to 30% by weight of phospholipids, 20 to 60% by weight (or 30 to 60% by weight) of cholesterol, and 0.1 to 10% by weight (or 0.25 to 10% by weight, 0.5 to 5% by weight) of lipid-PEG conjugate. As a further alternative, LNPs may contain, based on the total nanoparticle weight, 25–50 wt% cationic lipids, 10–20 wt% phospholipids, 35–55 wt% cholesterol, and 0.1–10 wt% (or 0.25–10 wt%, 0.5–5 wt%) lipid-PEG conjugates.
[0278] In some embodiments, the LNPs of this disclosure have wavelengths of approximately 20-200nm, 20-180nm, 20-170nm, 20-150nm, 20-120nm, 20-100nm, 20-90nm, 30-200nm, 30-180nm, 30-170nm, 30-150nm, 30-120nm, 30-100nm, 30-90nm, 40-200nm, 40-180nm, 40-170nm, 40-150nm, 40-120nm, 40-100nm, 40-90nm, 40-80nm, 40-70nm, 50-200nm, 50-180nm, 50-170nm, 50-150nm, 50-120nm, 50-100nm, It has an average diameter in the following ranges: 50-90nm, 60-200nm, 60-180nm, 60-170nm, 60-150nm, 60-120nm, 60-100nm, 60-90nm, 70-200nm, 70-180nm, 70-170nm, 70-150nm, 70-120nm, 70-100nm, 70-90nm, 80-200nm, 80-180nm, 80-170nm, 80-150nm, 80-120nm, 80-100nm, 80-90nm, 90-200nm, 90-180nm, 90-170nm, 90-150nm, 90-120nm, or 90-100nm, or an intermediate range of any of the above.
[0279] In some embodiments, the LNPs and / or LNP compositions of this disclosure are positively charged at an acidic pH and can encapsulate a payload (e.g., a therapeutic drug) through electrostatics produced by the negative charge of the payload (e.g., a therapeutic drug, e.g., an LTRP:gRNA system, or a polynucleotide encoding the same). The term “encapsulation” refers to a mixture of lipids that surrounds, embeds, and forms an LNP around a payload (e.g., a therapeutic drug) under physiological conditions. The term “encapsulation efficiency” is the amount of payload (e.g., a therapeutic drug) encapsulated by the LNP, divided by the total amount of payload (e.g., a therapeutic drug) used to fill the LNP with the payload (e.g., a therapeutic drug). The encapsulation efficiency of the LNPs and / or LNP compositions may be 70% or higher, 75% or higher, 80% or higher, 85% or higher, 90% or higher, 91% or higher, 92% or higher, 94% or higher, or 95% or higher. In other embodiments, the encapsulation efficiency of LNPs and / or LNP compositions is about 80%–99%, about 85%–98%, about 88%–95%, and about 90%–95%, or the payload (e.g., nucleic acids of the system) can be completely encapsulated within the lipid portion of the LNP composition, thereby protecting it from enzymatic degradation. In some embodiments, the payload (e.g., therapeutic agents) remains substantially undegraded after exposure of the LNPs and / or LNP compositions to nucleases at 37°C for at least about 20, 30, 45, or 60 minutes, or for at least about 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, or 36 hours. In some embodiments, the payload (e.g., nucleic acids of the system) is complexed with the lipid portion of the LNPs and / or LNP compositions. The LNPs and / or LNP compositions of this disclosure are nontoxic to mammals, such as humans.
[0280] The term “fully encapsulated” indicates that the payload (e.g., nucleic acids in the system) in the LNP and / or LNP composition is not significantly degraded after exposure to conditions that significantly degrade free DNA, RNA, or proteins. In a fully encapsulated system, less than about 25%, more preferably less than about 10%, and most preferably less than 5% of the payload (e.g., nucleic acids in the system) in the LNP and / or LNP composition is degraded by conditions that can degrade 100% of an unencapsulated payload. “Fully encapsulated” also indicates that the LNP and / or LNP composition is stable to serum and is not degraded into their constituent parts upon in vivo administration.
[0281] In some embodiments, the amount of LNP and / or LNP composition having a payload (e.g., therapeutic agent) encapsulated therein is about 30% to about 100%, about 40% to about 100%, about 50% to about 100%, about 60% to about 100%, about 70% to about 100%, about 80% to about 100%, about 90% to about 100%, about 30% to about 95%, about 40% to about 95%, about 50% to about 95%, about 60% to about 95%, %, about 70% to about 95%, about 80% to about 95%, about 85% to approximately 95%, approximately 90% to approximately 95%, approximately 30% to approximately 90%, approximately 40% to approximately 90%, approximately 50% to approximately 90%, approximately 60% to approximately 90%, approximately 70% to approximately 90%, approximately 80% to approximately 90%, or at least approximately 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or any intermediate range as described above.
[0282] In some embodiments, the amount of payload (e.g., nucleic acid) encapsulated within the LNP and / or LNP composition is approximately 30% to 100%, approximately 40% to 100%, approximately 50% to 100%, approximately 60% to 100%, approximately 70% to 100%, approximately 80% to 100%, approximately 90% to 100%, approximately 30% to 95%, approximately 40% to 95%, approximately 50% to 95%, approximately 60% to 95%, approximately 70% to 95%, approximately 80% to 95%, approximately 85% Approximately 95%, approximately 90% to approximately 95%, approximately 30% to approximately 90%, approximately 40% to approximately 90%, approximately 50% to approximately 90%, approximately 60% to approximately 90%, approximately 70% to approximately 90%, approximately 80% to approximately 90%, or at least approximately 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or any intermediate range as described above.
[0283] In some embodiments, the nucleic acids of the Disclosure, such as mRNA and / or gRNA encoding repressor fusion proteins, may be provided in a solution that is mixed with a lipid solution so that the nucleic acids can be encapsulated in lipid nanoparticles. A suitable nucleic acid solution may be any aqueous solution containing nucleic acids to be encapsulated at various concentrations. For example, a suitable nucleic acid solution may contain nucleic acids (or nucleic acids) at concentrations of about 0.01 mg / ml, 0.05 mg / ml, 0.06 mg / ml, 0.07 mg / ml, 0.08 mg / ml, 0.09 mg / ml, 0.1 mg / ml, 0.15 mg / ml, 0.2 mg / ml, 0.3 mg / ml, 0.4 mg / ml, 0.5 mg / ml, 0.6 mg / ml, 0.7 mg / ml, 0.8 mg / ml, 0.9 mg / ml, 1.0 mg / ml, 1.25 mg / ml, 1.5 mg / ml, 1.75 mg / ml, or 2.0 mg / ml or higher. In some embodiments, the nucleic acid includes mRNA encoding a repressor fusion protein, and a suitable mRNA solution is approximately 0.01-2.0 mg / ml, 0.01-1.5 mg / ml, 0.01-1.25 mg / ml, 0.01-1.0 mg / ml, 0.01-0.9 mg / ml, 0.01-0.8 mg / ml, 0.01-0.7 mg / ml, 0.01-0.6 mg / ml, 0.01-0.5 mg / ml, 0.01-0.4 mg / ml, 0.01-0.3 mg / ml, 0.01-0.2 mg / ml, 0.01-0.1 It may contain mRNA at concentrations ranging from mg / ml, 0.05-1.0 mg / ml, 0.05-0.9 mg / ml, 0.05-0.8 mg / ml, 0.05-0.7 mg / ml, 0.05-0.6 mg / ml, 0.05-0.5 mg / ml, 0.05-0.4 mg / ml, 0.05-0.3 mg / ml, 0.05-0.2 mg / ml, 0.05-0.1 mg / ml, 0.1-1.0 mg / ml, 0.2-0.9 mg / ml, 0.3-0.8 mg / ml, 0.4-0.7 mg / ml, or 0.5-0.6 mg / ml.In some embodiments, a suitable mRNA solution may contain mRNA at concentrations up to approximately 5.0 mg / ml, 4.0 mg / ml, 3.0 mg / ml, 2.0 mg / ml, 1.0 mg / ml, 0.9 mg / ml, 0.8 mg / ml, 0.7 mg / ml, 0.6 mg / ml, 0.5 mg / ml, 0.4 mg / ml, 0.3 mg / ml, 0.2 mg / ml, 0.1 mg / ml, 0.05 mg / ml, 0.04 mg / ml, 0.03 mg / ml, 0.02 mg / ml, 0.01 mg / ml, or 0.05 mg / ml. In some embodiments, a suitable gRNA solution may contain gRNA at concentrations up to approximately 5.0 mg / ml, 4.0 mg / ml, 3.0 mg / ml, 2.0 mg / ml, 1.0 mg / ml, 0.9 mg / ml, 0.8 mg / ml, 0.7 mg / ml, 0.6 mg / ml, 0.5 mg / ml, 0.4 mg / ml, 0.3 mg / ml, 0.2 mg / ml, 0.1 mg / ml, 0.05 mg / ml, 0.04 mg / ml, 0.03 mg / ml, 0.02 mg / ml, 0.01 mg / ml, or 0.05 mg / ml.
[0284] In some embodiments, LNPs are available in the following wavelength ranges: 20nm-200nm, 20-180nm, 20nm-170nm, 20nm-150nm, 20nm-120nm, 20nm-100nm, 20nm-90nm, 30nm-200nm, 30-180nm, 30nm-170nm, 30nm-150nm, 30nm-120nm, 30nm-100nm, 30nm-90nm, and 40nm- 200nm, 40~180nm, 40nm~170nm, 40nm~150nm, 40nm~120nm, 40nm~100nm, 40nm~90nm, 40nm~80nm, 40nm~ 70nm, 50nm~200nm, 50~180nm, 50nm~170nm, 50nm~150nm, 50nm~120nm, 50nm~100nm, 50nm~90nm, 60nm~ 200nm, 60~180nm, 60nm~170nm, 60nm~150nm, 60nm~120nm, 60nm~100nm, 60nm~90nm, 70nm~200nm, 70~1 80nm, 70nm~170nm, 70nm~150nm, 70nm~120nm, 70nm~100nm, 70nm~90nm, 80nm~200nm, 80~180nm, 80nm~ LNPs may have average diameters of 170 nm, 80 nm to 150 nm, 80 nm to 120 nm, 80 nm to 100 nm, 80 nm to 90 nm, 90 nm to 200 nm, 90 nm to 180 nm, 90 nm to 170 nm, 90 nm to 150 nm, 90 nm to 120 nm, or 90 nm to 100 nm for easy delivery into liver tissue, hepatocytes, and / or LSECs (hepatic sinusoidal endothelial cells). LNPs may be sized for easy delivery into organs or tissues, including, but not limited to, the liver, lungs, heart, spleen, and tumors. If the size of the LNP is smaller than the above range, it becomes difficult to maintain stability as the surface area of the LNP increases excessively, and thus delivery to the target tissue and / or drug effect may be reduced. LNPs may specifically target liver tissue. One mechanism by which therapeutic agents can be delivered using LNPs without being constrained by theory is through the mimicry of the metabolic behavior of natural lipoproteins; therefore, LNPs can be usefully delivered to the target through lipid metabolic processes carried out by the liver.During the delivery of therapeutic agents to hepatocytes and / or LSECs (hepatic sinusoidal endothelial cells), the diameter of the window leading from the sinusoidal lumen to the hepatocytes and LSECs is approximately 140 nm in mammals and approximately 100 nm in humans. Therefore, LNP compositions for therapeutic agent delivery having LNPs with diameters within the above range may have superior delivery efficiency to hepatocytes and LSECs when compared with LNPs having diameters outside the above range.
[0285] For example, the LNPs in an LNP composition may contain cationic lipids:phospholipids:cholesterol:lipid-PEG conjugates within the ranges described above, or in molar ratios of 20-50:10-30:30-60:0.5-5, 25-45:10-25:40-50:0.5-3, 25-45:10-20:40-55:0.5-3, or 25-45:10-20:40-55:1.0-1.5. LNPs containing components in molar ratios within the above ranges may have excellent delivery efficiency of therapeutic agents specific to cells of target organs.
[0286] In certain embodiments, LNPs exhibit a positive charge under acidic pH conditions by having a pKa of 5–8, 5.5–7.5, 6–7, or 6.5–7, and can efficiently encapsulate nucleic acids by readily forming complexes with them through electrostatic interactions. In such cases, LNPs can be usefully used as compositions for intracellular or in vivo delivery of therapeutic agents (e.g., nucleic acids).
[0287] In this specification, “encapsulate” or “embed” refers to the efficient delivery of a therapeutic agent by surrounding it with the particle surface and / or embedding it within the particle. Encapsulation efficiency means the therapeutic agent content encapsulated in the LNP relative to the total therapeutic agent content used for the preparation of the LNP.
[0288] The encapsulation of nucleic acids in the composition within the LNP may be 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 91% or more, 92% or more, 94% or more, or 95% or more of the LNP in the composition encapsulating the nucleic acids. In some embodiments, the encapsulation of nucleic acids in the composition within the LNP may be between 80% and 99%, between 80% and 97%, between 80% and 95%, between 85% and 95%, between 87% and 95%, between 90% and 95%, between 91% and 94%, between 91% and 95%, between 92% and 99%, between 92% and 97%, or between 92% and 95%. In some embodiments, the mRNA encoding the LTRP and the gRNA of any embodiment of this disclosure are completely encapsulated within the LNP.
[0289] The target organs to which nucleic acids are delivered by LNPs are, but are not limited to, the liver, lungs, heart, spleen, and tumors. LNPs according to one example are liver tissue specific, possess excellent biocompatibility, and can deliver nucleic acids of the composition with high efficiency; thus, they can be usefully used in relevant technical fields, such as lipid nanoparticle-mediated gene therapy. In certain embodiments, the target cells to which nucleic acids are delivered by LNPs according to one example may be hepatocytes and / or LSECs in vivo. In other embodiments, the present disclosure provides LNPs formulated for ex vivo delivery of nucleic acids of the embodiment to cells.
[0290] The present invention provides a pharmaceutical composition comprising a plurality of LNPs, including nucleic acids such as mRNA encoding long-term repressor fusion proteins and / or gRNA variants described herein, as well as a pharmaceutically acceptable carrier, diluent, or excipient.
[0291] In certain embodiments, LNPs containing nucleic acids have an electron-density core.
[0292] This disclosure provides an LNP comprising one or more nucleic acids, including (a) mRNA encoding a repressor fusion protein and / or a gRNA variant as described herein; (b) one or more cationic or ionizable cationic lipids or salts thereof comprising about 50 mol% to about 85 mol% of the total lipids present in the LNP; (c) one or more non-cationic lipids comprising about 13 mol% to about 49.5 mol% of the total lipids present in the LNP; and (d) one or more conjugate lipids that inhibit aggregation of the LNP comprising about 0.5 mol% to about 2 mol% of the total lipids present in the particles. In another embodiment, the Disclosure provides an LNP comprising (a) mRNA encoding a repressor fusion protein and / or a gRNA variant as described herein; (b) one or more cationic or ionizable cationic lipids or salts thereof comprising about 22 mol% to about 85 mol% of the total lipids present in the LNP; (c) one or more non-cationic / phospholipids comprising about 10 mol% to about 70 mol% of the total lipids present in the LNP; (d) 15 mol% to about 50 mol% sterols, and (d) one or more nucleic acids comprising about 1 mol% to about 5 mol% lipid-PEG or lipid-PEG-peptides in the particles. In certain embodiments, the long-term repressor fusion protein mRNA and gRNA may be present in the same LNP, or they may be present in different LNPs.
[0293] This disclosure provides LNPs comprising one or more nucleic acids, including (a) mRNA encoding a long-term repressor fusion protein as described herein; (b) a cationic lipid or salt thereof comprising about 52 mol% to about 62 mol% of the total lipids present in the LNP; (c) a mixture of phospholipids and cholesterol or derivatives thereof comprising about 36 mol% to about 47 mol% of the total lipids present in the LNP; and (d) a PEG-lipid conjugate comprising about 1 mol% to about 2 mol% of the total lipids present in the LNP. In certain embodiments, the formulation is a four-component system comprising about 1.4 mol% PEG-lipid conjugate (e.g., PEG2000-C-DMA), about 57.1 mol% cationic lipid (e.g., DLin-K-C2-DMA) or a salt thereof, about 7.1 mol% DPPC (or DSPC), and about 34.3 mol% cholesterol (or a derivative thereof). In some embodiments, the LNP comprises mRNA and gRNA encoding CasX as described herein.
[0294] In other embodiments, the LNP comprising one or more nucleic acids comprises (a) mRNA encoding a long-term repressor fusion protein and / or gRNA of any of the embodiments described herein; (b) a cationic lipid or salt thereof comprising about 46.5 mol% to about 66.5 mol% of the total lipids present in the LNP; (c) cholesterol or a derivative thereof comprising about 31.5 mol% to about 42.5 mol% of the total lipids present in the LNP; and (d) a PEG-lipid conjugate comprising about 1 mol% to about 2 mol% of the total lipids present in the LNP. In certain embodiments, the formulation is phospholipid-free and is a three-component system comprising about 1.5 mol% PEG-lipid conjugate (e.g., PEG2000-C-DMA), about 61.5 mol% cationic lipid (e.g., DLin-K-C2-DMA) or a salt thereof, and about 36.9 mol% cholesterol (or a derivative thereof). In some embodiments, the LNP comprises mRNA and gRNA encoding CasX as described herein.
[0295] Additional formulations are described in PCT Publication WO09 / 127060 and U.S. Patent Application Publications 2011 / 0071208A1 and 2011 / 0076335A1, and their disclosures are incorporated herein by reference in their entirety.
[0296] In other embodiments, the LNP comprising one or more nucleic acids comprises (a) mRNA encoding a long-term repressor fusion protein and gRNA of any embodiment described herein; (b) one or more cationic lipids or ionizable cationic lipids or salts thereof comprising about 2 mol% to about 50 mol% of the total lipids present in the LNP; (c) one or more non-cationic lipids or ionizable cationic lipids comprising about 5 mol% to about 90 mol% of the total lipids present in the LNP; and (d) one or more conjugate lipids that inhibit particle aggregation comprising about 0.5 mol% to about 20 mol% of the total lipids present in the LNP. In some embodiments, the LNP comprises mRNA and gRNA encoding CasX as described herein.
[0297] In other embodiments, the LNP comprising one or more nucleic acids comprises (a) mRNA encoding a long-term repressor fusion protein and gRNA of any embodiment described herein; (b) a cationic lipid or salt thereof comprising about 30 mol% to about 50 mol% of the total lipids present in the LNP; (c) a mixture of phospholipids and cholesterol or derivatives thereof comprising about 47 mol% to about 69 mol% of the total lipids present in the LNP; and (d) a PEG-lipid conjugate comprising about 1 mol% to about 3 mol% of the total lipids present in the LNP. In certain embodiments, the formulation is a four-component system comprising about 2 mol% PEG-lipid conjugate (e.g., PEG2000-C-DMA), about 40 mol% cationic lipid (e.g., DLin-K-C2-DMA) or a salt thereof, about 10 mol% DPPC (or DSPC), and about 48 mol% cholesterol (or a derivative thereof). In some embodiments, the LNP comprises mRNA and gRNA encoding CasX as described herein.
[0298] In other embodiments, the LNP comprising one or more nucleic acids comprises (a) mRNA encoding a long-term repressor fusion protein and gRNA of any embodiment described herein; (b) one or more cationic lipids or ionizable cationic lipids or salts thereof comprising about 50 mol% to about 65 mol% of the total lipids present in the LNP; (c) one or more non-cationic lipids or ionizable cationic lipids comprising about 25 mol% to about 45 mol% of the total lipids present in the LNP; and (d) one or more conjugate lipids that inhibit particle aggregation comprising about 5 mol% to about 10 mol% of the total lipids present in the LNP. In some embodiments, the LNP comprises mRNA and gRNA encoding CasX as described herein.
[0299] In other embodiments, the LNP comprising one or more nucleic acids comprises (a) mRNA encoding a long-term repressor fusion protein and gRNA of any of the embodiments described herein; (b) cationic lipids or salts thereof comprising about 50 mol% to about 60 mol% of the total lipids present in the LNP; (c) a mixture of phospholipids and cholesterol or derivatives thereof comprising about 35 mol% to about 45 mol% of the total lipids present in the LNP; and (d) PEG-lipid conjugates comprising about 5 mol% to about 10 mol% of the total lipids present in the LNP.
[0300] In certain embodiments, the noncationic lipid mixture in the formulation comprises (i) about 10 mol% to about 70 mol% of the total lipids present in the LNP, phospholipids; (ii) about 15 mol% to about 50 mol% of the total lipids present in the LNP, cholesterol or its derivatives; and 1 to 5% lipid-PEG or lipid-PEG-peptides. In certain embodiments, the formulation is a four-component system comprising about 7 mol% PEG-lipid conjugate (e.g., PEG750-C-DMA), about 54 mol% cationic lipid (e.g., DLin-K-C2-DMA) or its salt, about 7 mol% DPPC (or DSPC), and about 32 mol% cholesterol (or its derivatives).
[0301] In other embodiments, the LNP comprising one or more nucleic acids comprises (a) mRNA encoding a long-term repressor fusion protein and / or gRNA of any of the embodiments described herein; (b) a cationic lipid or salt thereof comprising about 55 mol% to about 65 mol% of the total lipids present in the LNP; (c) cholesterol or a derivative thereof comprising about 30 mol% to about 40 mol% of the total lipids present in the LNP; and (d) a PEG-lipid conjugate comprising about 5 mol% to about 10 mol% of the total lipids present in the LNP. In certain embodiments, the formulation is phospholipid-free and is a three-component system comprising about 7 mol% PEG-lipid conjugate (e.g., PEG750-C-DMA), about 58 mol% cationic lipid (e.g., DLin-K-C2-DMA) or a salt thereof, and about 35 mol% cholesterol (or a derivative thereof). In some embodiments, the LNP comprises mRNA and gRNA encoding CasX as described herein.
[0302] In other embodiments, the LNP comprising one or more nucleic acids comprises (a) mRNA encoding a long-term repressor fusion protein and / or gRNA of any embodiment described herein; (b) a cationic lipid or salt thereof comprising about 48 mol% to about 62 mol% of the total lipids present in the LNP; (c) a mixture of phospholipids and cholesterol or derivatives thereof, wherein the phospholipids comprise about 7 mol% to about 17 mol% of the total lipids present in the LNP, and the cholesterol or derivatives comprise about 25 mol% to about 40 mol% of the total lipids present in the LNP; and (d) a PEG-lipid conjugate comprising about 0.5 mol% to about 3.0 mol% of the total lipids present in the LNP. In some embodiments, the LNP comprises mRNA and gRNA encoding CasX as described herein.
[0303] IX. Methods for suppressing target nucleic acids In another embodiment, the Disclosure relates to a method for repressing or silencing the transcription of a target nucleic acid sequence of a gene in a population of cells using the LTRP:gRNA system of the Disclosure, either in vitro, ex vivo, or in vivo in a subject. The programmable nature of the system provided herein allows for precise targeting to achieve a desired effect in one or more regions of a given purpose within the target nucleic acid of the gene. In some embodiments, it may be desirable to repress or silencing a gene in cells containing mutations that cause disease or impairment in a subject.
[0304] In some embodiments, the method involves introducing into a cell the long-term repressor fusion protein of the Disclosure and one or more gRNAs with a targeting sequence complementary to a target nucleic acid, wherein the long-term repressor fusion protein is capable of complexing with the gRNA to form an RNP, and wherein the RNP is a cable for binding to the target nucleic acid and repressing or silencing the transcription of the gene in the cell (it is understood that additional cellular factors may be recruited and involved in the repression). In some embodiments, the method involves introducing into a cell the mRNA encoding the long-term repressor fusion protein of the Disclosure and one or more gRNAs with a targeting sequence complementary to a target nucleic acid, wherein the long-term repressor fusion protein is expressed and capable of complexing with the gRNA to form an RNP, and wherein the RNP is capable of binding to the target nucleic acid and repressing or silencing the transcription of the gene in the cell. In some embodiments, the long-term repressor fusion protein and the mRNA encoding the gRNA may be co-formulated in nanoparticles for delivery to cells of a population. In some embodiments, the mRNA encoding the long-term repressor fusion protein and gRNA may be formulated in separate nanoparticles for delivery to population cells. In some embodiments, the nanoparticles are lipid nanoparticles (LNPs) as described herein.
[0305] In some embodiments of the methods for repressing gene transcription, the LTRP:gRNA system of this disclosure can be designed to target any region of the gene or region of the gene for which transcriptional repression is desired, or a region thereproximal thereto. When repressing an entire gene, the use of a guide with a targeting sequence complementary to the transcription start site (TSS) or a sequence proximal thereto is envisioned in this disclosure. The core promoter serves as a binding platform for the transcription mechanism, and it includes Pol II and its associated general transcription factors (GTFs) (Haberle, V. et al. Eukaryotic core promoters and the functional basis of transcription initiation (Nat Rev Mol Cell)). Biol. 19(10):621(2018)). Variability in TSS selection has been proposed to include DNA "scrunching" and "anti-scrunching," characterized by (i) forward and backward movement of the RNA polymerase leading edge, but not the trailing edge, relative to the DNA, and (ii) expansion and contraction of the transcription bubble. In some embodiments, the targeting sequence of the gRNA in the LTRP:gRNA system is complementary to the target nucleic acid sequence located within 1 kb of the transcription start site (TSS) in the gene targeted for repression. Part of this method In some embodiments of this method, the targeting sequence of the LTRP:gRNA system gRNA is complementary to a target nucleic acid sequence located within 20 bp, 50 bp, 100 bp, 150 bp, 200 bp, 250 bp, 500 bp, 1 kb, or 1.5 kb upstream of the TSS of the gene targeted for repression. In some embodiments of this method, the targeting sequence of the LTRP:gRNA system gRNA is complementary to a target nucleic acid sequence located within 20 bp, 50 bp, 100 bp, 150 bp, 200 bp, 250 bp, 500 bp, 1 kb, or 1.5 kb downstream of the TSS of the gene targeted for repression.In some embodiments of this method, the targeting sequence of the LTRP:gRNA system gRNA is complementary to a target nucleic acid sequence located within 700 bp upstream to 700 bp downstream, 500 bp upstream to 500 bp downstream, 300 bp upstream to 300 bp downstream, or 100 bp upstream to 100 bp downstream of the gene's TSS. In some embodiments of this method, the targeting sequence of the LTRP:gRNA system gRNA is complementary to a target nucleic acid sequence located within 20 bp, 50 bp, 100 bp, 150 bp, 200 bp, 250 bp, 500 bp, or 1 kb of the enhancer of the gene targeted for repression. In some embodiments of this method, the targeting sequence of the LTRP:gRNA system gRNA is complementary to a target nucleic acid sequence located within 1 kb of the 3' to 5' untranslated region of the gene targeted for repression. In some embodiments of this method, the targeting sequence of the gRNA in the LTRP:gRNA system is complementary to the target nucleic acid sequence located within the open reading frame of the gene targeted for repression. In some embodiments of this method, the targeting sequence of the gRNA in the LTRP:gRNA system is complementary to the target nucleic acid sequence of an exon of the gene targeted for repression. In certain embodiments, the targeting sequence of the gRNA in the system of this disclosure is complementary to the target nucleic acid sequence of exon 1 of the gene targeted for repression. In other embodiments of this method, the targeting sequence of the gRNA in the system of this disclosure is complementary to the target nucleic acid sequence of an intron of the gene targeted for repression. In other embodiments of this method, the targeting sequence of the gRNA in the system of this disclosure is complementary to the target nucleic acid sequence of the intron-exon junction of the gene targeted for repression. In other embodiments of this method, the targeting sequence of the gRNA in the system of this disclosure is complementary to the target nucleic acid sequence of a regulatory element of the gene targeted for repression. In other embodiments of the Method, the targeting sequence of the gRNA in the System of the Disclosure is complementary to the sequence of the intergenetic region of the gene targeted for repression. In other embodiments of the Method, the targeting sequence of the gRNA in the System of the Disclosure is complementary to the junction, intron, or regulatory element of the exon of the gene targeted for repression.When the targeting sequence is complementary to a regulatory element, such regulatory elements include, but are not limited to, regions containing promoter regions, enhancer regions, intergenetic regions, 5' untranslated regions (5'UTR), 3' untranslated regions (3'UTR), conserved elements, and cis-regulatory elements. In some embodiments of the Method, the targeting sequence of the system's gRNA is complementary to a target nucleic acid sequence within 1 kb of the enhancer of the gene targeted for repression. In some embodiments of the Method, the targeting sequence of the system's gRNA of the Disclosure is complementary to a target nucleic acid sequence within the 3' untranslated region of the gene targeted for repression. The promoter region is intended to contain nucleotides within 5 kb of the start of the coding sequence, or, in the case of a gene enhancer element or conserved element, may be thousands of bp, hundreds of thousands of bp, or even millions of bp away from the coding sequence of the gene targeted for repression. As described above, the target is a target whose coding gene is intended to be repressed and / or epigenetically modified so that the gene product is not expressed in the cell or is expressed at a lower level. In some embodiments, upon binding of the RNP of the System to the binding site of the target nucleic acid, the System can repress the transcription of the gene at the 5' position relative to the RNP binding site. In other embodiments, upon binding of the RNP of the System to the binding site of the target nucleic acid, the System can repress the transcription of the gene at the 3' position relative to the RNP binding site.
[0306] This disclosure provides a method for transcriptional repression or silencing of a target gene in a population of cells. In some embodiments, the method involves contacting a population of cells with an LTRP:gRNA system containing mRNA having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with the sequence selected from the group consisting of SEQ ID NOs: 2409-18636, or the sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%, wherein the gRNA is SEQ ID NOs: 1744-1746, 2 The LTRP:gRNA system comprises a scaffold containing a sequence selected from the group consisting of 136-2144, 2146-2154, and 2156-2164, or a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity, wherein the gRNA contains a ligated targeting sequence complementary to the target nucleic acid of the gene to be repressed. In some embodiments, the LTRP:gRNA system comprises a gRNA variant containing the sequence of SEQ ID NO: 1744. In some embodiments, the LTRP:gRNA system comprises a gRNA variant containing the sequence of SEQ ID NO: 1745. In some embodiments, the LTRP:gRNA system comprises a gRNA variant containing the sequence of SEQ ID NO: 1746. In some embodiments, the LTRP:gRNA system includes one or more chemically modified gRNA variants, including gRNA variants containing the sequences of SEQ ID NOs. 2136-2144; 2146-2154; or 2156-2164, and is accompanied by a targeting sequence complementary to a target nucleic acid, with 20 nucleotides substituted on the 3' end of the gRNA of the enumerated sequences. In certain embodiments, the LTRP:gRNA system includes mRNA encoding an LTRP, containing a sequence selected from the group consisting of SEQ ID NOs. 2411, 2421, 2467, and 2477, and the gRNA includes a ligated targeting sequence complementary to the sequence of a gene targeted for repression or silencing, with 20 nucleotides substituted on the 3' end of SEQ ID NOs. 2156.
[0307] In some embodiments, the method for transcriptional repression or silencing of target nucleic acids involves contacting a population of cells with an LNP comprising an LTRP:gRNA system, which comprises an mRNA containing a sequence selected from the group consisting of sequences SEQ ID NOs. 2409 to 18636, or a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with that mRNA, and a gRNA containing a scaffold containing a sequence selected from the group consisting of sequences SEQ ID NOs. 1744 to 1746. Alternatively, the gRNA comprises a sequence selected from the group consisting of 2136-2144, 2146-2154, and 2156-2164, wherein the targeting sequence complementary to the target nucleic acid is substituted for 20 nucleotides on the 3' end of the gRNA of the listed sequence, or for a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity, and the gRNA comprises a targeting sequence complementary to the target nucleic acid of the gene to be repressed. In some embodiments of the method, the LNP comprises an LTRP:gRNA system comprising a gRNA variant comprising a scaffold containing the sequence of SEQ ID NO: 1744. In some embodiments of the method, the LNP comprising an LTRP:gRNA system comprises a gRNA variant comprising a scaffold containing the sequence of SEQ ID NO: 1745. In some embodiments of this method, the LNP containing the LTRP:gRNA system includes a gRNA variant containing a scaffold containing the sequence of SEQ ID NO: 1746. In some embodiments of this method, the LNP containing the LTRP:gRNA system includes a chemically modified gRNA variant containing the sequences of SEQ ID NOs: 2136-2144; 2146-2154; and 2156-2164, and includes a targeting sequence complementary to the target nucleic acid, with 20 nucleotides substituted on the 3' end of the gRNA of the listed sequences.In a particular embodiment of this method, the LNP comprising the LTRP:gRNA system comprises an mRNA encoding an LTRP containing a sequence selected from the group consisting of SEQ ID NOs: 2411, 2421, 2467, and 2477, and a gRNA containing the scaffold portion of SEQ ID NO: 2156, accompanied by a ligated targeting sequence complementary to the sequence of the gene targeted for repression or silencing, with 20 nucleotides substituted on the 3' end of SEQ ID NO: 2156.
[0308] In some embodiments of the Method, contacting cells with the LTRP:gRNA system of the Disclosure results in transcriptional repression of the targeted gene in at least about 1%, at least about 2%, at least about 3%, at least about 4%, at least about 5%, at least about 6%, at least about 7%, at least about 8%, at least about 9%, or at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, or at least about 80% of cells in a population targeted by the LTRP:gRNA system. In some embodiments of the Method, the gene in cells targeted by the LTRP:gRNA system is repressed or silenced so that the expression of the gene-encoded protein is reduced by at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, or at least about 90% compared to cells in which the gene is not targeted. In some embodiments, the repression of gene transcription in cells persists for at least about 8 hours, at least about 1 day, at least about 7 days, at least about 1 month, at least about 2 months, at least about 3 months, at least about 4 months, at least about 5 months, or at least about 6 months or longer. In some embodiments, the repression of transcription in cells treated with the LTRP:gRNA system of the embodiment is heritable and stable through one or more cell divisions. In some embodiments, the repression of transcription is stable through 1, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 cell divisions, or more. In some embodiments, the repression is determined in an in vitro assay. In some embodiments, the repression is determined in the subject by assay of cells removed from the subject, or by assay of proteins or markers in samples obtained from the subject.
[0309] The systems and methods described herein can be used in various cells associated with disease, such as liver, intestine, lung, heart, bone, kidney, eye, central nervous system, smooth muscle cells, macrophages, or arterial wall cells, in which gene products contributing to the disease or symptoms are repressed or silenced. In some embodiments, the cells targeted for transcriptional repression are eukaryotic cells. In some embodiments, the eukaryotic cells are selected from the group consisting of rodent cells, mouse cells, rat cells, primate cells, and non-human primate cells. In some embodiments, the eukaryotic cells are human cells. In some embodiments of this method, the cells are embryonic stem cells, induced pluripotent stem cells, germ cells, fibroblasts, oligodendrocytes, glial cells, hematopoietic stem cells, neuronal progenitor cells, neurons, astrocytes, muscle cells, osteocytes, hepatocytes, pancreatic cells, lung cells, kidney cells, retinal cells, cancer cells, T cells, B cells, NK cells, fetal cardiomyocytes, myofibroblasts, mesenchymal stem cells, autologous transplanted dilated cardiomyocytes, adipocytes, totipotent cells, pluripotent cells, hematopoietic stem cells, myoblasts, bone marrow cells, mesenchymal cells, parenchymal cells, epithelial cells, endothelial cells, mesothelial cells, fibroblasts, osteoblasts, chondrocytes, bone marrow-derived progenitor cells, cardiomyocytes, skeletal cells, fetal cells, undifferentiated cells, multipotent progenitor cells, unipotent progenitor cells, monocytes, cardiomyocytes, skeletal myoblasts, macrophages, capillary endothelial cells, xenogeneic cells, allogeneic cells, or postnatal stem cells.
[0310] This disclosure provides a method for reversing transcriptional repression resulting from the LTRP:gRNA system. In some embodiments, the transcriptional repression is reversible by the use of a DNMT inhibitor. In some embodiments of the method, the transcriptional repression is reversible by the use of a cytidine analog inhibitor of DNMT. In some embodiments, the transcriptional repression is reversible by treating cells with an inhibitor selected from the group consisting of azacitidine, decitabine, clofarabine, and zebralin. In some embodiments of the method, the reversal of transcriptional repression occurs in a subject treated with the system of the disclosure, and the method includes the administration of a therapeutically effective dose of a DNMT inhibitor.
[0311] X. Treatment method This disclosure provides a method for treating a disease or disorder in a subject requiring its use, using the LTRP:gRNA system of this disclosure. In some embodiments, the method of this disclosure can prevent, treat, and / or induce remission of a disease or disorder in a subject by administering a therapeutically effective dose of the composition of this disclosure to the subject. This approach can therefore be used for application in subjects with diseases or disorders, such as, but not limited to, autosomal dominant hypercholesterolemia (ADH), hypercholesterolemia, elevated total cholesterol levels, dyslipidemia, elevated low-density lipoprotein (LDL) levels, elevated LDL cholesterol levels, decreased high-density lipoprotein levels, fatty liver, coronary artery disease, ischemia, stroke, peripheral vascular disease, thrombosis, type 2 diabetes, high elevated blood pressure, atherosclerosis, obesity, Alzheimer's disease, neurodegeneration, or age-related macular degeneration (AMD). 【03...
Claims
1. A system for repressing gene transcription, wherein the system is (a) mRNA encoding a long-term repressor fusion protein (LTRP), The LTRP comprises mRNA containing a DNA methyltransferase (DNMT) 3A catalytic domain (DNMT3A), a DNMT3-like interaction domain (DNMT3L), a DNA-binding protein containing a non-catalyzed CasX (dCasX), and a first repressor domain (RD1) from the N-terminus to the C-terminus, (b) A system comprising a guide ribonucleic acid (gRNA) containing a targeting sequence complementary to the target nucleic acid sequence of a gene in a cell.
2. The system according to claim 1, wherein the LTRP includes a DNMT3A ATRX-DNMT3-DNMT3L domain (ADD) ligated to the DNMT3A at its N-terminus.
3. The system according to claim 1 or claim 2, wherein the mRNA comprises a sequence encoding the dCasX selected from the group consisting of SEQ ID NOs. 1948, 2405, and 2406, or a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto.
4. The system according to claim 3, wherein the mRNA comprises a sequence encoding the dCasX including sequence number 2406, or a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto.
5. The system according to any one of claims 1 to 3, wherein the mRNA comprises a sequence encoding the DNMT3A, including SEQ ID NO: 1955 or SEQ ID NO: 21878, or a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto.
6. The system according to any one of claims 1 to 4, wherein the mRNA comprises a sequence encoding the DNMT3L including SEQ ID NO: 1945 or SEQ ID NO: 21879, or a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto.
7. The system according to any one of claims 1 to 6, wherein the mRNA comprises a sequence encoding the RD1 selected from the group consisting of SEQ ID NOs. 18637 to 21830, or a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto.
8. The system according to claim 7, wherein the mRNA comprises a sequence encoding the RD1 selected from the group consisting of SEQ ID NOs. 18637 to 18646 and 20234 to 20243, or a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto.
9. The system according to claim 7 or claim 8, wherein the mRNA comprises a sequence encoding the RD1 selected from the group consisting of SEQ ID NOs. 18642 and 20239, or a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto.
10. The system according to claim 7 or claim 8, wherein the mRNA comprises a sequence encoding the RD1 selected from the group consisting of SEQ ID NOs. 18637 and 20234, or a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto.
11. The system according to claim 7 or claim 8, wherein the mRNA comprises a sequence encoding the RD1 selected from the group consisting of SEQ ID NOs. 18638 and 20235, or a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto.
12. The system according to any one of claims 1 to 11, wherein the mRNA comprises one or more sequences encoding nuclear localization sequences (NLS).
13. The system according to claim 12, wherein the mRNA comprises a sequence encoding one or more NLSs including sequence number 21875.
14. The system according to any one of claims 1 to 13, wherein the mRNA comprises one or more sequences encoding a linker peptide.
15. The system according to any one of claims 1 to 14, wherein the mRNA comprises a sequence encoding the LTRP selected from the group consisting of SEQ ID NOs. 2410 to 2428 and 2466 to 2484, or a sequence having at least about 70%, at least about 80%, at least about 90%, or at least about 95%, or at least about 95%, or at least about 96%, or at least about 97%, or at least about 98%, or at least about 99% sequence identity thereto.
16. The system according to claim 15, wherein the mRNA comprises a sequence encoding the LTRP selected from the group consisting of SEQ ID NOs: 2411, 2421, 2467, and 2477, or a sequence having at least about 70%, at least about 80%, at least about 90%, or at least about 95%, or at least about 95%, or at least about 96%, or at least about 97%, or at least about 98%, or at least about 99% sequence identity thereto.
17. The system according to claim 15, wherein the mRNA comprises a sequence encoding the LTRP selected from the group consisting of SEQ ID NOs: 2410, 2420, 2466, and 2476, or a sequence having at least about 70%, at least about 80%, at least about 90%, or at least about 95%, or at least about 95%, or at least about 96%, or at least about 97%, or at least about 98%, or at least about 99% sequence identity thereto.
18. The system according to claim 15, wherein the mRNA comprises a sequence encoding the LTRP selected from the group consisting of SEQ ID NOs. 2412, 2422, 2468, and 2478, or a sequence having at least about 70%, at least about 80%, at least about 90%, or at least about 95%, or at least about 95%, or at least about 96%, or at least about 97%, or at least about 98%, or at least about 99% sequence identity thereto.
19. (a) The dCasX contains the amino acid sequences of SEQ ID NOs. 4 to 29, or a sequence having at least about 70%, at least about 80%, at least about 90%, at least about 95%, or at least about 99% sequence identity thereto. (b) The RD1 includes an amino acid sequence selected from the group consisting of SEQ ID NOs: 130 to 224, or a sequence having at least about 70%, at least about 80%, at least about 90%, at least about 95%, or at least about 99% sequence identity thereto. (c) The DNMT3A contains the amino acid sequence of SEQ ID NO: 126, or a sequence having at least about 70%, at least about 80%, at least about 90%, at least about 95%, or at least about 99% sequence identity thereto, and / or (d) The system according to any one of claims 15 to 18, wherein the DNMT3L comprises the amino acid sequence of SEQ ID NO: 127, or a sequence having at least about 70%, at least about 80%, at least about 90%, at least about 95%, or at least about 99% sequence identity thereto.
20. (a) The dCasX contains the amino acid sequence of SEQ ID NOs: 4 to 29, (b) The RD1 comprises an amino acid sequence selected from the group consisting of SEQ ID NOs: 130 to 224, and optionally the RD1 comprises an amino acid sequence selected from the group consisting of SEQ ID NOs: 130, 131, and 135. (c) The DNMT3A comprises the amino acid sequence of SEQ ID NO: 126, and / or (d) The system according to any one of claims 15 to 18, wherein the DNMT3L comprises the amino acid sequence of SEQ ID NO:
127.
21. The system according to any one of claims 15 to 20, wherein the mRNA sequence encoding the LTRP is codon-optimized.
22. The system according to any one of claims 1 to 21, wherein the targeting sequence of the gRNA is complementary to a target nucleic acid sequence within 1 kb of the transcription start site (TSS) in the gene.
23. The system according to claim 22, wherein the gRNA includes a scaffold containing the sequence of Sequence ID No. 1746, or a sequence having at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto.
24. The system according to claim 23, wherein the gRNA is chemically modified.
25. The system according to claim 24, wherein the chemical modification comprises the addition of a 2'O-methyl group to one or more nucleotides of the gRNA.
26. The system according to claim 24 or 25, wherein one or more nucleotides located one, two, three, or four nucleotides from the 5' end, 3' end, or both ends of the gRNA are modified by the addition of a 2'O-methyl group.
27. The system according to any one of claims 24 to 26, wherein the chemical modification to the gRNA includes substitution of phosphorothioate bonds between two or more nucleotides of the gRNA.
28. The system according to claim 27, wherein the chemical modification includes the substitution of a phosphorothioate bond between two or more nucleotides located one, two, three, or four nucleotides from the 5' end, 3' end, or both ends of the gRNA.
29. The system according to any one of claims 24 to 28, wherein the gRNA comprises a sequence selected from sequence numbers 2156 to 2164, and comprises a targeting sequence complementary to a target nucleic acid substituted for 20 nucleotides on the 3' end of sequence numbers 2156 to 2164.
30. The system according to any one of claims 1 to 29, wherein the mRNA comprises a 5'UTR, a 3'UTR, a poly(A) sequence, and / or a 5' cap.
31. Lipid nanoparticles (LNPs) comprising the system according to any one of claims 1 to 30.
32. A pharmaceutical composition comprising the system described in any one of claims 1 to 30 or the LNP described in claim 31 and a pharmaceutically acceptable carrier, diluent, or excipient.
33. A method for suppressing gene transcription in a population of cells, the method comprising contacting the cells of the population with the system described in any one of claims 1 to 30, the LNP described in claim 31, or the pharmaceutical composition described in claim 32, wherein the contact results in suppression of gene transcription in the population of cells.
34. A composition for use in treating a disease in a subject requiring it, wherein the composition comprises a therapeutically effective dose of the system described in any one of claims 1 to 30, the LNP described in claim 31, or the pharmaceutical composition described in claim 32, wherein the transcription of a target gene in the subject is suppressed by the LTRP, thereby treating the disease.
35. A composition for use in the manufacture of a pharmaceutical product for treating a disease in a subject that requires it, wherein the composition comprises a therapeutically effective dose of the system described in any one of claims 1 to 30 or the LNP described in claim 31, wherein the transcription of a target gene in the subject is suppressed by the LTRP, thereby treating the disease.
36. A kit comprising the system according to any one of claims 1 to 30, the LNP according to claim 31, or the pharmaceutical composition according to claim 32, and instructions for use.