Compositions, systems and methods for modulating hepatitis b virus by targeting gene repression
By using an epigenetic DNA-targeting system, fusion proteins and gRNAs are used to target the HBV gene or its regulatory elements, thereby inhibiting HBV gene transcription. This addresses the efficiency and stability issues of existing treatment methods, resulting in reduced HBV replication and lower protein levels.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TUNE THERAPEUTICS INC
- Filing Date
- 2023-08-18
- Publication Date
- 2026-04-17
AI Technical Summary
Existing standard care methods, such as the administration of nucleoside analogs, pegylated interferon, antisense oligonucleotides, and siRNA, present challenges in terms of efficiency and stability in treating hepatitis B virus (HBV) infection, necessitating new and improved approaches to inhibit HBV gene transcription.
The epigenetic DNA-targeting system comprises a fusion protein and multiple guide RNAs (gRNAs). The fusion protein includes a DNA-binding domain and a transcriptional repressor effector domain. It represses the transcription of the HBV gene by targeting multiple target sites of the HBV gene or its regulatory elements.
It effectively reduces HBV replication and lowers HBV protein levels, avoiding gene damage or DNA breakage, and providing more stable treatment results.
Smart Images

Figure CN121874158A_ABST
Abstract
Description
This application is a divisional application of Chinese Patent Application No. 202380073064.1, filed on August 18, 2023, entitled "Composition, System and Method for Regulating Hepatitis B Virus by Targeted Gene Repression". Cross-reference to related applications
[0001] This application claims priority to U.S. Provisional Application No. 63 / 399,634, filed August 19, 2022, entitled “Compositions, systems, and methods for regulating hepatitis B virus by targeting gene inhibition”; U.S. Provisional Application No. 63 / 472,236, filed June 9, 2023, entitled “Compositions, systems, and methods for regulating hepatitis B virus by targeting gene inhibition”; and U.S. Provisional Application No. 63 / 531,309, filed August 7, 2023, entitled “Compositions, systems, and methods for regulating hepatitis B virus by targeting gene inhibition”, the contents of which are incorporated herein by reference in their entirety. By referencing and incorporating into the sequence list
[0002] This application is submitted together with an electronic sequence list. The sequence list is provided as a file with the title 224742002040SeqList.xml, created on August 18, 2023, and has a size of 1,595,298 bytes. Information from the electronic sequence list is incorporated herein by reference in its entirety. Technical Field
[0003] In some aspects, this disclosure relates to epigenetically modified DNA targeting systems, such as CRISPR-Cas / guide RNA (gRNA) systems, for transcriptional repression of hepatitis B virus (HBV) genes to promote cellular phenotypes leading to reduced HBV infection. In some embodiments, the epigenetically modified DNA targeting system binds to or targets a target site of at least one gene or regulatory element thereof in a cellular HBV DNA sequence. In some embodiments, the system is a multi-gene system that binds to or targets target sites of at least two genes or regulatory elements thereof. In some aspects, the systems of this disclosure relate to transcriptional repression of one or more HBV genes. In some aspects, this disclosure relates to methods and uses associated with the provided compositions, such as in the repression of HBV replication and expression in relation to the treatment of HBV infection. Background Technology
[0004] A large patient population (estimated at 1 million in the United States alone and 250 million worldwide) suffers from chronic hepatitis B infection. However, current standard of care (including methods to inhibit viral DNA transcription such as nucleoside analogs, pegylated interferon, antisense oligonucleotides, and siRNA) faces challenges in terms of efficacy and stability. Therefore, new and improved approaches are needed to overcome these challenges. This publication addresses these and other needs. Summary of the Invention
[0005] This document provides an epigenetically modified DNA targeting system comprising at least one DNA targeting module for repressing transcription of one or more hepatitis B virus (HBV) genes, wherein each of the at least one DNA targeting module comprises a fusion protein comprising: (a) a DNA-binding domain for targeting a target site in a hepatitis B virus DNA sequence; and (b) at least one transcriptional repressor effector domain. In some embodiments, the at least one DNA-binding domain comprises a clustered regularly spaced short palindromic repeat-associated (Cas)-guide RNA (gRNA) combination comprising (a) a Cas protein or a variant thereof and (b) at least one gRNA; a zinc finger protein (ZFP); a transcription activator-like effector (TALE); a meganuclease; a homing endonuclease; or an I-SceI enzyme or a variant thereof, optionally wherein the DNA-binding domain comprises a non-catalytically inactive variant of any of the foregoing. In some embodiments, the hepatitis B virus DNA sequence is an HBV gene or a regulatory element thereof. In some embodiments, the at least one DNA targeting module comprises a plurality of DNA targeting modules for targeting multiple target sites of one or more genes or their regulatory elements. In some embodiments, the plurality of DNA targeting modules comprises at least a first DNA targeting module and a second DNA targeting module, wherein: (1) the first DNA targeting module represses transcription of a first HBV gene, wherein the first DNA targeting module comprises a first fusion protein, the first fusion protein comprising (a) a DNA-binding domain for targeting a target site of the first gene or its regulatory DNA element; and (b) at least one transcriptional repressor domain; and (2) the second DNA targeting module represses transcription of a second HBV gene, wherein the second DNA targeting module comprises a second fusion protein, the second fusion protein comprising (a) a DNA-binding domain for targeting a target site of the second gene or its regulatory DNA element; and (b) At least one transcriptional repressor domain, optionally wherein: the first DNA targeting module and the second DNA targeting module have the same fusion protein, such that the first fusion protein and the second fusion protein are identical, and wherein the DNA-binding domain of the fusion protein is a clustered regularly spaced short palindromic repeat (Cas) protein or a variant thereof; and the first DNA targeting module contains a first guide RNA (gRNA) targeting a target site of a first HBV gene or a regulatory element thereof, and the second DNA targeting module contains a second gRNA targeting a target site of a second HBV gene or a regulatory element thereof.
[0006] This document also provides an epigenetically modified DNA targeting system for repressing transcription of one or more hepatitis B virus (HBV) genes, wherein the DNA targeting system comprises: (a) a fusion protein comprising a clustered regularly spaced short palindromic repeat (Cas) protein or a variant thereof and at least one transcriptional repressor effector domain; and (b) multiple guide RNAs (gRNAs), wherein the multiple guide RNAs comprise at least a first gRNA and a second gRNA, wherein the first gRNA targets a target site of a first HBV gene or a regulatory element thereof, and the second gRNA targets a target site of a second HBV gene or a regulatory element thereof, wherein the first gene and the second gene or regulatory element thereof regulate HBV replication and / or HBV transcription. In some embodiments, the DNA targeting system further comprises a third gRNA targeting a target site of a third gene or a regulatory element thereof, wherein the third gene or regulatory element thereof regulates HBV replication and / or HBV transcription. In some embodiments, the system further comprises a fourth gRNA targeting a target site of a fourth gene or its regulatory element, a fifth gRNA optionally targeting a fifth gene or its regulatory element, and / or a sixth gRNA optionally targeting a target site of a sixth gene or its regulatory element, wherein said gene or its regulatory element regulates hepatitis B virus replication and / or HBV transcription. In some embodiments, the first gene, the second gene, the third gene, the fourth gene, the fifth gene, and / or the sixth gene or its regulatory element are different.
[0007] This article also provides an epigenetically modified DNA targeting system for repressing transcription of one or more hepatitis B virus (HBV) genes, wherein the DNA targeting system comprises: (a) a fusion protein comprising a clustered regularly spaced short palindromic repeat (Cas) protein or a variant thereof and at least one transcriptional repressor effector domain; and (b) multiple guide RNAs (gRNAs) targeting multiple target sites of multiple genes or regulatory elements thereof, wherein the multiple genes or regulatory elements thereof regulate hepatitis B virus replication and / or HBV transcription.
[0008] This article also provides an epigenetic modified DNA targeting system comprising a single DNA targeting module for repressing transcription of more than one hepatitis B virus (HBV) gene, wherein the DNA targeting module comprises: (a) a fusion protein comprising a clustered regularly spaced short palindromic repeat (Cas) protein or a variant thereof and at least one transcriptional repressor effector domain; and (b) a guide RNA (gRNA) targeting multiple target sites of multiple genes or their regulatory elements, wherein the multiple genes or their regulatory elements regulate hepatitis B virus replication and / or HBV transcription.
[0009] In any implementation thereof, transcriptional repression results in reduced HBV replication and / or decreased HBV protein levels.
[0010] In any embodiment described herein, the DNA targeting system does not introduce gene damage or DNA breakage.
[0011] In any embodiment herein, the at least one DNA binding module comprises a plurality of DNA binding modules that together target a plurality of target sites in the HBV DNA sequence, optionally wherein each DNA binding module targets a different target site in the HBV DNA sequence.
[0012] In any embodiment of this document, the plurality of target sites is 2, 3, 4, 5, or 6 different target sites. In any embodiment of this document, the plurality of target sites are each located in a different HBV gene or its regulatory element.
[0013] In any implementation thereof, each target site is located in the same HBV gene or its regulatory element.
[0014] In any embodiment of this document, the system comprises 2 to 10 DNA targeting modules.
[0015] In any embodiment herein, any two or more of the DNA targeting modules have the same fusion protein, or any two or more of the DNA targeting modules contain different fusion proteins.
[0016] In any embodiment herein, the DNA-binding domain of each DNA-targeting module comprises a fusion protein comprising a clustered regularly spaced short palindromic repeat (Cas) protein or a variant thereof and at least one transcriptional repressor effector domain, and wherein each DNA-targeting module comprises a unique gRNA.
[0017] In any embodiment herein, the target site, or each of the target sites, is present in the form of covalently closed circular DNA (cccDNA), relaxed circular DNA (rcDNA), and / or in HBV viral DNA integrated into human genomic DNA. In any embodiment herein, the target site, or each of the target sites, is located at or near a gene or regulatory element thereof involved in controlling HBV replication and / or HBV transcription.
[0018] In any embodiment herein, the gene involved in controlling HBV replication and / or HBV transcription encodes a polymerase, envelope protein, capsid protein, transcription factor, or transcriptional transactivator. In any embodiment herein, the gene involved in controlling HBV replication and / or HBV transcription is a polymerase gene, an S-family gene, an X-gene, or a core family gene.
[0019] In any embodiment of this document, at least one target site is located in the gene encoding the X gene of the hepatitis B virus X protein (HBx) or a regulatory element thereof.
[0020] In any embodiment herein, the target site, or each of the target sites, is located at or near a regulatory element of the HBV gene involved in controlling HBV replication and / or HBV transcription. In some embodiments, the regulatory element is a promoter region. In some embodiments, the promoter region is a pre-S1 promoter, a pre-S2 promoter, an X promoter, or a basic core promoter. In some embodiments, the regulatory element is an enhancer region. In some embodiments, the enhancer region is an Enh1 or Enh2 enhancer region. In some embodiments, the regulatory element is a transcript processing control region.
[0021] In any embodiment herein, the target site, or each of the target sites, is located within the coding region of the HBV gene. In any embodiment herein, the target site, or each of the target sites, is located within 500 base pairs (bp), 1000 bp, or 1500 bp of the transcription start site. In any embodiment herein, the target site, or each of the target sites, is located within a target region located between 0 and 3300 base pairs (bp) of the HBV genome, optionally between 0 and 3182 bp corresponding to the position in the HBV genome shown in reference SEQ ID NO: 650. In any embodiment herein, the target site, or each of the target sites, is located within a target region, which is located at a base pair between 43 bp-490 bp, 1033 bp-1749 bp, 1800 bp-1950 bp, or 2953 bp-3182 bp of the HBV genome corresponding to the position shown in reference SEQ ID NO: 650. In any embodiment herein, the target site, or each of the target sites, is located within a target region, which is located at a base pair between 1 bp-42 bp, 491 bp-1032 bp, 1750 bp-1799 bp, or 1951 bp-2952 bp of the HBV genome corresponding to the position shown in reference SEQ ID NO: 650. In any embodiment herein, the target site, or each of the target sites, is located within a CpG island of the HBV genome. In any embodiment herein, the target site, or each of the target sites, is located within a target region located at a base pair between 67 bp-392 bp, 1033 bp-1749 bp, or 2215 bp-2490 bp of the HBV genome corresponding to the position shown in reference SEQ ID NO: 650. In any embodiment herein, the target site, or each of the target sites, is located within a target region located at a base pair between 1033 bp-1749 bp of the hepatitis B virus sequence at the nucleotide position shown in reference SEQ ID NO: 650. In any embodiment herein, the target site, or each of the target sites, is located within a target region located within 300 base pairs upstream of the start codon of the hepatitis B virus X protein (HBx).
[0022] This article also provides an epigenetic modified DNA targeting system comprising at least one DNA targeting module for repressing transcription of one or more hepatitis B virus (HBV) genes, wherein each of the at least one DNA targeting module comprises a fusion protein comprising: (a) a DNA-binding domain for targeting a target site within a target region spanning 300 base pairs upstream of the start codon of the hepatitis B virus X protein (HBx); and (b) at least one transcriptional repressor effector domain.
[0023] In any embodiment herein, the target site, or each of the target sites, is located within the HBx basal core promoter region. In any embodiment herein, the target site, or each of the target sites, is located within the HBx promoter / enhancer region.
[0024] In any embodiment herein, the target site, or each of the target sites, is located within a target region spanning 250 base pairs upstream of the start codon of the hepatitis B virus X protein (HBx). In any embodiment herein, the target site, or each of the target sites, is located within a target region having a sequence corresponding to a sequence located between 1060 and 1480 bp of the HBV genome as shown in reference SEQ ID NO: 650. In any embodiment herein, the target site, or each of the target sites, is located within a target region spanning 150 base pairs upstream of the start codon of the hepatitis B virus X protein (HBx). In any embodiment herein, the target site, or each of the target sites, is located within a target region spanning 120 base pairs upstream of the start codon of the hepatitis B virus X protein (HBx). In any embodiment herein, the target site, or each of the target sites, is within a target region sequence corresponding to a 1250-1374 bp sequence of the HBV genome spanning the HBV genome shown in reference SEQ ID NO: 650. In any embodiment herein, the target site, or each of the target sites, is within a target region sequence corresponding to a 1255-1302 bp sequence of the HBV genome spanning the HBV genome shown in reference SEQ ID NO: 650. In any embodiment herein, the target site, or each of the target sites, is within a target region sequence corresponding to a 1260-1300 bp sequence of the HBV genome spanning the HBV genome shown in reference SEQ ID NO: 650.
[0025] In any embodiment herein, the target site, or each of the target sites, is at least 70% homologous to all hepatitis B virus genomes. In any embodiment herein, the target site, or each of the target sites, is at least 70% homologous to at least 1000 hepatitis B virus genomes. In any embodiment herein, the target site, or each of the target sites, is at least 70% homologous to at least 1000 hepatitis B virus genomes and contains at most two mismatches.
[0026] In any embodiment herein, the target site or each of the target sites comprises the sequence shown in any one of SEQ ID NO: 1-195, a continuous portion of at least 14 nucleotides (nt) thereof, or a complementary sequence of any one of the foregoing. In any embodiment herein, the target site or each of the target sites comprises the sequence shown in any one of SEQ ID NO: 175, 138, 192, 152, 118, 125, 185, 63, 116, 124, 35, 82, a continuous portion of at least 14 nucleotides (nt) thereof, or a complementary sequence of any one of the foregoing. In any embodiment herein, the target site or each of the target sites comprises the sequence shown in any one of SEQ ID NO: 175, 138, 192, 152, 118, 125, 185, 63, 116, 124, 35, 82. In any embodiment herein, the target site or each of the target sites comprises the sequence shown in any one of SEQ ID NO: 5, 6, 12, 18, 22, 26, 29, 38, 42, 43, 51, 56, 61, 63, 68, 72, 75, 79, 82, 84, 88, 89, 98, 99, 113, 116, 121, 124, 125, 118, 130, 133, 135, 138, 143, 150, 152, 155, 158, 164, 165, 175, 176, 182, 185, 189, 190, 192, a continuous portion of at least 14 nucleotides (nt) thereof, or a complementary sequence to any one of the foregoing. In any embodiment herein, the target site or each of the target sites comprises the sequence shown in any one of SEQ ID NO: 5, 6, 12, 18, 22, 26, 29, 38, 42, 43, 51, 56, 61, 63, 68, 72, 75, 79, 82, 84, 88, 89, 98, 99, 113, 116, 121, 124, 125, 118, 130, 133, 135, 138, 143, 150, 152, 155, 158, 164, 165, 175, 176, 182, 185, 189, 190, 192. In any embodiment herein, the target site, or each of the target sites, is as shown in any one of SEQ ID NO: 12, 18, 20, 22, 26, 27, 46, 50, 63, 66, 73, 79, 185, 192, a continuous portion of at least 14 nucleotides thereof, or a complementary sequence of any of the foregoing. In any embodiment herein, the target site, or each of the target sites, is as shown in any one of SEQ ID NO: 12, 18, 20, 22, 26, 27, 46, 50, 63, 66, 73, 79, 185, 192.In any embodiment herein, the target site, or each of the target sites, comprises the sequence shown in SEQ ID NO: 22, a continuous portion of at least 14 nucleotides thereof, or a complementary sequence of any of the foregoing, optionally wherein the target site is as shown in SEQ ID NO: 22. In any embodiment herein, the target site, or each of the target sites, comprises the sequence shown in SEQ ID NO: 63, a continuous portion of at least 14 nucleotides thereof, or a complementary sequence of any of the foregoing, optionally wherein the target site is as shown in SEQ ID NO: 63.
[0027] In any embodiment herein, the gRNA or each of the gRNAs comprises a gRNA spacer sequence comprising any one of the sequences shown in SEQ ID NO: 196-390. In any embodiment herein, the gRNA or each of the gRNAs further comprises the sequence shown in SEQ ID NO: 587. In any embodiment herein, the gRNA or each of the gRNAs comprises the sequence shown in any one of SEQ ID NO: 196-390. In any embodiment herein, the gRNA or each of the gRNAs is as shown in any one of SEQ ID NO: 391-585. In any embodiment herein, the gRNA or each of the gRNAs comprises the sequence shown in any one of SEQ ID NO: 370, 333, 387, 347, 313, 320, 380, 256, 258, 311, 319, 230, 272, a continuous portion of at least 14 nucleotides thereof, or a complementary sequence of any of the foregoing, optionally wherein the gRNA or each of the gRNAs is shown in any one of SEQ ID NO: 565, 528, 542, 508, 515, 575, 515, 453, 506, 514, 425, or 472. In any embodiment herein, the gRNA or each of the gRNAs comprises the sequence shown in any one of SEQ ID NO: 370, 333, 387, 347, 313, 320, 380, 256, 258, 311, 319, 230, 272, optionally wherein the gRNA or each of the gRNAs is shown in any one of SEQ ID NO: 565, 528, 542, 508, 515, 575, 515, 453, 506, 514, 425, or 472.In any embodiment herein, the gRNA or each of the gRNAs comprises the sequence shown in any one of SEQ ID NO: 200, 201, 207, 217, 221, 224, 233, 237, 238, 246, 251, 256, 258, 263, 267, 274, 270, 277, 279, 283, 284, 293, 294, 308, 311, 313, 316, 319, 320, 325, 328, 330, 333, 338, 345, 347, 350, 353, 359, 360, 370, 371, 377, 380, 384, 385, 387, a continuous portion of at least 14 nucleotides thereof, or a complementary sequence to any one of the foregoing, optionally wherein the gRNA or each of the gRNAs is as shown in SEQ ID NO: 200, 201, 207, 217, 221, 224, 233, 237, 238, 246, 251, 256, 258, 263, 267, 274, 270, 277, 380, 384, 385, 387, attributable to the gRNA ... NO: Any one of 369, 395, 402, 408, 412, 416, 419, 428, 432, 433, 441, 446, 451, 453, 458, 462, 465, 469, 472, 474, 478, 479, 488, 489, 503, 506, 508, 511, 514, 515, 520, 523, 525, 575, 528, 533, 540, 542, 545, 548, 554, 555, 565, 566, 572, 579, 580, or 582. In any embodiment herein, the gRNA or each of the gRNAs comprises the sequence shown in any one of SEQ ID NO: 200, 201, 207, 217, 221, 224, 233, 237, 238, 246, 251, 256, 258, 263, 267, 274, 270, 277, 279, 283, 284, 293, 294, 308, 311, 313, 316, 319, 320, 325, 328, 330, 333, 338, 345, 347, 350, 353, 359, 360, 370, 371, 377, 380, 384, 385, or 387, optionally wherein the gRNA or each of the gRNAs is as shown in SEQ ID NO: Any one of 369, 395, 402, 408, 412, 416, 419, 428, 432, 433, 441, 446, 451, 453, 458, 462, 465, 469, 472, 474, 478, 479, 488, 489, 503, 506, 508, 511, 514, 515, 520, 523, 525, 575, 528, 533, 540, 542, 545, 548, 554, 555, 565, 566, 572, 579, 580, or 582 is shown.In any embodiment herein, the gRNA or each of the gRNAs comprises the sequence shown in any one of SEQ ID NO: 207, 213, 215, 217, 221, 222, 241, 245, 258, 261, 268, 274, 380, 387, a continuous portion of at least 14 nucleotides thereof, or a complementary sequence of any of the foregoing, optionally wherein the gRNA or each of the gRNAs is shown in any one of SEQ ID NO: 402, 408, 410, 412, 416, 417, 436, 440, 453, 456, 463, 469, 575, 582. In any embodiment herein, the gRNA or each of the gRNAs comprises the sequence shown in any one of SEQ ID NO: 207, 213, 215, 217, 221, 222, 241, 245, 258, 261, 268, 274, 380, 387. In any embodiment herein, the gRNA or each of the gRNAs comprises the sequence shown in SEQ ID NO: 402, 408, 410, 412, 416, 417, 436, 440, 453, 456, 463, 469, 575, 582. In any embodiment herein, the gRNA or each of the gRNAs comprises the sequence shown in SEQ ID NO: 217, a continuous portion of at least 14 nucleotides thereof, or a complementary sequence to any of the foregoing. In any embodiment herein, the gRNA or each of the gRNAs comprises the sequence shown in SEQ ID NO: 217. In any embodiment herein, the gRNA or each of the gRNAs is as shown in any of SEQ ID NO: 412. In any embodiment herein, the gRNA or each of the gRNAs comprises the sequence shown in SEQ ID NO: 258, a continuous portion of at least 14 nucleotides thereof, or a complementary sequence to any of the foregoing. In any embodiment herein, the gRNA or each of the gRNAs comprises the sequence shown in SEQ ID NO: 258. In any embodiment herein, the gRNA or each of the gRNAs is as shown in any of SEQ ID NO: 453.
[0028] In any embodiment herein, the target site, or each of the target sites, is at least 90% homologous to all hepatitis B virus genomes. In any embodiment herein, the target site, or each of the target sites, is at least 90% homologous to at least 1000 hepatitis B virus genomes. In any embodiment herein, the target site, or each of the target sites, is at least 90% homologous to at least 1000 hepatitis B virus genomes and contains at most two mismatches, optionally one or both.
[0029] In any embodiment herein, the target site or each of the target sites comprises the sequence shown in any of SEQ ID NO:35-100, a continuous portion of at least 14 nucleotides (nt) thereof, or a complementary sequence of any of the foregoing.
[0030] In any embodiment herein, the gRNA or each of the gRNAs comprises a gRNA spacer sequence comprising any one of the sequences shown in SEQ ID NO: 230-295. In any embodiment herein, the gRNA or each of the gRNAs further comprises the sequence shown in SEQ ID NO: 587. In any embodiment herein, the gRNA or each of the gRNAs comprises the sequence shown in any one of SEQ ID NO: 230-295, optionally wherein the gRNA or each of the gRNAs is shown in any one of SEQ ID NO: 425-490.
[0031] In any embodiment of this document, the at most two mismatches are located in the first 12 nt of the prototype spacer 5' end.
[0032] In any embodiment herein, the target site, or each of the target sites, is at least 90% homologous to at least 1,000 hepatitis B virus genomes and contains zero mismatches.
[0033] In any embodiment herein, the target site comprises the sequence shown in any one of SEQ ID NO: 1-34, a continuous portion of at least 14 nucleotides (nt) therein, or a complementary sequence to any one of the foregoing. In any embodiment herein, the gRNA or each of the gRNAs comprises a gRNA spacer sequence comprising the sequences shown in SEQ ID NO: 196-229.
[0034] In any embodiment herein, the length of the gRNA spacer sequence is between 14 nt and 24 nt or between 16 nt and 22 nt. In any embodiment herein, the length of the gRNA spacer sequence is 18 nt, 19 nt, 20 nt, 21 nt, or 22 nt.
[0035] In any embodiment herein, the gRNA spacer sequence contains modified nucleotides to increase stability.
[0036] In any embodiment herein, the at least one gRNA further comprises the sequence shown in SEQ ID NO: 587. In any embodiment herein, the gRNA or each of the gRNAs comprises the sequence shown in any one of SEQ ID NO: 196-229, optionally wherein the gRNA or each of the gRNAs is shown in any one of SEQ ID NO: 391-424. In any embodiment herein, the gRNA or each of the gRNAs comprises the sequence shown in any one of SEQ ID NO: 207, 213, 215, 217, 221, 222, 241, 245, 258, 261, 268, 274, 380, 387, optionally wherein the gRNA or each of the gRNAs is shown in any one of SEQ ID NO: 402, 408, 410, 412, 416, 417, 436, 440, 453, 456, 463, 469, 575, 582. In any embodiment herein, the gRNA comprises the sequence shown in SEQ ID NO: 217, optionally wherein the gRNA is shown in SEQ ID NO: 412.
[0037] In any embodiment herein, the Cas protein or a variant thereof is a Cas9 protein or a variant thereof. In any embodiment herein, the Cas protein or a variant thereof is a Cas12 protein or a variant thereof. In any embodiment herein, the Cas protein or a variant thereof is a variant Cas protein, wherein the variant Cas protein lacks nuclease activity or is an inactivated Cas (dCas) protein. In any embodiment herein, the variant Cas protein is a variant Cas9 protein lacking nuclease activity or being an inactivated Cas9 (dCas9) protein. In any embodiment herein, the Cas9 protein or a variant thereof is a Staphylococcus aureus Cas9 (SaCas9) protein or a variant thereof. In any embodiment herein, the variant Cas9 is a Staphylococcus aureus dCas9 protein (dSaCas9) containing at least one amino acid mutation selected from D10A and N580A, referenced to the position number of SEQ ID NO: 596. In any embodiment herein, the variant Cas9 protein comprises the sequence shown in SEQ ID NO: 597, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with SEQ ID NO: 597. In any embodiment herein, the Cas9 protein or a variant thereof is a Streptococcus pyogenes Cas9 (SpCas9) protein or a variant thereof. In any embodiment herein, the variant Cas9 is a Streptococcus pyogenes dCas9 (dSpCas9) protein comprising at least one amino acid mutation selected from D10A and H840A, referring to the position number in SEQ ID NO: 598. In any embodiment herein, the variant Cas9 protein comprises the sequence shown in SEQ ID NO: 599, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it.
[0038] In any embodiment herein, the at least one DNA-binding domain comprises an engineered zinc finger protein (eZFP). In any embodiment herein, the at least one DNA-binding domain is an eZFP. In any embodiment herein, the target site comprises a nucleotide sequence shown in any one of SEQ ID NO: 1045, 1046, 1052, a continuous portion thereof of at least 12 nt, or a complementary sequence to any one of the foregoing. In any embodiment herein, the target site comprises a nucleotide sequence shown in any one of SEQ ID NO: 1045, 1046, 1052.
[0039] In any embodiment of this document, the zinc finger protein comprises six zinc fingers, denoted as F1 to F6, sequentially from the N-terminus to the C-terminus, and wherein the amino acid sequence of the recognition region of each zinc finger is as follows: 1) F1: SEADRSR (SEQ ID NO: 720) F2: DRSNLTR (SEQ ID NO: 721) F3: QSSDLSR (SEQ ID NO: 722) F4: YHWYLKK (SEQ ID NO: 723) F5: RSDSLSV (SEQ ID NO: 724) F6: QNANRKT (SEQ ID NO: 725); 2) F1: RSDVLST (SEQ ID NO: 726) F2: DNSSRTR (SEQ ID NO: 727) F3: RPYTLRL (SEQ ID NO: 728) F4: DSSHRTR (SEQ ID NO: 729) F5: RSDHLSQ (SEQ ID NO: 730) F6: DSSHRTR (SEQ ID NO: 729) 731); 3) F1: RSDHLSQ (SEQ ID NO: 732) F2: QSADRTK (SEQ ID NO: 733) F3: RSDHLSQ (SEQ ID NO: 734) F4: RRSDLKR (SEQ ID NO: 735) F5: RSDHLSR (SEQ ID NO: 736) F6: QSSDLRR (SEQ ID NO: 5) F1: RSDHLSE (SEQ ID NO: 744) F2: QYSGRYY (SEQ ID NO: 745) F3: HGQTLNE (SEQ ID NO: 746) F4: QSGNLAR (SEQ ID NO: 747) F5: RSDSLLR (SEQ ID NO: 748) F6: CREYRGK (SEQ ID NO: 749);6) F1: QSANRTT (SEQ ID NO: 750) F2: RSANLTR (SEQ ID NO: 751) F3: RSDVLSE (SEQ ID NO: 752) F4: TSGHLSR (SEQ ID NO: 753) F5: QSSDLSR (SEQ ID NO: 754) F6: QWSTRKR (SEQ ID NO: 755) ; 7) F1: QSGNLAR (SEQ ID NO: 756) F2: ATCCLAH (SEQ ID NO: 757) F3: RWQYLPT (SEQ ID NO: 758) F4: DRSALAR (SEQ ID NO: 759) F5: RSDNLSE (SEQ ID NO:760) F6: KRCNLRC (SEQ ID NO: 761) ;8) F1: NPANLTR (SEQ ID NO: 762) F2: QNATRTK (SEQ IDNO: 763) F3: QSGHLAR (SEQ ID NO: 764) F4: NRHDRAK (SEQ ID NO: 765) F5: RSDHLSE (SEQ IDNO: 766) F6: QRRSRYK (SEQ ID NO: 767) ;9) F1: QSSDLSR (SEQ ID NO: 768) F2: HRSTRNR (SEQ ID NO: 769) F3: RSDVLSA (SEQ ID NO: 770)F4:DSRTRKN(SEQ ID NO: 771)F5:QSGSLTR(SEQ ID NO: 772)F6:DQSGLAH(SEQ ID NO: 773);10) F1:QNPAQWR(SEQ ID NO:774)F2:RSADLSR(SEQ ID NO: 775)F3:TSGSLSR(SEQ ID NO: 776)F4:RSDHLSR(SEQ ID NO:777)F5:RSDSLLR(SEQ ID NO: 778)F6:QSYDRFQ(SEQ ID NO: 779);11) F1:TSGSLSR(SEQID NO: 780) F2: RSDHLSR (SEQ ID NO: 781) F3: RDSLLR (SEQ ID NO: 782) F4: QSYDRFQ (SEQ ID NO: 783) F5: RSDNLST (SEQ ID NO: 784) F6: DNRDRIK (SEQ ID NO: 785)12) F1: DRSNLSR (SEQ ID NO: 786) F2: LRQNLIM (SEQ ID NO: 787) F3: ERGTLAR (SEQ ID NO: 788) F4: RSDALTQ (SEQ ID NO: 789) F5: RSDSLSQ (SEQ ID NO: 790) F6: RKADRTR (SEQ ID NO: 791) 13) F1: QYCCLTN (SEQ ID NO: 792) F2: TSGNLTR (SEQ ID NO: 793) F3: QSSDLSR (SEQ ID NO: 794) F4: FRYYLKR (SEQ ID NO: 795) F5: QSGDLTR (SEQ ID NO: 796) F6: DKGNLTK (SEQ ID NO: 797) ;14) F1: TSGSLSR (SEQ ID NO: 798) F2: RSDNLTT (SEQ ID NO: 799) F3: QSGNLAR (SEQ ID NO: 800) F4: DRTTLMR (SEQ ID NO: 801) F5: QSGHLAR (SEQ ID NO: 802) F6: QLTHLNS (SEQ ID NO: 803) ;15) F1: IKHDLHR (SEQ ID NO: 804) F2: RSANLTR (SEQ ID NO: 805) F3: RSDNLAR (SEQ ID NO: 806)F4:QNVSRPR(SEQ ID NO: 807)F5:RSDDLSK(SEQ IDNO: 808)F6:DSSHRTR(SEQ ID NO: 809);16) F1:RSDNLAR(SEQ ID NO: 810)F2:QNVSRPR(SEQ ID NO: 811)F3:RSDDLSK(SEQ ID NO: 812)F4:DSSHRTR(SEQ ID NO: 813)F5:TSSNRKT(SEQ ID NO: 814)F6:AQWTRAC(SEQ ID NO: 815);17) F1:RSDDLSK(SEQ ID NO:816) F2: DSSHRTR (SEQ ID NO: 817) F3: TSSNRKT (SEQ ID NO: 818) F4: AQWTRAC (SEQ ID NO: 819) F5: RKQTRTT (SEQ ID NO: 820) F6: HRSSLRR (SEQ ID NO: 821)18) F1: QSAHRKN (SEQID NO: 822) F2: TSSNRKT (SEQ ID NO: 823) F3: RSDNLSA (SEQ ID NO: 824) F4: RNNDRKT (SEQ ID NO: 825) F5: TSGSLSR (SEQ ID NO: 826) F6: QAGHLAK (SEQ ID NO: 827) 19) F1: RSDHLSQ (SEQ ID NO: 828) F2: ASSTRTK (SEQ ID NO: 829) F3: RSDDLTR (SEQ ID NO: 830) F4: QKSNLSS (SEQ ID NO: 831) F5: QSANRTT (SEQ ID NO: 832) F6: QNATRTK (SEQ ID NO: 833) ;20) F1: RSDTLSE (SEQ ID NO: 834) F2: RRWTLVG (SEQ ID NO: 835) F3: DRSNLSR (SEQ ID NO: 836) F4: QSGDLTR (SEQ ID NO: 837) F5: QSSDLSR (SEQ ID NO: 838) F6: YHWYLKK (SEQ ID NO: 839) ;21) F1: RSANLAR (SEQ ID NO: 840) F2: RSDNLRE (SEQ ID NO: 841) F3: RPYTLRL (SEQ ID NO: 842)F4:HRSNLNK(SEQ ID NO: 843)F5:QSGSLTR(SEQ ID NO: 844)F6:TSANLSR(SEQ ID NO: 845);22) F1:RSDDLVR(SEQ ID NO: 846)F2:TSGSLVR(SEQ IDNO: 847)F3:RSDKLVR(SEQ ID NO: 848)F4:RSDELVR(SEQ ID NO: 849)F5:TSHSLTE(SEQ IDNO: 850)F6:RADNLTE(SEQ ID NO: 851);23) F1:ERSHLRE(SEQ ID NO: 852) F2: THSHSLTE (SEQ ID NO: 853) F3: QAGHLAS (SEQ ID NO: 854) F4: THSHSLTE (SEQ ID NO: 855) F5: DPGHLVR (SEQ ID NO: 856) F6: TGSGNLVR (SEQ ID NO: 857)24) F1: RADNLTE (SEQ ID NO:858), F2: TSGSLVR (SEQ ID NO: 859), F3: RKDNLKN (SEQ ID NO: 860), F4: QSSSLVR (SEQ ID NO:861), F5: RSDKLVR (SEQ ID NO: 862), F6: DSGNLRV (SEQ ID NO: 863); 25) F1: QSSSLVR (SEQID NO: 864), F2: QSGDLRR (SEQ ID NO: 865), F3: RSDERKR (SEQ ID NO: 866), F4: HRTTLTN (SEQID NO: 867), F5: RSDHLTN (SEQ ID NO: 868), F6: TSGELVR (SEQ ID NO: 869); 26) F1: QSGDLRR (SEQ ID NO: 870), F2: RSDERKR (SEQ ID NO: 871), F3: HRTTLTN (SEQ ID NO: 872), F4: RSDHLTN (SEQ ID NO: 873), F5: TSGELVR (SEQ ID NO: 874), F6: RSDDLVR (SEQ ID NO:875); 27) F1: QRAHLER (SEQ ID NO: 876), F2: QLAHLRA (SEQ ID NO: 877), F3: DPGHLVR (SEQID NO: 878), F4: RRSACRR (SEQ ID NO: 879), F5: RSDHLTT (SEQ ID NO: 880), F6: QSSSLVR (SEQID NO: 881); and 28) F1: QSSNLVR (SEQ ID NO: 882), F2: RSDDLVR (SEQ ID NO: 883), F3: THLDLIR (SEQ ID NO: 884), F4: TSGNLTE (SEQ ID NO: 885), F5: RRSACRR (SEQ ID NO: 886), F6: RNDTLTE (SEQ ID NO: 887).;
[0040] In any embodiment of this document, the zinc finger protein comprises six zinc fingers, denoted as F1 to F6, sequentially from the N-terminus to the C-terminus, and wherein the amino acid sequence of the recognition region of each zinc finger is as follows: F1: QSAHRKN (SEQ ID NO: 822) F2: TSSNRKT (SEQ ID NO: 823) F3: RSDNLSA (SEQ ID NO: 824) F4: RNNDRKT (SEQ ID NO: 825) F5: TSGSLSR (SEQ ID NO: 826) F6: QAGHLAK (SEQ ID NO: 827).
[0041] In any embodiment of this document, the zinc finger protein comprises six zinc fingers, denoted as F1 to F6, sequentially from the N-terminus to the C-terminus, and wherein the amino acid sequence of the recognition region of each zinc finger is as follows: F1: RSDHLSQ (SEQ ID NO: 828) F2: ASSTRTK (SEQ ID NO: 829) F3: RSDDLTR (SEQ ID NO: 830) F4: QKSNLSS (SEQ ID NO: 831) F5: QSANRTT (SEQ ID NO: 832) F6: QNATRTK (SEQ ID NO: 833).
[0042] In any embodiment herein, the zinc finger protein comprises, sequentially from the N-terminus to the C-terminus, six zinc fingers denoted as F1 to F6, wherein the amino acid sequence of the recognition region of each zinc finger is as follows: F1: QSSSLVR (SEQ ID NO: 864) F2: QSGDLRR (SEQ ID NO: 865) F3: RSDERKR (SEQ ID NO: 866) F4: HRTTLTN (SEQ ID NO: 867) F5: RSDHLTN (SEQ ID NO: 868) F6: TSGELVR (SEQ ID NO: 869).
[0043] This article also provides an epigenetic DNA-targeting system comprising: a) an eZFP that binds to a target site in one or more HBV genes or their regulatory elements; and b) at least one effector domain that represses transcription of one or more HBV genes, wherein the zinc finger protein comprises, sequentially from the N-terminus to the C-terminus, six zinc fingers denoted as F1 to F6, and wherein the amino acid sequence of the recognition region of each zinc finger is as follows: 1) F1: SEADRSR (SEQ ID NO: 720) F2: DRSNLTR (SEQ ID NO: 721) F3: QSSDLSR (SEQ ID NO: 722) F4: YHWYLKK (SEQ ID NO: 723) F5: RSDSLSV (SEQ ID NO: 724) F6: QNANRKT (SEQ ID NO: 725); 2) F1: RSDVLST (SEQ ID NO: 726) F2: DNSSRTR (SEQ ID NO: 725) 727) F3: RPYTLRL (SEQ ID NO: 728) F4: DSSHRTR (SEQ ID NO: 729) F5: RSDHLSQ (SEQ ID NO: 730) F6: DSSHRTR (SEQ ID NO: 731); 3) F1: RSDHLSQ (SEQ ID NO: 732) F2: QSADRTK (SEQ ID NO: 733) F3: RSDHLSQ (SEQ ID NO: 734) F4: RRSDLKR (SEQ ID NO: 735) F5: RSDHLSR (SEQ ID NO: 736) F6: QSSDLRR (SEQ ID NO: 737); 4) F1: RSDNLSE (SEQ ID NO: 738) F2: TSSNRKT (SEQ ID NO: 739) F3: DRSHLTR (SEQ ID NO: 740) F4: RSDALTQ (SEQ ID NO: 741) F5: DRSALAR (SEQ ID NO: 742) F6: RRFTLSK (SEQ ID NO: 743); 5) F1: RSDHLSE (SEQ ID NO: 744) F2: QYSGRYY (SEQ ID NO: 745) F3: HGQTLNE (SEQ ID NO: 746) F4: QSGNLAR (SEQ ID NO: 747) F5: RSDSLLR (SEQ ID NO: 748) F6: CREYRGK (SEQ ID NO: 749);6) F1: QSANRTT (SEQ ID NO: 750) F2: RSANLTR (SEQ ID NO: 751) F3: RSDVLSE (SEQ ID NO: 752) F4: TSGHLSR (SEQ ID NO: 753) F5: QSSDLSR (SEQ ID NO: 754) F6: QWSTRKR (SEQ ID NO: 755) ; 7) F1: QSGNLAR (SEQ ID NO: 756) F2: ATCCLAH (SEQ IDNO: 757) F3: RWQYLPT (SEQ ID NO: 758) F4: DRSALAR (SEQ ID NO: 759) F5: RSDNLSE (SEQ IDNO: 760) F6: KRCNLRC (SEQ ID NO: 761) ;8) F1: NPANLTR (SEQ ID NO: 762) F2: QNATRTK (SEQ ID NO: 763) F3: QSGHLAR (SEQ ID NO: 764) F4: NRHDRAK (SEQ ID NO: 765) F5: RSDHLSE (SEQ ID NO: 766) F6: QRRSRYK (SEQ ID NO: 767) ;9) F1: QSSDLSR (SEQ ID NO: 768) F2: HRSTRNR (SEQ ID NO: 769) F3: RSDVLSA (SEQ ID NO: 770)F4:DSRTRKN(SEQ ID NO:771)F5:QSGSLTR(SEQ ID NO: 772)F6:DQSGLAH(SEQ ID NO: 773);10) F1:QNPAQWR(SEQID NO: 774)F2:RSADLSR(SEQ ID NO: 775)F3:TSGSLSR(SEQ ID NO: 776)F4:RSDHLSR(SEQID NO: 777)F5:RSDSLLR(SEQ ID NO: 778)F6:QSYDRFQ(SEQ ID NO: 779);11) F1:TSGSLSR(SEQ ID NO: 780) F2: RSDHLSR (SEQ ID NO: 781) F3: RDSLLR (SEQ ID NO: 782) F4: QSYDRFQ (SEQ ID NO: 783) F5: RSDNLST (SEQ ID NO: 784) F6: DNRDRIK (SEQ ID NO: 785)12) F1: DRSNLSR (SEQ ID NO: 786) F2: LRQNLIM (SEQ ID NO: 787) F3: ERGTLAR (SEQ ID NO: 788) F4: RSDALTQ (SEQ ID NO: 789) F5: RSDSLSQ (SEQ ID NO: 790) F6: RKADRTR (SEQ ID NO: 791) 13) F1: QYCCLTN (SEQ ID NO: 792) F2: TSGNLTR (SEQ ID NO: 793) F3: QSSDLSR (SEQ ID NO: 794) F4: FRYYLKR (SEQ ID NO: 795) F5: QSGDLTR (SEQ ID NO: 796) F6: DKGNLTK (SEQ ID NO: 797) ;14) F1: TSGSLSR (SEQ ID NO: 798) F2: RSDNLTT (SEQ IDNO: 799) F3: QSGNLAR (SEQ ID NO: 800) F4: DRTTLMR (SEQ ID NO: 801) F5: QSGHLAR (SEQ IDNO: 802) F6: QLTHLNS (SEQ ID NO: 803) ;15) F1: IKHDLHR (SEQ ID NO: 804) F2: RSANLTR (SEQ ID NO: 805) F3: RSDNLAR (SEQ ID NO: 806)F4:QNVSRPR(SEQ ID NO: 807)F5:RSDDLSK(SEQ ID NO: 808)F6:DSSHRTR(SEQ ID NO: 809);16) F1:RSDNLAR(SEQ ID NO:810)F2:QNVSRPR(SEQ ID NO: 811)F3:RSDDLSK(SEQ ID NO: 812)F4:DSSHRTR(SEQ ID NO:813)F5:TSSNRKT(SEQ ID NO: 814)F6:AQWTRAC(SEQ ID NO: 815);17) F1: RSDDLSK(SEQID NO: 816) F2: DSSHRTR (SEQ ID NO: 817) F3: TSNRKT (SEQ ID NO: 818) F4: AQWTRAC (SEQ ID NO: 819) F5: RKQTRTT (SEQ ID NO: 820) F6: HRSSLRR (SEQ ID NO: 821)18) F1: QSAHRKN (SEQ ID NO: 822) F2: TSSNRKT (SEQ ID NO: 823) F3: RSDNLSA (SEQ ID NO: 824) F4: RNNDRKT (SEQ ID NO: 825) F5: TSGSLSR (SEQ ID NO: 826) F6: QAGHLAK (SEQ ID NO: 827) 19) F1: RSDHLSQ (SEQ ID NO: 828) F2: ASSTRTK (SEQ ID NO: 829) F3: RSDDLTR (SEQ ID NO: 830) F4: QKSNLSS (SEQ ID NO: 831) F5: QSANRTT (SEQ ID NO: 832) F6: QNATRTK (SEQ ID NO: 833) ;20) F1: RSDTLSE (SEQ ID NO: 834) F2: RRWTLVG (SEQ ID NO: 835) F3: DRSNLSR (SEQ ID NO: 836) F4: QSGDLTR (SEQ ID NO: 837) F5: QSSDLSR (SEQ ID NO: 838) F6: YHWYLKK (SEQ ID NO: 839) ;21) F1: RSANLAR (SEQ ID NO: 840) F2: RSDNLRE (SEQ ID NO: 841) F3: RPYTLRL (SEQ ID NO: 842)F4:HRSNLNK(SEQ ID NO: 843)F5:QSGSLTR(SEQ IDNO: 844)F6:TSANLSR(SEQ ID NO: 845);22) F1:RSDDLVR(SEQ ID NO: 846)F2:TSGSLVR(SEQ ID NO: 847)F3:RSDKLVR(SEQ ID NO: 848)F4:RSDELVR(SEQ ID NO: 849)F5:TSHSLTE(SEQ ID NO: 850)F6:RADNLTE(SEQ ID NO: 851);23) F1:ERSHLRE(SEQ ID NO:852) F2: THSHSLTE (SEQ ID NO: 853) F3: QAGHLAS (SEQ ID NO: 854) F4: THSHSLTE (SEQ ID NO: 855) F5: DPGHLVR (SEQ ID NO: 856) F6: TGSGNLVR (SEQ ID NO: 857)24) F1: RADNLTE (SEQ ID NO: 858) F2: TSGSLVR (SEQ ID NO: 859) F3: RKDNLKN (SEQ ID NO: 860) F4: QSSSLVR (SEQ ID NO: 861) F5: RSDKLVR (SEQ ID NO: 862) F6: DSG NLRV (SEQ ID NO: 863); 25) F1: QSSSLVR (SEQ ID NO: 864) F2: QSGDLRR (SEQ ID NO: 865) F3: RSDERKR (SEQ ID NO: 866) F4: HRTTLTN (SEQ ID NO: 867) F5: RSDHLTN (SEQ ID NO: 868) F6: TSGELVR (SEQ ID NO: 869); 26) F1: QSGDLRR (SEQ ID NO: 870) F2: RSDERKR (SEQ ID NO: 871) F3: HRTTLTN (SEQ ID NO: 872) F4: RSDHLTN (SEQ ID NO: 873) F5: TSGELVR (SEQ ID NO: 874) F6: RSDDLVR (SEQ ID NO: 875); 27) F1: QRAHLER (SEQ ID NO: 876) F2: QLAHLRA (SEQ ID NO: 877) F3: DPGHLVR (SEQ ID NO: 878) F4: RRSACRR (SEQ ID NO: 879) F5: RSDHLTT (SEQ ID NO: 880) F6: QSSSLVR (SEQ ID NO: 881); and 28) F1: QSSNLVR (SEQ ID NO: 882) F2: RSDDLVR (SEQ ID NO: 883) F3: THLDLIR (SEQ ID NO: 884) F4: TSGNLTE (SEQ ID NO: 885) F5: RRSACRR (SEQ ID NO: 886) F6: RNDTLTE (SEQ ID NO: 887).;
[0044] In any embodiment herein, the engineered zinc finger protein comprises a sequence or a portion thereof shown in any of SEQ ID NO: 692-719, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the engineered zinc finger protein is encoded by a sequence or a portion thereof shown in any of SEQ ID NO: 888-915, or a nucleotide sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it.
[0045] This article also provides an epigenetically modified DNA targeting system comprising: a) an engineered zinc finger protein that binds to a target site in one or more HBV genes or their regulatory elements; and b) at least one effector domain that represses transcription of one or more HBV genes, wherein the zinc finger protein comprises, in sequence from the N-terminus to the C-terminus, six zinc fingers denoted as F1 to F6, and wherein the amino acid sequence of each zinc finger recognition region is as follows: F1: QSAHRKN (SEQ ID NO: 822) F2: TSSNRKT (SEQ ID NO: 823) F3: RSDNLSA (SEQ ID NO: 824) F4: RNNDRKT (SEQ ID NO: 825) F5: TSGSLSR (SEQ ID NO: 826) F6: QAGHLAK (SEQ ID NO: 827).
[0046] In any embodiment herein, the eZFP comprises the sequence shown in SEQ ID NO: 709 or a portion thereof, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the engineered zinc finger protein comprises the sequence shown in any of SEQ ID NO: 709.
[0047] In any embodiment herein, the engineered zinc finger protein is encoded by the sequence shown in SEQ ID NO: 905 or a portion thereof, or by a nucleotide sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In any embodiment herein, the engineered zinc finger protein is encoded by the sequence shown in any of SEQ ID NO: 905.
[0048] This article also provides an epigenetically modified DNA targeting system comprising: a) an engineered zinc finger protein that binds to a target site in one or more HBV genes or their regulatory elements; and b) at least one effector domain that represses transcription of one or more HBV genes, wherein the zinc finger protein comprises, in sequence from the N-terminus to the C-terminus, six zinc fingers denoted as F1 to F6, and wherein the amino acid sequence of each zinc finger recognition region is as follows: F1: RSDHLSQ (SEQ ID NO: 828) F2: ASSTRTK (SEQ ID NO: 829) F3: RSDDLTR (SEQ ID NO: 830) F4: QKSNLSS (SEQ ID NO: 831) F5: QSANRTT (SEQ ID NO: 832) F6: QNATRTK (SEQ ID NO: 833). In some embodiments, the engineered zinc finger protein comprises the sequence shown in SEQ ID NO: 710 or a portion thereof, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity therewith. In some embodiments, the engineered zinc finger protein comprises the sequence shown in any one of SEQ ID NO: 710. In any embodiment herein, the engineered zinc finger protein is encoded by the sequence shown in SEQ ID NO: 906 or a portion thereof, or a nucleotide sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity therewith. In any embodiment herein, the engineered zinc finger protein is encoded by the sequence shown in any one of SEQ ID NO: 906.
[0049] This article also provides an epigenetically modified DNA targeting system comprising: a) an engineered zinc finger protein that binds to a target site in one or more HBV genes or their regulatory elements; and b) at least one effector domain that represses transcription of one or more HBV genes, wherein the zinc finger protein comprises, in sequence from the N-terminus to the C-terminus, six zinc fingers denoted as F1 to F6, and wherein the amino acid sequence of each zinc finger recognition region is as follows: F1: QSSSLVR (SEQ ID NO: 864) F2: QSGDLRR (SEQ ID NO: 865) F3: RSDERKR (SEQ ID NO: 866) F4: HRTTLTN (SEQ ID NO: 867) F5: RSDHLTN (SEQ ID NO: 868) F6: TSGELVR (SEQ ID NO: 869). In some embodiments, the engineered zinc finger protein comprises the sequence shown in SEQ ID NO: 716 or a portion thereof, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity therewith. In some embodiments, the engineered zinc finger protein comprises the sequence shown in any one of SEQ ID NO: 716. In some embodiments, the engineered zinc finger protein is encoded by the sequence shown in SEQ ID NO: 912 or a portion thereof, or a nucleotide sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity therewith. In some embodiments, the engineered zinc finger protein is encoded by the sequence shown in any one of SEQ ID NO: 912.
[0050] In any embodiment of this document, the at least one effector domain induces transcriptional repression. In any embodiment of this document, the at least one effector domain is a DNA methyltransferase. In any embodiment of this document, the at least one effector domain comprises a DNA methyltransferase and a repressor domain capable of recruiting a heterochromatin-inducing factor, or optionally wherein the heterochromatin-inducing factor comprises a histone methyltransferase. In any embodiment of this document, the at least one effector domain comprises a DNA methyltransferase and a histone methyltransferase. In any embodiment of this document, the at least one effector domain is selected from the KRAB repressor domain, ERF repressor domain, Mxi1 repressor domain, SID4X repressor domain, Mad-SID repressor domain, LSD1 repressor domain, or DNMT3A, DNMT3A-3L, DNMT3A / L-KRAB fusion repressor domain, DNMT3B domain-binding protein, EZH2 repressor domain, or LSD1 repressor domain, or a variant of any of the foregoing. In any embodiment herein, at least one effector domain comprises a sequence or domain thereof selected from any one of SEQ ID NO: 590 or 600-608, 651, 661, 664, 665, 666, 668, and 669, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with any of the foregoing. In any embodiment herein, the at least one effector domain comprises a KRAB domain or a variant thereof. In any embodiment herein, the at least one effector domain comprises the sequence shown in SEQ ID NO: 590, a portion thereof, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with any of the foregoing. In any embodiment herein, the at least one effector domain comprises a DNMT3A / L domain or a variant thereof. In any embodiment herein, the at least one effector domain comprises the sequence shown in SEQ ID NO: 604 and 607, a portion thereof, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with any of the foregoing; or the at least one effector domain comprises the sequence shown in SEQ ID NO: 651, a portion thereof, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with SEQ ID NO: 651.
[0051] In any embodiment herein, the fusion protein comprises the DNMT3A / 3L-dSpCas9-KRAB fusion protein. In any embodiment herein, the fusion protein comprises the sequence shown in SEQ ID NO: 645, a portion thereof, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with any of the foregoing.
[0052] In any embodiment herein, the at least one effector domain is fused with the N-terminus, C-terminus, or both the N-terminus and C-terminus of the DNA binding domain or a component thereof.
[0053] In any embodiment herein, the fusion protein is encoded by the sequence shown in SEQ ID NO: 680, a portion thereof, or a nucleic acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with any of the foregoing. In any embodiment herein, the fusion protein is encoded by the sequence shown in SEQ ID NO: 680.
[0054] In any embodiment herein, the fusion protein is encoded by a nucleic acid sequence represented by any one of SEQ ID NO: 916-943, a portion thereof, or having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with any of the foregoing. In any embodiment herein, the fusion protein comprises a sequence represented by any one of SEQ ID NO: 944-971, a portion thereof, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with any of the foregoing. In any embodiment herein, the fusion protein comprises a sequence represented by any one of SEQ ID NO: 961, 962, or 968.
[0055] In any embodiment herein, the fusion protein comprises the DNMT3A / 3L-eZFP-KRAB fusion protein.
[0056] In any embodiment of this document, the fusion protein is encoded by a nucleic acid sequence represented by any one of SEQ ID NO: 972-999, a portion thereof, or having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with any of the aforementioned sequences. In any embodiment of this document, the fusion protein is encoded by a nucleic acid sequence represented by any one of SEQ ID NO: 933, 934, or 940, a portion thereof, or having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with any of the aforementioned sequences. In any embodiment of this document, the fusion protein comprises an amino acid sequence represented by any one of SEQ ID NO: 1000-1027, a portion thereof, or having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with any of the aforementioned sequences. In any embodiment herein, the fusion protein comprises the sequence, a portion thereof, or an amino acid sequence that is at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with any of the foregoing.
[0057] In any embodiment herein, the fusion protein further comprises one or more nuclear localization signals (NLS).
[0058] In any embodiment herein, the fusion protein further comprises one or more adapters connecting two or more of the following: the DNA-binding domain, the at least one effector domain, and the one or more nuclear localization signals.
[0059] In any embodiment herein, the DNA targeting system targets all hepatitis B virus genomes.
[0060] In any embodiment herein, the DNA targeting system targets at least 70% of the entire hepatitis B virus genome. In any embodiment herein, the DNA targeting system targets at least 60% of the entire hepatitis B virus genome. In any embodiment herein, the DNA targeting system targets at least 50% of the entire hepatitis B virus genome.
[0061] In any embodiment of this document, the DNA targeting system shall not introduce gene damage or DNA breakage at or near the target site.
[0062] In any embodiment of this document, repression of transcription of one or more HBV genes results in a reduction in RNA and / or protein levels derived from HBV DNA sequences. In any embodiment of this document, transcription repression includes a reduction in total hepatitis B virus RNA transcript levels. In any embodiment of this document, transcription repression includes a reduction in hepatitis B virus precore (“pre-C”), pregenomic (“pgRNA”), preS1, preS2 / S, and HBx levels. In any embodiment of this document, transcription repression includes a reduction in HBx levels. In any embodiment of this document, transcription repression includes a reduction in hepatitis B surface antigen (HBsAg) and / or hepatitis B core-associated antigen (HbcrAg) protein levels. In any embodiment of this document, transcription repression includes a reduction of HBsAg transcript and / or protein levels by at least 90%. In any embodiment of this document, transcription repression includes a reduction of HbcrAg transcript and / or protein levels derived from cccDNA by at least 50%.
[0063] This document also provides a guide RNA (gRNA) that binds to a target site in a hepatitis B virus DNA sequence. In some embodiments, the hepatitis B virus DNA sequence is a hepatitis B virus (HBV) gene or a regulatory element thereof. In some embodiments, the target site is present as covalently closed circular DNA (cccDNA), relaxed circular DNA (rcDNA), and / or integrated into human genomic DNA. In any embodiment herein, the target site is located at or near a gene or regulatory element thereof involved in controlling HBV replication and / or HBV transcription. In any embodiment herein, the gene involved in controlling HBV replication and / or HBV transcription encodes a polymerase, envelope protein, capsid protein, transcription factor, or transcription transactivator. In any embodiment herein, the gene involved in controlling HBV replication and / or HBV transcription is a polymerase gene, an S family gene, an X gene, or a core family gene. In any embodiment herein, the target site is located in a gene or regulatory element thereof encoding the hepatitis B virus X protein (HBx) X gene. In any embodiment of this document, the target site is located at or near a regulatory element involved in controlling HBV replication and / or HBV transcription.
[0064] In any embodiment of this document, the regulatory element is a promoter region. In any embodiment of this document, the promoter region is a pre-S1 promoter, a pre-S2 promoter, an X promoter, or a basic core promoter. In any embodiment of this document, the regulatory element is an enhancer region. In any embodiment of this document, the enhancer region is an Enh1 or Enh2 enhancer region. In any embodiment of this document, the regulatory element is a transcript processing control region.
[0065] In any embodiment herein, the target site is a coding region. In any embodiment herein, the target site is located within 500 bp, 1000 bp, or 1500 bp of the transcription start site. In any embodiment herein, the target site is located within a target region, which is located between 0 and 3300 base pairs (bp) of the HBV genome, optionally between 0 and 3189 bp at the position corresponding to the HBV genome shown in reference SEQ ID NO: 650. In any embodiment herein, the target site is located within a target region, which is located between 43 bp-490 bp, 1033 bp-1749 bp, 1800 bp-1950 bp, or 2953 bp-3182 bp of the HBV genome at the position corresponding to the HBV genome shown in reference SEQ ID NO: 650. In any embodiment herein, the target site is located within a target region located at base pairs between 1 bp-42 bp, 491 bp-1032 bp, 1750 bp-1799 bp, or 1951 bp-2952 bp of the HBV genome corresponding to the position of the HBV genome shown in reference SEQ ID NO: 650.
[0066] In any embodiment herein, the target site, or each of the target sites, is located within a CpG island of the HBV genome. In any embodiment herein, the target site is located within a target region situated at a base pair between 67 bp and 392 bp, 1033 bp and 1749 bp, or 2215 bp and 2490 bp of the HBV genome corresponding to the location shown in reference SEQ ID NO: 650. In any embodiment herein, the target site is located within a target region situated at a base pair between 1033 bp and 1749 bp of the HBV genome corresponding to the location shown in reference SEQ ID NO: 650. In any embodiment herein, the target site, or each of the target sites, is located within a target region spanning 300 base pairs upstream of the start codon of the hepatitis B virus X protein (HBx).
[0067] This document also provides a gRNA (gRNA) that binds to a target site within a target region spanning 300 base pairs upstream of the start codon of the hepatitis B virus X protein (HBx). In any embodiment of this document, the target site is located in the HBx basic core promoter region. In any embodiment of this document, the target site is located within the HBx promoter / enhancer region. In any embodiment of this document, the target site is within a target region spanning 250 base pairs upstream of the start codon of the hepatitis B virus X protein (HBx). In any embodiment of this document, the target site is within a target region spanning 1060–1480 bp of the HBV genome corresponding to the position shown in reference SEQ ID NO: 650. In any embodiment of this document, the target site is within a target region spanning 150 base pairs upstream of the start codon of the hepatitis B virus X protein (HBx). In any embodiment of this document, the target site is within a target region spanning 120 base pairs upstream of the start codon of the hepatitis B virus X protein (HBx). In any embodiment herein, the target site is within a target region sequence corresponding to a 1250-1374 bp sequence of the HBV genome spanning the HBV genome shown in reference SEQ ID NO: 650. In any embodiment herein, the target site is within a target region sequence corresponding to a 1255-1302 bp sequence of the HBV genome spanning the HBV genome shown in reference SEQ ID NO: 650. In any embodiment herein, the target site is within a target region sequence corresponding to a 1260-1300 bp sequence of the HBV genome spanning the HBV genome shown in reference SEQ ID NO: 650.In any embodiment herein, the gRNA comprises the sequence shown in any one of SEQ ID NO: 200, 201, 207, 217, 221, 224, 233, 237, 238, 246, 251, 256, 258, 263, 267, 274, 270, 277, 279, 283, 284, 293, 294, 308, 311, 313, 316, 319, 320, 325, 328, 330, 333, 338, 345, 347, 350, 353, 359, 360, 369, 370, 371, 377, 380, 384, 385, or 387, a continuous portion of at least 14 nucleotides thereof, or a complementary sequence to any one of the foregoing, optionally wherein the gRNA or each of the gRNAs as shown in SEQ ID NO: 200, 201, 207, 217, 221, 224, 233, 237, 238, 246, 251, 256, 258, 263, 267, 274, 275, 386, 387, 385, or 387, a continuous portion of at least 14 nucleotides thereof, or a complementary sequence to any one of the foregoing, optionally wherein the gRNA or each of the gRNAs as shown in SEQ ID NO: 200, 201, 207, 217, 221, 224, 233, 237, 238, 246 NO: Any one of 395, 402, 408, 412, 416, 419, 428, 432, 433, 441, 446, 451, 453, 458, 462, 465, 469, 472, 474, 478, 479, 488, 489, 503, 506, 508, 511, 514, 515, 520, 523, 525, 575, 528, 533, 540, 542, 545, 548, 554, 555, 565, 566, 572, 579, 580, or 582 is shown. In any embodiment herein, the gRNA is as shown in any of SEQ ID NO: 395, 402, 408, 412, 416, 419, 428, 432, 433, 441, 446, 451, 453, 458, 462, 465, 469, 472, 474, 478, 479, 488, 489, 503, 506, 508, 511, 514, 515, 520, 523, 525, 575, 528, 533, 540, 542, 545, 548, 554, 555, 565, 566, 572, 579, 580, or 582. In any embodiment herein, the gRNA comprises the sequence shown in any one of SEQ ID NO: 207, 213, 215, 217, 221, 222, 241, 245, 258, 261, 268, 274, 380, 387, a continuous portion of at least 14 nucleotides thereof, or a complementary sequence of any of the foregoing, optionally wherein the gRNA or each of the gRNAs is shown in any one of SEQ ID NO: 402, 408, 410, 412, 416, 417, 436, 440, 453, 456, 463, 469, 575, or 582.In any embodiment herein, the gRNA comprises the sequence shown in any one of SEQ ID NO: 402, 408, 410, 412, 416, 417, 436, 440, 453, 456, 463, 469, 575, or 582. In any embodiment herein, the gRNA comprises the sequence shown in SEQ ID NO: 217, a continuous portion of at least 14 nucleotides thereof, or a complementary sequence to any of the foregoing, optionally wherein the gRNA or each of the gRNAs is shown in any one of SEQ ID NO: 412. In any embodiment herein, the gRNA comprises the sequence shown in SEQ ID NO: 412.
[0068] This document also provides a CRISPR Cas-guide RNA (gRNA) assembly comprising: (a) a clustered regularly spaced short palindromic repeat-associated (Cas) protein or a variant thereof; and (b) at least one gRNA according to any one of claims 165-202, said gRNA targeting the Cas protein or a variant thereof to a target site in a hepatitis B virus DNA sequence. In some embodiments, the Cas protein or a variant thereof is a Cas9 protein or a variant thereof. In any embodiment herein, the Cas protein or a variant thereof is a variant Cas protein, wherein the variant Cas protein lacks nuclease activity or is an inactivated Cas (dCas) protein. In any embodiment herein, the variant Cas protein is a variant Cas9 protein lacking nuclease activity or being an inactivated Cas9 (dCas9) protein. In any embodiment herein, the Cas9 protein or a variant thereof is Staphylococcus aureus Cas9 (SaCas9) protein or a variant thereof. In some embodiments, the variant Cas9 is a Staphylococcus aureus dCas9 protein (dSaCas9) containing at least one amino acid mutation selected from D10A and N580A, referenced to position number SEQ ID NO: 596. In some embodiments, the variant Cas9 protein contains the sequence shown in SEQ ID NO: 597, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity therewith. In some embodiments, the Cas9 protein or a variant thereof is a Streptococcus pyogenes Cas9 (SpCas9) protein or a variant thereof. In some embodiments, the variant Cas9 is a Streptococcus pyogenes dCas9 (dSpCas9) protein containing at least one amino acid mutation selected from D10A and H840A, referenced to position number SEQ ID NO: 598. In some embodiments, the variant Cas9 protein comprises the sequence shown in SEQ ID NO: 599, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it.
[0069] A polynucleotide is also provided that encodes the epigenetic modification DNA targeting system or the DNA targeting system fusion protein disclosed herein, the gRNA disclosed herein, the CRISPR Cas-gRNA combination disclosed herein, or a portion or component of any of the foregoing.
[0070] It also provides a variety of polynucleotides that encode the epigenetic modification DNA targeting system or the DNA targeting system fusion protein disclosed herein, the gRNA disclosed herein, the CRISPR Cas-gRNA combination disclosed herein, or a portion or component of any of the foregoing.
[0071] A vector is also provided that comprises the polynucleotides disclosed herein. A vector is also provided that comprises a variety of polynucleotides disclosed herein.
[0072] In any embodiment of this document, the vector is a viral vector. In some embodiments of this document, the vector is an adeno-associated virus (AAV) vector. In some embodiments of this document, the vector is selected from AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, or AAV9. In some embodiments of this document, the vector is a lentiviral vector. In some embodiments of this document, the vector is a non-viral vector. In some embodiments of this document, the non-viral vector is selected from lipid nanoparticles, liposomes, exosomes, or cell-penetrating peptides. In any embodiment of this document, the vector exhibits tropism for hepatitis B virus-infected cells. In any embodiment of this document, the vector comprises one or two or more vectors.
[0073] This article also provides a method for promoting epigenetic modifications within a target region of a hepatitis B virus sequence, the method comprising introducing an epigenetic modification DNA targeting system that targets a target site within the target region into HBV-infected cells containing the hepatitis B virus sequence.
[0074] This article also provides a method for increasing CpG methylation in a target region of a hepatitis B virus sequence, the method comprising introducing an epigenetic modified DNA targeting system that targets a target site in the target region into HBV-infected cells containing the hepatitis B virus sequence.
[0075] This document also provides a method for promoting epigenetic modification of target regions in hepatitis B virus sequences, the method comprising introducing into HBV-infected cells containing hepatitis B virus sequences the epigenetic modification DNA targeting system disclosed herein, the gRNA disclosed herein, the CRISPR Cas-gRNA combination disclosed herein, the polynucleotide disclosed herein, a variety of polynucleotides disclosed herein, the vector disclosed herein, or a portion or component of any of the foregoing.
[0076] This document also provides a method for increasing CpG methylation of a target region in a hepatitis B virus sequence, the method comprising introducing into HBV-infected cells containing a hepatitis B virus sequence the epigenetic modification DNA targeting system disclosed herein, the gRNA disclosed herein, the CRISPR Cas-gRNA combination disclosed herein, the polynucleotide disclosed herein, a variety of polynucleotides disclosed herein, the vector disclosed herein, or a portion or component of any of the foregoing.
[0077] In any embodiment herein, the target region comprises a continuous nucleotide sequence within 1033 bp-1749 bp of the hepatitis B virus sequence corresponding to the nucleotide position of reference SEQ ID NO: 650.
[0078] This article also provides a method for reducing the transcription of one or more genes in HBV-infected cells containing hepatitis B virus sequences, the method comprising introducing an epigenetic modified DNA targeting system into the cells, the epigenetic modified DNA targeting system inducing targeted CpG methylation within the hepatitis B virus sequence at a nucleotide position referenced in SEQ ID NO: 650.
[0079] This article also provides a method for reducing hepatitis B virus infection in HBV-infected cells, the method comprising introducing an epigenetically modified DNA targeting system into cells containing a hepatitis B virus sequence, the epigenetically modified DNA targeting system inducing targeted CpG methylation within a target region of the hepatitis B virus sequence at a nucleotide position referenced in SEQ ID NO: 650.
[0080] In any embodiment herein, the epigenetic modified DNA targeting system comprises at least one DNA targeting module comprising a fusion protein comprising (a) a DNA binding domain for targeting a target site in a hepatitis B virus DNA sequence; and (b) at least one effector domain comprising a DNA methyltransferase effector domain.
[0081] In any embodiment herein, the CpG methylation region is within 500 base pairs of the target region. In any embodiment herein, the introduction occurs in vivo or ex vivo.
[0082] In any embodiment herein, the cell is a mammalian cell. In any embodiment herein, the cell is a human cell. In any embodiment herein, the cell contains integrated HBV DNA. In any embodiment herein, the cell is a hepatocyte containing a pool of free HBV cccDNA. In some embodiments, the hepatocyte expresses HBV protein, wherein the HBV protein is HBsAg, HBeAg, or HBcrAg, or a combination thereof.
[0083] This article also provides a method for reducing hepatitis virus infection in subjects, the method comprising administering to a subject infected with hepatitis B an epigenetically modified DNA targeting system that increases CpG methylation in a target region of a hepatitis B virus sequence, wherein the epigenetically modified DNA targeting system comprises (a) a DNA-binding domain for targeting a target site in a hepatitis B virus DNA sequence; and (b) at least one effector domain comprising a DNA methyltransferase effector domain.
[0084] In any embodiment of this document, the target region is a region in the HBV genome containing CpG. In any embodiment of this document, the target region comprises a continuous nucleotide sequence within 67 bp-392 bp, 1033 bp-1749 bp, or 2215 bp-2490 bp of the hepatitis B virus sequence corresponding to the nucleotide position of reference SEQ ID NO: 650. In any embodiment of this document, the target region comprises a continuous nucleotide sequence within 1033 bp-1749 bp of the hepatitis B virus sequence corresponding to the nucleotide position of reference SEQ ID NO: 650. In any embodiment of this document, the target region is located within 300 base pairs upstream of the start codon of the hepatitis B virus X protein (HBx). In any embodiment of this document, the target region is within the HBx basic core promoter region. In any embodiment of this document, the target region is within the HBx promoter / enhancer region. In any embodiment herein, the target region is located within 250 base pairs upstream of the start codon of the hepatitis B virus X protein (HBx). In any embodiment herein, the target region comprises a continuous nucleotide sequence within 1060 bp–1480 bp of the hepatitis B virus sequence corresponding to the nucleotide position of reference SEQ ID NO: 650. In any embodiment herein, the target region is located within 150 base pairs upstream of the start codon of the hepatitis B virus X protein (HBx). In any embodiment herein, the target region is located within 120 base pairs upstream of the start codon of the hepatitis B virus X protein (HBx). In any embodiment herein, the target region comprises a continuous nucleotide sequence within 1250 bp–1374 bp of the hepatitis B virus sequence corresponding to the nucleotide position of reference SEQ ID NO: 650. In some embodiments, the target region has the sequence shown in SEQ ID NO: 1068. In any embodiment herein, the target region comprises a continuous nucleotide sequence within a 1260 bp-1300 bp segment of the hepatitis B virus sequence corresponding to the nucleotide position of reference SEQ ID NO: 650. In some embodiments, the target region has the sequence shown in SEQ ID NO: 1070.
[0085] In any embodiment herein, the DNA-binding domain comprises a clustered regularly spaced short palindromic repeat-associated (Cas)-guide RNA (gRNA) combination comprising (a) a Cas protein or a variant thereof and (b) at least one gRNA; a zinc finger protein (ZFP); a transcription activator-like effector (TALE); a meganuclease; a homing endonuclease; or an I-SceI enzyme or a variant thereof, optionally wherein the DNA-binding domain comprises a non-catalytically inactive variant of any of the foregoing.
[0086] In any embodiment of this document, the method comprises a CRISPR Cas-guide RNA (gRNA) combination comprising: (a) a clustered regularly spaced short palindromic repeat sequence-associated (Cas) protein or a variant thereof; and (b) at least one gRNA according to any one of claims 165-202, the gRNA targeting the Cas protein or a variant thereof to a target site in a hepatitis B virus DNA sequence.
[0087] In any embodiment herein, the target site or each of the target sites comprises the sequence shown in any one of SEQ ID NO: 5, 6, 12, 18, 22, 26, 29, 38, 42, 43, 51, 56, 61, 63, 68, 72, 75, 79, 82, 84, 88, 89, 98, 99, 113, 116, 121, 124, 125, 118, 130, 133, 135, 138, 143, 150, 152, 155, 158, 164, 165, 175, 176, 182, 185, 189, 190, 192, a continuous portion of at least 14 nucleotides (nt) thereof, or a complementary sequence to any one of the foregoing. In any embodiment herein, the target site or each of the target sites comprises the sequence shown in any one of SEQ ID NO: 12, 18, 20, 22, 26, 27, 46, 50, 63, 66, 73, 79, 185, 192, a continuous portion of at least 14 nucleotides (nt) thereof, or a complementary sequence of any of the foregoing. In any embodiment herein, the target site or each of the target sites comprises the sequence shown in SEQ ID NO: 22, a continuous portion of at least 14 nucleotides (nt) thereof, or a complementary sequence of any of the foregoing. In any embodiment herein, the gRNA, wherein the gRNA or each of the gRNAs comprises SEQ ID NO: The sequence shown in any one of 200, 201, 207, 217, 221, 224, 233, 237, 238, 246, 251, 256, 258, 263, 267, 274, 270, 277, 279, 283, 284, 293, 294, 308, 311, 313, 316, 319, 320, 325, 328, 330, 333, 338, 345, 347, 350, 353, 359, 360, 370, 371, 377, 380, 384, 385, 387, a continuous portion of at least 14 nucleotides thereof, or a complementary sequence of any one of the foregoing, optionally wherein the gRNA or each of the gRNAs as shown in SEQ ID NO: Any one of 395, 402, 408, 412, 416, 419, 428, 432, 433, 441, 446, 451, 453, 458, 462, 465, 469, 472, 474, 478, 479, 488, 489, 503, 506, 508, 511, 514, 515, 520, 523, 525, 575, 528, 533, 540, 542, 545, 548, 554, 555, 565, 566, 572, 579, 580, or 582 is shown.In any embodiment herein, the gRNA or each of the gRNAs comprises the sequence shown in any one of SEQ ID NO: 207, 213, 215, 217, 221, 222, 241, 245, 258, 261, 268, 274, 380, 387, a continuous portion of at least 14 nucleotides thereof, or a complementary sequence of any of the foregoing, optionally wherein the gRNA or each of the gRNAs is shown in any one of SEQ ID NO: 402, 408, 410, 412, 416, 417, 436, 440, 453, 456, 463, 469, 575, 582. In any embodiment herein, the gRNA or each of the gRNAs comprises the sequence shown in any of SEQ ID NO: 217, a continuous portion of at least 14 nucleotides thereof, or a complementary sequence to any of the foregoing, optionally wherein the gRNA or each of the gRNAs is shown in any of SEQ ID NO: 412.
[0088] In any embodiment herein, the at least one DNA-binding domain comprises an engineered zinc finger protein (eZFP). In any embodiment herein, the target site comprises the nucleotide sequence shown in any of SEQ ID NO: 1045, 1046, 1052, a continuous portion thereof of at least 12 nt, or a complementary sequence to any of the foregoing. In any embodiment herein, the target site comprises the nucleotide sequence shown in any of SEQ ID NO: 1045, 1046, 1052.
[0089] In any embodiment herein, at least one effector domain is a DNA methyltransferase. In any embodiment herein, at least one effector domain comprises a DNA methyltransferase and a repressor domain capable of recruiting a heterochromatin-inducing factor, or optionally wherein the heterochromatin-inducing factor comprises a histone methyltransferase. In any embodiment herein, the at least one effector domain comprises a DNA methyltransferase and a histone methyltransferase. In any embodiment herein, the at least one effector domain comprises a DNMT3A / L domain or a variant thereof. In any embodiment herein, the at least one effector domain comprises an effector domain comprising the sequence shown in SEQ ID NO: 604 and 607, a portion thereof, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with any of the foregoing. In any embodiment herein, the at least one effector domain further comprises a KRAB domain or a variant thereof. In any embodiment herein, the at least one effector domain further comprises the sequence shown in SEQ ID NO: 590, a portion thereof, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with any of the foregoing. In any embodiment herein, the DNA targeting system comprises the DNMT3A / 3L-dSpCas9-KRAB domain or a variant thereof. In any embodiment herein, the DNA targeting system comprises the sequence shown in SEQ ID NO: 645, a portion thereof, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with any of the foregoing.
[0090] In any embodiment herein, the DNA targeting system comprises the sequence shown in SEQ ID NO: 680, a portion thereof, or a nucleic acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with the foregoing. In any embodiment herein, the DNA targeting system comprises the sequence shown in SEQ ID NO: 680.
[0091] This document also provides a method for inhibiting the transcription of one or more genes in hepatitis B virus-infected cells, the method comprising introducing into hepatitis B virus-infected cells an epigenetically modified DNA-targeting system disclosed herein, a gRNA disclosed herein, a CRISPR Cas-gRNA combination disclosed herein, a polynucleotide disclosed herein, a plurality of polynucleotides disclosed herein, a vector disclosed herein, or a portion or component of any of the foregoing. In some embodiments, the DNA-targeting system epigenetically modifies the one or more genes. In some embodiments, the transcription of the one or more genes is reduced compared to comparable cells not subjected to the method.
[0092] In any embodiment herein, transcription of the one or more genes is reduced by at least about 1.25-fold, 1.5-fold, 1.75-fold, 2.0-fold, 2.5-fold, 2.75-fold, 3.0-fold, 3.5-fold, 3.75-fold, 4.0-fold, 4.5-fold, 4.75-fold, 5.0-fold, 5.25-fold, 5.5-fold, 5.75-fold, or 6-fold. In any embodiment herein, repression of transcription of the one or more genes results in HBV replication and / or reduced HBV transcription. In any embodiment herein, the HBV-infected cells are mammalian cells.
[0093] In any embodiment herein, the HBV-infected cell is a human cell. In any embodiment herein, the cell contains integrated HBV DNA. In any embodiment herein, the cell is a hepatocyte containing a pool of free HBV cccDNA. In any embodiment herein, the hepatocyte expresses HBV protein, wherein the HBV protein is HBsAg and / or HBeAg. In any embodiment herein, the HBV-infected cell is present in a subject.
[0094] In any embodiment of this document, the subject is a human. In any embodiment of this document, the subject has HBV virus infection. In any embodiment of this document, the subject has hepatocytes containing integrated HBV DNA. In any embodiment of this document, the subject has hepatocytes containing a pool of free HBV cccDNA. In any embodiment of this document, the subject has hepatocytes expressing HBV protein, wherein the HBV protein is HBsAg, HBeAg, or HBcrAg, or a combination thereof.
[0095] In any embodiment of this document, the subject suffers from a disease, condition, or disorder associated with the HBV virus infection. In any embodiment of this document, the disease, condition, or disorder is liver disease or cancer. In any embodiment of this document, the disease, condition, or disorder is acute hepatitis, chronic hepatitis, liver failure, or cirrhosis. In any embodiment of this document, the disease, condition, or disorder is cancer, optionally wherein the cancer is hepatocellular carcinoma.
[0096] This document also provides a pharmaceutical composition comprising a carrier disclosed herein. In any embodiment herein, the carrier is conjugated to an aminoglycoside derivative of galactose, optionally wherein the carrier is partially conjugated to N-acetylgalactosamine (GalNAc).
[0097] This document also provides a pharmaceutical composition comprising the epigenetic DNA-targeting system or the fusion protein disclosed herein, the gRNA disclosed herein, the CRISPR Cas-gRNA combination disclosed herein, the polynucleotide disclosed herein, a plurality of polynucleotides disclosed herein, the vector disclosed herein, or a portion or component of any of the foregoing.
[0098] This document also provides a pharmaceutical composition for use in treating a subject with HBV virus infection. In any embodiment of this document, the subject suffers from a disease, condition, or disorder associated with the HBV virus infection.
[0099] This article also provides a pharmaceutical composition for use in treating a subject with a disease, disorder, or condition associated with HBV infection.
[0100] This document also provides a pharmaceutical composition for use in the preparation of a medicament for treating HBV virus infection in a subject. In some embodiments, the HBV virus infection is associated with a disease, disability, or condition.
[0101] This article also provides a pharmaceutical composition for use in the preparation of a medicament for treating a subject with a disease, condition, or disorder associated with HBV virus infection.
[0102] In any embodiment of this document, the disease, condition, or obstacle is liver disease or cancer. In any embodiment of this document, the disease, condition, or obstacle is acute hepatitis, chronic hepatitis, liver failure, or cirrhosis. In any embodiment of this document, the disease, condition, or obstacle is cancer, optionally hepatocellular carcinoma. In any embodiment of this document, the pharmaceutical composition will be administered to the subject in vivo.
[0103] In any embodiment herein, upon administration of the pharmaceutical composition, transcription of one or more HBV genes in the subject's cells is repressed. In any embodiment herein, the one or more HBV genes are involved in controlling HBV replication and / or HBV transcription. In any embodiment herein, the one or more genes are polymerase genes, S-family genes, X-genes, or core family genes.
[0104] This document also provides a method for treating a disease, condition, or disorder in a subject in need, the method comprising administering to the subject a disclosed epigenetic modified DNA targeting system, a disclosed gRNA, a disclosed CRISPR Cas-gRNA combination, a disclosed polynucleotide, a disclosed plurality of disclosed polynucleotides, a disclosed vector, a disclosed pharmaceutical composition, or a portion or component of any of the foregoing.
[0105] This document also provides a method for reducing hepatitis B virus infection in a subject, the method comprising administering to a subject suffering from hepatitis B virus infection the epigenetic modified DNA targeting system disclosed herein, the gRNA disclosed herein, the CRISPR Cas-gRNA combination disclosed herein, the polynucleotide disclosed herein, multiple polynucleotides disclosed herein, the vector disclosed herein, the pharmaceutical composition disclosed herein, or any part or component of the foregoing.
[0106] This document also provides an engineered zinc finger protein (eZFP) that binds to a target site in one or more HBV genes or their regulatory elements, wherein the target site is located within a 1033 bp–1749 bp target region of the HBV genome at a location corresponding to the HBV genome shown in reference SEQ ID NO: 650. In any embodiment herein, the target site is located within a 300-base pair upstream of the start codon of the hepatitis B virus X protein (HBx). In any embodiment herein, the target site is located in the HBx basic core promoter region. In any embodiment herein, the target site is located within the HBx promoter / enhancer region. In any embodiment herein, the target site is located within a 250-base pair upstream of the start codon of the hepatitis B virus X protein (HBx). In any embodiment herein, the target site is located within a 1060–1480 bp target region of the HBV genome at a location corresponding to the HBV genome shown in reference SEQ ID NO: 650. In any embodiment herein, the target site is located within a target region spanning 150 base pairs upstream of the start codon of the hepatitis B virus X protein (HBx). In any embodiment herein, the target site is located within a target region spanning 120 base pairs upstream of the start codon of the hepatitis B virus X protein (HBx). In any embodiment herein, the target site is located within a target region sequence corresponding to a 1250-1374 bp sequence of the HBV genome spanning the HBV genome shown in reference SEQ ID NO: 650. In some embodiments, the target region has the sequence shown in SEQ ID NO: 1068. In any embodiment herein, the target site is located within a target region sequence corresponding to a 1255-1302 bp sequence of the HBV genome spanning the HBV genome shown in reference SEQ ID NO: 650. In some embodiments, the target region has the sequence shown in SEQ ID NO: 1069. In any embodiment herein, the target site is within a target region sequence corresponding to a 1260-1300 bp sequence of the HBV genome spanning the HBV genome shown in reference SEQ ID NO: 650. In some embodiments, the target region has the sequence shown in SEQ ID NO: 1070. In any embodiment herein, the target site is within a target region sequence corresponding to a 1255-1290 bp sequence of the HBV genome spanning the HBV genome shown in reference SEQ ID NO: 650. In any embodiment herein, the target site comprises the nucleotide sequence shown in any one of SEQ ID NO: 1028-1055, a continuous portion thereof of at least 12 nt, or a complementary sequence to any of the foregoing.In any embodiment herein, the target site comprises the nucleotide sequence shown in any of SEQ ID NO: 1028-1055.
[0107] In any embodiment herein, the target site comprises the nucleotide sequence shown in any one of SEQ ID NO: 1045, 1046, or 1052, a continuous portion thereof of at least 12 nt, or a complementary sequence to any one of the foregoing. In any embodiment herein, the target site comprises the nucleotide sequence shown in any one of SEQ ID NO: 1045, 1046, or 1052.
[0108] In any embodiment of this document, the zinc finger protein comprises six zinc fingers, denoted as F1 to F6, sequentially from the N-terminus to the C-terminus, and wherein the amino acid sequence of the recognition region of each zinc finger is as follows: 1) F1: SEADRSR (SEQ ID NO: 720) F2: DRSNLTR (SEQ ID NO: 721) F3: QSSDLSR (SEQ ID NO: 722) F4: YHWYLKK (SEQ ID NO: 723) F5: RSDSLSV (SEQ ID NO: 724) F6: QNANRKT (SEQ ID NO: 725); 2) F1: RSDVLST (SEQ ID NO: 726) F2: DNSSRTR (SEQ ID NO: 727) F3: RPYTLRL (SEQ ID NO: 728) F4: DSSHRTR (SEQ ID NO: 729) F5: RSDHLSQ (SEQ ID NO: 730) F6: DSSHRTR (SEQ ID NO: 729) 731); 3) F1: RSDHLSQ (SEQ ID NO: 732) F2: QSADRTK (SEQ ID NO: 733) F3: RSDHLSQ (SEQ ID NO: 734) F4: RRSDLKR (SEQ ID NO: 735) F5: RSDHLSR (SEQ ID NO: 736) F6: QSSDLRR (SEQ ID NO: 5) F1: RSDHLSE (SEQ ID NO: 744) F2: QYSGRYY (SEQ ID NO: 745) F3: HGQTLNE (SEQ ID NO: 746) F4: QSGNLAR (SEQ ID NO: 747) F5: RSDSLLR (SEQ ID NO: 748) F6: CREYRGK (SEQ ID NO: 749);6) F1: QSANRTT (SEQ ID NO: 750) F2: RSANLTR (SEQ ID NO: 751) F3: RSDVLSE (SEQ ID NO: 752) F4: TSGHLSR (SEQ ID NO: 753) F5: QSSDLSR (SEQ ID NO: 754) F6: QWSTRKR (SEQ ID NO: 755) ; 7) F1: QSGNLAR (SEQ ID NO: 756) F2: ATCCLAH (SEQ ID NO: 757) F3: RWQYLPT (SEQ ID NO: 758) F4: DRSALAR (SEQ ID NO: 759) F5: RSDNLSE (SEQ ID NO:760) F6: KRCNLRC (SEQ ID NO: 761) ;8) F1: NPANLTR (SEQ ID NO: 762) F2: QNATRTK (SEQ IDNO: 763) F3: QSGHLAR (SEQ ID NO: 764) F4: NRHDRAK (SEQ ID NO: 765) F5: RSDHLSE (SEQ IDNO: 766) F6: QRRSRYK (SEQ ID NO: 767) ;9) F1: QSSDLSR (SEQ ID NO: 768) F2: HRSTRNR (SEQ ID NO: 769) F3: RSDVLSA (SEQ ID NO: 770)F4:DSRTRKN(SEQ ID NO: 771)F5:QSGSLTR(SEQ ID NO: 772)F6:DQSGLAH(SEQ ID NO: 773);10) F1:QNPAQWR(SEQ ID NO:774)F2:RSADLSR(SEQ ID NO: 775)F3:TSGSLSR(SEQ ID NO: 776)F4:RSDHLSR(SEQ ID NO:777)F5:RSDSLLR(SEQ ID NO: 778)F6:QSYDRFQ(SEQ ID NO: 779);11) F1:TSGSLSR(SEQID NO: 780) F2: RSDHLSR (SEQ ID NO: 781) F3: RDSLLR (SEQ ID NO: 782) F4: QSYDRFQ (SEQ ID NO: 783) F5: RSDNLST (SEQ ID NO: 784) F6: DNRDRIK (SEQ ID NO: 785)12) F1: DRSNLSR (SEQ ID NO: 786) F2: LRQNLIM (SEQ ID NO: 787) F3: ERGTLAR (SEQ ID NO: 788) F4: RSDALTQ (SEQ ID NO: 789) F5: RSDSLSQ (SEQ ID NO: 790) F6: RKADRTR (SEQ ID NO: 791) 13) F1: QYCCLTN (SEQ ID NO: 792) F2: TSGNLTR (SEQ ID NO: 793) F3: QSSDLSR (SEQ ID NO: 794) F4: FRYYLKR (SEQ ID NO: 795) F5: QSGDLTR (SEQ ID NO: 796) F6: DKGNLTK (SEQ ID NO: 797) ;14) F1: TSGSLSR (SEQ ID NO: 798) F2: RSDNLTT (SEQ ID NO: 799) F3: QSGNLAR (SEQ ID NO: 800) F4: DRTTLMR (SEQ ID NO: 801) F5: QSGHLAR (SEQ ID NO: 802) F6: QLTHLNS (SEQ ID NO: 803) ;15) F1: IKHDLHR (SEQ ID NO: 804) F2: RSANLTR (SEQ ID NO: 805) F3: RSDNLAR (SEQ ID NO: 806)F4:QNVSRPR(SEQ ID NO: 807)F5:RSDDLSK(SEQ IDNO: 808)F6:DSSHRTR(SEQ ID NO: 809);16) F1:RSDNLAR(SEQ ID NO: 810)F2:QNVSRPR(SEQ ID NO: 811)F3:RSDDLSK(SEQ ID NO: 812)F4:DSSHRTR(SEQ ID NO: 813)F5:TSSNRKT(SEQ ID NO: 814)F6:AQWTRAC(SEQ ID NO: 815);17) F1:RSDDLSK(SEQ ID NO:816) F2: DSSHRTR (SEQ ID NO: 817) F3: TSSNRKT (SEQ ID NO: 818) F4: AQWTRAC (SEQ ID NO: 819) F5: RKQTRTT (SEQ ID NO: 820) F6: HRSSLRR (SEQ ID NO: 821)18) F1: QSAHRKN (SEQID NO: 822) F2: TSSNRKT (SEQ ID NO: 823) F3: RSDNLSA (SEQ ID NO: 824) F4: RNNDRKT (SEQ ID NO: 825) F5: TSGSLSR (SEQ ID NO: 826) F6: QAGHLAK (SEQ ID NO: 827) 19) F1: RSDHLSQ (SEQ ID NO: 828) F2: ASSTRTK (SEQ ID NO: 829) F3: RSDDLTR (SEQ ID NO: 830) F4: QKSNLSS (SEQ ID NO: 831) F5: QSANRTT (SEQ ID NO: 832) F6: QNATRTK (SEQ ID NO: 833) ;20) F1: RSDTLSE (SEQ ID NO: 834) F2: RRWTLVG (SEQ ID NO: 835) F3: DRSNLSR (SEQ ID NO: 836) F4: QSGDLTR (SEQ ID NO: 837) F5: QSSDLSR (SEQ ID NO: 838) F6: YHWYLKK (SEQ ID NO: 839) ;21) F1: RSANLAR (SEQ ID NO: 840) F2: RSDNLRE (SEQ ID NO: 841) F3: RPYTLRL (SEQ ID NO: 842)F4:HRSNLNK(SEQ ID NO: 843)F5:QSGSLTR(SEQ ID NO: 844)F6:TSANLSR(SEQ ID NO: 845);22) F1:RSDDLVR(SEQ ID NO: 846)F2:TSGSLVR(SEQ IDNO: 847)F3:RSDKLVR(SEQ ID NO: 848)F4:RSDELVR(SEQ ID NO: 849)F5:TSHSLTE(SEQ IDNO: 850)F6:RADNLTE(SEQ ID NO: 851);23) F1:ERSHLRE(SEQ ID NO: 852) F2: THSHSLTE (SEQ ID NO: 853) F3: QAGHLAS (SEQ ID NO: 854) F4: THSHSLTE (SEQ ID NO: 855) F5: DPGHLVR (SEQ ID NO: 856) F6: TGSGNLVR (SEQ ID NO: 857)24) F1: RADNLTE (SEQ ID NO:858), F2: TSGSLVR (SEQ ID NO: 859), F3: RKDNLKN (SEQ ID NO: 860), F4: QSSSLVR (SEQ ID NO:861), F5: RSDKLVR (SEQ ID NO: 862), F6: DSG NLRV (SEQ ID NO: 863); 25) F1: QSSSLVR (SEQID NO: 864), F2: QSGDLRR (SEQ ID NO: 865), F3: RSDERKR (SEQ ID NO: 866), F4: HRTTLTN (SEQID NO: 867), F5: RSDHLTN (SEQ ID NO: 868), F6: TSGELVR (SEQ ID NO: 869); 26) F1: QSGDLRR (SEQ ID NO: 870), F2: RSDERKR (SEQ ID NO: 871), F3: HRTTLTN (SEQ ID NO: 872), F4: RSDHLTN (SEQ ID NO: 873), F5: TSGELVR (SEQ ID NO: 874), F6: RSDDLVR (SEQ ID NO:875); 27) F1: QRAHLER (SEQ ID NO: 876), F2: QLAHLRA (SEQ ID NO: 877), F3: DPGHLVR (SEQID NO: 878), F4: RRSACRR (SEQ ID NO: 879), F5: RSDHLTT (SEQ ID NO: 880), F6: QSSSLVR (SEQID NO: 881); and 28) F1: QSSNLVR (SEQ ID NO: 882), F2: RSDDLVR (SEQ ID NO: 883), F3: THLDLIR (SEQ ID NO: 884), F4: TSGNLTE (SEQ ID NO: 885), F5: RRSACRR (SEQ ID NO: 886), F6: RNDTLTE (SEQ ID NO: 887).;
[0109] In any embodiment of this document, the zinc finger protein comprises six zinc fingers, denoted as F1 to F6, sequentially from the N-terminus to the C-terminus, and wherein the amino acid sequence of the recognition region of each zinc finger is as follows: F1: QSAHRKN (SEQ ID NO: 822) F2: TSSNRKT (SEQ ID NO: 823) F3: RSDNLSA (SEQ ID NO: 824) F4: RNNDRKT (SEQ ID NO: 825) F5: TSGSLSR (SEQ ID NO: 826) F6: QAGHLAK (SEQ ID NO: 827).
[0110] In any embodiment of this document, the zinc finger protein comprises six zinc fingers, denoted as F1 to F6, sequentially from the N-terminus to the C-terminus, and wherein the amino acid sequence of the recognition region of each zinc finger is as follows: F1: RSDHLSQ (SEQ ID NO: 828) F2: ASSTRTK (SEQ ID NO: 829) F3: RSDDLTR (SEQ ID NO: 830) F4: QKSNLSS (SEQ ID NO: 831) F5: QSANRTT (SEQ ID NO: 832) F6: QNATRTK (SEQ ID NO: 833).
[0111] In any embodiment herein, the zinc finger protein comprises, sequentially from the N-terminus to the C-terminus, six zinc fingers denoted as F1 to F6, wherein the amino acid sequence of the recognition region of each zinc finger is as follows: F1: QSSSLVR (SEQ ID NO: 864) F2: QSGDLRR (SEQ ID NO: 865) F3: RSDERKR (SEQ ID NO: 866) F4: HRTTLTN (SEQ ID NO: 867) F5: RSDHLTN (SEQ ID NO: 868) F6: TSGELVR (SEQ ID NO: 869).
[0112] This article also provides an eZFP that binds to a target site in one or more HBV genes or their regulatory elements, wherein the zinc finger protein comprises six zinc fingers, denoted as F1 to F6, sequentially from the N-terminus to the C-terminus, and wherein the amino acid sequence of the recognition region of each zinc finger is as follows: 1) F1: SEADRSR (SEQ ID NO: 720) F2: DRSNLTR (SEQ ID NO: 721) F3: QSSDLSR (SEQ ID NO: 722) F4: YHWYLKK (SEQ ID NO: 723) F5: RSDSLSV (SEQ ID NO: 724) F6: QNANRKT (SEQ ID NO: 725); 2) F1: RSDVLST (SEQ ID NO: 726) F2: DNSSRTR (SEQ ID NO: 727) F3: RPYTLRL (SEQ ID NO: 728) F4: DSSHRTR (SEQ ID NO: 729) F5: RSDHLSQ (SEQ ID NO: 729) F6 ...6: RPYTLRL (SEQ ID NO: 729) F7: RPYTLRL (SEQ ID NO: 728) F6: RPYTLRL (SEQ ID NO: 729) F7: RPYTLRL (SEQ ID NO: 728) F7: RPYTLRL (SEQ ID NO: 729) F6: RPYTLRL (SEQ ID NO: 728) F7: RPYTLRL (SEQ ID NO: 729) F6: RPYTLRL (SEQ ID NO: 728) F7: RPYTLRL (SEQ ID NO: 72 730) F6: DSSHRTR (SEQ ID NO: 731); 3) F1: RSDHLSQ (SEQ ID NO: 732) F2: QSADRTK (SEQ ID NO: 733) F3: RSDHLSQ (SEQ ID NO: 734) F4: RRSDLKR (SEQ ID NO: 735) F5: RSDHLSR (SEQ ID NO: 736) F6: QSSDLRR (SEQ ID NO: 737); 4) F1: RSDNLSE (SEQ ID NO: 738) F2: TSSNRKT (SEQ ID NO: 739) F3: DRSHLTR (SEQ ID NO: 740) F4: RSDALTQ (SEQ ID NO: 741) F5: DRSALAR (SEQ ID NO: 742) F6: RRFTLSK (SEQ ID NO: 743); 5) F1: RSDHLSE (SEQ ID NO: 744) F2: QYSGRYY (SEQ ID NO: 745) F3: HGQTLNE (SEQ ID NO: 746) F4: QSGNLAR (SEQ ID NO: 747) F5: RSDSLLR (SEQ ID NO: 748) F6: CREYRGK (SEQ ID NO: 749);6) F1: QSANRTT (SEQ ID NO: 750) F2: RSANLTR (SEQ ID NO: 751) F3: RSDVLSE (SEQ ID NO: 752) F4: TSGHLSR (SEQ ID NO: 753) F5: QSSDLSR (SEQ ID NO: 754) F6: QWSTRKR (SEQ ID NO: 755) ; 7) F1: QSGNLAR (SEQ ID NO: 756) F2: ATCCLAH (SEQ ID NO: 757) F3: RWQYLPT (SEQ ID NO: 758) F4: DRSALAR (SEQ ID NO: 759) F5: RSDNLSE (SEQ ID NO: 760) F6: KRCNLRC (SEQ ID NO: 761) ;8) F1: NPANLTR (SEQ ID NO: 762) F2: QNATRTK (SEQ ID NO: 763) F3: QSGHLAR (SEQ IDNO: 764) F4: NRHDRAK (SEQ ID NO: 765) F5: RSDHLSE (SEQ ID NO: 766) F6: QRRSRYK (SEQ IDNO: 767) ;9) F1: QSSDLSR (SEQ ID NO: 768) F2: HRSTRNR (SEQ ID NO: 769) F3: RSDVLSA (SEQ ID NO: 770)F4:DSRTRKN(SEQ ID NO: 771)F5:QSGSLTR(SEQ ID NO: 772)F6:DQSGLAH(SEQ ID NO: 773);10) F1:QNPAQWR(SEQ ID NO: 774)F2:RSADLSR(SEQ ID NO:775)F3:TSGSLSR(SEQ ID NO: 776)F4:RSDHLSR(SEQ ID NO: 777)F5:RSDSLLR(SEQ ID NO:778)F6:QSYDRFQ(SEQ ID NO: 779);11) F1:TSGSLSR(SEQ ID NO: 780) F2: RSDHLSR (SEQID NO: 781) F3: RDSLLR (SEQ ID NO: 782) F4: QSYDRFQ (SEQ ID NO: 783) F5: RSDNLST (SEQID NO: 784) F6: DNRDRIK (SEQ ID NO: 785)12) F1: DRSNLSR (SEQ ID NO: 786) F2: LRQNLIM (SEQ ID NO: 787) F3: ERGTLAR (SEQ ID NO: 788) F4: RSDALTQ (SEQ ID NO: 789) F5: RSDSLSQ (SEQ ID NO: 790) F6: RKADRTR (SEQ ID NO: 791) 13) F1: QYCCLTN (SEQ IDNO: 792) F2: TSGNLTR (SEQ ID NO: 793) F3: QSSDLSR (SEQ ID NO: 794) F4: FRYYLKR (SEQ IDNO: 795) F5: QSGDLTR (SEQ IDNO: 795) ID NO: 796) F6: DKGNLTK (SEQ ID NO: 797) ;14) F1: TSGSLSR (SEQ ID NO: 798) F2: RSDNLTT (SEQ ID NO: 799) F3: QSGNLAR (SEQ ID NO: 800) F4: DRTTLMR (SEQ ID NO: 801) F5: QSGHLAR (SEQ ID NO: 802) F6: QLTHLNS (SEQ ID NO: 803) ;15) F1: IKHDLHR (SEQ ID NO: 804) F2: RSANLTR (SEQ ID NO: 805) F3: RSDNLAR (SEQ ID NO: 806)F4: 810)F2:QNVSRPR(SEQ ID NO: 811)F3:RSDDLSK(SEQID NO: 812)F4:DSSHRTR(SEQ ID NO: 813)F5:TSSNRKT(SEQ ID NO: 814)F6:AQWTRAC(SEQID NO: 815);17) F1: RSDDLSK(SEQ ID NO: 816) F2: DSSHRTR (SEQ ID NO: 817) F3: TSNRKT (SEQ ID NO: 818) F4: AQWTRAC (SEQ ID NO: 819) F5: RKQTRTT (SEQ ID NO: 820) F6: HRSSLRR (SEQ ID NO: 821)18) F1: QSAHRKN (SEQ ID NO: 822) F2: TSSNRKT (SEQ ID NO: 823) F3: RSDNLSA (SEQ ID NO: 824) F4: RNNDRKT (SEQ ID NO: 825) F5: TSGSLSR (SEQ ID NO: 826) F6: QAGHLAK (SEQ ID NO: 827) 19) F1: RSDHLSQ (SEQ ID NO: 828) F2: ASSTRTK (SEQ ID NO: 829) F3: RSDDLTR (SEQ ID NO: 830) F4: QKSNLSS (SEQ ID NO: 831) F5: QSANRTT (SEQ ID NO: 832) F6: QNATRTK (SEQ ID NO: 833) ;20) F1: RSDTLSE (SEQ ID NO: 834) F2: RRWTLVG (SEQ ID NO: 835) F3: DRSNLSR (SEQ ID NO: 836) F4: QSGDLTR (SEQ ID NO: 837) F5: QSSDLSR (SEQ ID NO: 838) F6: YHWYLKK (SEQ ID NO: 839) ;21) F1: RSANLAR (SEQ ID NO: 840) F2: RSDNLRE (SEQ ID NO: 841) F3: RPYTLRL (SEQ ID NO: 842)F4:HRSNLNK(SEQID NO: 843)F5:QSGSLTR(SEQ ID NO: 844)F6:TSANLSR(SEQ ID NO: 845);22) F1:RSDDLVR(SEQ ID NO: 846)F2:TSGSLVR(SEQ ID NO: 847)F3:RSDKLVR(SEQ ID NO: 848)F4:RSDELVR(SEQ ID NO: 849)F5:TSHSLTE(SEQ ID NO: 850)F6:RADNLTE(SEQ ID NO:851);23) F1:ERSHLRE(SEQ ID NO: 852) F2: THSHSLTE (SEQ ID NO: 853) F3: QAGHLAS (SEQ ID NO: 854) F4: THSHSLTE (SEQ ID NO: 855) F5: DPGHLVR (SEQ ID NO: 856) F6: TGSGNLVR (SEQ ID NO: 857)24) F1: RADNLTE (SEQ ID NO: 858), F2: TSGSLVR (SEQ ID NO: 859), F3: RKDNLKN (SEQ ID NO: 860), F4: QSSSLVR (SEQ ID NO: 861), F5: RSDKLVR (SEQ ID NO: 862), F6: DSG NLRV (SEQ ID NO: 863); 25) F1: QSSSLVR (SEQ ID NO: 864), F2: QSGDLRR (SEQ ID NO: 865), F3: RSDERKR (SEQ ID NO: 866), F4: HRTTLTN (SEQ ID NO: 867), F5: RSDHLTN (SEQ ID NO: 868), F6: TSGELVR (SEQ ID NO: 869); 26) F1: QSGDLRR (SEQ ID NO: 870), F2: RSDERKR (SEQ ID NO: 871), F3: HRTTLTN (SEQ ID NO: 872), F4: RSDHLTN (SEQ ID NO: 873), F5: TSGELVR (SEQ ID NO: 874), F6: RSDDLVR (SEQ ID NO: 875); 27) F1: QRAHLER (SEQ ID NO: 876), F2: QLAHLRA (SEQ ID NO: 877), F3: DPGHLVR (SEQ ID NO: 878), F4: RRSACRR (SEQ ID NO: 879), F5: RSDHLTT (SEQ ID NO: 880), F6: QSSSLVR (SEQ ID NO: 881); and 28) F1: QSSNLVR (SEQ ID NO: 882), F2: RSDDLVR (SEQ ID NO: 883), F3: THLDLIR (SEQ ID NO: 884), F4: TSGNLTE (SEQ ID NO: 885), F5: RRSACRR (SEQ ID NO: 886), F6: RNDTLTE (SEQ ID NO: 887).;
[0113] In any embodiment herein, the engineered zinc finger protein comprises a sequence or a portion thereof shown in any of SEQ ID NO: 692-719, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In any embodiment herein, the engineered zinc finger protein is encoded by a sequence or a portion thereof shown in any of SEQ ID NO: 888-915, or a nucleotide sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it.
[0114] This document also provides an engineered zinc finger protein that binds to a target site in one or more HBV genes or their regulatory elements, wherein the zinc finger protein comprises, sequentially from the N-terminus to the C-terminus, six zinc fingers denoted as F1 to F6, namely F1: QSAHRKN (SEQ ID NO: 822), F2: TSSNRKT (SEQ ID NO: 823), F3: RSDNLSA (SEQ ID NO: 824), F4: RNNDRKT (SEQ ID NO: 825), F5: TSGSLSR (SEQ ID NO: 826), and F6: QAGHLAK (SEQ ID NO: 827). In some embodiments, the engineered zinc finger protein comprises the sequence shown in SEQ ID NO: 709 or a portion thereof, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the engineered zinc finger protein comprises any one of the sequences shown in SEQ ID NO: 709. In some embodiments, the engineered zinc finger protein is encoded by the sequence shown in SEQ ID NO: 905 or a portion thereof, or by a nucleotide sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the engineered zinc finger protein is encoded by the sequence shown in any one of SEQ ID NO: 905.
[0115] This document also provides an engineered zinc finger protein that binds to a target site in one or more HBV genes or their regulatory elements, wherein the zinc finger protein comprises six zinc fingers, designated F1 to F6, sequentially from the N-terminus to the C-terminus, and wherein the amino acid sequence of each zinc finger recognition region is as follows: F1: RSDHLSQ (SEQ ID NO: 828) F2: ASSTRTK (SEQ ID NO: 829) F3: RSDDLTR (SEQ ID NO: 830) F4: QKSNLSS (SEQ ID NO: 831) F5: QSANRTT (SEQ ID NO: 832) F6: QNATRTK (SEQ ID NO: 833). In some embodiments, the engineered zinc finger protein comprises the sequence shown in SEQ ID NO: 710 or a portion thereof, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the engineered zinc finger protein comprises the sequence shown in any of SEQ ID NO: 710. In some embodiments, the engineered zinc finger protein is encoded by the sequence shown in SEQ ID NO: 906, or a portion thereof, or a nucleotide sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the engineered zinc finger protein is encoded by the sequence shown in any of SEQ ID NO: 906.
[0116] This document also provides an engineered zinc finger protein that binds to a target site in one or more HBV genes or their regulatory elements, wherein the zinc finger protein comprises six zinc fingers, designated F1 to F6, sequentially from the N-terminus to the C-terminus, and wherein the amino acid sequence of the recognition region of each zinc finger is as follows: F1: QSSSLVR (SEQ ID NO: 864) F2: QSGDLRR (SEQ ID NO: 865) F3: RSDERKR (SEQ ID NO: 866) F4: HRTTLTN (SEQ ID NO: 867) F5: RSDHLTN (SEQ ID NO: 868) F6: TSGELVR (SEQ ID NO: 869). In some embodiments, the engineered zinc finger protein comprises the sequence shown in SEQ ID NO: 716 or a portion thereof, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the engineered zinc finger protein comprises the sequence shown in any of SEQ ID NO: 716. In some embodiments, the engineered zinc finger protein is encoded by the sequence shown in SEQ ID NO: 912, or a portion thereof, or a nucleotide sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the engineered zinc finger protein is encoded by the sequence shown in any of SEQ ID NO: 912. Attached Figure Description
[0117] Figure 1 The fold change in total HBV RNA mediated by each guide RNA and dCas9-KRAB effector fusion protein was depicted. The fold change was correlated with the median target location of each gRNA across all HBV genotypes.
[0118] Figure 2 The conservation of guide RNA in HBV genome types (HBV0-HBV12) was depicted.
[0119] Figures 3A-3B This describes the action on total HBV RNA mediated by the selected top gRNA candidates (HBVg_192 (SEQ ID NO: 582), HBVg_17 (SEQ ID NO: 407), HBVg_63 (SEQ ID NO: 453)) on day 5 post-transfection. Figure 3A ) and HBsAg ( Figure 3B () hinders.
[0120] Figure 4The fold changes in total HBV RNA mediated by each guide RNA and the DNMT3A / L-dCas9-KRAB effector fusion protein were depicted. The fold changes were plotted against the median target location of each gRNA in the HBV genome.
[0121] Figure 5 The inhibition of total HBV RNA mediated by the selected top gRNA candidates (HBVg_142 (SEQ ID NO: 532), HBVg_138 (SEQ ID NO: 528), HBVg_185 (SEQ ID NO: 575), HBVg_152 (SEQ ID NO: 542)) was described at days 41 and 61 post-transfection.
[0122] Figures 6A-6B The durable and stable repression of total HBV RNA mediated by gRNAs HBVg_63 (SEQ ID NO: 453), HBVg_185 (SEQ ID NO: 575), and HBVg_56 (SEQ ID NO: 446) was described.
[0123] Figure 7 Normalized total HBV RNA expression is shown after re-dose with a combination of exemplary gRNA and epigenetic editor.
[0124] Figure 8 The fold change in 3.5 kb HBV RNA mediated by each guide RNA and the DNMT3A / L-dSpCas9-KRAB effector fusion protein was depicted in a real infection model. The fold change was plotted against the median target location of each gRNA in the HBV genome.
[0125] Figure 9 The correlation plot between the two cccDNA screenings shows the consistency between infections.
[0126] Figure 10 The gRNAs that inhibit integration of HBV, cccDNA, or both HBV and cccDNA are shown.
[0127] Figure 11 The study lists single gRNAs that target HBV cccDNA and integrated DNA, and achieve multiple HBV repressions at multiple genes and their regulatory elements.
[0128] Figure 12 The sequences targeted by single gRNAs that achieve multiple HBV repressions on multiple genes and their regulatory elements are listed.
[0129] Figure 13The fold change in total HBV RNA mediated by each guide RNA and exemplary epigenetic editor was depicted. The fold change was plotted as a correlation with the median target location of each gRNA in the HBV genome.
[0130] Figure 14 The study demonstrated the repression of the targeted HBV cccDNA transcript by the gRNA alone.
[0131] Figures 15A-15B The inhibition of cccDNA in a primary human hepatocyte (PHH) infection model was described. Figure 15A The inhibition of HBV RNA from two PHH donors was demonstrated by HBVg_22 (SEQ ID NO: 412). Figure 15B This study demonstrates inhibition mediated by HBVg_22 (SEQ ID NO: 412) in PHH cells infected with two doses of HBV.
[0132] Figures 16A-16B The HepG2.NTCP model is shown. Figure 16A ) and PxB PHH model ( Figure 16B In the comparison of repression mediated by HBVg_22 (SEQ ID NO: 412) with either dSpCas9-KRAB alone (SEQ ID NO: 595) or DNMT3A / L-dSpCas9-KRAB (“D3AL-K”; SEQ ID NO: 645) fusion.
[0133] Figure 17 This study describes multiple targeted transcriptional repressions of different regions within HBV RNA in Hep3B cells.
[0134] Figure 18 This demonstrates repression mediated by a single gRNA or multiple gRNAs in the PLC / PRF / 5 (Alexander) cell model.
[0135] Figure 19 This demonstrates a multiplexing approach using a combination of two gRNAs with an exemplary dSpCas9-effectant in a PXB primary human hepatocyte (PHH) cell model.
[0136] Figure 20A and Figure 20B Methyl capture sequencing analysis of cccDNA and integrated HBV DNA is shown.
[0137] Figure 21 The increased methylation pattern following delivery of mRNA encoding DNMT3A / L-dSpCas9-KRAB along with various gRNAs is shown.
[0138] Figure 22 The durable CpG island 2-methylation pattern was demonstrated after delivery of mRNA encoding DNMT3A / L-dCas9-KRAB and HBVg_22 (SEQ ID NO: 412) to HBV-infected PHH donors.
[0139] Figures 23A to 23D : Figure 23A and Figure 23B RNA sequencing analysis revealed gRNA-dependent and epigenetic editor-dependent gene expression changes. Figure 23C This study showed minimal changes in differentially expressed genes mediated by gRNA compared to a lipid-only control. Figure 23D Genes for which there was no differential expression between non-targeted gRNA and HBVg_22 (SEQ ID NO: 412) at any dose or time point were shown.
[0140] Figure 24 This is a schematic diagram illustrating an in vivo study of human chimeric livers in mice.
[0141] Figure 25 Repression mediated by the combination of the exemplary DNMT3A / L-dCas9-KRAB fusion protein and gRNA HBVg_22 (SEQ ID NO: 412) was demonstrated five days after administration (D5). Repression was measured by monitoring HBsAg protein levels and HBV DNA two days before administration (pre-dose) and five days after administration (D5).
[0142] Figure 26 It describes the fold change in indicators after administration compared to before administration.
[0143] Figures 27A-27D It shows the encoding dCas9-KRAB ( Figures 27A-27C ) or DNMT3A / L-dCas9-KRAB ( Figure 27D Combining the mRNA of HBVg_22 (SEQ ID NO: 412) alone or multiplexed with HBVg_185 (SEQ ID NO: 575) for re-administration, or via delivery via LNP conjugated with GalNAc, resulted in repression in FRG mice.
[0144] Figures 28A-28C This study shows a human chimeric FRG mouse procedure following delivery of GalNAc-conjugated PEG LNPs. Figure 28A Stable repression was demonstrated after a single administration of an LNP containing mRNA encoding DNMT3A / 3L-dSpCas9-KRAB and HBVg_22 (SEQ ID NO: 412) on day 33. Figure 28B The inhibition level trajectory of a single mouse is shown. Figure 28C Tissue samples showing significantly reduced pgRNA signaling in mice delivered HBVg_22 (SEQ ID NO: 412) compared to mice receiving non-targeted gRNA were presented.
[0145] Figures 29A-29B The fold change in total HBV RNA mediated by each fusion protein containing the KRAB epigenetic editor and zinc finger protein (ZFP) was depicted. ZFP-KRAB fusion proteins were screened for HBV repression in Hep3B cells in two batches. Figure 29A and Figure 29B ).
[0146] Figure 30 This study presents a preliminary comparison between repression induced by the ZFP-KRAB fusion protein and the combination of the dCas9-KRAB fusion protein and gRNA in Hep3B cells.
[0147] Figures 31A-31B The target location of each ZFP along the X promoter (HBx) site is shown. Figure 31A The highlighted regions represent the areas within the HBx promoter targeted by the most effective gRNAs (HBVg_22 and HBVg_63) and the most effective ZFP-KRAB fusion proteins (eZFP_18, eZFP_19, and eZFP_25). Figure 31B The sites in the HBx region targeted by gRNAs HBVg_22 and HBVg_63, as well as the ZFP-KRAB fusion proteins eZFP_18, eZFP_19, and eZFP_25, were depicted.
[0148] Figures 32A-32B The fold change in total HBV RNA mediated by each eZFP-KRAB fusion protein was depicted in HepG2.NTCP cells.
[0149] Figure 33 The fold change in total HBV RNA mediated by each eZFP-KRAB fusion protein and the conservation of the fusion protein in HBV subtypes are shown.
[0150] Figure 34 The changes in total HBV RNA fold mediated by each DNMT3A / L-eZFP-KRAB fusion protein were shown on days 4 and 15 post-transfection.
[0151] Figure 35 This study showed minimal changes in differentially expressed genes mediated by the eZFP-KRAB fusion protein compared to the lipid-only control.
[0152] Figures 36A-36B Revealed the relationship with GFP ( Figure 36A ) or non-targeted (NT) control ( Figure 36B Compared to RNA sequencing analysis of changes in gene expression dependent on the eZFP-KRAB fusion protein, this study also included a comparison of the results. Detailed Implementation
[0153] Hepatitis B is a potentially life-threatening liver infection caused by the hepatitis B virus (HBV). HBV infection is a global public health problem that causes chronic liver infection and increases the risk of cirrhosis and liver cancer. The WHO estimates that in 2019, 296 million people worldwide had chronic hepatitis B infection (including 1 million in the United States), with 1.5 million new infections each year. In 2019, hepatitis B caused an estimated 820,000 deaths, primarily due to cirrhosis and hepatocellular carcinoma (the primary form of liver cancer).
[0154] HBV belongs to the Hepatoviridae family, a small enveloped hepatoviridae family (Wei L. and Ploss A. Nature communications 12(1591) 1-13 (2021)). The HBV virion contains a compact, partially double-stranded, approximately 3.2 kb relaxed circular DNA (rcDNA) genome. The genome contains four breaks: a covalently linked HBV polymerase and a 10-nucleotide (nt) DNA flap at the 5′ end of the negative strand; and a 5′ capped RNA primer and a single-stranded DNA (ssDNA) nick on the positive strand. The HBV genome is a double-stranded DNA molecule of approximately 3.2 kbp, but can be longer or shorter depending on the specific HBV strain (e.g., up to 3300 bp or larger). An exemplary HBV genome is the hepatitis B virus genome (hepatitis B virus subtype ayw, complete genome, GenBank: U95551.1), SEQ ID NO: 650. At least 10 genotypes (A to J) have been identified, with differences between genotypes not exceeding 8%. Genotype subtypes also exist, including those classified as HBV genotype A (A1-A7), genotype B (B1-B9), genotype C (C1-16), genotype D (D1-D8), and genotype F (F1-F4) (Zhang et al., World J Gastroenterol., 2015, 22:126-144). Within each genotype, sequence identity differences are only about 4%. It should be understood that the systems and methods presented are applicable to a variety of HBV genomes, particularly considering high sequence similarity. For the purposes of this article, the nucleotide position numbers mentioned refer to the nucleotide (base pair) numbers of the HBV DNA sequence described under GenBank accession number U95551.1, as shown in SEQ ID NO: 650. Those skilled in the art will understand that target sites or base pair positions (as described herein) in another HBV genome may not be in the same location, but can still be homologous or substantially homologous sequences (e.g., 1, 2, or 3 mismatches), as can be determined by aligning the HBV genome sequence with the sequence shown in SEQ ID NO: 650. Therefore, one or more corresponding positions can be readily identified by aligning the HBV genome sequence with the reference sequence shown in SEQ ID NO: 650. Hepatitis B virus contains a circular genome; therefore, for the purposes of this document, the numbering of nucleotide positions referring to linear HBV DNA sequences may be offset by a few nucleotides, such as depending on the start of the linear sequence. For example, the sequences shown in SEQ ID NO: 650 and SEQ ID NO: 1071 are identical sequences, but the start of the linear sequence is offset by two nucleotides due to the difference in the start of the linear sequence.Identifying the corresponding sequence regions between different sequences in the HBV genome is entirely within the capabilities of a skilled technician.
[0155] The HBV life cycle includes processes such as viral entry, cccDNA formation, transcription, replication, assembly, secretion, and integration. After the virus enters a hepatocyte via the bile acid transporter NTCP11, the viral nucleocapsid carrying HBV rcDNA is transported to the nucleus. The rcDNA is released, and four damages on the rcDNA are fully repaired to form a supercoiled cccDNA molecule (also known as a mini-chromosome). Viral repair factors are optional for repair, and cccDNA generally relies on the host DNA repair machinery, including TDP2, DNA polymerase (POL)κ, POLα, DNA ligases 1 and 3, and valve endonuclease 1. HBV hijacks liver-enriched transcription factors that are ubiquitous in the host for cccDNA transcriptional regulation. cccDNA is a key viral reservoir driving chronic HBV infection and serves as a template for all HBV viral transcripts. Another form of HBV DNA in the host is HBV DNA stably integrated into the host genome (Zhao K. et al., Cell Press-The Innovation 1(2): 1-10 (2020)). Double-stranded linear DNA (dslDNA) is the primary substrate for integration into the host genome. Due to the minimal sequence homology between viral and cellular DNA, the NHEJ DNA repair pathway is considered the mechanism for HBV DNA integration. HBV DNA integration occurs throughout the host genome at double-strand breaks, with end deletions of up to 200 bp from the integrated HBV DNA being common. No specific chromosomal hotspots or common recurrence sites have been observed among patients. There is some evidence of enrichment at specific genomic sites in tumor tissues (Sung W. et al., Nature Genetics 44(7):765-9 (2012)). Although it does not produce progeny viruses, integrated HBV DNA can generate viral RNA and proteins. HBV DNA integration occurs more frequently in hepatocellular carcinoma cells (84%) than in normal liver tissue (30%).
[0156] Current standard of care includes nucleoside analogues (e.g., lamivudine) and pegylated interferon therapy. Nucleoside analogues work by inhibiting HBV polymerase activity, leading to reduced viral replication. However, prolonged treatment duration, increased viral resistance, and the emergence of mutant strains have reduced the effectiveness of nucleoside therapy (Papatheodoridis GV et al., Am. J. Gastroenterol 97(7):1618-28 (2002). Pegylated interferon therapy has been tested alone or in combination with nucleoside analogs (e.g., lamivudine) to inhibit viral DNA transcription. Pegylated interferon therapy has been shown to mediate different effects on the innate and adaptive parts of the immune system, with a strong depletion effect on CD8 T cells, limiting the efficacy of the therapy (Micco L et al., Journal of Hepatology 58(2): 225-233 (2013); Stelma F et al., Journal of Infectious Disease 212(7):1042-51 (2015); marcellin P et al., New England Journal of Medicine 351(12):1206-17). (2004)). Nucleotide analogues and pegylation therapy have failed to eliminate or inhibit the production of HBV surface antigen (HBsAg), which has been associated with poor prognosis in HBV infection. Other therapies, including antisense oligonucleotides (ASO) and siRNA approaches centered on reducing HBsAg to achieve functional cure (Billioud G. et al., Journal of Hepatology 64(4):781-9 (2015); Gane E. et al., Hepatology 74(4):1795-1808 (2021); Flisiak R. et al., Expert Opinion on Biology Therapy 18(6):609-617), have shown promise in inhibiting HBsAg, HBeAg, and HBV DNA synthesis. However, the functional benefit of any of these therapies for liver tissue regeneration remains unclear.
[0157] Current antiviral therapies rarely achieve a cure because they inhibit cytoplasmic HBV genome replication and do not directly target the cccDNA form—a form that acts as an intermediate in HBV replication and a reservoir for viral persistence (Yang G. et al., Theranostics 9(24):7345-58 (2019)). Genome engineering approaches, such as nucleases or base editors, target the removal or mutagenesis of the cccDNA pool to functionally cure the infection. However, such nuclease-based therapies have the potential to produce chromosomal abnormalities and are therefore not preferred, highlighting the need for better HBV treatments.
[0158] The persistent presence of free cccDNA pools in infected hepatocytes remains a key obstacle to complete eradication by anti-HBV therapy. cccDNA accumulates in the cell nucleus as chromatin-like miniature chromosomes assembled from histones and non-histones. Due to its unnatural state, cccDNA exhibits unusual chromatin regulation. For example, changes in the epigenetic state of cccDNA have been found to determine its transcriptional activity (Yang G. et al., Theranostics 9(24):7345-58 (2019)). For instance, the host nucleosome assembly machinery (HAT1 / CAF-1) acetylates histone H4 at H4K5 and H4K12 sites, facilitating cccDNA assembly. Acetylation markers on histones of cccDNA, in turn, promote HBV replication and cccDNA accumulation. This transcriptional activity is largely driven by the presence or absence of activating epigenetic markers on cccDNA; repressive histone markers (e.g., H3K27me3 and H3K9me3) are scarce, indicating limited repression in cccDNA (Tropberger P. et al., PNAS, 112(42):E5715-E5724 (2015); Riviere L. et al., JHepatol 15(00450):S0168-8278 (2015)).
[0159] The expected clinical outcomes have been linked to key epigenetic features within the cccDNA mini-chromosome. Studies have found that cccDNA contains readily methylated CpG islands associated with HBV behavior (Zhang Y. et al., PlosOne 9(10):e110442 (2014); Vivekanandan P et al., Journal of infectious diseases, 199(9):1286-1291 (2009); Vivekanandan P et al., Journal of Virology, 84(9):4321-4329 (2010); Vivekanandan P et al., Journal of Viral hepatitis 15(2):103-107 (2008); Jain S. et al., Scientific Reports 5: 10478 (2015)). Methylation of CpG islands II and III has been associated with low levels of serum HBV DNA and HBsAg titers in patients (Zhang Y. et al., PlosOne 9(10):e110442(2014)). HBV genotype, HBeAg positivity, patient age, and stage of liver fibrosis have been found to be associated with cccDNA CpG methylation status. In vitro methylation studies have further confirmed that CpG island II methylation can significantly reduce cccDNA transcription and subsequent viral core DNA replication (Zhang Y. et al., PlosOne 9(10):e110442(2014)), establishing the importance of chromatin for cccDNA regulation and as a potential target for the treatment of chronic HBV infection. Antiviral agents and a wide range of epigenetic modifiers, such as IFNα, have been attributed to reducing post-translational modifications of active histones, thereby downregulating the transcription of cccDNA at the transcriptional level (Tropberger P. et al., PNAS, 112(42):E5715-E5724 (2015); Belloni L et al., Journal of Clinical Investigation 122L529-537 (2012); Allweiss L et al., Journal of Hepatology 60:500-507 (2014); Lucifora J et al., Science 343: 1221-1228 (2014)).
[0160] The provided implementations are based on the understanding that epigenetic silencing of one or more HBV viral genes (including those present on cccDNA) may be a feasible therapeutic approach for curing HBV infection. This document discloses methods for improving HBV infection and, in some cases, potentially achieving functional cure, through precise epigenetic silencing of HBV in cccDNA form, relaxed circular DNA (rcDNA) form, and integrated into human genomic DNA. The methods described herein demonstrate high efficacy, safety, and stability. In some embodiments, the methods target all forms of HBV in the same manner, utilizing a non-mutagenic platform and targeting the transcriptional source rather than downstream transcripts. The persistence of the epigenetic editing methods offers promise for treating HBV infection because methylation can be inherited by cell progeny. In some embodiments, the methods described herein target multiple locations on the viral genome to ensure a deep and durable response across HBV variants. In some embodiments, the epigenetic approach results in the silencing of HBV replication, HBV transcription, and protein production from HBV DNA. The provided implementations do not rely on immune reactivation and clearance of infected hepatocytes but are based on direct epigenetic silencing (e.g., HBV repression).
[0161] The embodiments provided herein include an epigenetically modified DNA targeting system comprising at least one DNA targeting module for repressing the transcription of one or more hepatitis B virus (HBV) genes and / or their regulatory elements, wherein each of the at least one DNA targeting module comprises a fusion protein comprising: (a) a DNA-binding domain for targeting a target site in a hepatitis B virus DNA sequence; and (b) at least one transcriptional repressor effector domain. In some embodiments, the provided epigenetically modified DNA targeting system is used for multi-target repression of multiple different genes or their regulatory elements that regulate hepatitis B virus (HBV) replication and / or HBV transcription. In some embodiments, the epigenetically modified DNA targeting system comprises multiple DNA targeting modules for repressing the transcription of multiple genes or their regulatory elements that regulate hepatitis B virus (HBV) replication and / or HBV transcription. In some embodiments, the DNA targeting module comprises (a) a fusion protein comprising a clustered regularly spaced short palindromic repeat (Cas) protein or a variant thereof and at least one transcriptional repressor effector domain; and (b) multiple guide RNAs (gRNAs), the multiple guide RNAs comprising at least a first gRNA and a second gRNA. In some embodiments, the first gRNA targets a target site of a first gene or a regulatory element thereof, and the second gRNA targets a target site of a second gene or a regulatory element thereof. The first gene and the second gene or a regulatory element thereof regulate hepatitis B virus replication and / or HBV transcription. This document also provides polynucleotides, vectors, and compositions containing the thereof encoding the DNA targeting system or the fusion protein of the DNA targeting system.
[0162] In some embodiments of the provided epigenetic modification DNA targeting system, the DNA-binding domain is a nuclease-free clustered regularly spaced short palindromic repeat (Cas) protein or a variant thereof, such as inactive Cas (dCas, e.g., dCas9), and the DNA targeting system further comprises at least one gRNA capable of complexing with the Cas. In some embodiments, the DNA-binding domain is a nuclease-free clustered regularly spaced short palindromic repeat (Cas) protein or a variant thereof that complexes with a guide RNA (gRNA). In such systems, the gRNA has a spacer sequence capable of hybridizing to a target site of the gene or its regulatory element. This document also provides related gRNAs (including Cas / gRNA combinations), polynucleotides, compositions, and methods relating to or associated with the epigenetic modification DNA targeting system.
[0163] In some embodiments of the provided epigenetically modified DNA targeting system, the DNA-binding domain is a protein domain engineered for specific binding to the target site sequence. For example, in some embodiments, the DNA-binding domain is a zinc finger (ZFN) based DNA-binding domain or a transcription activator-like effector DNA-binding domain as described herein.
[0164] This article also provides methods for regulating transcription or phenotype of liver cells using the described epigenetic DNA-targeting system. This article also provides methods for inhibiting HBV replication and / or protein levels using the described epigenetic DNA-targeting system. In some embodiments, the methods can be used as a therapy for treating HBV infection, such as hepatitis.
[0165] In some embodiments, the target site exists as covalently closed circular DNA (cccDNA), relaxed circular DNA (rcDNA), and / or integrated into genomic DNA. In some embodiments, the target site is located at or near genes involved in HBV replication and / or HBV transcription, or their regulatory elements (such as regulatory elements or coding regions). This document also provides epigenetically modified DNA targeting systems multiplexed with multiple DNA targeting modules, enabling the system to target combinations of such genes or their regulatory elements. In some embodiments, each module of the DNA targeting system represses the transcription of a different gene. This document also provides methods for reducing HBV replication and / or transcription using the epigenetically modified DNA targeting system. In some embodiments, the method can be used to treat liver diseases (e.g., hepatitis), cancers (e.g., hepatocellular carcinoma), or HBV infection (acute or chronic hepatitis).
[0166] Therefore, in some embodiments, the DNA targeting system comprises synthetic transcription factors capable of regulating (e.g., reducing or repressing) gene transcription in a targeted manner. In the provided embodiments, the provided epigenetically modified DNA targeting system reduces the transcription of said genes and / or their regulatory elements, or multiple said genes and / or their regulatory elements, thereby promoting HBV replication and / or silencing transcription. The provided embodiments can be used to target multiple genetic mechanisms to treat HBV in infected patients while avoiding viral resistance, the costs associated with prolonged treatment, and the poor efficacy of current combination therapies. This method provides a substantial clinical solution for the treatment of HBV infection by reducing viral replication and transcription from cccDNA and integrated HBV DNA and avoiding problems associated with current therapies.
[0167] All publications (including patent documents, scientific articles, and databases) mentioned in this application are incorporated herein by reference in their entirety for all purposes, as if each individual publication were incorporated separately by reference. Where the definitions described herein are contrary to or otherwise inconsistent with those described in patents, applications, published applications, and other publications incorporated herein by reference, the definitions described herein shall prevail over those incorporated herein by reference.
[0168] The chapter titles used in this article are for organizational purposes only and should not be construed as limiting the topics described. I. DNA Targeting System
[0169] In some embodiments, a DNA targeting system is provided that is capable of specifically targeting a target site in at least one gene (also referred to herein as a target gene) or its DNA regulatory element (e.g., a regulatory element) and reducing transcription of said at least one gene. In the provided embodiments, for each targeted target gene or its regulatory element, the DNA targeting system comprises a DNA-binding domain that binds to the target site in the gene or its regulatory element. In some embodiments, the DNA targeting system further comprises at least one effector domain capable of epigenetically modifying one or more DNA bases of the gene or its regulatory element, wherein the epigenetic modification results in a reduction in gene transcription (e.g., repression or reduction of gene transcription compared to the absence of a DNA targeting system). Therefore, the terms DNA targeting system and epigenetically modified DNA targeting system are used interchangeably herein. In some embodiments, the DNA targeting system comprises a fusion protein comprising: (a) at least one DNA-binding domain capable of targeting a target site; and (b) at least one effector domain capable of reducing gene transcription. For example, said at least one effector domain is a transcriptional repressor domain.
[0170] In some embodiments, the DNA targeting system contains at least one DNA targeting module, wherein each DNA targeting module of the system is a component of the DNA targeting system capable of independently targeting a target site, such as a target gene or its regulatory element. In some embodiments, each DNA targeting module includes (a) a DNA-binding domain capable of targeting a target site of a target gene or its regulatory element that regulates HBV replication and / or HBV transcription; and (b) an effector domain capable of reducing gene transcription.
[0171] In some embodiments, the DNA targeting system includes a single DNA targeting module for targeting and repressing a single gene. In some embodiments, the DNA targeting module includes (a) a DNA-binding domain capable of targeting a target site of a target gene or its regulatory element that regulates HBV replication and / or HBV transcription; and (b) an effector domain capable of reducing gene transcription.
[0172] In some embodiments, the DNA targeting system comprises a single DNA targeting module for targeting and repressing more than one gene or its regulatory element. Thus, in some embodiments, the single DNA targeting module provides a multiple epigenetic modification DNA targeting system that targets, regulates (e.g., represses) more than one gene or its regulatory element. In some embodiments, the DNA targeting module comprises (a) a DNA-binding domain capable of targeting target sites of more than one target gene or its regulatory element that regulates HBV replication and / or HBV transcription; and (b) an effector domain capable of reducing gene transcription. In some embodiments, the DNA targeting system comprises a single DNA targeting module for targeting and repressing at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, or 30 genes or their regulatory elements. In a particular embodiment, the DNA targeting module is cross-reactive to each of the target sites of more than one gene. In some embodiments, a single DNA targeting module provides a multiple epigenetic modification DNA targeting system that represses transcription of more than one gene or its regulatory element. In some embodiments, the DNA targeting module represses transcription of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, or 30 genes or their regulatory elements.
[0173] In some embodiments, the DNA targeting system is a multiplex DNA targeting system comprising multiple DNA targeting modules, each targeting a different target site of one or more genes or their regulatory elements. In some embodiments, the different target sites are located in the same region of the gene or its regulatory element. In some embodiments, the different target sites are present in a regulatory element such as a promoter. In some embodiments, the target sites overlap, such that any two or more DNA targeting modules bind to the overlapping target sites.
[0174] In some embodiments, the DNA targeting system comprises multiple DNA targeting modules, each of which is used to target and repress a different gene. In some embodiments, the DNA targeting system is a multiplex DNA targeting system, i.e., it targets target sites in more than one gene or its regulatory element. The term DNA targeting system can include a multiplex epigenetic modification DNA targeting system comprising more than one DNA targeting module. In some embodiments, each DNA targeting module within the multiplex epigenetic modification DNA targeting system targets target sites in a different gene or its regulatory element to repress a different gene, in contrast to other DNA targeting modules in the system. In some embodiments, each DNA targeting module within the multiplex epigenetic modification DNA targeting system targets target sites in more than one gene or its regulatory element to repress the transcription of said more than one gene. In some embodiments, each DNA targeting module represses the transcription of at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, or at least 30 genes.
[0175] A multiplexed epigenetic modification DNA targeting system comprises at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30 DNA targeting modules, or any value between the foregoing. In some embodiments, the multiplexed epigenetic modification DNA targeting system represses the transcription of at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, or at least 30 genes.
[0176] In some embodiments, any two DNA targeting modules of the DNA targeting system contain separate (i.e., non-overlapping) components. In some embodiments, each DNA targeting module of the DNA targeting system contains separate (i.e., non-overlapping) components. For example, the DNA targeting system may include: a first DNA targeting module containing a first fusion protein containing a DNA-binding domain (e.g., a ZFN- or TALE-based DNA-binding domain) targeting a first target site; and a second DNA targeting module containing a second fusion protein containing a second DNA-binding domain (e.g., a ZFN- or TALE-based DNA-binding domain) targeting a second target site.
[0177] In some embodiments, any two DNA targeting modules of the DNA targeting system may contain shared (i.e., overlapping) components. In some embodiments, each DNA targeting module of the DNA targeting system contains shared (i.e., overlapping) components. For example, the DNA targeting system may include: a first DNA targeting module comprising (a) a fusion protein containing a Cas protein and a transcriptional repressor domain, and (b) a first gRNA that complexes with the Cas protein and targets a first target site of a first HBV gene or its regulatory element; and a second DNA targeting module comprising (a) the fusion protein of the first DNA targeting module, and (b) a second gRNA that complexes with the Cas protein and targets a second target site of a second HBV gene or its regulatory element. It should be understood that providing two or more different gRNAs for a given Cas protein enables different molecules of the same Cas protein to target the target sites of two or more gRNAs. Conversely, as described herein, different Cas protein variants (e.g., SpCas9 and SaCas9) are compatible with different gRNA scaffold sequences and PAMs. Therefore, a single DNA targeting system comprising multiple CRISPR / Cas-based non-overlapping DNA targeting modules can be engineered.
[0178] In some aspects, this document provides an epigenetic modification DNA targeting system comprising multiple DNA targeting modules for repressing the transcription of multiple genes regulating HBV replication and / or HBV transcription. In some embodiments, the multiple DNA targeting modules comprise a first DNA targeting module for repressing the transcription of a first gene or its regulatory element among the multiple genes, and a second DNA targeting module for repressing the transcription of a second gene among the multiple genes. In some embodiments, each DNA targeting module comprises a fusion protein comprising: (a) a DNA-binding domain for targeting a target site of one of the multiple genes, and (b) at least one transcriptional repressor domain. In some embodiments, the target site is located at or near an HBV gene or its regulatory element. In some embodiments, the HBV gene or its regulatory element is involved in controlling HBV replication and / or HBV transcription. Regulatory elements may be promoter regions (e.g., pre-S1 promoter, pre-S2 promoter, X promoter, or basic core promoter), enhancer regions (e.g., Enh1 or Enh2 enhancer regions), or any other transcript processing control region (e.g., regions involved in 5' capping, splicing, and / or 3' polyadenylation).
[0179] In some respects, this document provides an epigenetically modified DNA targeting system comprising at least one DNA targeting module for repressing transcription of one or more hepatitis B virus (HBV) genes; wherein each of the at least one DNA targeting module comprises a fusion protein comprising: (a) a DNA-binding domain for targeting a target site in a hepatitis B virus DNA sequence (such as the HBV gene or its regulatory elements); and (b) at least one transcriptional repressor effector domain.
[0180] In some aspects, this document provides an epigenetically modified DNA targeting system comprising at least one DNA targeting module for repressing transcription of one or more hepatitis B virus (HBV) genes, wherein each of the at least one DNA targeting module comprises: (a) a fusion protein comprising a clustered regularly spaced short palindromic repeat-associated (Cas) protein or a variant thereof and at least one transcriptional repressor effector domain; and (b) multiple guide RNAs (gRNAs) targeting multiple target sites of multiple genes or regulatory elements thereof, wherein the multiple genes or regulatory elements thereof regulate hepatitis B virus replication and / or HBV transcription. In aspects of the provided embodiments, the multiple target sites are 2, 3, 4, 5, or 6 different target sites. In aspects of the provided embodiments, the multiple target sites are each located in a different HBV gene or regulatory element thereof.
[0181] In aspects of the provided embodiments, the DNA targeting system described herein targets genes or regulatory elements thereof to reduce the transcription of one or more HBV genes in cells infected with hepatitis B virus (HBV), wherein the reduced transcription regulates one or more activities or functions of the HBV-infected cells, such as the expression of HBV RNA and / or HBV proteins. In some embodiments, the reduction in gene transcription leads to a reduction in gene expression in the infected cells, i.e., reduced gene expression. In some embodiments, the reduction in gene transcription, such as reduced gene expression, leads to a reduction in protein expression in the infected cells, i.e., reduced protein expression.
[0182] In some embodiments, the cell is a hepatocyte, such as hepatocytes, hepatic stellate cells (HSCs), Kupffer cells, and hepatic sinusoidal endothelial cells. For example, this document provides a DNA targeting system that targets a gene or its regulatory element to reduce transcription of the HBV gene in a target cell, wherein the reduced transcription regulates one or more activities or functions of HBV, such as HBV transcription and protein expression. In some embodiments, the reduction in gene transcription leads to a reduction in gene expression in the target cell, i.e., reduced gene expression. In some embodiments, the cell is a hepatocyte.
[0183] In some respects, the cells are derived from human subjects. In other respects, the cells are cells within the subject (i.e., cells in the body).
[0184] In some embodiments, the DNA-binding domain comprises or is derived from a CRISPR-associated (Cas) protein, a zinc finger protein (ZFP), a transcription activator-like effector (TALE), a megnuclease, a homing endonuclease, an I-SceI enzyme, or a variant thereof. In some embodiments, the DNA-binding domain comprises a catalytically inactivated (e.g., nuclease-inactive or nuclease-inactivated) variant of any of the foregoing. In some embodiments, the DNA-binding domain comprises an inactivated Cas9 (dCas9) protein or a variant thereof, which is catalytically inactivated and therefore lacks nuclease activity and cannot cleave DNA.
[0185] In some embodiments, the DNA-binding domain comprises or is derived from a Cas protein or a variant thereof, such as a nuclease-free Cas or dCas (e.g., dCas9), and the DNA-targeting system comprises one or more guide RNAs (gRNAs), such as a combination of two or three gRNAs. In some embodiments, the gRNA comprises a spacer sequence capable of targeting and / or hybridizing to a target site. In some embodiments, the gRNA is capable of complexing with a Cas protein or a variant thereof. In some aspects, the gRNA guides or recruits a Cas protein or a variant thereof to a target site. In some embodiments, the effector domain comprises a transcriptional repressor domain and / or is capable of reducing gene transcription. In some embodiments, the effector domain directly or indirectly causes a reduction in gene transcription. In some embodiments, the effector domain induces, catalyzes, or causes transcriptional repression. In some embodiments, the effector domain induces transcriptional repression. In some aspects, the effector domain is selected from the KRAB domain, ERF repressor domain, MXI1 domain, SID4X domain, MAD-SID domain, DNMT family protein domains (e.g., DNMT3A or DNMT3B), fusions of one or more DNMT family proteins or their domains (e.g., DNMT3A / L, which comprises a fusion of the DNMT3A and DNMT3L domains), LSD1, SunTag domain, EZH2 domain, partial or complete functional fragments or domains of any of the foregoing, or combinations of any of the foregoing. In some embodiments, the effector domain is KRAB. In some embodiments, the effector domain is DNMT3A / L.
[0186] In some embodiments, the fusion protein of the DNA targeting system includes a dCas9-KRAB fusion protein. In some embodiments, the fusion protein of the DNA targeting system includes a DNMT3A / L-dCas9-KRAB fusion protein. In some embodiments, the fusion protein of the DNA targeting system includes a KRAB-dCas9-DNMT3A / L fusion protein.
[0187] The following subsections provide exemplary components and characteristics of DNA targeting systems. A. Target site and target location
[0188] In any embodiment herein, the target site is a gene and / or its regulatory element in the hepatitis B virus (HBV) genome. In some embodiments, the target site is present in covalently closed circular DNA (cccDNA), relaxed circular DNA (rcDNA), and / or integrated into human genomic DNA. In some embodiments, the target site is located within a hepatitis B virus DNA sequence. In some embodiments, the hepatitis B virus DNA sequence is an HBV gene or its regulatory element. In some embodiments, the target site is located at or near a gene involved in HBV replication and / or HBV transcription. In some embodiments, the epigenetic modification DNA targeting system comprises at least one DNA targeting module for repressing the transcription of one or more HBV genes by targeting the target site. In some aspects, repressing the transcription of HBV genes, such as reduced gene expression, leads to silencing of HBV replication (e.g., reduced HBV replication) and / or silencing of HBV transcription.
[0189] Referring to the provided disclosure, it should be understood that HBV-positive (+) cells (e.g., HBV-infected cells) mean that the cells express any of the HBV markers described herein (e.g., HBV RNA transcripts and / or proteins). Similarly, it should be understood that cells negative (-) for a specific marker are cells that do not express the marker, or whose expression level is undetectable. Antibodies and other binding entities can be used to detect the expression level of the marker protein to identify or detect a given cell surface marker. Suitable antibodies may include polyclonal, monoclonal, fragment (e.g., Fab fragments), single-chain antibodies, and other forms of specific binding molecules. Antibody reagents for the aforementioned cell surface markers are readily available to those skilled in the art. Many well-known methods for assessing the expression level of surface markers or proteins can be used, such as by affinity-based methods, such as immunoaffinity-based methods (e.g., in the case of surface markers, such as detection by flow cytometry). In some embodiments, the marker is a fluorophore, and flow cytometry is used to detect or identify cell surface markers on cells (e.g., hepatocytes). In some implementations, a different label is used for each of the different labels by multicolor flow cytometry. In some implementations, surface expression can be determined by flow cytometry, for example, by staining with an antibody that specifically binds to the label and detecting the binding of the antibody to the label.
[0190] In some embodiments, a cell (e.g., hepatocytes) is positive (pos or +) for a specific marker if it is detectably present on or within the cell (which may be an intracellular marker or a surface marker, such as HBeAg or HBsAg). In some embodiments, surface expression is positive if staining by flow cytometry is detected at a level significantly higher than that detected by the same procedure with a type-matched control under otherwise identical conditions, and / or at a level substantially similar to or, in some cases, higher than, that of cells known to be positive for the marker, and / or at a level higher than that of cells known to be negative for the marker. In some embodiments, cells (e.g., hepatocytes) contacted by a DNA-targeting system express less of a specific marker (e.g., HBeAg) if staining is significantly lower compared to similar cells not contacted by the DNA-targeting system described herein.
[0191] In some embodiments, a cell (e.g., a hepatocyte) is negative (neg or -) for a specific marker if the specific marker (which may be an intracellular or surface marker) is not detectably present on or within the cell. In some embodiments, surface expression is negative if staining is not detectable by flow cytometry at a level significantly higher than that detected by the same procedure with a type-matched control under otherwise identical conditions, and / or at a level significantly lower than that of cells known to be positive for the marker, and / or at a level substantially similar to that of cells known to be negative for the marker.
[0192] In some embodiments, the phenotype of the infected cells and / or individuals is functionally characterized. In some aspects, the phenotype can be characterized by the presence of HBV RNA transcripts in the infected cells. In some aspects, the phenotype can be characterized by the presence of any or a combination of HBV proteins in the infected cells. In some aspects, the phenotype can be characterized by the presence of antibodies against any of the markers described herein. In some aspects, antibodies include, but are not limited to, anti-HBc-IgM, total anti-HBc antibodies, and antibodies against HBeAg. In some embodiments, RNA transcripts, proteins, and / or antibodies are measured, detected, and / or quantified using any suitable technique known in the art. For example, real-time PCR technology can be used to measure, detect, and / or quantify RNA transcripts. Enzyme-linked immunosorbent assay (ELISA) can be used to measure, detect, and / or quantify HBV proteins (e.g., HBsAg, HBeAg, and / or HBcrAg).
[0193] Target genes and / or regulatory elements regulated by the provided DNA targeting system (including the multiplex epigenetic modification DNA targeting system described herein) include any genes and / or regulatory elements whose transcription and expression are reduced in cells (e.g., HBV-infected cells). Various methods can be used to characterize the transcriptional or expression levels of genes in cells (e.g., hepatocytes), such as after the cells have been exposed to or introduced with the provided DNA targeting system. In some embodiments, the transcriptional activity or expression of genes can be analyzed by RNA analysis. In some embodiments, RNA analysis includes RNA quantification. In some embodiments, RNA quantification is performed by reverse transcription quantitative PCR (RT-qPCR), multiplex qRT-PCR, fluorescence in situ hybridization (FISH), RNA sequencing (RNA-seq), or a combination thereof.
[0194] In some implementations, the gene or transcript is a gene or transcript whose gene expression or transcript presence is reduced in a cell (e.g., HBV-infected cells, such as hepatocytes) upon contact with or introduction of a provided DNA targeting system (such as a multiplex epigenetic DNA targeting system). In some aspects, multiple genes or transcripts are targeted by the multiplex epigenetic DNA targeting system, such as through one or more of its DNA targeting modules. In such a system, each gene or transcript of the multiplex DNA targeting system is a gene or transcript whose gene expression or transcript presence is reduced in a cell upon contact with or introduction of the provided multiplex epigenetic DNA targeting system. In some implementations, the reduction in gene expression or change in transcript levels in cells (e.g., HBV-infected cells, such as hepatocytes) compared to gene levels in control cells is approximately a log2 change of the following folds: at least 1.25, 1.5, 1.75, 2.0, 2.5, 2.75, 3.0, 3.25, 3.5, 3.75, 4.0, 4.25, 4.5, 4.75, 5.0, 5.25, 5.5, 5.75, 6.25, 6.50, 6.75, 7.0, 7.25, 7.50, 7.75, 8.0, 8.25, 8.5, 8.75, 9.0, or any value between any of the foregoing values. In some embodiments, the reduction in gene expression or change in transcript levels in cells (e.g., HBV-infected cells, such as hepatocytes) compared to gene levels in control cells is about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, about 100%, or any value between any of the foregoing values. In some embodiments, the reduction in gene expression or change in transcript levels in cells compared to gene levels in control cells is greater than 90%. In some embodiments, the guide RNA is shown in SEQ ID NO: 565, 528, 542, 508, 515, 575, 515, 453, 506, 514, 425, or 472. In some embodiments, the reduction in gene expression or change in transcript levels in cells compared to gene levels in control cells is greater than 75%.In some implementations, the guide RNA is shown in SEQ ID NO: 565 (HBVg_175), 528 (HBVg_138), 582 (HBVg_192), 542 (HBVg_152), 508 (HBVg_118), 515 (HBVg_125), 575 (HBVg_185), 453 (HBVg_63), 506 (HBVg_116), 514 (HBVg_124), 395 (HBVg_5), 472 (HBVg_8). 2), 451 (HBVg_61), 488 (HBVg_98), 540 (HBVg_150), 533 (HBVg_143), 572 (HBVg_182), 566 (HBVg_1 76), 489 (HBVg_99), 469 (HBVg_79), 408 (HBVg_18), 465 (HBVg_75), 402 (HBVg_12), 474 (HBVg_84) , 525 (HBVg_135), 416 (HBVg_26), 396 (HBVg_6), 554 (HBVg_164), 419 (HBVg_29), 545 (HBVg_155), 446 (HBVg_56), 580 (HBVg_190), 555 (HBVg_165), 412 (HBVg_22), 428 (HBVg_38), 458 (HBVg_68), 5 Among 48 (HBVg_158), 511 (HBVg_121), 432 (HBVg_42), 441 (HBVg_51), 433 (HBVg_43), 579 (HBVg_189), 479 (HBVg_89), 478 (HBVg_88), 520 (HBVg_130), 462 (HBVg_72), 523 (HBVg_133), and 503 (HBVg_113).
[0195] In the provided implementation, cccDNA is transcribed into five HBV RNAs (0.7 kb, 2.1 kb, 2.4 kb, longer, and shorter 3.5 kb RNAs) by host RNA polymerase. Transcription of cccDNA is controlled by four promoters (basic core promoter, pre-S1 promoter, pre-S2 promoter, and X promoter) and two enhancers (enhancer I and enhancer II). Figure 10.7-kb RNA is translated into HBV X protein (HBx), which functions as a transcriptional regulator. 2.1-kb RNA is translated into HBV small surface protein (S) and medium surface protein (M). 2.4-kb RNA is translated into HBV large surface protein (L). L, M, and S can self-assemble to form empty subviral particles (SVPs) (including globular SVPs and filamentous SVPs), of which only filamentous SVPs and viral particles contain a large amount of L protein. Globular SVPs are secreted via a constitutive secretion pathway. Filamentous SVPs are secreted via multivesicular bodies (MVBs) by the endosomal sorting complex (ESCRT) mechanism required for transport. The longer 3.5-kb RNA is called precore RNA (pre-C RNA) and can be translated into precore protein, more widely known as HBV e antigen (HBeAg). The shorter 3.5-kb RNA is pregenomic RNA (pgRNA), which has two functions: first, as a template for the translation of HBV polymerase (Pol) and the core protein; and second, as a template for intracapsular reverse transcription (formed by the core protein) to form HBV rcDNA via Pol. These nucleocapsids are then coated with HBV surface proteins (L, M, and S) to form mature viral particles, which are secreted via the ESCRT / MVB pathway. Alternatively, these nucleocapsids can be transported into the nucleus to form cccDNA. In some implementations, transcription and / or translation of HBV genes are repressed, such as by reduced gene expression, resulting in the silencing of any of the following HBV markers: HBV X protein (HBx), hepatitis B surface antigen (HBsAg) such as small surface protein (S), medium surface protein (M), or large surface protein (L), or HBV e antigen (HBeAg). In some embodiments, repression of HBV gene transcription and / or translation, such as reduced gene expression, leads to silencing of HB core-associated antigen (HBcrAg). HBcrAg comprises three pre-core / core protein products, including hepatitis B virus core antigen (HBcAg), HBeAg, and the 22-kDA pre-core protein (p22cr). In some aspects, cccDNA, total HBV DNA, serum HBcrAg, HBsAg, HBeAg, hepatitis B virus core antibody (anti-HBc), HBV DNA, and HBV RNA are quantified as readouts to measure reduced HBV transcription and / or translation. In some embodiments, the target site is located in a gene encoding any HBV protein. In some embodiments, the target site is located in a regulatory element (e.g., a promoter or enhancer) of a gene encoding any HBV protein.
[0196] In some embodiments, the target site for the epigenetic modification DNA targeting system is located in a gene involved in HBV replication and / or HBV transcription. In some aspects, the target site for the epigenetic modification DNA targeting system is located in or near a gene or regulatory element involved in controlling HBV replication and / or HBV transcription. In some embodiments, the gene involved in HBV replication and / or HBV transcription is a polymerase gene, an S-family gene, an X-gene, and / or a core family gene. In some embodiments, the gene involved in HBV replication and / or transcription encodes a polymerase, an envelope protein, a capsid protein, a transcription factor, or a transcriptional transactivator. In some embodiments, the regulatory element involved in HBV replication and / or HBV transcription is a promoter region, an enhancer region, and / or any transcript processing control region. In some embodiments, the promoter region is a pre-S1 promoter, a pre-S2 promoter, an X promoter, or a basic core promoter. In some embodiments, the enhancer region is an Enh1 enhancer and / or an Enh2 enhancer region. In some implementations, the transcript processing control region is the region that encodes signals for 5' capping, splicing, and / or 3' polyadenylation.
[0197] In some embodiments, the target site is a sequence within a target region having a sequence corresponding to the base pair (bp) position 1 bp-42 bp, 491 bp-1032 bp, 1750 bp-1799 bp, or 1951 bp-2952 bp of the reference hepatitis B virus genome (hepatitis B virus subtype ayw, complete genome, GenBank: U95551.1). In some embodiments, the target site is within a target region of the HBV genome, wherein the target region has the sequence shown in SEQ ID NO: 1056 or its complementary sequence. In some embodiments, the target site is within a target region of the HBV genome, wherein the target region has the sequence shown in SEQ ID NO: 1058 or its complementary sequence. In some embodiments, the target site is within a target region of the HBV genome, wherein the target region has the sequence shown in SEQ ID NO: 1060 or its complementary sequence. In some implementations, the target site is located within the target region of the HBV genome, wherein the target region has the sequence shown in SEQ ID NO: 1062 or its complementary sequence.
[0198] In some embodiments, the target site is a sequence within a target region having a sequence corresponding to the base pair (bp) position 43 bp-490 bp, 1033 bp-1749 bp, 1800 bp-1950 bp, or 2953 bp-3182 bp of the reference hepatitis B virus genome (hepatitis B virus subtype ayw, complete genome, GenBank: U95551.1). In some embodiments, the target site is within a target region of the HBV genome, wherein the target region has the sequence shown in SEQ ID NO: 1057 or its complementary sequence. In some embodiments, the target site is within a target region of the HBV genome, wherein the target region has the sequence shown in SEQ ID NO: 1059 or its complementary sequence. In some embodiments, the target site is within a target region of the HBV genome, wherein the target region has the sequence shown in SEQ ID NO: 1061 or its complementary sequence. In some implementations, the target site is located within a target region of the HBV genome, wherein the target region has the sequence shown in SEQ ID NO: 1063 or its complementary sequence.
[0199] In some embodiments, the target site is a sequence within a target region having a sequence corresponding to the sequence in the reference hepatitis B virus genome (hepatitis B virus subtype ayw, complete genome, GenBank: U95551.1) SEQ ID NO: 650 located at base pair (bp) positions between 67 bp-392 bp (CpG island 1), 1033 bp-1749 bp (CpG island 2), or 2215 bp-2490 bp (CpG island 3). In some embodiments, the target site is within a target region of the HBV genome, wherein the target region has the sequence shown in SEQ ID NO: 1064 or its complementary sequence. In some embodiments, the target site is within a target region of the HBV genome, wherein the target region has the sequence shown in SEQ ID NO: 1059 or its complementary sequence. In some embodiments, the target site is within a target region of the HBV genome, wherein the target region has the sequence shown in SEQ ID NO: 1066 or its complementary sequence.
[0200] In some embodiments, the target site is located in the polymerase gene or its regulatory elements. The polymerase gene (also known as the P gene) encodes a multifunctional enzyme (also known as P polymerase, HBVgp1, or DNA-guided DNA polymerase) that converts the viral RNA genome into dsDNA within the viral cytoplasmic capsid. The polymerase exhibits DNA polymerase activity capable of replicating DNA or RNA templates, as well as ribonuclease H (RNase H) activity that cleaves the RNA strand of the RNA-DNA heteroduplex in a partially advancing 3' to 5' endonuclease pattern. The polymerase gene ORF completely overlaps with the pre-S / S ORF and partially overlaps with the ORFs of the core family and the X gene. In some implementations, the target site is a sequence within a target region having a sequence corresponding to the sequence in the reference hepatitis B virus genome (hepatitis B virus subtype ayw, complete genome, GenBank: U95551.1) SEQ ID NO: 650 located between 1 bp-1621 bp, 1374 bp-1838 bp, or 2307 bp-3182 bp in the HBV genome.
[0201] In some embodiments, the target site is in the S family genes or their regulatory elements. In some embodiments, the target site is in the S gene, the pre-S1 promoter, and / or the pre-S2 promoter region. The S family genes encode three different structurally related envelope proteins synthesized from different start codons, referred to as large (L), medium (M), and small (S) hepatitis B (HB) virus proteins (also referred to as L-HBs, M-HBs, and s-HBs, respectively). These three proteins have the same C-terminus but different N-terminal extensions. In some embodiments, the target site is a sequence within a target region having a sequence corresponding to the sequence in the reference hepatitis B virus genome (hepatitis B virus subtype ayw, complete genome, GenBank: U95551.1) SEQ ID NO: 650 located between 1 bp-837 bp, 1 bp-155 bp, or 2854 bp-3182 bp in the HBV genome.
[0202] In some embodiments, the target site is in the X gene or its regulatory elements. The X gene (also known as HBx, HBVgp3, peptide X, pX) is a gene encoding a multifunctional protein that regulates transcriptional regulation, protein degradation pathways, apoptosis, signal transduction, cell cycle progression, and genetic stability through direct or indirect interaction with host factors. The X gene protein regulates protein degradation pathways, apoptosis, transcription, signal transduction, cell cycle progression, and genetic stability through direct or indirect interaction with host factors. In some embodiments, the target site is a sequence within a target region having a sequence corresponding to the sequence between 1374 bp and 1838 bp of the HBV genome in SEQ ID NO: 650 of the reference hepatitis B virus genome (hepatitis B virus subtype ayw, complete genome, GenBank: U95551.1). In some embodiments, the start codon encoding the HBx protein (HBx start codon) is located at the 1376th residue pair of the HBV genome at the position corresponding to the HBV genome shown in SEQ ID NO: 650. In some embodiments, the start codon encoding the HBx protein (HBx start codon) is located at the 1374th residue pair of the HBV genome, corresponding to the position shown in reference SEQ ID NO: 1071. This study found that targeting the target site in the upstream region of the X gene start codon using the provided epigenetic modification DNA targeting system exhibits high activity against viral replication and transcription in cells that inhibit HBV infection. In some embodiments, the target region is located within a CpG island of the HBV genome. In some embodiments, the target site is a sequence within the target region having a sequence corresponding to the sequence in the HBV genome shown in reference SEQ ID NO: 650, located between 1033 and 1749 bp. In some embodiments, the target site is in the HBx promoter / enhancer #1 region, such as within a target region having a sequence corresponding to the sequence in the HBV genome shown in reference SEQ ID NO: 650, located between 1100 and 1350 bp. In some implementations, the target site is located in the basic core promoter region, such as in a target region having a sequence that corresponds to the HBV genome sequence between 1600-1750 bp shown in reference SEQ ID NO: 650.
[0203] In some embodiments, the target site is within a target region that spans within 300 base pairs (bp), 250 bp, 200 bp, 150 bp, 140 bp, 130 bp, 120 bp, 110 bp, or 100 bp upstream of the HBx start codon. In some embodiments, the target site is within a target region having a sequence corresponding to the HBV genome sequence shown in SEQ ID NO: 650, located between 1250 and 1374 bp. In some embodiments, the target site is within a target region of the HBV genome, wherein the target region has the sequence shown in SEQ ID NO: 1068 or its complementary sequence. In some embodiments, the target site is a sequence within a target region having a sequence corresponding to the HBV genome sequence shown in SEQ ID NO: 650, located between 1255 and 1302 bp. In some embodiments, the target site is within a target region of the HBV genome, wherein the target region has the sequence shown in SEQ ID NO: 1069 or its complementary sequence. In some embodiments, the target site is a sequence within a target region having a sequence corresponding to the HBV genome sequence shown in reference SEQ ID NO: 650, located between 1260 and 1300 bp. In some embodiments, the target site is within a target region of the HBV genome, wherein the target region has the sequence shown in SEQ ID NO: 1070 or its complementary sequence. In some embodiments, the target site, or each of the target sites, is within a target region located at a base pair between 1060 and 1480 bp of the HBV genome corresponding to the position shown in reference SEQ ID NO: 650. In some embodiments, the target site is within a target region of the HBV genome, wherein the target region has the sequence shown in SEQ ID NO: 1067. This document provides exemplary DNA binding systems for targeting target sites in such regions, including systems having various DNA binding domains, including CRISPR / Cas systems and ZFP.
[0204] In some implementations, the target site is in a core family gene or its regulatory element. In some implementations, the regulatory element is the Enh2 promoter. In some implementations, the regulatory element is the basic core promoter (BCP). The core promoter (CP) region of the viral genome plays a crucial role in viral replication and morphogenesis (Quarleri J, World Journal of Gastroenterology 20(2): 425-435 (2014)). The core promoter region guides the initiation of transcription to synthesize both precore mRNA and pregenomic RNA (pgRNA). The CP region consists of the basic core promoter (BCP), which initiates the transcription of precore mRNA (also known as pre-C, C gene, HBVgp4) and pgRNA, and consists of an upstream regulatory region (URR) containing positive and negative regulatory elements that regulate promoter activity. Several transcription factors bind to regulatory sequence elements of the nucleocapsid (CP), such as C / EBP, HNF1, HNF3 / 4, and COUP-TF1, to differentially regulate the synthesis of pre-C mRNA and pgRNA. The presence of AT-rich regions or TATA-like cassettes within the CP is also attributed to pgRNA transcription. The pre-core mRNA encodes the external core antigen (also known as the capsid protein, precapsid protein, HBeAg, pre-core protein, p25), which self-assembles to form an icosahedral capsid to package the viral genome. pgRNA is translated into a polymerase, the nucleocapsid protein HBcAg, and soluble secreted HBeAg protein. Additionally, pgRNA is integrated into the progeny nucleocapsid and reverse transcribed into DNA by a co-assembled viral polymerase, forming new HBV viral particles. These mature nucleocapsids containing relaxed circular DNA (rcDNA) can re-deliver their genome to the nucleus of the same cell to build pools of 10-100 copies of cccDNA molecules, or they can interact with envelope proteins at the endoplasmic reticulum / Golgi apparatus and be secreted as novel infectious viral particles (Pollicino T., et al., Journal of hepatology, 61(2):P408-417 (2014)). In some embodiments, the target site is in a core family gene or its regulatory element. In some aspects, targeting one or more sites within a core family gene or its regulatory element includes repressing pgRNA transcripts. In some aspects, repression of pgRNA transcripts includes silencing HBV replication.In some embodiments, the target site is a sequence within a target region having a sequence corresponding to the sequence in the reference hepatitis B virus genome (hepatitis B virus subtype ayw, complete genome, GenBank: U95551.1) SEQ ID NO: 650 located between 1590 bp-1815 bp, 1636 bp-1744, 1751 bp-1769, 1814 bp-1900 bp, 1816 bp-2455, and 1800 bp-1950 bp. In some embodiments, transcriptional reduction includes a reduction in the total level of hepatitis B virus RNA transcripts. In some embodiments, transcriptional reduction includes a reduction in the level of hepatitis B virus precore (“pre-C”) and / or pregenomic (“pgRNA”) RNA.
[0205] In some aspects, the target site is a coding region. In some embodiments, genes involved in HBV replication and / or HBV transcription encode HBV X protein (HBx), S family proteins (HBsAg) (such as small surface proteins (S-HBs), medium surface proteins (M-HBs), or HBV large surface proteins (L-HBs)), precore protein (HBeAg), HBV core-associated antigen (HBcrAg), polymerase, core protein, and precore protein. In some aspects, the target site is a sequence within a target region having a sequence corresponding to the sequence in the reference hepatitis B virus genome (hepatitis B virus subtype ayw, complete genome, GenBank: U95551.1) SEQ ID NO: 650 located between 1 bp-42 bp, 43 bp-1090 bp, 1091 bp-1849 bp, or 1850 bp-2455 bp, or 2455 bp-3182 bp in the HBV genome.
[0206] In some embodiments, transcriptional repression includes a reduction in the levels of hepatitis B surface antigen (HBsAg) and / or hepatitis B core-associated antigen (HBcrAg) proteins. In some embodiments, transcriptional repression includes a reduction of HBsAg transcript and / or protein levels by at least 90%. In some embodiments, transcriptional repression includes a reduction of HBcrAg transcript and / or protein levels from cccDNA by at least 50%.
[0207] In some implementations, transcriptional repression includes a reduction in the levels of the hepatitis B virus precore (“pre-C”), pregenome (“pgRNA”), pre-S1, pre-S2 / S, and HBx.
[0208] In some embodiments, the multiple epigenetic modification DNA targeting system targets or binds to target sites in a gene, such as any of the target sites described above. In some embodiments, the target site is located in a regulatory DNA element of a gene in a cell (e.g., a hepatocyte). In some embodiments, the regulatory DNA element is a sequence in which a gene regulatory protein can bind and influence gene transcription. In some embodiments, the regulatory DNA element is a cis-, trans-, distal, proximal, upstream, or downstream regulatory DNA element of a gene. In some embodiments, the regulatory DNA element is a promoter or enhancer of a gene. In some embodiments, the target site is located within a promoter, enhancer, exon, intron, untranslated region (UTR), 5' UTR, or 3' UTR of a gene. In some embodiments, the promoter is a nucleotide sequence to which an RNA polymerase binds to initiate gene transcription. In some embodiments, the promoter is a nucleotide sequence typically located between 100 bp and 1000 bp from the gene transcription start site, such as a nucleotide sequence located within approximately 100 bp, approximately 500 bp, or approximately 1000 bp from the gene transcription start site. In some implementations, the target site is located within a sequence that has an unknown or known function suspected of controlling gene expression.
[0209] In some implementations, the target site is located within approximately 50 base pairs (bp) of the transcription start site, approximately 100 bp, approximately 150 bp, approximately 200 bp, approximately 250 bp, approximately 300 bp, approximately 350 bp, approximately 400 bp, approximately 450 bp, approximately 500 bp, approximately 600 bp, approximately 650 bp, approximately 700 bp, approximately 750 bp, approximately 800 bp, approximately 850 bp, approximately 900 bp, approximately 1000 bp, approximately 1050 bp, approximately 1100 bp, approximately 1200 bp, approximately 1250 bp, approximately 1300 bp, approximately 1350 bp, approximately 1400 bp, approximately 1450 bp, and approximately 1500 bp.
[0210] In some embodiments, the target site is located within a target region, which is located at a base pair between 1 bp and 3300 bp in the HBV genome. In some embodiments, the target site is a sequence within the target region having a sequence corresponding to the sequence in the reference hepatitis B virus genome (hepatitis B virus subtype ayw, complete genome, GenBank: U95551.1) SEQ ID NO 650 located between 43 bp-490 bp, 1033 bp-1749 bp, 1800 bp-1950 bp, or 2953 bp-3182 in the HBV genome. In some implementations, the target site is a sequence within a target region having a sequence corresponding to SEQ ID NO 650 of the reference hepatitis B virus genome (hepatitis B virus subtype ayw, complete genome, GenBank: U95551.1) located between 1 bp-42 bp, 491 bp-1032 bp, 1750 bp-1799 bp, or 1951 bp-2952 bp or 3198 bp-3182 bp in the HBV genome. Based on phylogenetic analysis and sequence divergence, HBV can be divided into 10 genotypes (A to J) based on an intergroup divergence of 8% or higher in the complete nucleotide sequence (Norder H, et al., Complete genomes, phylogenetic relatedness, and structural proteins of six strains of the hepatitis B virus, four of which represent two new genotypes. Virology. Feb. 1994; 198(2): 489-503; Stuyver L, et al., A new genotype of hepatitis B virus: complete genome and phylogenetic relatedness. J Gen Virol. Jan. 2000; 81(Pt 1): 67-74; Arauz-Ruiz P, et al., Genotype H: a new American genotype of hepatitis B virus revealed in Central America. J GenVirol. Aug. 2002; 83(Pt 8): 2059-2073). Evidence suggests that HBV genotype influences clinical outcomes, mutation patterns in the pre-core and core promoter regions, HBeAg seroconversion rates, and response to interferon therapy.Most genotypes have specific geographical distributions; genotypes A and D are prevalent in Western Europe and North America, while genotypes B and C are prevalent in East Asia and Oceania.
[0211] In some embodiments, the target site is at least 70% homologous to all hepatitis B virus genotypes (e.g., genomes). In some embodiments, the target site is at least 70% homologous to at least 500, at least 1000, at least 1500, at least 2000, at least 2500, at least 3000, at least 3500, at least 4000, at least 4500, at least 5000, at least 5500, at least 6000, at least 6500, or at least 7000 hepatitis B virus genomes.
[0212] In some embodiments, the target site is at least 70% homologous to at least 1000 hepatitis B virus genomes and contains at most two mismatches. In some embodiments, the target site comprises a sequence shown in any one of SEQ ID NO: 1-195, a continuous portion of at least 14 nucleotides (nt) of any one of SEQ ID NO: 1-195, or a complementary sequence of any one of the foregoing. In some embodiments, the target site is a continuous portion of any one of SEQ ID NO: 1-195 of length 15, 16, 17, 18, or 19 nucleotides, or a complementary sequence of any one of the foregoing. In some embodiments, the target site is a sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity with all or continuous portions of the target site sequence described above. In some embodiments, the target site is a sequence shown in any one of SEQ ID NO: 1-195.
[0213] In some embodiments, the target site is a sequence of 14 to 22 nucleotides. In some embodiments, the target site is a sequence of 14 to 19 nucleotides. In some embodiments, the target site is a sequence of 14 nucleotides. In some embodiments, the target site is a sequence of 15 nucleotides. In some embodiments, the target site is a sequence of 16 nucleotides. In some embodiments, the target site is a sequence of 17 nucleotides. In some embodiments, the target site is a sequence of 18 nucleotides. In some embodiments, the target site is a sequence of 19 nucleotides.
[0214] In any of the embodiments provided herein, the target site is complementary to the reference sequence (i.e., the specific sequence shown in the reference sequence listing as SEQ ID NO). In some of any embodiments, the complementary sequence is the inverse complementary sequence of the reference sequence.
[0215] In any of the embodiments provided herein, the target site comprises a reference sequence (i.e., a specific sequence shown in SEQ ID NO in the reference sequence listing). In any of the embodiments provided herein, the target site is the sequence shown by the reference sequence (i.e., a specific sequence shown in SEQ ID NO in the reference sequence listing).
[0216] In any of the embodiments provided herein, the target site is a continuous portion of at least 14 nucleotides (14 nt) of the reference sequence (i.e., the specific sequence shown in the reference sequence listing as SEQ ID NO). In some embodiments, the continuous portion is 15 nucleotides. In some embodiments, the continuous portion is 16 nucleotides. In some embodiments, the continuous portion is 17 nucleotides. In some embodiments, the continuous portion is 18 nucleotides. In some embodiments, the continuous portion is 19 nucleotides.
[0217] In some implementations, the target site is at least 90% homologous to all hepatitis B virus genomes. In some implementations, the target site is at least 90% homologous to the following:
[0218] At least 500, at least 1000, at least 1500, at least 2000, at least 2500, at least 3000, at least 3500, at least 4000, at least 4500, at least 5000, at least 5500, at least 6000, at least 6500, at least 7000 hepatitis B virus genomes.
[0219] In some embodiments, the target site is at least 90% homologous to at least 1000 hepatitis B virus genomes and contains one or two mismatches. In some embodiments, the target site comprises a sequence shown in any one of SEQ ID NO: 35-100, a continuous portion of at least 14 nt of any one of SEQ ID NO: 35-100, or a complementary sequence of any one of the foregoing. In some embodiments, the target site is a continuous portion of any one of SEQ ID NO: 35-100 having a length of 15, 16, 17, 18, or 19 nucleotides, or a complementary sequence of any one of the foregoing. In some embodiments, the target site is a sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity with all or continuous portions of the target site sequence described above. In some embodiments, the target site is a sequence shown in any one of SEQ ID NO: 35-100.
[0220] In some implementations, the mismatch is located in the first 12 nts of the 5' end of the prototype spacer neighbor motif (PAM), as indicated by 'n' in 'nnnnnnnnnnnnNNNNNNNN-NGG'.
[0221] In some embodiments, the target site is at least 90% homologous to at least 1000 hepatitis B virus genomes and contains zero mismatches. In some embodiments, the target site comprises a sequence shown in any one of SEQ ID NO: 1-34, a continuous portion of at least 14 nt of any one of SEQ ID NO: 1-34, or a complementary sequence of any one of the foregoing. In some embodiments, the target site is a continuous portion of any one of SEQ ID NO: 1-34 with a length of 15, 16, 17, 18, or 19 nucleotides, or a complementary sequence of any one of the foregoing. In some embodiments, the target site is a sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity with all or continuous portions of the target site sequence described above. In some embodiments, the target site is a sequence shown in any one of SEQ ID NO: 1-34.
[0222] In any embodiment herein, the target site or each of the target sites comprises the sequence shown in any one of SEQ ID NO: 175, 138, 192, 152, 118, 125, 185, 63, 116, 124, 35, 82, a continuous portion of at least 14 nucleotides (nt) therein, or a complementary sequence of any one of the foregoing. In some embodiments, the target site is a continuous portion of any one of SEQ ID NO: 175, 138, 192, 152, 118, 125, 185, 63, 116, 124, 35, 82 of length 15, 16, 17, 18, or 19 nucleotides, or a complementary sequence of any one of the foregoing. In some embodiments, the target site is a sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity with all or a continuous portion of the target site sequence described above. In some embodiments, the target site is a sequence shown in any one of SEQ ID NO: 175, 138, 192, 152, 118, 125, 185, 63, 116, 124, 35, 82. In some embodiments, compared to gene levels in control cells, there is a reduction in gene expression or a change in transcript levels greater than 90%.
[0223] In any embodiment herein, the target site or each of the target sites comprises the sequence shown in any one of SEQ ID NO: 5, 10, 12, 18, 22, 26, 29, 38, 56, 61, 62, 63, 68, 72, 79, 80, 82, 84, 98, 99, 116, 118, 121, 124, 125, 135, 138, 143, 150, 152, 158, 164, 175, 176, 182, 185, 189, 190, 192, a continuous portion of at least 14 nucleotides (nt) thereof, or a complementary sequence to any one of the foregoing. In some implementations, the target site is a continuous portion of any one of SEQ ID NO: 5, 10, 12, 18, 22, 26, 29, 38, 56, 61, 62, 63, 68, 72, 75, 79, 80, 82, 84, 98, 99, 116, 118, 121, 124, 125, 135, 138, 143, 150, 152, 158, 164, 175, 176, 182, 185, 189, 190, 192, of length 15, 16, 17, 18, or 19 nucleotides, or a complementary sequence of any one of the foregoing. In some embodiments, the target site is a sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity with all or a continuous portion of the target site sequence described above. In some embodiments, the target site is a sequence represented by any one of SEQ ID NO: 5, 10, 12, 18, 22, 26, 29, 38, 56, 61, 62, 63, 68, 72, 75, 79, 80, 82, 84, 98, 99, 116, 118, 121, 124, 125, 135, 138, 143, 150, 152, 158, 164, 175, 176, 182, 185, 189, 190, or 192. In some implementations, the reduction in gene expression or the change in transcript levels in cells is greater than 80% compared to the gene levels in control cells.
[0224] In any embodiment herein, the target site or each of the target sites comprises the sequence shown in any one of SEQ ID NO: 5, 6, 12, 18, 22, 26, 29, 38, 42, 43, 51, 56, 61, 63, 68, 72, 75, 79, 82, 84, 88, 89, 98, 99, 113, 116, 121, 124, 125, 118, 130, 133, 135, 138, 143, 150, 152, 155, 158, 164, 165, 175, 176, 182, 185, 189, 190, 192, a continuous portion of at least 14 nucleotides (nt) thereof, or a complementary sequence to any one of the foregoing. In some implementations, the target site is a continuous portion of any one of SEQ ID NO: 5, 6, 12, 18, 22, 26, 29, 38, 42, 43, 51, 56, 61, 63, 68, 72, 75, 79, 82, 84, 88, 89, 98, 99, 113, 116, 121, 124, 125, 118, 130, 133, 135, 138, 143, 150, 152, 155, 158, 164, 165, 175, 176, 182, 185, 189, 190, 192, with a length of 15, 16, 17, 18, or 19 nucleotides, or a complementary sequence of any one of the foregoing. In some implementations, the target site is a sequence that has at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity with all or a continuous portion of the target site sequence described above. In some embodiments, the target site is any one of the sequences shown in SEQ ID NO: 5, 6, 12, 18, 22, 26, 29, 38, 42, 43, 51, 56, 61, 63, 68, 72, 75, 79, 82, 84, 88, 89, 98, 99, 113, 116, 121, 124, 125, 118, 130, 133, 135, 138, 143, 150, 152, 155, 158, 164, 165, 175, 176, 182, 185, 189, 190, or 192. In some embodiments, compared to gene levels in control cells, there is a reduction in gene expression or a change in transcript levels greater than 75%.
[0225] In any embodiment herein, the target site or each of the target sites comprises the sequence shown in any one of SEQ ID NO: 22, 63, 75, 99, 116, 124, 138, 143, 150, 152, 175, 176, 192, a continuous portion of at least 14 nucleotides (nt) therein, or a complementary sequence to any one of the foregoing. In some embodiments, the target site comprises a sequence of length 15, 16, 17, 18, or 19 nucleotides, or a complementary sequence to any one of the foregoing. In some embodiments, the target site is a sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity with all or a continuous portion of the target site sequence described above. In some embodiments, the target site is a sequence shown in any one of SEQ ID NO: 22, 63, 75, 99, 116, 124, 138, 143, 150, 152, 175, 176, or 192.
[0226] In any embodiment herein, the target site or each of the target sites comprises the sequence shown in any one of SEQ ID NO: 12, 18, 20, 22, 26, 27, 46, 50, 63, 66, 73, 79, 185, 192, a continuous portion of at least 14 nucleotides (nt) therein, or a complementary sequence of any one of the foregoing. In some embodiments, the target site is a continuous portion of any one of SEQ ID NO: 12, 18, 20, 22, 26, 27, 46, 50, 63, 66, 73, 79, 185, 192 of length 15, 16, 17, 18, or 19 nucleotides, or a complementary sequence of any one of the foregoing. In some embodiments, the target site is a sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity with all or a continuous portion of the target site sequence described above. In some embodiments, the target site is a sequence shown in any one of SEQ ID NO: 12, 18, 20, 22, 26, 27, 46, 50, 63, 66, 73, 79, 185, or 192.
[0227] In some embodiments, the target site, or each of the target sites, comprises a nucleotide sequence shown in any one of SEQ ID NO: 1028-1055, a continuous portion of at least 12 nt thereof, or a complementary sequence of any one of the foregoing. In some embodiments, the target site is a continuous portion of any one of SEQ ID NO: 1028-1055 having a length of 13, 14, 16, 16, 17, or 18 nucleotides, or a complementary sequence of any one of the foregoing. In some embodiments, the target site is a sequence shown in any one of SEQ ID NO: 1028-1055.
[0228] In some embodiments, the target site or each of the target sites comprises the sequence shown in SEQ ID NO: 22, a continuous portion of at least 14 nucleotides (nt) thereof, or a complementary sequence thereof. In any embodiment herein, the target site or each of the target sites comprises a continuous portion of the sequence shown in SEQ ID NO: 22 of 14-19 nucleotides (nt) in length, or a complementary sequence thereof. In some embodiments, the target site is the sequence shown in SEQ ID NO: 22. In some embodiments, the target site can be targeted using the DNA targeting system provided herein. In some embodiments, the DNA-binding domain is dCas9, specifically dSpCas9, and is used in combination with a complementary gRNA to target the target site. In some embodiments, the gRNA has the spacer sequence shown in SEQ ID NO: 217 or a continuous portion thereof complementary to the target site. In some embodiments, the gRNA also comprises the scaffold sequence of dSpCas9 shown in SEQ ID NO: 587. In some embodiments, the DNA targeting system comprises a dSpCas9 fusion protein having the effector domain described herein, and the gRNA shown in SEQ ID NO: 22 (e.g., HBVg_22).
[0229] In some embodiments, the target site or each of the target sites comprises the sequence shown in SEQ ID NO: 63, a continuous portion of at least 14 nucleotides (nt) thereof, or a complementary sequence thereof. In any embodiment herein, the target site or each of the target sites comprises a continuous portion of the sequence shown in SEQ ID NO: 63 of 14-20 nucleotides (nt) in length, or a complementary sequence thereof. In some embodiments, the target site is the sequence shown in SEQ ID NO: 63. In some embodiments, the target site can be targeted using the DNA targeting system provided herein. In some embodiments, the DNA-binding domain is dCas9, specifically dSpCas9, and is used in combination with a complementary gRNA to target the target site. In some embodiments, the gRNA has the spacer sequence shown in SEQ ID NO: 217 or a continuous portion thereof complementary to the target site. In some embodiments, the gRNA also comprises the scaffold sequence of SpCas9 shown in SEQ ID NO: 587. In some embodiments, the DNA targeting system comprises a dSpCas9 fusion protein having the effector domain described herein, and the gRNA shown in SEQ ID NO: 63 (e.g., HBVg_63).
[0230] In any embodiment herein, the target site or each of the target sites comprises the sequence shown in SEQ ID NO: 1045, a continuous portion of at least 12 nucleotides (nt) thereof, or a complementary sequence to the foregoing. In any embodiment herein, the target site or each of the target sites comprises a continuous portion of the sequence shown in SEQ ID NO: 1045 of length 12-18 nucleotides (nt) or a complementary sequence to the foregoing. In some embodiments, the target site is the sequence shown in SEQ ID NO: 1045. In some embodiments, the DNA-binding domain is an eZFP for targeting the target site. In some embodiments, the ZFP comprises the recognition motifs shown in SEQ ID NO: 822, 823, 824, 825, 826, and 827. In some embodiments, the eZFP has the sequence shown in SEQ ID NO: 709. In some embodiments, the eZFP is an eZFP designated as eZFP_18.
[0231] In any embodiment herein, the target site or each of the target sites comprises the sequence shown in SEQ ID NO: 1046, a continuous portion of at least 12 nucleotides (nt) thereof, or a complementary sequence to the foregoing. In any embodiment herein, the target site or each of the target sites comprises a continuous portion of the sequence shown in SEQ ID NO: 1046 of length 12-18 nucleotides (nt) or a complementary sequence to the foregoing. In some embodiments, the target site is the sequence shown in SEQ ID NO: 1046. In some embodiments, the DNA-binding domain is an eZFP targeting the target site. In some embodiments, the ZFP comprises the recognition motifs shown in SEQ ID NO: 828, 829, 830, 831, 832, and 833. In some embodiments, the eZFP has the sequence shown in SEQ ID NO: 710. In some embodiments, the eZFP is the eZFP designated as eZFP_19.
[0232] In any embodiment herein, the target site or each of the target sites comprises the sequence shown in SEQ ID NO: 1052, a continuous portion of at least 12 nucleotides (nt) thereof, or a complementary sequence to the foregoing. In any embodiment herein, the target site or each of the target sites comprises a continuous portion of the sequence shown in SEQ ID NO: 1052 of 12-18 nucleotides (nt) in length, or a complementary sequence to the foregoing. In some embodiments, the target site is the sequence shown in SEQ ID NO: 1052. In some embodiments, the DNA-binding domain is an eZFP for targeting the target site. In some embodiments, the ZFP comprises the recognition motifs shown in SEQ ID NO: 864, 865, 866, 867, 868, and 869. In some embodiments, the eZFP has the sequence shown in SEQ ID NO: 716. In some embodiments, the eZFP is the eZFP designated as eZFP_25.
[0233] In some embodiments, the target site is present in covalently closed circular DNA (cccDNA), relaxed circular DNA (rcDNA), and / or integrated into human genomic DNA. In some embodiments, targeting the target site leads to silencing of HBV replication (e.g., reduced HBV replication) and / or HBV transcription.
[0234] In some embodiments, this document provides a multiplex epigenetic modification DNA targeting system that targets at least two target genes or combinations of their regulatory DNA elements as described herein. In some embodiments, the multiplex epigenetic modification DNA targeting system targets two, three, four, five, six or more target genes or their regulatory DNA elements as described herein.
[0235] In some implementations, in the provided multiple epigenetic modification DNA targeting system, the target sites are located in different HBV genes. In some implementations, the target sites are located in the same HBV gene.
[0236] In some implementations, this document provides a multiple epigenetic modification DNA targeting system that targets any combination of the genes described herein and / or their regulatory elements.
[0237] In some embodiments, this document provides a multiple epigenetic modification DNA targeting system that targets a first gene or its regulatory element and a second gene or its regulatory element. In some embodiments, the first gene or its regulatory element is selected from polymerase genes, S family genes, X genes, core family genes, pre-S1 promoters, pre-S2 promoters, X promoters, basic core promoters, Enh1 enhancers, Enh2 enhancers, transcript processing control regions, and any coding regions within the HBV genome; the second gene or its regulatory element is selected from polymerase genes, S family genes, X genes, core family genes, pre-S1 promoters, pre-S2 promoters, X promoters, basic core promoters, Enh1 enhancers, Enh2 enhancers, transcript processing control regions, and any coding regions within the HBV genome, and the first gene or its regulatory element differs from the second gene or its regulatory element. The first and second target sites can be any target sites as described above.
[0238] In some embodiments, this document provides a multiplex epigenetic modification DNA targeting system that targets a first regulatory element and a second regulatory element. In some embodiments, the first and second regulatory elements are selected from combinations listed in Table 1. Table 1. Combinations of the first and second regulatory elements targeted by the multiple epigenetic modification DNA targeting system provided in this paper.
[0239] In some implementations, this document provides a multiple epigenetic modification DNA targeting system that targets a first gene or its regulatory element, a second gene or its regulatory element, and a third gene or its regulatory element. In some implementations, the first gene or its regulatory element is selected from the polymerase gene, S family gene, X gene, core family gene, pre-S1 promoter, pre-S2 promoter, X promoter, basic core promoter, Enh1 enhancer, Enh2 enhancer, transcript processing control region, and any coding region within the HBV genome; the second gene or its regulatory element is selected from the polymerase gene, S family gene, X gene, core family gene, pre-S1 promoter, pre-S2 promoter, X promoter, basic core promoter, Enh1 enhancer, Enh2 enhancer, transcript processing control region, and any coding region within the HBV genome; and the third gene or its regulatory element is selected from the polymerase gene, S family gene, X gene, core family gene, pre-S1 promoter, pre-S2 promoter, X promoter, basic core promoter, Enh1 enhancer, Enh2 enhancer, and transcript processing control region, and the first gene or its regulatory element, the second gene or its regulatory element, and the third gene or its regulatory element are different from each other. The first target site, the second target site, and the third target site can be any target site as described above.
[0240] In some embodiments, this document provides a multiple epigenetic modification DNA targeting system that targets a first regulatory element, a second regulatory element, and a third regulatory element. In some embodiments, the first regulatory element is selected from L-HBs promoters, M-HBs promoters, S-HBs promoters, X promoters, basic core promoters, Enh1 enhancers, and Enh2 enhancers; the second regulatory element is selected from L-HBs promoters, M-HBs promoters, S-HBs promoters, X promoters, basic core promoters, Enh1 enhancers, and Enh2 enhancers; and the third regulatory element is selected from L-HBs promoters, M-HBs promoters, S-HBs promoters, X promoters, basic core promoters, Enh1 enhancers, and Enh2 enhancers, and the first, second, and third regulatory elements are different. In some embodiments, the first regulating element is the Enh1 enhancer, the second regulating element is selected from the L-HBs promoter, M-HBs promoter, S-HBs promoter, X promoter, basic core promoter, and Enh2 enhancer, and the third regulating element is selected from the L-HBs promoter, M-HBs promoter, S-HBs promoter, X promoter, basic core promoter, and Enh2 enhancer, and the second and third regulating elements are different. In some embodiments, the first regulating element is the Enh1 enhancer, the second regulating element is L-HBs, and the third regulating element is selected from the M-HBs promoter, S-HBs promoter, X promoter, basic core promoter, and Enh2 enhancer.
[0241] In some implementations, the first regulating element, the second regulating element, and the third regulating element are selected from the combinations listed in Table 2. Table 2. Combinations of the first, second, and third regulatory elements targeted by the multiple epigenetic modification DNA targeting system provided in this paper.
[0242] In some embodiments, this document provides a multiple epigenetic modification DNA targeting system that targets a first gene or its regulatory element, a second gene or its regulatory element, a third gene or its regulatory element, and a fourth gene or its regulatory element. In some embodiments, the first gene or its regulatory element is selected from polymerase genes, S family genes, X genes, core family genes, pre-S1 promoters, pre-S2 promoters, X promoters, basic core promoters, Enh1 enhancers, Enh2 enhancers, transcript processing control regions, and any coding regions within the HBV genome; the second gene or its regulatory element is selected from polymerase genes, S family genes, X genes, core family genes, pre-S1 promoters, pre-S2 promoters, X promoters, basic core promoters, Enh1 enhancers, Enh2 enhancers, transcript processing control regions, and any coding regions within the HBV genome; the third gene or its regulatory element... The elements are selected from polymerase genes, S family genes, X genes, core family genes, pre-S1 promoters, pre-S2 promoters, X promoters, basic core promoters, Enh1 enhancers, Enh2 enhancers, and transcript processing control regions. The fourth gene or its regulatory element is selected from polymerase genes, S family genes, X genes, core family genes, pre-S1 promoters, pre-S2 promoters, X promoters, basic core promoters, Enh1 enhancers, Enh2 enhancers, and transcript processing control regions. Furthermore, the first gene or its regulatory element, the second gene or its regulatory element, the third gene or its regulatory element, and the fourth gene or its regulatory element are distinct from each other. The first target site, the second target site, the third target site, and the fourth target site can be any target site as described above.
[0243] In some embodiments, this document provides multiple epigenetic modification DNA targeting systems that target the same gene or its regulatory elements. For example, two or more multiple epigenetic modification DNA targeting systems target the same or common gene or its regulatory elements. In some embodiments, the gene or its regulatory element is selected from polymerase genes, S family genes, X genes, core family genes, pre-S1 promoters, pre-S2 promoters, X promoters, basic core promoters, Enh1 enhancers, Enh2 enhancers, transcript processing control regions, and any coding regions within the HBV genome, and the second gene or its regulatory element is selected from polymerase genes, S family genes, X genes, core family genes, pre-S1 promoters, pre-S2 promoters, X promoters, basic core promoters, Enh1 enhancers, Enh2 enhancers, transcript processing control regions, and any coding regions within the HBV genome. In some implementations, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, or at least 30 multiple epigenetic modification DNA targeting systems target the same gene or its regulatory elements.
[0244] In some embodiments, the target site for multiple editing (e.g., via the multiple epigenetic modification DNA targeting system described herein) is any one of the sequences shown in SEQ ID NO: 12, 18, 20, 22, 26, 27, 46, 50, 63, 66, 73, 79, 185, 192. In some embodiments, the target site for multiple editing (e.g., via the multiple epigenetic modification DNA targeting system described herein) is the sequence shown in SEQ ID NO: 22. The target site can be any target site as described above. B. CRISPR-based DNA targeting systems
[0245] This article provides an epigenetic DNA targeting system based on the CRISPR / Cas system, namely, a CRISPR / Cas-based DNA targeting system capable of binding to target sites in target genes or their regulatory elements. In some embodiments, the provided epigenetic DNA targeting system is a multiplex epigenetic DNA targeting system based on the CRISPR / Cas system, namely, a CRISPR / Cas-based DNA targeting system capable of targeting target sites in a combination of target genes or their regulatory elements.
[0246] In some embodiments, the CRISPR / Cas DNA-binding domain is nuclease-free, such as including dCas (e.g., dCas9), such that the system binds to a target site in a target gene or its regulatory element without mediating nucleic acid cleavage at the target site. In some embodiments, the DNA targeting system does not introduce gene damage or DNA breaks. CRISPR / Cas-based DNA targeting systems can be used to regulate the expression of target genes in cells such as hepatocytes. In some embodiments, the target gene or its regulatory element may include any target gene or its regulatory element described herein, including any target gene or its regulatory element described in Section IA above. In some embodiments, the target site of the target gene or its regulatory element may include any target site described herein, including any target site described in Section IA above. In some embodiments, the CRISPR / Cas-based DNA targeting system may include any known Cas enzyme, and is generally nuclease-free or dCas. In some embodiments, the CRISPR / Cas-based DNA targeting system comprises: a fusion protein of a nuclease-free Cas protein or a variant thereof and an effector domain that reduces gene transcription (e.g., a transcriptional repressor), and at least one gRNA.
[0247] The CRISPR system (also known as the CRISPR / Cas system or CRISPR-Cas system) refers to a conserved microbial nuclease system found in the genomes of bacteria and archaea, providing an acquired form of immunity against invading bacteriophages and plasmids. Clustered regularly spaced short palindromic repeats (CRISPR) are loci containing multiple repetitive DNA elements separated by non-repetitive DNA sequences called spacers. Spacers are short foreign DNA sequences incorporated into the genome between CRISPR repeat sequences as a "memory" of past exposures. Spacers encode DNA-targeting portions of RNA molecules that confer the specificity of the CRISPR system for nucleic acid cleavage. CRISPR loci contain or are adjacent to: one or more CRISPR-associated (Cas) genes, which can act as RNA-guided nuclease-mediated cleavage; and non-protein-coding DNA elements encoding RNA molecules capable of programming the specificity of CRISPR-mediated nucleic acid cleavage.
[0248] In a type II CRISPR / Cas system containing the Cas protein Cas9, two RNA molecules and the Cas9 protein form a ribonucleoprotein (RNP) complex to guide Cas9 nuclease activity. The CRISPR RNA (crRNA) contains a spacer sequence complementary to the target nucleic acid sequence (target site) and encodes the sequence specificity of the complex. Trans-activating crRNA (tracrRNA) pairs with a portion of the crRNA, forming a structure that complexes with the Cas9 protein, thus forming the Cas / RNA RNP complex.
[0249] Naturally occurring CRISPR / Cas systems, such as those with Cas9, have been engineered to allow for efficient programming of Cas / RNA RNPs to target desired sequences in target cells for gene editing and regulation of gene expression. tracrRNA and crRNA have been engineered to form a single chimeric guide RNA molecule, commonly referred to as guide RNA (gRNA), as described, for example, in: WO 2013 / 176772 A1, WO 2014 / 093661 A2, WO 2014 / 093655 A2, Jinek, M. et al., Science 337(6096):816-21 (2012) or Cong, L. et al., Science 339(6121):819-23 (2013). The spacer sequence of the gRNA can be selected by the user to target the Cas / gRNA RNP complex to a desired locus, such as a target gene and / or a desired target site in its regulatory elements.
[0250] Cas proteins are also engineered to allow targeting of Cas / gRNA RNPs without inducing cleavage at the target site. Mutations in Cas proteins can reduce or eliminate the nuclease activity of Cas proteins, thereby rendering them non-catalytic. Cas proteins with reduced or eliminated nuclease activity are referred to as inactive Cas (dCas) or Cas proteins without nuclease activity (iCas), and these terms are used interchangeably in this document. Exemplary inactive Cas9 (dCas9) derived from Streptococcus pyogenes contains silent mutations (D10A and H840A) in the RuvC and HNH nuclease domains, as described, for example, in: WO 2013 / 176772 A1, WO 2014 / 093661A2, Jinek, M. et al. Science 337(6096):816-21 (2012), and Qi, L. et al. Cell 152(5):1173-83 (2013). Exemplary dCas variants derived from the Cas12 system (i.e., Cpf1) are described, for example, in: WO 2017 / 189308 A1 and Zetsche, B. et al. Cell 163(3):759-71 (2015). Conserved domains mediating nucleic acid cleavage, such as the RuvC and HNH endonuclease domains, are readily identifiable in Cas orthogonal homologs and can be mutated to produce inactivating variants, for example, as described in: Zetsche, B. et al. Cell 163(3):759-71 (2015).
[0251] dCas fusion proteins with transcriptional and / or epigenetic regulators have been used as multifunctional platforms for ectopic regulation of gene expression in target cells. These include fusions of Cas with effector domains such as transcriptional activators or transcriptional repressors. For example, fusing dCas9 with a transcriptional activator (such as VP64, a polypeptide consisting of four tandem copies of VP16, the 16-amino acid transcriptional activation domain of herpes simplex virus) can lead to robust induction of gene expression. Alternatively, fusing dCas9 with a transcriptional repressor (such as KRAB (Krüppel-related box)) can lead to robust repression of gene expression. Various dCas fusion proteins with transcriptional and epigenetic regulatory factors can be engineered for the regulation of gene expression, as described below: WO 2014 / 197748, WO 2016 / 130600, WO 2017 / 180915, WO2021 / 226555, WO 2013 / 176772, WO 2014 / 152432, WO 2014 / 093661, WO 2021 / 247570, Adli, M. Nat. Commun. 9, 1911 (2018), Perez-Pinera, P. et al. Nat. Methods 10, 973-976 (2013), Mali, P. et al. Nat. Biotechnol. 31, 833-838 (2013), Maeder, ML et al. Nat. Methods 10, 977-979 (2013), Gilbert, LA et al. Cell 154(2):442-451 (2013), and Nuñez, JK et al. Cell 184(9):2503-2519 (2021).
[0252] In some aspects, a DNA targeting system is provided comprising a fusion protein containing a DNA-binding domain and an effector domain, the DNA-binding domain comprising a nuclease-free Cas protein or a variant thereof, the effector domain being used to reduce transcription or induce transcriptional repression (i.e., a transcriptional repressor) when targeting a target gene or its regulatory element in a cell (e.g., a hepatocyte). In such embodiments, the DNA targeting system further comprises one or more gRNAs provided in combination with or as a complex of a dCas protein or a variant thereof for targeting the DNA targeting system to a target site of a target gene or its regulatory element. In some embodiments, the fusion protein is guided by a guide RNA to a specific target site sequence of a target gene or its regulatory element, wherein the effector domain mediates the targeting of epigenetic modifications to reduce or repress transcription of the target gene. In some embodiments, a combination of gRNAs guides the fusion protein to a combination of target site sequences in a combination of genes or their regulatory elements, wherein the effector domain mediates the targeting of epigenetic modifications to reduce or repress transcription of the combination of target genes. As further described below, any of a variety of effector domains that reduce or repress transcription may be used. 1. CRISPR-based DNA binding domain
[0253] In some respects, the DNA-binding domain contains a CRISPR-associated (Cas) protein or a variant thereof, or is derived from a Cas protein or a variant thereof, and is non-nuclease-active (i.e., it is a dCas protein).
[0254] In some embodiments, the Cas protein is derived from a type 1 CRISPR system (i.e., a multi-Cas protein system), such as type I, III, or IV CRISPR systems. In some embodiments, the Cas protein is derived from a type 2 CRISPR system (i.e., a single-Cas protein system), such as type II, V, or VI CRISPR systems. In some embodiments, the Cas protein is derived from a type V CRISPR system. In some embodiments, the Cas protein is derived from the Cas12 protein (i.e., Cpf1) or a variant thereof, for example, as described in: WO 2017 / 189308 A1 and Zetsche, B. et al. Cell. 163(3):759-71 (2015). In some embodiments, the Cas protein is derived from a type II CRISPR system. In some implementations, the Cas protein is derived from the Cas9 protein or a variant thereof, as described below: WO 2013 / 176772 A1, WO 2014 / 152432 A2, WO 2014 / 093661 A2, WO 2014 / 093655 A2, Jinek, M. et al. Science 337(6096):816-21 (2012), Mali, P. et al. Science 339(6121):823-6 (2013), Cong, L. et al. Science 339(6121):819-23 (2013), Perez-Pinera, P. et al. Nat. Methods 10, 973-976 (2013), or Mali, P. et al. Nat. Biotechnol. 31, 833-838 (2013). Various CRISPR / Cas systems and associated Cas proteins used in gene editing and regulation have been described, for example, in Moon, SB et al. Exp. Mol. Med. 51, 1-11 (2019), Zhang, FQ Rev. Biophys. 52, E6 (2019), and Makarova KS et al. Methods Mol. Biol. 1311:47-75 (2015).
[0255] In some embodiments, the dCas9 protein may comprise a sequence derived from a naturally occurring Cas9 molecule or a variant thereof. In some embodiments, the dCas9 protein may comprise a sequence derived from a naturally occurring Cas9 molecule or a variant thereof, said naturally occurring Cas9 molecule being derived from *Streptococcus pyogenes*, *Streptococcus thermophilus*, *Staphylococcus aureus*, *Campylobacter jejuni*, *Neisseria meningitidis*, *F. novicida*, *Streptococcus canis*, or *Staphylococcus auricularis*. In some embodiments, the dCas9 protein comprises a sequence derived from a naturally occurring *Staphylococcus aureus* Cas9 molecule. In some embodiments, the dCas9 protein comprises a sequence derived from a naturally occurring *Streptococcus pyogenes* Cas9 molecule.
[0256] Non-limiting examples of Cas9 orthologs from other bacterial strains include, but are not limited to, the Cas proteins identified in: the deep-sea single-celled cyanobacterium *Acaryochloris marina* MBIC11017; *Acetohalobium arabaticum* DSM 5501; *Acidithiobacillus caldus*; *Acidithiobacillus ferrooxidans* ATCC 23270; *Alicyclobacillus acidocaldarius* LAA1; *Acidocaldarius* subspecies DSM 446; *Allochromatium vinosum* DSM 180; *Ammonifex degensii* KC4; and *Anabaena variabilis* ATCC. 29413; *Arthrospira maxima* CS-328; *Arthrospira platensis* Paraca strain; *Arthrospira* sp. PCC 8005; *Bacillus pseudomycoides* DSM 12442; *Bacillus selenitireducens* MLS10; Burkholderiales bacteria 1_1_47; *Caldicelulosiruptor becscii* DSM 6725; *Candidatus desulforudis audaxviator* MP104C; *Caldicellulosiruptor hydrothermalis* 108; Clostridium phage c-st; *Clostridium botulinum* A3Loch Maree strain; Clostridium botulinum Ba4657; Clostridium difficile QCD-63q42; Crocosphaera watsonii WH 8501; Cyanothece sp.ATCC 51142; *Cymbidium* species CCY0110; *Cymbidium* species PCC 7424; *Cymbidium* species PCC 7822; *Exiguobacterium sibiricum* 255-15; *Finegoldia magna* ATCC 29328; *Ktedonobacter racemifer* DSM 44963; *Lactobacillus delbrueckii* Bulgarian subspecies PB2003 / 044-T3-4; *Lactobacillus salivarius* ATCC 11741; *Listeria innocua*; *Lyngbya* sp. PCC 8106; *Marinobacter* species sp.) ELB17; Methanohalobium evestigatum Z-7303; Microcystis bacteriophage Ma-LMM01; Microcystis aeruginosa NIES-843; Microscilla marina ATCC23134; Microcoleus chthonoplastes PCC 7420; Neisseria meningitidis; Nitrosococcus halophilus Nc4; Nocardiopsis dassonvillei (DSM 43111); Nodularia spumigena CCY9414; Nostoc sp. PCC 7120; Oscillatoriasp. PCC 6506; Pelotomaculum thermopropionicum SI; Petrotoga mobilis SJ95; Polaromonas naphthalenivorans CJ2; Species of the genus Polaromonas (Polaromonas sp.)JS666; Pseudoalteromonas haloplanktis TAC125; Streptomyces pristinaespiralis ATCC 25486; Streptomyces ATCC 25486; Streptococcus thermophilus; Streptomyces viridochromogenes DSM 40736; Streptosporangium roseum DSM 43021; Synechococcus sp. PCC 7335; and Thermosiphoafricanus TCF52B (Chylinski et al., RNA Biol., 2013; 10(5): 726-737).
[0257] In some embodiments, the Cas protein is a variant lacking nuclease activity (i.e., the dCas protein). In some embodiments, the Cas protein is mutated to reduce or eliminate nuclease activity. Such Cas proteins are interchangeably referred to herein as inactive Cas or non-active Cas (dCas) or Cas protein lacking nuclease activity (iCas). In some embodiments, the variant Cas protein is a variant Cas9 protein lacking nuclease activity or acting as an inactive Cas9 protein (dCas9 or iCas9).
[0258] In some embodiments, the Cas9 protein or a variant thereof is derived from the Staphylococcus aureus Cas9 (SaCas9) protein or a variant thereof. In some embodiments, the variant Cas9 is a Staphylococcus aureus dCas9 protein (dSaCas9) containing at least one amino acid mutation selected from D10A and N580A, referenced to the position number of SEQ ID NO: 596. In some embodiments, the variant Cas9 protein comprises the sequence shown in SEQ ID NO: 597, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity therewith.
[0259] In some embodiments, the Cas9 protein or a variant thereof is derived from the Streptococcus pyogenes Cas9 (SpCas9) protein or a variant thereof. In some embodiments, the variant Cas9 is a Streptococcus pyogenes dCas9 (dSpCas9) protein containing at least one amino acid mutation selected from D10A and H840A, referenced to the position number of SEQ ID NO: 598. In some embodiments, the variant Cas9 protein comprises the sequence shown in SEQ ID NO: 599, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity therewith. 2. Guide RNA
[0260] In some embodiments, the Cas protein (e.g., dCas9) is provided in combination with one or more guide RNAs (gRNAs) or as a complex. In some aspects, the gRNA is a nucleic acid that facilitates the specific targeting or homing of the gRNA / Cas RNP complex to target sites (as described above) of target genes and / or their regulatory elements. In some embodiments, the target site of the gRNA may be referred to as a prototypical spacer.
[0261] This document provides gRNAs, such as gRNAs that target or bind to a target site or its DNA regulatory elements (as described in Section IA above). In some embodiments, the gRNA is capable of complexing with a Cas protein or a variant thereof. In some embodiments, the gRNA comprises a gRNA spacer sequence (i.e., a spacer sequence or guide sequence) capable of hybridizing with or being complementary to a target site, such as any target site as described in Section IA or further below. In some embodiments, the gRNA comprises a scaffold sequence that complexes with or binds to a Cas protein.
[0262] In some embodiments, the gRNAs provided herein are chimeric gRNAs. Typically, gRNAs can be monomolecules (i.e., composed of a single RNA molecule) or modular (containing more than one, and typically two, separate RNA molecules). Modular gRNAs can be engineered to be monomolecules, wherein sequences from individual modular RNA molecules are contained within a single gRNA molecule; these are sometimes referred to as chimeric gRNAs, synthetic gRNAs, or monomolecules. In some embodiments, a chimeric gRNA is a fusion of two non-coding RNA sequences, namely a crRNA sequence and a tracrRNA sequence, as described, for example, in WO 2013 / 176772 A1 or Jinek, M. et al., Science 337(6096):816-21 (2012). In some embodiments, the chimeric gRNA mimics the naturally occurring crRNA:tracrRNA duplex involved in type II effector systems, where the naturally occurring crRNA:tracrRNA duplex acts as a guide for the Cas9 protein.
[0263] In some aspects, the spacer sequence of the gRNA is a polynucleotide sequence comprising at least a portion having sufficient complementarity with the target site or its DNA regulatory element (e.g., any of those described in Section IA) to hybridize with the target site in the target gene and / or its regulatory element and to guide the CRISPR complex to sequence-specific binding to the target site. Perfect complementarity is not necessarily required if sufficient complementarity exists to induce hybridization and promote the formation of the CRISPR complex. In some embodiments, the gRNA comprises a spacer sequence that is, for example, at least 80%, 85%, 90%, 95%, 98%, 99%, or 100% complementary (e.g., perfectly complementary) to the target site. The strand of the target nucleic acid containing the target site sequence may be referred to as the “complementary strand” of the target nucleic acid.
[0264] In some embodiments, the length of the gRNA spacer sequence is between approximately 14 nucleotides (nt) and approximately 26 nt, or between 16 nt and 22 nt. In some embodiments, the length of the gRNA spacer sequence is 14 nt, 15 nt, 16 nt, 17 nt, 18 nt, 19 nt, 20 nt, 21 nt, or 22 nt, 23 nt, 24 nt, 25 nt, or 26 nt. In some embodiments, the length of the gRNA spacer sequence is 18 nt, 19 nt, 20 nt, 21 nt, or 22 nt. In some embodiments, the length of the gRNA spacer sequence is 19 nt.
[0265] The target site of gRNA may be referred to as the prototype spacer. In some aspects, the spacer is designed to target a prototype spacer having a specific prototype spacer neighbor motif (PAM), which is the sequence immediately adjacent to the prototype spacer, contributing to and / or being essential for Cas binding specificity. Different CRISPR / Cas systems have different PAM requirements for targeting. For example, in some embodiments, *Streptococcus pyogenes* Cas9 uses PAM 5'-NGG-3' (SEQ ID NO: 629), where N is any nucleotide. *Staphylococcus aureus* Cas9 uses PAM 5'-NNGRRT-3' (SEQ ID NO: 630), where N is any nucleotide and R is G or A. *Neisseria meningitidis* Cas9 uses PAM 5'-NNNNGATT-3' (SEQ ID NO: 631), where N is any nucleotide. Campylobacter jejuni Cas9 uses PAM 5′-NNNNRYAC-3′ (SEQ ID NO: 632), where N is any nucleotide, R is G or A, and Y is C or T. Streptococcus thermophilus uses PAM 5′-NNAGAAW-3′ (SEQ ID NO: 633), where N is any nucleotide and W is A or T. The novel culprit Francisella Cas9 uses PAM 5′-NGG-3′ (SEQ ID NO: 634), where N is any nucleotide. Treponema denticola Cas9 uses PAM 5′-NAAAAC-3′ (SEQ ID NO: 635), where N is any nucleotide. Cas12a (also known as Cpf1) from various species uses PAM 5′-TTTV-3′ (SEQ ID NO: 636). Cas proteins can be used or engineered to use PAMs different from those listed above. For example, the mutant SpCas9 protein can use PAM 5'-NGG-3' (SEQ ID NO: 629), 5'-NGAN-3' (SEQ ID NO: 637), 5'-NGNG-3' (SEQ ID NO: 638), 5'-NGAG-3' (SEQ ID NO: 639), or 5'-NGCG-3' (SEQ ID NO: 640). In some embodiments, the prototype spacer for the gRNA used to complex with *Streptococcus pyogenes* Cas9 or a variant thereof is shown in SEQ ID NO: 588. In some embodiments, the prototype spacer for the gRNA used to complex with *Staphylococcus aureus* Cas9 or a variant thereof is shown in SEQ ID NO: 589.
[0266] Spacer sequences can be selected to reduce the degree of secondary structure within the spacer sequence. Secondary structures can be determined using any suitable polynucleotide folding algorithm.
[0267] In some implementations, the gRNA (including the guide sequence) will contain the base uracil (U), while the DNA encoding the gRNA molecule will contain the base thymine (T). Although not wishing to be bound by theory, in some implementations it is believed that the complementarity of the guide and target sequences contributes to the specificity of the interaction between the gRNA / Cas molecule complex and the target nucleic acid. It should be understood that in the guide and target sequence pair, the uracil base in the guide sequence will pair with the adenine base in the target sequence.
[0268] In some implementations, one, more than one, or all nucleotides of the gRNA may be modified, for example, to make the gRNA less susceptible to degradation and / or to improve biocompatibility. As a non-limiting example, the backbone of the gRNA may be modified with phosphate thioesters or one or more other modifications. In some cases, the nucleotides of the gRNA may contain 2' modifications, such as 2-acetylation, 2' methylation, or one or more other modifications.
[0269] Methods for designing gRNAs and exemplary targeting domains may include, for example, those described in the following: International PCT Publications WO 2014 / 197748 A2, WO 2016 / 130600 A2, WO 2017 / 180915 A2, WO 2021 / 226555A2, WO 2013 / 176772 A1, WO 2014 / 152432 A2, WO 2014 / 093661 A2, WO 2014 / 093655 A2, WO2015 / 089427 A1, WO 2016 / 049258 A2, WO 2016 / 123578 A1, WO 2021 / 076744 A1, WO 2014 / 191128 A1, WO 2015 / 161276 A2, WO 2017 / 193107 A2 and WO 2017 / 093969 A1.
[0270] In some implementations, the gRNAs provided herein target sites present in covalently closed circular DNA (cccDNA) and relaxed circular DNA (rcDNA). In some aspects, the target site can be any HBV genomic sequence most suitable for establishing DNA methylation that leads to silencing of HBV RNA transcription.
[0271] In some implementations, the target site is located at or near a gene or regulatory element involved in controlling HBV replication and / or HBV transcription. In some aspects, the target site is located at or near a promoter. In some aspects, the target site is located near an enhancer region. In some aspects, the target site is located at or near a transcript processing control region. In some aspects, the target site can be any gene most suitable for establishing DNA methylation to silence HBV transcription.
[0272] In some implementations, the gRNA targets the following target sites, which include: a sequence selected from any one of SEQ ID NO: 1-34 as shown in Table 3, a continuous portion of at least 14 nucleotides thereof, a complementary sequence of any one of the aforementioned, or a sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity with any one of the aforementioned.
[0273] In some implementations, the gRNA targets the following target sites, which include: a sequence selected from any one of SEQ ID NO: 35-100 as shown in Table 4, or a continuous portion thereof of at least 14 nt, or a sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity with any of the aforementioned sequences.
[0274] In some implementations, the gRNA targets the following target sites, which include: a sequence selected from any one of SEQ ID NOs: 101-195 as shown in Table 5, a continuous portion of at least 14 nucleotides thereof, a complementary sequence of any one of the aforementioned, or a sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity with any one of the aforementioned.
[0275] In some embodiments, the gRNA further comprises the scaffold sequence shown in SEQ ID NO: 587. In some embodiments, the gRNA further comprises a scaffold sequence. In some embodiments, the scaffold sequence comprises the sequence shown in SEQ ID NO: 587 (GUUUAAGAGCUAUGCUGGAAACAGCAUAGCAAGUUUAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC), or a sequence having all or part of it at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity. In some embodiments, the scaffold sequence is shown in SEQ ID NO: 587.
[0276] In some embodiments, the gRNA comprises a sequence selected from any one of SEQ ID NO: 391-585 as shown in Table 6, or a sequence having sequence identity of 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% with any one of SEQ ID NO: 391-585. In some embodiments, the gRNA is shown in any one of SEQ ID NO: 391-585.
[0277] In some embodiments, the gRNA comprises a sequence selected from any one of SEQ ID NO: 391-424, or a sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity with any of the aforementioned sequences.
[0278] In some embodiments, the gRNA comprises a sequence selected from any one of SEQ ID NO: 425-490, or a sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity with any of the aforementioned sequences.
[0279] In some embodiments, the gRNA comprises a sequence selected from any one of SEQ ID NO: 490-585, or a sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity with any of the aforementioned sequences.
[0280] In some embodiments, any one of the provided gRNA sequences is provided in combination with or in conjugate with Cas9. In some embodiments, Cas9 is dCas9. In some embodiments, dCas9 is dSpCas9, such as dSpCas9 shown in SEQ ID NO: 599. Table 3. Target site sequences and gRNA spacers (HBV0) with >90% HBV genome conservation and no mismatches. Table 4. Target site sequences and gRNA spacers (HBV1) with >90% HBV genome conservation and 1-2 mismatches. Table 5. Target site sequences and gRNA spacers (HBV2) with 70%-90% HBV genome conservation and at most 2 mismatches. Table 6. gRNAs targeting genes
[0281] In some embodiments, the gRNA provided herein targets a target site in the hepatitis B virus genome. In some embodiments, the gRNA targets a site located between 0 and 3300 bp in the HBV genome. In some embodiments, the gRNA targets sites located between 43 bp-490 bp, 1033 bp-1749 bp, 1800 bp-1950 bp, or 2953 bp-3182 bp in the HBV genome corresponding to the position of SEQ ID NO: 650 in the reference hepatitis B virus genome (hepatitis B virus subtype ayw, complete genome, GenBank: U95551.1). In some embodiments, the gRNA targets sites located between 1 bp-42 bp, 491 bp-1032 bp, 1750 bp-1799 bp, or 1951 bp-2952 bp in the HBV genome corresponding to the position of SEQ ID NO: 650 in the reference hepatitis B virus genome (hepatitis B virus subtype ayw, complete genome, GenBank: U95551.1). In some embodiments, the gRNA targets sites at or near regulatory elements involved in HBV replication and / or transcription. In some embodiments, the gRNA targets polymerase genes, S family genes, X genes, or core family genes. In some embodiments, the gRNA targets the M / S-HBs promoter region, X promoter region, basic core promoter region, or L-HBs promoter region. In some embodiments, the gRNA targets the Enh1 or Enh2 enhancer region. In some embodiments, the gRNA targets the HBV coding region.
[0282] In some embodiments, the gRNA provided herein comprises: a sequence selected from any one of SEQ ID NO: 196-229, a continuous portion of at least 14 nucleotides (e.g., 14, 15, 16, 17, 18, or 19 nucleotides), a complementary sequence of any one of the foregoing, or a sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity with any one of the foregoing. In some embodiments, the gRNA comprises a spacer sequence comprising: a sequence selected from any one of SEQ ID NO: 230-295, a continuous portion of at least 14 nt (e.g., 14, 15, 16, 17, 18, or 19 nucleotides), or a sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity with any of the foregoing. In some embodiments, the gRNA comprises a spacer sequence comprising: a sequence selected from any one of SEQ ID NO: 296-390, a continuous portion of at least 14 nt (e.g., 14, 15, 16, 17, 18, or 19 nucleotides), or a sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity with any of the foregoing. In some embodiments, the gRNA further comprises a scaffold sequence. In some embodiments, the scaffold sequence comprises the sequence shown in SEQ ID NO: 587, or a sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity with SEQ ID NO: 587. In some embodiments, the gRNA comprising the spacer sequence and the scaffold sequence comprises a sequence selected from any one of SEQ ID NO: 391-585, or a sequence having all or part of it having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity.
[0283] In some embodiments, the gRNA comprises a spacer sequence comprising: a sequence selected from any one of SEQ ID NO: 370, 333, 387, 347, 313, 320, 380, 256, 258, 311, 319, 230, 272, a continuous portion of at least 14 nucleotides (e.g., 14, 15, 16, 17, 18, or 19 nucleotides), a complementary sequence to any one of the aforementioned sequences, or a sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity with any one of the aforementioned sequences. In some embodiments, the gRNA also comprises a scaffold sequence. In some embodiments, the scaffold sequence comprises the sequence shown in SEQ ID NO: 587, or a sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity with SEQ ID NO: 587. In some embodiments, the gRNA comprising the spacer sequence and the scaffold sequence comprises a sequence selected from any one of SEQ ID NO: 565, 528, 542, 508, 515, 575, 515, 453, 506, 514, 425, or 472, or a sequence having all or part of it having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity with it. In some implementations, the gRNA is shown in SEQ ID NO: 565, 528, 542, 508, 515, 575, 515, 453, 506, 514, 425 or 472.
[0284] In some embodiments, the gRNA comprises a spacer sequence, said spacer sequence comprising: selected from SEQ ID NO: 200, 205, 207, 213, 217, 221, 224, 233, 251, 256, 257, 258, 263, 267, 274, 275, 277, 279, 293, 294, 311, 313, 316, 319, 320, 330, 333, 338, 345, 347, 353, 359, 370, 371, 377, 380, 384, The sequence of any one of SEQ ID NO: 385, 387, a continuous portion of at least 14 nucleotides (e.g., 14, 15, 16, 17, 18, or 19 nucleotides), a complementary sequence of any one of the foregoing, or a sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity with any one of the foregoing. In some embodiments, the gRNA further comprises a scaffold sequence. In some embodiments, the scaffold sequence comprises the sequence shown in SEQ ID NO: 587, or a sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity with SEQ ID NO: 587. In some embodiments, the gRNA comprising the spacer sequence and the scaffold sequence comprises a sequence selected from SEQ ID NO: The sequence of any one of 395, 400, 402, 408, 412, 416, 419, 428, 446, 451, 452, 453, 458, 462, 465, 469, 470, 472, 474, 488, 489, 506, 508, 511, 514, 515, 525, 528, 533, 540, 542, 548, 554, 565, 566, 572, 575, 579, 580, 582, or any sequence that has all or part of it being 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity.In some embodiments, the gRNA comprises a spacer sequence, said spacer sequence comprising: selected from SEQ ID NO: 200, 201, 207, 217, 221, 224, 233, 237, 238, 246, 251, 256, 258, 263, 267, 274, 270, 277, 279, 283, 284, 293, 294, 308, 311, 313, 316, 319, 320, 325, 328, 330, 333, 338, 345, 347, 350, 353, 359, 360, 370, 37 The sequence of any one of SEQ ID NO: 587, or a continuous portion of at least 14 nucleotides (e.g., 14, 15, 16, 17, 18, or 19 nucleotides), a complementary sequence to any one of the preceding sequences, or a sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity with any one of the preceding sequences. In some embodiments, the gRNA further comprises a scaffold sequence. In some embodiments, the scaffold sequence comprises the sequence shown in SEQ ID NO: 587, or a sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity with SEQ ID NO: 587. In some implementations, the gRNA comprising the spacer sequence and the scaffold sequence comprises a sequence selected from SEQ ID NO: 369, 395, 402, 408, 412, 416, 419, 428, 432, 433, 441, 446, 451, 453, 458, 462, 465, 469, 472, 474, 478, 479, 488, 489, 503, 506, 508, 511, 514, 515, 520, 523, 525, 575, 528, 533. The sequence of any one of 540, 542, 545, 548, 554, 555, 565, 566, 572, 579, 580, 582, or the sequence which has all or part of it the sequence identity of 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9% or 100%.In some implementations, the gRNA is shown in SEQ ID NO: 369, 395, 402, 408, 412, 416, 419, 428, 432, 433, 441, 446, 451, 453, 458, 462, 465, 469, 472, 474, 478, 479, 488, 489, 503, 506, 508, 511, 514, 515, 520, 523, 525, 575, 528, 533, 540, 542, 545, 548, 554, 555, 565, 566, 572, 579, 580, or 582.
[0285] In some embodiments, the gRNA comprises a spacer sequence comprising: a sequence selected from any one of SEQ ID NO: 207, 213, 215, 217, 221, 222, 241, 245, 258, 261, 268, 274, 380, 387, a continuous portion of at least 14 nucleotides (e.g., 14, 15, 16, 17, 18, or 19 nucleotides), a complementary sequence to any one of the aforementioned sequences, or a sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity with any one of the aforementioned sequences. In some embodiments, the gRNA comprises a spacer sequence comprising: a continuous portion of sequence SEQ ID NO: 217, of which at least 14 nucleotides (e.g., 14, 15, 16, 17, 18, or 19 nucleotides), a complementary sequence to any of the foregoing, or a sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity with SEQ ID NO: 217. In some embodiments, the gRNA further comprises a scaffold sequence. In some embodiments, the scaffold sequence comprises the sequence shown in SEQ ID NO: 587, or a sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity with SEQ ID NO: 587. In some embodiments, the gRNA comprising the spacer sequence and the scaffold sequence comprises a sequence selected from any one of SEQ ID NO: 402, 408, 410, 412, 416, 417, 436, 440, 453, 456, 463, 469, 575, 582, or a sequence having all or part of that sequence identity of at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100%. In some embodiments, the provided multiplex epigenetic modification DNA targeting system for epigenetic modification of at least two genes and / or their regulatory elements comprises any of the aforementioned gRNAs complexed with a Cas protein (such as the Cas9 protein). In some embodiments, Cas9 is dCas9. In some embodiments, dCas9 is dSpCas9 (such as the dSpCas9 shown in SEQ ID NO: 599) or a variant and / or fusion thereof.
[0286] In some implementations, this document provides combinations of gRNAs. In some implementations, this document provides multiple epigenetic modification DNA targeting systems comprising combinations of gRNAs.
[0287] In some embodiments, the combination of gRNAs comprises at least two gRNAs targeting at least two different genes or their regulatory elements. In some embodiments, the combination of gRNAs comprises a first gRNA targeting a first gene or its regulatory element and a second gRNA targeting a second gene or its regulatory element. In some embodiments, the first gRNA targets genes or their regulatory elements selected from: polymerase genes, S family genes, X genes, core family genes, pre-S1 promoters, pre-S2 promoters, X promoters, basic core promoters, Enh1 enhancers, Enh2 enhancers, transcript processing control regions, and any coding regions within the HBV genome; the second gRNA targets genes or their regulatory elements selected from: polymerase genes, S family genes, X genes, core family genes, pre-S1 promoters, pre-S2 promoters, X promoters, basic core promoters, Enh1 enhancers, Enh2 enhancers, transcript processing control regions, and any coding regions within the HBV genome, and the first gRNA and the second gRNA target different genes or their regulatory elements. In some embodiments, the first gRNA targets the Enh1 enhancer, while the second gRNA targets a gene selected from the following: L-HBs promoter, M-HBs promoter, S-HBs promoter, X promoter, basic core promoter, S gene promoter, and Enh2 enhancer. In some embodiments, the first gRNA targets the Enh1 enhancer, while the second gRNA targets L-HBs. In some embodiments, the first and second gRNAs target a combination of two genes selected from the combinations of genes listed in Table 1. In some embodiments, the first and second gRNAs are each independently selected from any gRNA described herein.
[0288] In some embodiments, the combination of gRNAs comprises at least three gRNAs targeting at least three different genes or their regulatory elements. In some embodiments, the combination of gRNAs comprises a first gRNA targeting a first gene, a second gRNA targeting a second gene, and a third gRNA targeting a third gene. In some embodiments, the combination of gRNAs comprises at least three gRNAs targeting at least three different genes or their regulatory elements. In some embodiments, the combination of gRNAs comprises a first gRNA targeting a first gene or its regulatory element, a second gRNA targeting a second gene or its regulatory element, and a third gRNA targeting a third gene or its regulatory element. In some implementations, the first gRNA targets genes or regulatory elements thereof selected from the following: polymerase genes, S family genes, X genes, core family genes, pre-S1 promoters, pre-S2 promoters, X promoters, basic core promoters, Enh1 enhancers, Enh2 enhancers, transcript processing control regions, and any coding regions within the HBV genome; the second gRNA targets genes or regulatory elements thereof selected from the following: polymerase genes, S family genes, X genes, core family genes, pre-S1 promoters, pre-S2 promoters, X promoters, basic core promoters, Enh1 enhancers, Enh2 enhancers, transcript processing control regions, and any coding regions within the HBV genome; and the third gRNA targets genes or regulatory elements thereof selected from the following: polymerase genes, S family genes, X genes, core family genes, pre-S1 promoters, pre-S2 promoters, X promoters, basic core promoters, Enh1 enhancers, Enh2 enhancers, transcript processing control regions, and any coding regions within the HBV genome, and the first gRNA, second gRNA, and third gRNA target different genes or regulatory elements thereof. In some embodiments, the first gRNA targets the Enh1 enhancer, the second gRNA targets a gene selected from the L-HBs promoter, M-HBs promoter, S-HBs promoter, X promoter, basic core promoter, S gene promoter, Enh1 enhancer, and Enh2 enhancer, and the third gRNA targets a gene selected from the L-HBs promoter, M-HBs promoter, S-HBs promoter, X promoter, basic core promoter, S gene promoter, Enh1 enhancer, and Enh2 enhancer, and the second and third gRNAs target different genes. In some embodiments, the first gRNA targets the Enh1 enhancer, the second gRNA targets the L-HBs promoter, and the third gRNA targets a gene selected from the M-HBs promoter, S-HBs promoter, X promoter, basic core promoter, S gene promoter, and Enh2 enhancer. In some embodiments, the first, second, and third gRNAs target a combination of three genes selected from the combinations of genes listed in Table 2.In some implementations, the first gRNA, the second gRNA, and the third gRNA are each independently selected from any gRNA described herein.
[0289] In some embodiments, the combination of gRNAs comprises at least four gRNAs targeting at least four different genes or their regulatory elements. In some embodiments, the combination of gRNAs comprises a first gRNA targeting a first gene, a second gRNA targeting a second gene, a third gRNA targeting a third gene, and a fourth gRNA targeting a fourth gene. In some embodiments, the combination of gRNAs comprises at least four gRNAs targeting at least four different genes or their regulatory elements. In some embodiments, the combination of gRNAs comprises a first gRNA targeting a first gene or its regulatory element, a second gRNA targeting a second gene or its regulatory element, a third gRNA targeting a third gene or its regulatory element, and a fourth gRNA targeting a third gene or its regulatory element. In some implementations, the first gRNA target is selected from the following genes or their regulatory elements: polymerase genes, S family genes, X genes, core family genes, pre-S1 promoters, pre-S2 promoters, X promoters, basic core promoters, Enh1 enhancers, Enh2 enhancers, transcript processing control regions, and any coding regions within the HBV genome; the second gRNA target is selected from the following genes or their regulatory elements: polymerase genes, S family genes, X genes, core family genes, pre-S1 promoters, pre-S2 promoters, X promoters, basic core promoters, Enh1 enhancers, Enh2 enhancers, transcript processing control regions, and any coding regions within the HBV genome; the third gRNA target is selected from the following genes or their regulatory elements. Regulatory elements: polymerase genes, S family genes, X genes, core family genes, pre-S1 promoter, pre-S2 promoter, X promoter, basic core promoter, Enh1 enhancer, Enh2 enhancer, transcript processing control region, and any coding region within the HBV genome. The fourth gRNA targets genes or regulatory elements thereof selected from the following: polymerase genes, S family genes, X genes, core family genes, pre-S1 promoter, pre-S2 promoter, X promoter, basic core promoter, Enh1 enhancer, Enh2 enhancer, transcript processing control region, and any coding region within the HBV genome, and the first gRNA, second gRNA, third gRNA, and fourth gRNA target different genes or regulatory elements thereof.In some embodiments, the first gRNA targets the Enh1 enhancer; the second gRNA targets a gene selected from the following: L-HBs promoter, M-HBs promoter, S-HBs promoter, X promoter, basic core promoter, S gene promoter, Enh1 enhancer, and Enh2 enhancer; the third gRNA targets a gene selected from the following: L-HBs promoter, M-HBs promoter, S-HBs promoter, X promoter, basic core promoter, S gene promoter, Enh1 enhancer, and Enh2 enhancer; and the fourth gRNA targets a gene selected from the following: L-HBs promoter, M-HBs promoter, S-HBs promoter, X promoter, basic core promoter, S gene promoter, Enh1 enhancer, and Enh2 enhancer. Furthermore, the first, second, third, and fourth gRNAs target different genes or their regulatory elements. In some embodiments, the first, second, third, and fourth gRNAs target four genes or combinations of their regulatory elements selected from combinations of genes or their regulatory elements listed in Table 2. In some implementations, the first gRNA, second gRNA, third gRNA, and fourth gRNA are each independently selected from any gRNA described herein.
[0290] In some embodiments, the combination of gRNAs comprises at least five gRNAs targeting at least five different genes or their regulatory elements. In some embodiments, the combination of gRNAs comprises at least six gRNAs targeting at least six different genes and / or their regulatory elements. In some embodiments, the first, second, third, fourth, fifth, and / or sixth genes or their regulatory elements are different. C. Engineered zinc finger protein (eZFP)
[0291] In some aspects, this document provides zinc finger proteins (ZFPs), such as engineered zinc finger proteins (eZFPs). In some embodiments, eZFPs are capable of binding to or to target sites in regulatory elements of genes within the HBV gene or hepatitis B virus sequence. In some aspects, eZFPs can facilitate specific targeting of effector domains to achieve transcriptional repression of genes or regulatory elements. In some embodiments, this document provides an epigenetically modified DNA targeting system comprising a fusion protein containing eZFPs and one or more other elements (such as effector domains for transcriptional repression). Thus, in some aspects, such as in combination with compositions and methods for treating HBV-related diseases or disorders (e.g., HBV infection, liver disease, or cancer), eZFPs contribute to a reduction in the expression of HBV genes or regulatory elements.
[0292] In some embodiments, zinc finger proteins (ZFPs), zinc finger DNA-binding proteins, or zinc finger DNA-binding domains are domains within proteins or larger proteins that bind DNA in a sequence-specific manner via one or more zinc fingers. A zinc finger is an amino acid sequence region within a binding domain that has a structure stabilized by zinc ion coordination. ZFPs include artificially created or engineered ZFPs (eZFPs) containing ZFP domains that target specific DNA sequences (typically 9-18 nucleotides long) and are generated through the assembly of individual zinc fingers. ZFPs include those where a single finger domain is approximately 30 amino acids long and contains an α-helix containing two invariant histidine residues coordinated to two cysteine residues via a single β-turn with zinc, and where the ZFP has two, three, four, five, or six fingers. Typically, the sequence specificity of a ZFP can be altered by amino acid substitutions at four helical positions (−1, 2, 3, and 6) on the zinc finger recognition helix (also known as the zinc finger recognition region). Therefore, for example, ZFP or molecules containing ZFP (such as fusion proteins) may not exist naturally, for example, and may be engineered to bind to selected target sites.
[0293] In some implementations, zinc fingers can be custom-designed (i.e., user-designed) and / or obtained from commercial sources. Various methods for designing zinc finger proteins are available. For example, methods for designing zinc finger proteins to bind to target DNA sequences are described, for example, in the following: Liu, Q. et al., PNAS, 94(11):5525-30 (1997); Wright, DA et al., Nat. Protoc., 1(3):1637-52 (2006); Gersbach, CA et al., Acc. Chem. Res., 47(8):2309-18 (2014); Bhakta MS et al., Methods Mol. Biol., 649:3-30 (2010); and Gaj et al., Trends Biotechnol, 31(7):397-405 (2013). Furthermore, various network-based tools for designing zinc finger proteins that bind to target DNA sequences are publicly available. See, for example, the zinc finger tool design website from Scripps, available at scripps.edu / barbas / zfdesign / zfdesignhome.php. Various commercial services for designing zinc finger proteins that bind to target DNA sequences are also available. See, for example, commercially available services or kits from CreativeBiolabs (creative-biolabs.com / Design-and-Synthesis-of-Artificial-Zinc-Finger-Proteins.html), the Zinc Finger Consortium Modular Assembly Kit from Addgene (addgene.org / kits / zfc-modular-assembly / ), or the CompoZr custom ZFN service from Sigma Aldrich (sigmaaldrich.com / life-science / zinc-finger-nuclease-technology / custom-zfn.html).
[0294] In some embodiments, this document provides an epigenetically modified DNA targeting system comprising a fusion protein containing an eZFP and one or more other elements (such as an effector domain for transcriptional repression). In some embodiments, at least one DNA-binding domain of the epigenetically modified DNA targeting system comprises an engineered zinc finger protein (eZFP). In some embodiments, the epigenetically modified DNA targeting system comprises an engineered zinc finger protein (eZFP) that binds to a target site in one or more HBV genes or their regulatory elements. The target site targeted by any provided eZFP can be any target site as described herein, such as any target site described in Section IA.
[0295] In some embodiments, the target site of the eZFP (e.g., the eZFP contained in a fusion protein of an epigenetic modification DNA targeting system) provided herein is in one or more HBV genes or their regulatory elements. In some embodiments, the target site is in a CpG island (e.g., CpG island 1, CpG island 2, CpG island 3) of the HBV genome. In some embodiments, the target site is in CpG island 2 of the HBV genome. In some embodiments, the target site is within a target region spanning 1033 bp–1749 bp across the HBV genome, corresponding to the location shown in reference SEQ ID NO: 650. In some embodiments, the target site is within a target region spanning 300 base pairs (bp), 250 bp, 200 bp, 150 bp, 140 bp, 130 bp, 120 bp, 110 bp, or 100 bp upstream of the HBx start codon. In some embodiments, the target site is within the target region sequence, which corresponds to a 1250-1374 bp sequence in the HBV genome shown in SEQ ID NO: 650. In some embodiments, the target region has the sequence shown in SEQ ID NO: 1068. In some embodiments, the target site is within the target region sequence, which corresponds to a 1255-1302 bp sequence in the HBV genome shown in SEQ ID NO: 650. In some embodiments, the target region has the sequence shown in SEQ ID NO: 1069. In some embodiments, the target site is within the target region sequence, which corresponds to a 1260-1300 bp sequence in the HBV genome shown in SEQ ID NO: 650. In some embodiments, the target region has the sequence shown in SEQ ID NO: 1070.
[0296] In some embodiments, the target site of the eZFP provided herein (e.g., the eZFP contained in a fusion protein of an epigenetic modification DNA targeting system) comprises a nucleotide sequence shown in any one of SEQ ID NO: 1028-1055, a continuous portion of at least 12 nt thereof, or a complementary sequence of any one of the foregoing. In some embodiments, the target site of the eZFP provided herein comprises a nucleotide sequence shown in any one of SEQ ID NO: 1028-1055. In some embodiments, the target site of the eZFP provided herein comprises a nucleotide sequence shown in any one of SEQ ID NO: 1045, 1046, or 1052, a continuous portion of at least 12 nt thereof, or a complementary sequence of any one of the foregoing. In some embodiments, the target site comprises a nucleotide sequence shown in any one of SEQ ID NO: 1045, 1046, or 1052.
[0297] In some embodiments, the target site is contained in double-stranded DNA, such as an HBV sequence integrated into human genomic DNA. In some embodiments, the target site is contained in a covalently closed circular (cccDNA) HBV sequence. In some embodiments, the target site is contained in a relaxed circular (rcDNA) HBV sequence. In some embodiments, the eZFP is capable of binding to the target site. In some embodiments, the eZFP binds to the target site. In some embodiments, the binding is target-specific. For example, in some embodiments, the eZFP binds to the target site and not to other sites containing different sequences. For example, in some embodiments, the individual eZFP disclosed herein binds to any one of the target sites shown in SEQ ID NO: 1028-1055 and not to different target sites. In some embodiments, the individual eZFP disclosed herein binds to any one of the target sites shown in SEQ ID NO: 1045, 1046, or 1052 and not to different target sites. In some embodiments, the target site of the eZFP provided herein contains the sequences shown in Table 7. Table 7. eZFP target sequences
[0298] In some embodiments, the target site of the eZFP provided herein (e.g., the eZFP contained in a fusion protein of an epigenetic modification DNA targeting system) comprises the nucleotide sequence shown in SEQ ID NO: 1045. In some embodiments, it comprises at least a continuous 12 nt portion, or a complementary sequence to any of the foregoing. In some embodiments, the target site of the eZFP provided herein comprises the sequence shown in SEQ ID NO: 1045.
[0299] In some embodiments, the target site of the eZFP provided herein (e.g., the eZFP contained in a fusion protein of an epigenetic modification DNA targeting system) comprises the nucleotide sequence shown in SEQ ID NO: 1046. In some embodiments, it comprises at least a continuous 12 nt portion, or a complementary sequence to any of the foregoing. In some embodiments, the target site of the eZFP provided herein comprises the sequence shown in SEQ ID NO: 1046.
[0300] In some embodiments, the target site of the eZFP provided herein (e.g., the eZFP contained in a fusion protein of an epigenetic modification DNA targeting system) comprises the nucleotide sequence shown in SEQ ID NO: 1052. In some embodiments, it comprises at least a continuous 12 nt portion, or a complementary sequence to any of the foregoing. In some embodiments, the target site of the eZFP provided herein comprises the sequence shown in SEQ ID NO: 1052.
[0301] In some embodiments, the eZFP comprises multiple zinc fingers. In some embodiments, each zinc finger comprises a recognition region. In some embodiments, the recognition regions together facilitate sequence-specific binding of the eZFP, such as sequence-specific binding to a specific target site. In some embodiments, the eZFP comprises six zinc fingers, numbered F1 to F6 sequentially from the N-terminus to the C-terminus, each containing a corresponding recognition region F1 to F6 that facilitates sequence-specific binding to a specific target site.
[0302] In some embodiments, the characteristics of the eZFP targeting the specific target sites provided herein are shown in Table E4. In some embodiments, the eZFP comprises six zinc fingers, sequentially designated F1 to F6 from the N-terminus to the C-terminus, each zinc finger containing a corresponding recognition region F1-F6, as shown in Table E4. In some embodiments, recognition regions F1-F6 facilitate specific binding to the target site sequence indicated in Table E4. In some embodiments, the eZFP comprises an amino acid sequence containing the recognition region, as shown in Table E4. In some embodiments, the eZFP may be encoded by a DNA sequence as shown in Table 8. Table 8. eZFP DNA Sequence
[0303] In some implementations, eZFP (e.g., eZFP contained in a fusion protein of an epigenetic modification DNA targeting system) comprises a zinc finger protein containing six zinc fingers, sequentially denoted as F1 to F6 from the N-terminus to the C-terminus. In some implementations, the amino acid sequence of each zinc finger recognition region is as follows: 1) F1: SEADRSR (SEQ ID NO: 720) F2: DRSNLTR (SEQ ID NO: 721) F3: QSSDLSR (SEQ ID NO: 722) F4: YHWYLKK (SEQ ID NO: 723) F5: RSDSLSV (SEQ ID NO: 724) F6: QNANRKT (SEQ ID NO: 725); 2) F1: RSDVLST (SEQ ID NO: 726) F2: DNSSRTR (SEQ ID NO: 727) F3: RPYTLRL (SEQ ID NO: 728) F4: DSSHRTR (SEQ ID NO: 729) F5: RSDHLSQ (SEQ ID NO: 730) F6: DSSHRTR (SEQ ID NO: 731); 3) F1: RSDHLSQ (SEQ ID NO: 729) F5: RSDHLSQ (SEQ ID NO: 720) F6: DSSHRTR (SEQ ID NO: 721); 732) F2: QSADRTK (SEQ ID NO: 733) F3: RSDHLSQ (SEQ ID NO: 734) F4: RRSDLKR (SEQ ID NO: 735) F5: RSDHLSR (SEQ ID NO: 736) F6: QSSDLRR (SEQ ID NO: 737); 4) F1: RSDNLSE (SEQ ID NO: 732) 738) F2: TSSNRKT (SEQ ID NO: 739) F3: DRSHLTR (SEQ ID NO: 740) F4: RSDALTQ (SEQ ID NO: 741) F5: DRSALAR (SEQ ID NO: 742) F6: RRFTLSK (SEQ ID NO: 743); 5) F1: RSDHLSE (SEQ ID NO: 743) 744) F2: QYSGRYY (SEQ ID NO: 745) F3: HGQTLNE (SEQ ID NO:746) F4: QSGNLAR (SEQ ID NO: 747) F5: RSDSLLR (SEQ ID NO: 748) F6: CREYRGK (SEQ ID NO: 749);6) F1: QSANRTT (SEQ ID NO: 750) F2: RSANLTR (SEQ ID NO: 751) F3: RSDVLSE (SEQ ID NO: 752) F4: TSGHLSR (SEQ ID NO: 753) F5: QSSDLSR (SEQ ID NO: 754) F6: QWSTRKR (SEQ ID NO: 755) 7) F1: QSGNLAR (SEQ ID NO: 756) F2: ATCCLAH (SEQ ID NO: 757) F3: RWQYLPT (SEQ ID NO: 758) F4: DRSALAR (SEQ ID NO: 759) F5: RSDNLSE (SEQ ID NO: 760) F6: KRCNLRC (SEQ ID NO: 761) ;8) F1: NPANLTR (SEQ ID NO: 762) F2: QNATRTK (SEQ ID NO: 763) F3: QSGHLAR (SEQ ID NO: 764) F4: NRHDRAK (SEQ ID NO: 765) F5: RSDHLSE (SEQ ID NO: 766) F6: QRRSRYK (SEQ ID NO: 767) ;9) F1: QSSDLSR (SEQ ID NO: 768) F2: HRSTRNR (SEQ ID NO: 769) F3: RSDVLSA (SEQ ID NO: 770)F4:DSRTRKN(SEQ ID NO: 771)F5:QSGSLTR(SEQID NO: 772)F6:DQSGLAH(SEQ ID NO: 773);10) F1:QNPAQWR(SEQ ID NO: 774)F2:RSADLSR(SEQ ID NO: 775)F3:TSGSLSR(SEQ ID NO: 776)F4:RSDHLSR(SEQ ID NO: 777)F5:RSDSLLR(SEQ ID NO: 778)F6:QSYDRFQ(SEQ ID NO: 779);11) F1:TSGSLSR(SEQ IDNO: 780) F2: RSDHLSR (SEQ ID NO: 781) F3: RDSLLR (SEQ ID NO: 782) F4: QSYDRFQ (SEQ ID NO: 783) F5: RSDNLST (SEQ ID NO: 784) F6: DNRDRIK (SEQ ID NO: 785)12) F1: DRSNLSR (SEQ ID NO: 786) F2: LRQNLIM (SEQ ID NO: 787) F3: ERGTLAR (SEQ ID NO: 788) F4: RSDALTQ (SEQ ID NO: 789) F5: RSDSLSQ (SEQ ID NO: 790) F6: RKADRTR (SEQ ID NO: 791) 13) F1: QYCCLTN (SEQ ID NO: 792) F2: TSGNLTR (SEQ ID NO: 793) F3: QSSDLSR (SEQ ID NO: 794) F4: FRYYLKR (SEQ ID NO: 795) F5: QSGDLTR (SEQ ID NO: 796) F6: DKGNLTK (SEQ ID NO: 797) ;14) F1: TSGSLSR (SEQ ID NO: 798) F2: RSDNLTT (SEQ ID NO: 799) F3: QSGNLAR (SEQ ID NO: 800) F4: DRTTLMR (SEQ ID NO: 801) F5: QSGHLAR (SEQ ID NO: 802) F6: QLTHLNS (SEQ ID NO: 803) ;15) F1: IKHDLHR (SEQ ID NO: 804) F2: RSANLTR (SEQ ID NO: 805) F3: RSDNLAR (SEQ ID NO: 806)F4:QNVSRPR(SEQ ID NO: 807)F5:RSDDLSK(SEQ ID NO: 808)F6:DSSHRTR(SEQ ID NO: 809);16) F1:RSDNLAR(SEQ ID NO: 810)F2:QNVSRPR(SEQ IDNO: 811)F3:RSDDLSK(SEQ ID NO: 812)F4:DSSHRTR(SEQ ID NO: 813)F5:TSSNRKT(SEQ IDNO: 814)F6:AQWTRAC(SEQ ID NO: 815);17) F1: RSDDLSK(SEQ ID NO: 816) F2: DSSHRTR (SEQ ID NO: 817) F3: TSNRKT (SEQ ID NO: 818) F4: AQWTRAC (SEQ ID NO: 819) F5: RKQTRTT (SEQ ID NO: 820) F6: HRSSLRR (SEQ ID NO: 821)18) F1: QSAHRKN (SEQ ID NO: 822) F2: TSSNRKT (SEQ ID NO: 823) F3: RSDNLSA (SEQ ID NO: 824) F4: RNNDRKT (SEQ ID NO: 825) F5: TSGSLSR (SEQ ID NO: 826) F6: QAGHLAK (SEQ ID NO: 827) 19) F1: RSDHLSQ (SEQ ID NO: 828) F2: ASSTRTK (SEQ ID NO: 829) F3: RSDDLTR (SEQ ID NO: 830) F4: QKSNLSS (SEQ ID NO: 831) F5: QSANRTT (SEQ ID NO: 832) F6: QNATRTK (SEQ ID NO: 833) ;20) F1: RSDTLSE (SEQ ID NO: 834) F2: RRWTLVG (SEQ ID NO: 835) F3: DRSNLSR (SEQ ID NO: 836) F4: QSGDLTR (SEQ ID NO: 837) F5: QSSDLSR (SEQ ID NO: 838) F6: YHWYLKK (SEQ ID NO: 839) ;21) F1: RSANLAR (SEQ ID NO: 840) F2: RSDNLRE (SEQ ID NO: 841) F3: RPYTLRL (SEQ ID NO: 842)F4:HRSNLNK(SEQ ID NO: 843)F5:QSGSLTR(SEQ ID NO: 844)F6:TSANLSR(SEQID NO: 845);m22) F1:RSDDLVR(SEQ ID NO: 846)F2:TSGSLVR(SEQ ID NO: 847)F3:RSDKLVR(SEQ ID NO: 848)F4:RSDELVR(SEQ ID NO: 849)F5:TSHSLTE(SEQ ID NO: 850)F6:RADNLTE(SEQ ID NO: 851);23) F1:ERSHLRE(SEQ ID NO: 852) F2: THSHSLTE (SEQ IDNO: 853) F3: QAGHLAS (SEQ ID NO: 854) F4: THSHSLTE (SEQ ID NO: 855) F5: DPGHLVR (SEQ IDNO: 856) F6: TGSGNLVR (SEQ ID NO: 857)24) F1:RADNLTE(SEQ ID NO: 858)F2:TSGSLVR(SEQ ID NO: 859)F3:RKDNLKN(SEQ ID NO: 860)F4:QSSSLVR(SEQ ID NO: 861)F5:RSDKLVR(SEQ ID NO: 862)F6:DSGNLRV(SEQ ID NO: 863);25) F1:QSSSLVR(SEQ ID NO:864)F2:QSGDLRR(SEQ ID NO: 865)F3:RSDERKR(SEQ ID NO: 866)F4:HRTTLTN(SEQ ID NO:867)F5:RSDHLTN(SEQ ID NO: 868) F6: TGSGELVR (SEQ ID NO: 869) ;26) F1: QSGDLRR (SEQ ID NO: 870) F2: RSDERKR (SEQ ID NO: 871) F3: HRTTLTN (SEQ ID NO: 872) F4: RSDHLTN (SEQ ID NO: 873) F5: TGSGELVR (SEQ ID NO: 874) F6: RSDDLVR (SEQ ID NO: 875) ;27) F1: QRAHLER (SEQ ID NO: 876) F2: QLAHLRA (SEQ ID NO: 877) F3: DPGHLVR (SEQ ID NO: F1: 882)F2:RSDDLVR(SEQ ID NO: 883)F3:THLDLIR(SEQ ID NO: 884)F4:TSGNLTE(SEQ ID NO: 885)F5:RRSACRR(SEQ ID NO: 886)F6:RNDTLTE(SEQ ID NO: 887)。;
[0304] In some embodiments, the eZFP (e.g., the eZFP contained in a fusion protein of an epigenetic modification DNA targeting system) comprises the sequence shown in any one of SEQ ID NO: 692-719, or a portion thereof, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP is encoded by the sequence shown in any one of SEQ ID NO: 888-915, or a portion thereof, or a nucleotide sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it.
[0305] In some embodiments, this document provides eZFP (e.g., eZFP contained in a fusion protein of an epigenetic modification DNA targeting system), such as eZFP_1 as described herein. In some embodiments, eZFP targets a target site comprising the nucleotide sequence shown in SEQ ID NO: 1028, a continuous portion thereof of at least 12 nt, or a complementary sequence to any of the foregoing. In some embodiments, eZFP targets a target site comprising the nucleotide sequence shown in SEQ ID NO: 1028. In some embodiments, the eZFP comprises six zinc fingers, numbered F1 to F6 sequentially from the N-terminus to the C-terminus, each zinc finger containing a corresponding zinc finger recognition region F1 to F6, and the amino acid sequence of each zinc finger recognition region is as follows: F1: SEADRSR (SEQ ID NO: 720) F2: DRSNLTR (SEQ ID NO: 721) F3: QSSDLSR (SEQ ID NO: 722) F4: YHWYLKK (SEQ ID NO: 723) F5: RSDSLSV (SEQ ID NO: 724) F6: QNANRKT (SEQ ID NO: 725). In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 692, or a portion thereof, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 692. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 888 or a nucleotide sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 888.
[0306] In some embodiments, this document provides eZFP (e.g., eZFP contained in a fusion protein of an epigenetic modification DNA targeting system), such as eZFP_2 as described herein. In some embodiments, eZFP targets a target site comprising the nucleotide sequence shown in SEQ ID NO: 1029, a continuous portion thereof of at least 12 nt, or a complementary sequence to any of the foregoing. In some embodiments, eZFP targets a target site comprising the nucleotide sequence shown in SEQ ID NO: 1029. In some embodiments, the eZFP comprises six zinc fingers, numbered F1 to F6 sequentially from the N-terminus to the C-terminus, each zinc finger containing a corresponding zinc finger recognition region F1 to F6, and the amino acid sequence of each zinc finger recognition region is as follows: F1: RSDVLST (SEQ ID NO: 726) F2: DNSSRTR (SEQ ID NO: 727) F3: RPYTLRL (SEQ ID NO: 728) F4: DSSHRTR (SEQ ID NO: 729) F5: RSDHLSQ (SEQ ID NO: 730) F6: DSSHRTR (SEQ ID NO: 731). In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 693, or a portion thereof, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 693. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 889 or a nucleotide sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 889.
[0307] In some embodiments, this document provides eZFPs (e.g., eZFPs contained in fusion proteins of epigenetic modification DNA targeting systems), such as eZFP_3 as described herein. In some embodiments, eZFPs target target sites comprising the nucleotide sequence shown in SEQ ID NO: 1030, a continuous portion thereof of at least 12 nt, or a complementary sequence to any of the foregoing. In some embodiments, eZFPs target target sites comprising the nucleotide sequence shown in SEQ ID NO: 1030. In some embodiments, the eZFP comprises six zinc fingers, numbered F1 to F6 sequentially from the N-terminus to the C-terminus, each zinc finger containing a corresponding zinc finger recognition region F1 to F6, and the amino acid sequence of each zinc finger recognition region is as follows: F1: RSDHLSQ (SEQ ID NO: 732) F2: QSADRTK (SEQ ID NO: 733) F3: RSDHLSQ (SEQ ID NO: 734) F4: RRSDLKR (SEQ ID NO: 735) F5: RSDHLSR (SEQ ID NO: 736) F6: QSSDLRR (SEQ ID NO: 737). In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 694, or a portion thereof, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 694. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 890 or a nucleotide sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 890.
[0308] In some embodiments, this document provides eZFPs (e.g., eZFPs contained in fusion proteins of epigenetic modification DNA targeting systems), such as eZFP_4 as described herein. In some embodiments, eZFPs target target sites comprising the nucleotide sequence shown in SEQ ID NO: 1031, a continuous portion thereof of at least 12 nt, or a complementary sequence to any of the foregoing. In some embodiments, eZFPs target target sites comprising the nucleotide sequence shown in SEQ ID NO: 1031. In some embodiments, the eZFP comprises six zinc fingers, numbered F1 to F6, sequentially from the N-terminus to the C-terminus. Each zinc finger contains a corresponding zinc finger recognition region F1 to F6, and the amino acid sequence of each zinc finger recognition region is as follows: F1: RSDNLSE (SEQ ID NO: 738) F2: TSSNRKT (SEQ ID NO: 739) F3: DRSHLTR (SEQ ID NO: 740) F4: RSDALTQ (SEQ ID NO: 741) F5: DRSALAR (SEQ ID NO: 742) F6: RRFTLSK (SEQ ID NO: 743). In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 695, or a portion thereof, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 695. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 891 or a nucleotide sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 891.
[0309] In some embodiments, this document provides eZFPs (e.g., eZFPs contained in fusion proteins of epigenetic modification DNA targeting systems), such as eZFP_5 as described herein. In some embodiments, eZFPs target target sites comprising the nucleotide sequence shown in SEQ ID NO: 1032, a continuous portion thereof of at least 12 nt, or a complementary sequence to any of the foregoing. In some embodiments, eZFPs target target sites comprising the nucleotide sequence shown in SEQ ID NO: 1032. In some embodiments, the eZFP comprises six zinc fingers, numbered F1 to F6 sequentially from the N-terminus to the C-terminus, each zinc finger containing a corresponding zinc finger recognition region F1 to F6, and the amino acid sequence of each zinc finger recognition region is as follows: F1: RSDHLSE (SEQ ID NO: 744) F2: QYSGRYY (SEQ ID NO: 745) F3: HGQTLNE (SEQ ID NO: 746) F4: QSGNLAR (SEQ ID NO: 747) F5: RSDSLLR (SEQ ID NO: 748) F6: CREYRGK (SEQ ID NO: 749). In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 696, or a portion thereof, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 696. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 892 or a nucleotide sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 892.
[0310] In some embodiments, this document provides eZFP (e.g., eZFP contained in a fusion protein of an epigenetic modification DNA targeting system), such as eZFP_6 as described herein. In some embodiments, eZFP targets a target site comprising the nucleotide sequence shown in SEQ ID NO: 1033, a continuous portion thereof of at least 12 nt, or a complementary sequence to any of the foregoing. In some embodiments, eZFP targets a target site comprising the nucleotide sequence shown in SEQ ID NO: 1033. In some embodiments, the eZFP comprises six zinc fingers, numbered F1 to F6 sequentially from the N-terminus to the C-terminus, each zinc finger containing a corresponding zinc finger recognition region F1 to F6, and the amino acid sequence of each zinc finger recognition region is as follows: F1: QSANRTT (SEQ ID NO: 750) F2: RSANLTR (SEQ ID NO: 751) F3: RSDVLSE (SEQ ID NO: 752) F4: TSGHLSR (SEQ ID NO: 753) F5: QSSDLSR (SEQ ID NO: 754) F6: QWSTRKR (SEQ ID NO: 755). In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 697, or a portion thereof, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 697. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 893 or a nucleotide sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 893.
[0311] In some embodiments, this document provides eZFPs (e.g., eZFPs contained in fusion proteins of epigenetic modification DNA targeting systems), such as eZFP_7 as described herein. In some embodiments, eZFPs target target sites comprising the nucleotide sequence shown in SEQ ID NO: 1034, a continuous portion thereof of at least 12 nt, or a complementary sequence to any of the foregoing. In some embodiments, eZFPs target target sites comprising the nucleotide sequence shown in SEQ ID NO: 1034. In some embodiments, the eZFP comprises six zinc fingers, numbered F1 to F6, sequentially from the N-terminus to the C-terminus. Each zinc finger contains a corresponding zinc finger recognition region F1 to F6, and the amino acid sequence of each zinc finger recognition region is as follows: F1: QSGNLAR (SEQ ID NO: 756) F2: ATCCLAH (SEQ ID NO: 757) F3: RWQYLPT (SEQ ID NO: 758) F4: DRSALAR (SEQ ID NO: 759) F5: RSDNLSE (SEQ ID NO: 760) F6: KRCNLRC (SEQ ID NO: 761). In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 698, or a portion thereof, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 698. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 894 or a nucleotide sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 894.
[0312] In some embodiments, this document provides eZFPs (e.g., eZFPs contained in fusion proteins of epigenetic modification DNA targeting systems), such as eZFP_8 as described herein. In some embodiments, eZFPs target target sites comprising the nucleotide sequence shown in SEQ ID NO: 1035, a continuous portion thereof of at least 12 nt, or a complementary sequence to any of the foregoing. In some embodiments, eZFPs target target sites comprising the nucleotide sequence shown in SEQ ID NO: 1035. In some embodiments, the eZFP comprises six zinc fingers, numbered F1 to F6, sequentially from the N-terminus to the C-terminus. Each zinc finger contains a corresponding zinc finger recognition region F1 to F6, and the amino acid sequence of each zinc finger recognition region is as follows: F1: NPANLTR (SEQ ID NO: 762) F2: QNATRTK (SEQ ID NO: 763) F3: QSGHLAR (SEQ ID NO: 764) F4: NRHDRAK (SEQ ID NO: 765) F5: RSDHLSE (SEQ ID NO: 766) F6: QRRSRYK (SEQ ID NO: 767). In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 699, or a portion thereof, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 699. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 895 or a nucleotide sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 895.
[0313] In some embodiments, this document provides eZFP (e.g., eZFP contained in a fusion protein of an epigenetic modification DNA targeting system), such as eZFP_9 as described herein. In some embodiments, eZFP targets a target site comprising the nucleotide sequence shown in SEQ ID NO: 1036, a continuous portion thereof of at least 12 nt, or a complementary sequence to any of the foregoing. In some embodiments, eZFP targets a target site comprising the nucleotide sequence shown in SEQ ID NO: 1036. In some embodiments, the eZFP comprises six zinc fingers, numbered F1 to F6 sequentially from the N-terminus to the C-terminus, each zinc finger containing a corresponding zinc finger recognition region F1 to F6, and the amino acid sequence of each zinc finger recognition region is as follows: F1: QSSDLSR (SEQ ID NO: 768) F2: HRSTRNR (SEQ ID NO: 769) F3: RSDVLSA (SEQ ID NO: 770) F4: DSRTRKN (SEQ ID NO: 771) F5: QSGSLTR (SEQ ID NO: 772) F6: DQSGLAH (SEQ ID NO: 773). In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 700, or a portion thereof, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 700. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 896 or a nucleotide sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 896.
[0314] In some embodiments, this document provides eZFPs (e.g., eZFPs contained in fusion proteins of epigenetic modification DNA targeting systems), such as eZFP_10 described herein. In some embodiments, eZFPs target target sites comprising the nucleotide sequence shown in SEQ ID NO: 1037, a continuous portion thereof of at least 12 nt, or a complementary sequence to any of the foregoing. In some embodiments, eZFPs target target sites comprising the nucleotide sequence shown in SEQ ID NO: 1037. In some embodiments, the eZFP comprises six zinc fingers, numbered F1 to F6 sequentially from the N-terminus to the C-terminus, each zinc finger containing a corresponding zinc finger recognition region F1 to F6, and the amino acid sequence of each zinc finger recognition region is as follows: F1: QNPAQWR (SEQ ID NO: 774) F2: RSADLSR (SEQ ID NO: 775) F3: TSGSLSR (SEQ ID NO: 776) F4: RSDHLSR (SEQ ID NO: 777) F5: RSDSLLR (SEQ ID NO: 778) F6: QSYDRFQ (SEQ ID NO: 779). In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 701, or a portion thereof, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 701. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 897 or a nucleotide sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 897.
[0315] In some embodiments, this document provides eZFPs (e.g., eZFPs contained in fusion proteins of epigenetic modification DNA targeting systems), such as eZFP_11 described herein. In some embodiments, eZFPs target target sites comprising the nucleotide sequence shown in SEQ ID NO: 1038, a continuous portion thereof of at least 12 nt, or a complementary sequence to any of the foregoing. In some embodiments, eZFPs target target sites comprising the nucleotide sequence shown in SEQ ID NO: 1038. In some embodiments, the eZFP comprises six zinc fingers, numbered F1 to F6 sequentially from the N-terminus to the C-terminus, each zinc finger containing a corresponding zinc finger recognition region F1 to F6, and the amino acid sequence of each zinc finger recognition region is as follows: F1: TSGSLSR (SEQ ID NO: 780) F2: RSDHLSR (SEQ ID NO: 781) F3: RSDSLLR (SEQ ID NO: 782) F4: QSYDRFQ (SEQ ID NO: 783) F5: RSDNLST (SEQ ID NO: 784) F6: DNRDRIK (SEQ ID NO: 785). In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 702, or a portion thereof, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 702. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 898 or a nucleotide sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 898.
[0316] In some embodiments, this document provides eZFPs (e.g., eZFPs contained in fusion proteins of epigenetic modification DNA targeting systems), such as eZFP_12 as described herein. In some embodiments, eZFPs target target sites comprising the nucleotide sequence shown in SEQ ID NO: 1039, a continuous portion thereof of at least 12 nt, or a complementary sequence to any of the foregoing. In some embodiments, eZFPs target target sites comprising the nucleotide sequence shown in SEQ ID NO: 1039. In some embodiments, the eZFP comprises six zinc fingers, numbered F1 to F6 sequentially from the N-terminus to the C-terminus, each zinc finger containing a corresponding zinc finger recognition region F1 to F6, and the amino acid sequence of each zinc finger recognition region is as follows: F1: DRSNLSR (SEQ ID NO: 786) F2: LRQNLIM (SEQ ID NO: 787) F3: ERGTLAR (SEQ ID NO: 788) F4: RSDALTQ (SEQ ID NO: 789) F5: RSDSLSQ (SEQ ID NO: 790) F6: RKADRTR (SEQ ID NO: 791). In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 703, or a portion thereof, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 703. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 899 or a nucleotide sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 899.
[0317] In some embodiments, this document provides eZFPs (e.g., eZFPs contained in fusion proteins of epigenetic modification DNA targeting systems), such as eZFP_13 described herein. In some embodiments, eZFPs target target sites comprising the nucleotide sequence shown in SEQ ID NO: 1040, a continuous portion thereof of at least 12 nt, or a complementary sequence to any of the foregoing. In some embodiments, eZFPs target target sites comprising the nucleotide sequence shown in SEQ ID NO: 1040. In some embodiments, the eZFP comprises six zinc fingers, numbered F1 to F6 sequentially from the N-terminus to the C-terminus, each zinc finger containing a corresponding zinc finger recognition region F1 to F6, and the amino acid sequence of each zinc finger recognition region is as follows: F1: QYCCLTN (SEQ ID NO: 792) F2: TSGNLTR (SEQ ID NO: 793) F3: QSSDLSR (SEQ ID NO: 794) F4: FRYYLKR (SEQ ID NO: 795) F5: QSGDLTR (SEQ ID NO: 796) F6: DKGNLTK (SEQ ID NO: 797). In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 704, or a portion thereof, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 704. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 900 or a nucleotide sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 900.
[0318] In some embodiments, this document provides eZFPs (e.g., eZFPs contained in fusion proteins of epigenetic modification DNA targeting systems), such as eZFP_14 described herein. In some embodiments, eZFPs target target sites comprising the nucleotide sequence shown in SEQ ID NO: 1041, a continuous portion thereof of at least 12 nt, or a complementary sequence to any of the foregoing. In some embodiments, eZFPs target target sites comprising the nucleotide sequence shown in SEQ ID NO: 1041. In some embodiments, the eZFP comprises six zinc fingers, numbered F1 to F6 sequentially from the N-terminus to the C-terminus, each zinc finger containing a corresponding zinc finger recognition region F1 to F6, and the amino acid sequence of each zinc finger recognition region is as follows: F1: TSGSLSR (SEQ ID NO: 798) F2: RSDNLTT (SEQ ID NO: 799) F3: QSGNLAR (SEQ ID NO: 800) F4: DRTTLMR (SEQ ID NO: 801) F5: QSGHLAR (SEQ ID NO: 802) F6: QLTHLNS (SEQ ID NO: 803). In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 705, or a portion thereof, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 705. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 901 or a nucleotide sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 901.
[0319] In some embodiments, this document provides eZFPs (e.g., eZFPs contained in fusion proteins of epigenetic modification DNA targeting systems), such as eZFP_15 as described herein. In some embodiments, eZFPs target target sites comprising the nucleotide sequence shown in SEQ ID NO: 1042, a continuous portion thereof of at least 12 nt, or a complementary sequence to any of the foregoing. In some embodiments, eZFPs target target sites comprising the nucleotide sequence shown in SEQ ID NO: 1042. In some embodiments, the eZFP comprises six zinc fingers, numbered F1 to F6 sequentially from the N-terminus to the C-terminus, each zinc finger containing a corresponding zinc finger recognition region F1 to F6, and the amino acid sequence of each zinc finger recognition region is as follows: F1: IKHDLHR (SEQ ID NO: 804) F2: RSANLTR (SEQ ID NO: 805) F3: RSDNLAR (SEQ ID NO: 806) F4: QNVSRPR (SEQ ID NO: 807) F5: RSDDLSK (SEQ ID NO: 808) F6: DSSHRTR (SEQ ID NO: 809). In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 706, or a portion thereof, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 706. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 902 or a nucleotide sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 902.
[0320] In some embodiments, this document provides eZFPs (e.g., eZFPs contained in fusion proteins of epigenetic modification DNA targeting systems), such as eZFP_16 as described herein. In some embodiments, eZFPs target target sites comprising the nucleotide sequence shown in SEQ ID NO: 1043, a continuous portion thereof of at least 12 nt, or a complementary sequence to any of the foregoing. In some embodiments, eZFPs target target sites comprising the nucleotide sequence shown in SEQ ID NO: 1043. In some embodiments, the eZFP comprises six zinc fingers, numbered F1 to F6 sequentially from the N-terminus to the C-terminus, each zinc finger containing a corresponding zinc finger recognition region F1 to F6, and the amino acid sequence of each zinc finger recognition region is as follows: F1: RSDNLAR (SEQ ID NO: 810) F2: QNVSRPR (SEQ ID NO: 811) F3: RSDDLSK (SEQ ID NO: 812) F4: DSSHRTR (SEQ ID NO: 813) F5: TSSNRKT (SEQ ID NO: 814) F6: AQWTRAC (SEQ ID NO: 815). In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 707, or a portion thereof, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 707. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 903 or a nucleotide sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 903.
[0321] In some embodiments, this document provides eZFPs (e.g., eZFPs contained in fusion proteins of epigenetic modification DNA targeting systems), such as eZFP_17 described herein. In some embodiments, eZFPs target target sites comprising the nucleotide sequence shown in SEQ ID NO: 1044, a continuous portion thereof of at least 12 nt, or a complementary sequence to any of the foregoing. In some embodiments, eZFPs target target sites comprising the nucleotide sequence shown in SEQ ID NO: 1044. In some embodiments, the eZFP comprises six zinc fingers, numbered F1 to F6, sequentially from the N-terminus to the C-terminus. Each zinc finger contains a corresponding zinc finger recognition region F1 to F6, and the amino acid sequence of each zinc finger recognition region is as follows: F1: RSDDLSK (SEQ ID NO: 816) F2: DSSHRTR (SEQ ID NO: 817) F3: TSSNRKT (SEQ ID NO: 818) F4: AQWTRAC (SEQ ID NO: 819) F5: RKQTRTT (SEQ ID NO: 820) F6: HRSSLRR (SEQ ID NO: 821). In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 708, or a portion thereof, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 708. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 904 or a nucleotide sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 904.
[0322] In some embodiments, this document provides eZFPs (e.g., eZFPs contained in fusion proteins of epigenetic modification DNA targeting systems), such as eZFP_18 as described herein. In some embodiments, eZFPs target target sites comprising the nucleotide sequence shown in SEQ ID NO: 1045, a continuous portion thereof of at least 12 nt, or a complementary sequence to any of the foregoing. In some embodiments, eZFPs target target sites comprising the nucleotide sequence shown in SEQ ID NO: 1045. In some embodiments, the eZFP comprises six zinc fingers, numbered F1 to F6 sequentially from the N-terminus to the C-terminus, each zinc finger containing a corresponding zinc finger recognition region F1 to F6, and the amino acid sequence of each zinc finger recognition region is as follows: F1: QSAHRKN (SEQ ID NO: 822) F2: TSSNRKT (SEQ ID NO: 823) F3: RSDNLSA (SEQ ID NO: 824) F4: RNNDRKT (SEQ ID NO: 825) F5: TSGSLSR (SEQ ID NO: 826) F6: QAGHLAK (SEQ ID NO: 827). In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 709, or a portion thereof, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 709. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 905 or a nucleotide sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 905.
[0323] In some embodiments, this document provides eZFPs (e.g., eZFPs contained in fusion proteins of epigenetic modification DNA targeting systems), such as eZFP_19 described herein. In some embodiments, eZFPs target target sites comprising the nucleotide sequence shown in SEQ ID NO: 1046, a continuous portion thereof of at least 12 nt, or a complementary sequence to any of the foregoing. In some embodiments, eZFPs target target sites comprising the nucleotide sequence shown in SEQ ID NO: 1046. In some embodiments, the eZFP comprises six zinc fingers, numbered F1 to F6, sequentially from the N-terminus to the C-terminus. Each zinc finger contains a corresponding zinc finger recognition region F1 to F6, and the amino acid sequence of each zinc finger recognition region is as follows: F1: RSDHLSQ (SEQ ID NO: 828) F2: ASSTRTK (SEQ ID NO: 829) F3: RSDDLTR (SEQ ID NO: 830) F4: QKSNLSS (SEQ ID NO: 831) F5: QSANRTT (SEQ ID NO: 832) F6: QNATRTK (SEQ ID NO: 833). In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 710, or a portion thereof, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 710. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 906 or a nucleotide sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 906.
[0324] In some embodiments, this document provides eZFPs (e.g., eZFPs contained in fusion proteins of epigenetic modification DNA targeting systems), such as eZFP_20 described herein. In some embodiments, eZFPs target target sites comprising the nucleotide sequence shown in SEQ ID NO: 1047, a continuous portion thereof of at least 12 nt, or a complementary sequence to any of the foregoing. In some embodiments, eZFPs target target sites comprising the nucleotide sequence shown in SEQ ID NO: 1047. In some embodiments, the eZFP comprises six zinc fingers, numbered F1 to F6, sequentially from the N-terminus to the C-terminus. Each zinc finger contains a corresponding zinc finger recognition region F1 to F6, and the amino acid sequence of each zinc finger recognition region is as follows: F1: RSDTLSE (SEQ ID NO: 834) F2: RRWTLVG (SEQ ID NO: 835) F3: DRSNLSR (SEQ ID NO: 836) F4: QSGDLTR (SEQ ID NO: 837) F5: QSSDLSR (SEQ ID NO: 838) F6: YHWYLKK (SEQ ID NO: 839). In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 711, or a portion thereof, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 711. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 907 or a nucleotide sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 907.
[0325] In some embodiments, this document provides eZFPs (e.g., eZFPs contained in fusion proteins of epigenetic modification DNA targeting systems), such as eZFP_21 described herein. In some embodiments, eZFPs target target sites comprising the nucleotide sequence shown in SEQ ID NO: 1048, a continuous portion thereof of at least 12 nt, or a complementary sequence to any of the foregoing. In some embodiments, eZFPs target target sites comprising the nucleotide sequence shown in SEQ ID NO: 1048. In some embodiments, the eZFP comprises six zinc fingers, numbered F1 to F6 sequentially from the N-terminus to the C-terminus, each zinc finger containing a corresponding zinc finger recognition region F1 to F6, and the amino acid sequence of each zinc finger recognition region is as follows: F1: RSANLAR (SEQ ID NO: 840) F2: RSDNLRE (SEQ ID NO: 841) F3: RPYTLRL (SEQ ID NO: 842) F4: HRSNLNK (SEQ ID NO: 843) F5: QSGSLTR (SEQ ID NO: 844) F6: TSANLSR (SEQ ID NO: 845). In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 712, or a portion thereof, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 712. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 908 or a nucleotide sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 908.
[0326] In some embodiments, this document provides eZFPs (e.g., eZFPs contained in fusion proteins of epigenetic modification DNA targeting systems), such as eZFP_22 as described herein. In some embodiments, eZFPs target target sites comprising the nucleotide sequence shown in SEQ ID NO: 1049, a continuous portion thereof of at least 12 nt, or a complementary sequence to any of the foregoing. In some embodiments, eZFPs target target sites comprising the nucleotide sequence shown in SEQ ID NO: 1049. In some embodiments, the eZFP comprises six zinc fingers, numbered F1 to F6 sequentially from the N-terminus to the C-terminus, each zinc finger containing a corresponding zinc finger recognition region F1 to F6, and the amino acid sequence of each zinc finger recognition region is as follows: F1: RSDDLVR (SEQ ID NO: 846) F2: TSGSLVR (SEQ ID NO: 847) F3: RSDKLVR (SEQ ID NO: 848) F4: RSDELVR (SEQ ID NO: 849) F5: TSHSLTE (SEQ ID NO: 850) F6: RADNLTE (SEQ ID NO: 851). In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 713, or a portion thereof, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 713. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 909 or a nucleotide sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 909.
[0327] In some embodiments, this document provides eZFPs (e.g., eZFPs contained in fusion proteins of epigenetic modification DNA targeting systems), such as eZFP_23 described herein. In some embodiments, eZFPs target target sites comprising the nucleotide sequence shown in SEQ ID NO: 1050, a continuous portion thereof of at least 12 nt, or a complementary sequence to any of the foregoing. In some embodiments, eZFPs target target sites comprising the nucleotide sequence shown in SEQ ID NO: 1050. In some embodiments, the eZFP comprises six zinc fingers, numbered F1 to F6, sequentially from the N-terminus to the C-terminus. Each zinc finger contains a corresponding zinc finger recognition region F1 to F6, and the amino acid sequence of each zinc finger recognition region is as follows: F1: ERSHLRE (SEQ ID NO: 852) F2: TSHSLTE (SEQ ID NO: 853) F3: QAGHLAS (SEQ ID NO: 854) F4: TSHSLTE (SEQ ID NO: 855) F5: DPGHLVR (SEQ ID NO: 856) F6: TSGNLVR (SEQ ID NO: 857). In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 714, or a portion thereof, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 714. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 910 or a nucleotide sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 910.
[0328] In some embodiments, this document provides eZFPs (e.g., eZFPs contained in fusion proteins of epigenetic modification DNA targeting systems), such as eZFP_24 described herein. In some embodiments, eZFPs target target sites comprising the nucleotide sequence shown in SEQ ID NO: 1051, a continuous portion thereof of at least 12 nt, or a complementary sequence to any of the foregoing. In some embodiments, eZFPs target target sites comprising the nucleotide sequence shown in SEQ ID NO: 1051. In some embodiments, the eZFP comprises six zinc fingers, numbered F1 to F6 sequentially from the N-terminus to the C-terminus, each zinc finger containing a corresponding zinc finger recognition region F1 to F6, and the amino acid sequence of each zinc finger recognition region is as follows: F1: RADNLT (SEQ ID NO: 858) F2: TSGSLVR (SEQ ID NO: 859) F3: RKDNLKN (SEQ ID NO: 860) F4: QSSSLVR (SEQ ID NO: 861) F5: RSDKLVR (SEQ ID NO: 862) F6: DSGNLRV (SEQ ID NO: 863). In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 715, or a portion thereof, or an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP comprises the amino acid sequence shown in SEQ ID NO: 715. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 911 or a nucleotide sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it. In some embodiments, the eZFP is encoded by the nucleotide sequence shown in SEQ ID NO: 911.
[0329] In some embodiments, this document provides eZFPs (e.g., eZFPs contained in fusion proteins of epigenetic modification DNA targeting systems), such as eZFP_25 as described herein. In some embodiments, eZFPs target target sites comprising the nucleotide sequence shown in SEQ ID NO: 1052, a continuous portion thereof of at least 12 nt, or a complementary sequence to any of the foregoing. In some embodiments, eZFPs target target sites comprising the nucleotide sequence shown in SEQ ID NO: 1052. In some embodiments, the eZFP comprises six zinc fingers, numbered F1 to F6 sequentially from the N-terminus to the C-terminus, each zinc finger containing a corresponding zinc finge...
Claims
1. An epigenetic DNA-targeting system, the epigenetic DNA-targeting system comprising at least one DNA-targeting module for blocking transcription of one or more hepatitis B virus (HBV) genes, wherein each of the at least one DNA-targeting module comprises a fusion protein, the fusion protein comprising: (a) A DNA-binding domain for targeting sites in the hepatitis B virus DNA sequence; and (b) At least one transcriptional repressor effector domain.
2. The epigenetic modification DNA targeting system of claim 1, wherein the at least one DNA-binding domain comprises a clustered regularly spaced short palindromic repeat-associated (Cas)-guide RNA (gRNA) combination, the combination comprising (a) a Cas protein or a variant thereof and (b) at least one gRNA capable of associating with the Cas protein to target the Cas protein to the target site; a zinc finger protein (ZFP); a transcription activator-like effector (TALE); a meganuclease; a homing endonuclease; or an I-SceI enzyme or a variant thereof, optionally wherein the DNA-binding domain comprises a non-catalytically inactive variant of any of the foregoing.
3. The epigenetic modified DNA targeting system according to claim 1 or 2, wherein the hepatitis B virus DNA sequence is the HBV gene or a regulatory element thereof.
4. The epigenetic modification DNA targeting system according to any one of claims 1-3, wherein the at least one DNA targeting module comprises a plurality of DNA targeting modules for targeting a plurality of target sites of one or more HBV genes or their regulatory elements.
5. The epigenome-modifying DNA-targeting system of claim 4, wherein the plurality of DNA- targeting modules comprises at least a first DNA-targeting module and a second DNA- targeting module, wherein: (1) The first DNA targeting module represses transcription of the first HBV gene, wherein the first DNA targeting module comprises a first fusion protein, the first fusion protein comprising (a) a DNA-binding domain for targeting a target site of the first gene or its regulatory DNA element; and (b) at least one transcriptional repressor domain; and (2) The second DNA targeting module represses transcription of the second HBV gene, wherein the second DNA targeting module comprises a second fusion protein, the second fusion protein comprising (a) a DNA-binding domain for targeting a target site of the second gene or its regulatory DNA element; and (b) at least one transcriptional repressor domain, optionally wherein: The first DNA targeting module and the second DNA targeting module have the same fusion protein, such that the first fusion protein and the second fusion protein are identical, and wherein the DNA-binding domain of the fusion protein is a clustered regularly spaced short palindromic repeat (Cas) protein or a variant thereof capable of associating with the first guide RNA (gRNA) and the second gRNA, wherein The first DNA targeting module contains the first gRNA targeting a target site of a first HBV gene or its regulatory element, and the second DNA targeting module contains the second gRNA targeting a target site of a second HBV gene or its regulatory element.
6. An epigenetic DNA-targeting system for blocking the transcription of one or more hepatitis B virus (HBV) genes, wherein the DNA-targeting system comprises: (a) A fusion protein comprising a clustered regularly spaced short palindromic repeat (Cas) protein or a variant thereof and at least one transcriptional repressor effector domain; and (b) Multiple guide RNAs (gRNAs), wherein the multiple guide RNAs include at least a first gRNA and a second gRNA. The first gRNA targets a target site of the first HBV gene or its regulatory element, and the second gRNA targets a target site of the second HBV gene or its regulatory element. The first gene and the second gene or their regulatory elements regulate hepatitis B virus replication and / or HBV transcription.
7. The epigenetic modification DNA targeting system of claim 6, wherein the DNA targeting system further comprises a third gRNA targeting a third gene or its regulatory element, the third gene or its regulatory element regulating hepatitis B virus replication and / or HBV transcription. Optionally, the system further comprises a fourth gRNA targeting a target site of a fourth gene or its regulatory element, a fifth gRNA optionally targeting a fifth gene or its regulatory element, and / or a sixth gRNA optionally targeting a target site of a sixth gene or its regulatory element. The gene or its regulatory element thereon regulates hepatitis B virus replication and / or HBV transcription.
8. The epigenetic DNA targeting system of claim 7, wherein the first gene, the second gene, the third gene, the fourth gene, the fifth gene, and / or the sixth gene or their regulatory elements are different.
9. An epigenetic DNA-targeting system for blocking the transcription of one or more hepatitis B virus (HBV) genes, wherein the DNA-targeting system comprises: (a) A fusion protein comprising a clustered regularly spaced short palindromic repeat (Cas) protein or a variant thereof and at least one transcriptional repressor effector domain; and (b) Multiple guide RNAs (gRNAs) that target multiple sites of multiple genes or their regulatory elements. The aforementioned genes or their regulatory elements regulate hepatitis B virus replication and / or HBV transcription.
10. An epigenetic DNA-targeting system, said epigenetic DNA-targeting system comprising a single DNA-targeting module for blocking transcription of more than one hepatitis B virus (HBV) gene, The DNA targeting module described herein includes: (a) A fusion protein comprising a clustered regularly spaced short palindromic repeat (Cas) protein or a variant thereof and at least one transcriptional repressor effector domain; and (b) Guide RNAs (gRNAs) that target multiple sites of multiple genes or their regulatory elements. The aforementioned genes or their regulatory elements regulate hepatitis B virus replication and / or HBV transcription.
Citation Information
Patent Citations
Lipids and lipid nanoparticle formulations for delivery of nucleic acids
US10723692B2
Method for gene editing
US10941395B2
Liposomal apparatus and manufacturing methods
US20040142025A1
Systems and methods for manufacturing liposomes
US20070042031A1
Crispr / CAS9-based repressors for silencing gene targets in vivo and methods of use
US20190127713A1