Inhibition of unintended mutations in gene editing
A nucleobase deaminase inhibitor system addresses the specificity issues of CRISPR/Cas-based base editors, enhancing precision in genome editing by minimizing off-target mutations and enabling effective therapeutic applications.
Patent Information
- Application Number
- JP2025170065
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-02-02
- Filing Date
- 2025-10-08
- Publication Date
- 2026-01-27
AI Technical Summary
Current genome editing tools, particularly CRISPR/Cas-based base editors, suffer from low specificity, leading to unintended mutations in single-stranded DNA regions, limiting their application in therapeutic contexts such as correcting T to C transitions associated with human diseases.
Development of a nucleobase deaminase inhibitor fused to a nucleobase deaminase via a protease cleavage site, allowing precise control over editing by inhibiting activity in off-target regions and enabling targeted base editing.
Enhances the specificity of base editing, reducing unintended mutations and facilitating clinical translation by ensuring accurate gene correction.
Smart Images

Figure 2026012727000025 
Figure 2026012727000026 
Figure 2026012727000027
Abstract
Description
[Background technology]
[0001] Genome editing uses engineered nucleases (molecular scissors) to alter how DNA is generated. Genome editing is a type of genetic engineering in which insertions, deletions, or substitutions are made in the genome of an organism. The use of tools to genetically manipulate the genomes of cells and organisms is a key step in life science research. research, biotechnology / agricultural technology development, and most importantly, pharmaceutical / clinical innovation. It has a wide range of applications. For example, genome editing can correct driver mutations that underlie genetic diseases. Genome editing can be used to lead to a complete cure for these diseases in living organisms. , manipulating the genome of crops to increase crop yields and improve resistance to environmental pollution or pathogen infection. It can also be used to confer resistance to crops. Biomolecular transformation is of great importance in the development of renewable bioenergy.
[0002] CRISPR / Cas(Clustered regularly interspa ced short palindromic repeats / CRISPR-ass The mediated protein (MOP) system offers unparalleled editing efficiency, convenience, and It is the most powerful genome editing tool since its conception and has potential applications in biology. Cas nucleases are directed by guide RNAs (gRNAs) and are involved in the transcription of various cellular DNA double-strand breaks (DS) at target genomic sites in cells (both cell lines and cells of biological origin). B) These DSBs can then be repaired by endogenous DNA repair systems. This can be used to perform desired genome editing.
[0003] Generally, the two main DNA repair pathways are DSB, non-homologous end joining (NHEJ) and NHEJ can be activated by homology-directed repair (HDR). Random insertions / deletions (indels) can be introduced into the NA region, and open reading frames can be inserted. This leads to an open reading frame (ORF) shift and ultimately to gene inactivation. In contrast, when HDR is triggered, the genomic DNA sequence at the target site is transformed into a homologous recombination site. The DNA is replaced by the sequence of an exogenous donor DNA template via a genetic alteration mechanism. However, the occurrence of homologous recombination is cell-type specific and cell cycle dependent. NHEJ is dependent on HDR, and NHEJ is triggered more frequently than HDR, so HDR-mediated inheritance The practical efficiency of correction is low (typically less than 5%). Therefore, the efficiency of HDR is relatively low. Therefore, CRISPR / Cas in the field of precision gene therapy (disease-driven gene correction) Translation of genome editing tools has been limited.
[0004] CRISPR / Cas system is used to transduce APOBEC (apolipoprotein B mRNA editing enzyme, catalytic polypepti de-like) AID (activation-induced cytidine deaminase Induced cytidine deaminase family and integrating base enzymes The recently developed BE (Beta) technology has significantly improved the efficiency of CRISPR / Cas-mediated gene correction. Cas9 nickase (nCas9) or catalytically insensitive Cpf1 (dCp f1, also known as dCas12a), which is a member of the APOBEC / AID family. The cytosine (C) deamination activity of the leukocyte member is intentionally directed at target bases in the genome. can be directed against thymine (T) and catalyze the substitution of C for thymine (T) in these bases. can.
[0005] However, APOBEC / AID family members do not recognize single-stranded DNA (ssDNA ) region, current base editing The specificity of the system is compromised, which limits the application, e.g., the use of BE for therapeutic purposes. This limits their use in restoring T to C mutations that cause human disease. It can specifically edit cytosines in the target region but not in other ssDNA regions. It is desirable to create a new BE that does not cause a C to T mutation. New BEs will enable more specific base editing in a variety of organisms. Importantly, this high specificity of BE is particularly important for disease-associated T to C transitions. This will facilitate potential clinical translation in gene therapy with the recovery of allergies. Deaf. Summary of the Invention
[0006] The present disclosure provides, in some embodiments, off-target base editing common to current base editors. Provides a base editor useful for genome editing that reduces or eliminates mutations In certain embodiments, nucleobase deaminase inhibitors inhibit the activity of a nucleic acid molecule involved in genome editing. The nucleobase deaminase inhibitor is cleavably fused to a nucleobase deaminase. In the presence of nucleotides, nucleobase deaminases cannot react with nucleotide molecules (reaction At the target editing site, nucleobase deaminase inhibitors are cut off. The nucleic acid salt can then be cleaved to release a fully active nucleobase deaminase. The base deaminase can perform editing as desired.
[0007] Thus, in one embodiment, a nucleic acid sequence comprising a nucleobase deaminase or catalytic domain thereof is provided. a first fragment, a second fragment containing a nucleobase deaminase inhibitor, and and a protease cleavage site between the first fragment and the second fragment. Thus, fusion proteins are provided.
[0008] In some embodiments, the nucleobase deaminase is an adenosine deaminase. In some embodiments, the adenosine deaminase is a tRNA-specific c adenosine deaminase (TadA), adenosine deaminase aminase tRNA specific 1(ADAT1), adenosine deaminase tRNA specific 2(ADAT2), adenos ine deaminase tRNA specific 3(ADAT3), ade nosine deaminase RNA specific B1(ADARB1) , adenosine deaminase RNA specific B2(ADA RB2), adenosine monophosphate deaminase 1 (AMPD1), adenosine monophosphate deaminas e 2(AMPD2), adenosine monophosphate deami nase 3(AMPD3), adenosine deaminase(ADA), a denosine deaminase 2(ADA2), adenosine dea minase like(ADAL), adenosine deaminase do main containing 1(ADAD1), adenosine deami nase domain containing 2(ADAD2), adenosin e deaminase RNA specific (ADAR) and adenosine ne deaminase RNA specific B1 (ADARB1) is selected from the group.
[0009] In some embodiments, the nucleobase deaminase is a cytidine deaminase. In some embodiments, the cytidine deaminase is APOBEC3B (A3B), APOBEC3C(A3C), APOBEC3D(A3D), APOBEC3F(A3F ), APOBEC3G(A3G), APOBEC3H(A3H), APOBEC1(A1 ), APOBEC3(A3), APOBEC2(A2), APOBEC4(A4) and In some embodiments, the cytidine is selected from the group consisting of: AICDA (AID). The cytidine deaminase is a human or mouse cytidine deaminase. In one embodiment, the catalytic domain is a mouse A3 cytidine deaminase domain 1 (CDA1). or human A3B cytidine deaminase domain 2 (CDA2).
[0010] In some embodiments, the nucleobase deaminase inhibitor is a nucleobase deaminase inhibitor. In some embodiments, the inhibitory domain of the nucleobase deaminase is In some embodiments, the inhibitor is an inhibitory domain of cytidine deaminase. The nucleobase deaminase inhibitor is an inhibitory domain of adenosine deaminase. In some embodiments, the nucleobase deaminase inhibitor is selected from the group consisting of SEQ ID NOs: 1-2. and an amino acid sequence selected from Tables 1 and 2 (SEQ ID NOs: 48 to 135), or At least one amino acid sequence selected from sequence numbers 1 to 2 and Tables 1 and 2 In some embodiments, the amino acid sequence has at least 85% sequence identity with the Nucleobase deaminase inhibitors include those having the amino acid sequence of SEQ ID NO: 1, the amino acid sequence of SEQ ID NO: 1, It comprises the amino acid residues AA76 to AA149, or the amino acid sequence of SEQ ID NO:2.
[0011] In some embodiments, the first fragments are clustered and regularly spaced. CRISPR-associated proteins (Cas) In some embodiments, the Cas protein is SpCas9, FnCas9, St1Cas9, St3Cas9, NmCas9, SaCas9, AsCpf1, LbC pf1, FnCpf1, VQR SpCas9, EQR SpCas9, VRER Sp Cas9, xSpCas9, SpCas9-NG, RHA FnCas9, KKH Sa Cas9, NmeCas9, StCas9, CjCas9, AsCpf1, FnCpf1 , SsCpf1, PcCpf1, BpCpf1, CmtCpf1, LiCpf1, PmC pf1, Pb3310Cpf1, Pb4417Cpf1, BsCpf1, EeCpf1, BhCas12b, AkCas12b, EbCas12b, LsCas12b, RfCa s13d, LwaCas13a, PspCas13b, PguCas13b and Ran Cas13b.
[0012] In some embodiments, the protease cleavage site is a cleavage site for the TuMV protease, PP V protease, PVY protease, ZIKV protease and WNV protease is a protease cleavage site of a protease selected from the group consisting of:
[0013] In some embodiments, the protease cleavage site is an autocleavage site. In some embodiments, the protease cleavage site is a TEV protease cleavage site. In some embodiments, the fusion protein comprises a TEV protease or a fragment thereof. In some embodiments, the third fragment further comprises a third fragment comprising the third flag. The cleavage site of the TEV protease is not cleaved by itself. Contains protease fragments.
[0014] In another embodiment, a cytidine deaminase or its catalytic domain, cluster CRISPR-associated (Cas) transcription factors a first fragment comprising a protein and a first TEV protease fragment; a second fragment comprising a cytidine deaminase inhibitor and a second fragment comprising the first fragment. a TEV protease cleavage site between said first fragment and said second fragment, The TEV protease fragment alone does not cleave the TEV protease cleavage site. Fusion proteins that cannot be synthesized are also provided.
[0015] In some embodiments, the fusion protein comprises a uracil glycosylase inhibitor. In some embodiments, the cytidine deaminase inhibitor further comprises a cytidine deaminase inhibitor (UGI). the TEV protease cleavage site, the cytidine deaminase or its catalytic domain the main, the Cas protein, and the first TEV protease fragment In some embodiments, the first TEV protease is arranged from the N-terminus to the C-terminus. The fragment is the N-terminal domain (SEQ ID NO: 3) or the C-terminal domain of the TEV protease. In some embodiments, the TEV protease cleavage domain is the cleavage domain (SEQ ID NO: 4). The truncated site has the amino acid sequence of SEQ ID NO:5.
[0016] Additionally, in one embodiment, there is provided a method for performing gene editing in a cell at a target site. The cells are then transfected with (a) a fusion protein of the present disclosure and (b) a fusion protein targeting the target site. A guide RNA or a target RNA that targets the target site and tracrRNA (c) a crRNA that further comprises a tag sequence; and (d) a crRNA that can bind to the tag sequence. This involves introducing a second TEV protease fragment linked to the A recognition peptide. Well, a method is also provided.
[0017] In some embodiments, one or more of the molecules comprises a polynucleotide encoding the molecule. In some embodiments, the first TEV protein is introduced into the cell by a the second TEV protease fragment and the second TEV protease fragment interact to form When the TEV protease cleavage site is present, the TEV protease cleavage site can be cleaved. In one embodiment, the second TEV protease fragment is fused to the RNA recognition peptide. are combined.
[0018] In some embodiments, the tag sequence comprises the MS2 sequence (SEQ ID NO: 16). In some embodiments, the RNA recognition peptide is MS2 coat protein (MCP, SEQ ID NO: 1). In some embodiments, the tag sequence comprises the PP7 sequence (SEQ ID NO: 18). wherein the RNA recognition peptide comprises PP7 coat protein (PCP, SEQ ID NO: 23). or the tag sequence comprises a boxB sequence (SEQ ID NO: 20), and the RNA recognition peptide The peptides include the boxB coat protein (N22p, SEQ ID NO: 24).
[0019] Also, in one embodiment, a fusion protein of the present disclosure and a fusion protein that binds to an RNA sequence are provided. a second TEV protease fragment bound to an RNA recognition peptide capable of Also provided is a kit or package for performing gene editing, comprising:
[0020] Yet another embodiment provides a first cytidine deaminase or a second cytidine deaminase comprising a catalytic domain thereof. a fragment of 1, and a second fragment containing the inhibitory domain of a second cytidine deaminase. a fragment of the first cytidine deaminase, wherein the second cytidine deaminase The fusion protein may be the same as or different from the above.
[0021] In another embodiment, a nucleobase deaminase or catalytic domain thereof, a nucleobase deaminase a nucleic acid base deaminase inhibitor, a first RNA recognition peptide, and said nucleic acid base deaminase or TEV protease between its catalytic domain and the nucleobase deaminase inhibitor A fusion protein is provided that includes a first fragment that includes a cleavage site.
[0022] In some embodiments, the fusion protein comprises the TEV protease cleavage site. a TEV protease fragment that cannot cleave alone, and a second RNA In some embodiments, the fusion protein further comprises a second fragment comprising a recognition peptide. The fusion protein has a self-cleavage site between the first fragment and the second fragment. It further includes rank.
[0023] In some embodiments, the fusion protein comprises a second TEV protease fragment. a third fragment containing the first TEV protease fragment, The TEV protease fragment is cleaved in the presence of the second TEV protease fragment. In some embodiments, the fusion protein can cleave the second a second self-cleavage site between the first fragment and the third fragment, Upon cleavage of the second self-cleavage site, the fusion protein is fused to an RNA recognition peptide. This releases the second TEV protease fragment.
[0024] Also, in one embodiment, a sequence A target single guide RNA comprising a first spacer having complementarity with a second PA A second spacer having sequence complementarity to a second nucleic acid sequence adjacent to the M site. A single guide RNA and clustered regularly interspaced short palindromic sequences Contains CRISPR-associated repeat (Cas) proteins and nucleobase deaminases The second PAM site is 34 to 91 bases away from the first PAM site. Also provided are guide RNA systems. In some embodiments, the second spacer In some embodiments, the second spacer is 8 to 15 bases in length. It is 12 bases long.
[0025] In one embodiment, in the 5' to 3' direction, a first stem loop portion, a second stem loop portion, a scaffold comprising a first portion, a third stem-loop portion, and a fourth stem-loop portion wherein the third stem loop contains five base pairs therein. In another embodiment, the present disclosure provides a guide RNA between bases 45 and 55. By introducing base pairing into the In some embodiments, the scaffold comprises a sequence similar to SEQ ID NO: 3. In some embodiments, the guide R comprises a sequence selected from the group consisting of: 2 to 43. The NA is at least 100, or 120 nucleotides in length.
[0026] Another embodiment is a method for performing gene editing in a cell at a target site, comprising: Clustered regularly interspaced short palindromic repeats (CRISP) are expressed in the memory cells. R) a first viral particle encapsulating a first construct encoding a related (Cas) protein , and a second construct encoding a reverse transcriptase fused to an RNA recognition peptide. and introducing a second viral particle comprising the second viral particle.
[0027] In some embodiments, the second construct comprises an R In some embodiments, the C further encodes a guide RNA comprising a guide RNA recognition site. The as protein was SpCas9-NG (SEQ ID NO: 46) or xSpCas9 (SEQ ID NO: No. 47).
[0028] including, but not limited to, a polynucleotide encoding a fusion protein of the present disclosure; A construct comprising said polynucleotide, a cell comprising said polynucleotide or said construct and compositions comprising any of the above are also provided. [Brief explanation of the drawings]
[0029] [Figure 1] Unintended base substitutions caused by the current BE in the Sa-SITE31 ssDNA region. Figure 1A: Schematic diagram showing co-expression of SaD10A nickase and Sa-sgSITE31, which triggers the formation of an ssDNA region at the Sa-sgSITE31 target site. Figure 1B: Schematic diagram showing co-transfection of Sa-sgSITE31-expressing plasmids and SaD10A nickase-expressing plasmids with BE3-expressing plasmids, hA3A-BE3-expressing plasmids, or empty vector. Figure 1C: Unintended base substitutions caused by BE3 and hA3A-BE3. The dashed box represents the location of the unintended base substitution in the Sa-sgSITE31 target site. [Figure 2]Unintended base substitutions caused by the current BE in the Sa-SITE42 ssDNA region. Figure 2A: Schematic diagram showing co-expression of Sa-sgSITE42 with SaD10A nickase, which triggers the formation of an ssDNA region at the Sa-sgSITE42 on-target site. Figure 2B: Schematic diagram showing co-transfection of Sa-sgSITE42-expressing plasmids and SaD10A nickase-expressing plasmids with BE3-expressing plasmids, hA3A-BE3-expressing plasmids, or empty vector. Figure 2C: Unintended base substitutions caused by BE3 and hA3A-BE3. The dashed box represents the location of the unintended base substitution at the Sa-sgSITE42 target site. [Figure 3] Unintended base substitutions caused by the current BE in the Sa-F1 ssDNA region. Figure 3A: Schematic diagram showing coexpression of SaD10A nickase and Sa-sgF1, which triggers the formation of an ssDNA region at the Sa-sgF1 on-target site. Figure 3B: Schematic diagram showing cotransfection of a plasmid expressing Sa-sgF1 and a plasmid expressing SaD10A nickase with a plasmid expressing BE3, a plasmid expressing hA3A-BE3, or an empty vector. Figure 3C: Unintended base substitutions caused by BE3 and hA3A-BE3. The dashed box represents the location of the unintended base substitution in the Sa-sgF1 target site. [Figure 4]mA3CDA2 inhibits C to T base editing activity in the TET1 region. Figure 4A: Schematic diagram showing the regions of the CDA domain in mA3, rA1, and hA3A. Figure 4B: Schematic diagram showing co-transfection of a plasmid expressing sgTET1 with a plasmid expressing mA3-BE3, a plasmid expressing mA3CDA1-BE3, a plasmid expressing mA3rev-BE3, a plasmid expressing mA3rev-2A-BE3, a plasmid expressing BE3, a plasmid expressing mA3CDA2-BE3, a plasmid expressing mA3CDA2-2A-BE3, a plasmid expressing hA3A-BE3, a plasmid expressing mA3CDA2-hA3A-BE3, or a plasmid expressing mA3CDA2-2A-hA3A-BE3. Figure 4C: mA3CDA2 inhibits C to T editing activity of mA3CDA1-BE3, BE3, and hA3A-BE3. The dashed box represents the position of the C to T base edit at the sgTET1 target site. [Figure 5] mA3CDA2 inhibits C to T base editing activity in the RNF2 region. Figure 5A: Schematic diagram showing the regions of the CDA domain in mA3, rA1, and hA3A. Figure 5B: Schematic diagram showing co-transfection of a plasmid expressing sgRNF2 with a plasmid expressing mA3-BE3, a plasmid expressing mA3CDA1-BE3, a plasmid expressing mA3rev-BE3, a plasmid expressing mA3rev-2A-BE3, a plasmid expressing BE3, a plasmid expressing mA3CDA2-BE3, a plasmid expressing mA3CDA2-2A-BE3, a plasmid expressing hA3A-BE3, a plasmid expressing mA3CDA2-hA3A-BE3, or a plasmid expressing mA3CDA2-2A-hA3A-BE3. Figure 5C: mA3CDA2 inhibits C to T editing activity of mA3CDA1-BE3, BE3, and hA3A-BE3. The dashed box represents the position of the C to T base edit at the sgRNF2 target site. [Figure 6]mA3CDA2 inhibits C to T base editing activity in the SITE3 region. Figure 6A: Schematic diagram showing the region of the CDA domain in mA3, rA1, and hA3A. Figure 6B: Schematic diagram showing cotransfection of a plasmid expressing sgSITE3 with a plasmid expressing mA3-BE3, a plasmid expressing mA3CDA1-BE3, a plasmid expressing mA3rev-BE3, a plasmid expressing mA3rev-2A-BE3, a plasmid expressing BE3, a plasmid expressing mA3CDA2-BE3, a plasmid expressing mA3CDA2-2A-BE3, a plasmid expressing hA3A-BE3, a plasmid expressing mA3CDA2-hA3A-BE3, or a plasmid expressing mA3CDA2-2A-hA3A-BE3. Figure 6C: mA3CDA2 inhibits C to T editing activity of mA3CDA1-BE3, BE3, and hA3A-BE3. The dashed box represents the position of the C to T base edit at the sgSITE3 target site. [Figure 7] hA3BCDA1 inhibits C to T base editing activity in the TET1 region. Figure 7A: Schematic showing the region of the CDA domain in hA3B. Figure 7B: Schematic showing co-transfection of a plasmid expressing sgTET1 with a plasmid expressing hA3B-BE3, a plasmid expressing hA3BCDA2-BE3, or a plasmid expressing hA3B-2A-BE3. Figure 7C: hA3BCDA1 inhibits the C to T editing activity of hA3BCDA2-BE3. The dashed box represents the position of the C to T base edit at the sgTET1 target site. [Figure 8]hA3BCDA1 inhibits C to T base editing activity in the RNF2 region. Figure 8A: Schematic showing the region of the CDA domain in hA3B. Figure 8B: Schematic showing co-transfection of a plasmid expressing sgRNF2 with a plasmid expressing hA3B-BE3, a plasmid expressing hA3BCDA2-BE3, or a plasmid expressing hA3B-2A-BE3. Figure 8C: hA3BCDA1 inhibits the C to T editing activity of hA3BCDA2-BE3. The dashed box represents the position of the C to T base edit in the sgRNF2 target site. [Figure 9] hA3BCDA1 inhibits C to T base editing activity in the SITE3 region. Figure 9A: Schematic diagram showing the region of the CDA domain in hA3B. Figure 9B: Schematic diagram showing co-transfection of a plasmid expressing sgSITE3 with a plasmid expressing hA3B-BE3, a plasmid expressing hA3BCDA2-BE3, or a plasmid expressing hA3B-2A-BE3. Figure 9C: hA3BCDA1 inhibits the C to T editing activity of hA3BCDA2-BE3. The dashed box represents the position of the C to T base edit at the sgSITE3 target site. [Figure 10]Mapping the split site of mA3 by examining base editing efficiency in the FANCF region (Figure 10A: Schematic showing the regions of the two CDA domains in mA3 and the sites used to split mA3 (AA196 / AA197, AA207 / AA208, AA215 / AA216, AA229 / AA230, AA237 / AA238). Figure 10B: Schematic diagram showing co-transfection of a sgFANCF-expressing plasmid with a mA3rev-BE3-196-expressing plasmid, a mA3rev-2A-BE3-196-expressing plasmid, a mA3rev-BE3-expressing plasmid, a mA3rev-2A-BE3-expressing plasmid, a mA3rev-BE3-215-expressing plasmid, a mA3rev-2A-BE3-215-expressing plasmid, a mA3rev-BE3-229-expressing plasmid, a mA3rev-BE3-237-expressing plasmid, or a mA3rev-2A-BE3-237-expressing plasmid. Figure 10C: Split sites spanning AA196 / AA197 to AA237 / AA238 generally maintain C-to-T editing efficiency. The dashed box represents the position of the C-to-T base edit at the sgFANCF target site. [Figure 11]Mapping the split site of mA3 by examining base editing efficiency in the SITE2 region. Figure 11A: Schematic showing the regions of the two CDA domains in mA3 and the sites used to split mA3 (AA196 / AA197, AA207 / AA208, AA215 / AA216, AA229 / AA230, AA237 / AA238). Figure 11B: Schematic diagram showing cotransfection of a plasmid expressing sgSITE2 with a plasmid expressing mA3rev-BE3-196, a plasmid expressing mA3rev-2A-BE3-196, a plasmid expressing mA3rev-BE3, a plasmid expressing mA3rev-2A-BE3, a plasmid expressing mA3rev-BE3-215, a plasmid expressing mA3rev-2A-BE3-215, a plasmid expressing mA3rev-BE3-229, a plasmid expressing mA3rev-2A-BE3-229, a plasmid expressing mA3rev-BE3-237, or a plasmid expressing mA3rev-2A-BE3-237. Figure 11C: Split sites spanning AA196 / AA197 to AA237 / AA238 generally maintain C-to-T editing efficiency. The dashed box represents the position of the C-to-T base edit at the sgSITE2 target site. [Figure 12]Mapping the split site of mA3 by examining base editing efficiency in the SITE4 region. Figure 12A: Schematic showing the regions of the two CDA domains in mA3 and the sites used to split mA3 (AA196 / AA197, AA207 / AA208, AA215 / AA216, AA229 / AA230, AA237 / AA238). Figure 12B: Schematic diagram showing co-transfection of a plasmid expressing sgSITE4 with a plasmid expressing mA3rev-BE3-196, a plasmid expressing mA3rev-2A-BE3-196, a plasmid expressing mA3rev-BE3, a plasmid expressing mA3rev-2A-BE3, a plasmid expressing mA3rev-BE3-215, a plasmid expressing mA3rev-2A-BE3-215, a plasmid expressing mA3rev-BE3-229, a plasmid expressing mA3rev-2A-BE3-229, a plasmid expressing mA3rev-BE3-237, or a plasmid expressing mA3rev-2A-BE3-237. Figure 12C: Split sites spanning AA196 / AA197 to AA237 / AA238 generally maintain C-to-T editing efficiency. The dashed box represents the position of the C-to-T base edit at the sgSITE4 target site. [Figure 13] Mapping of the minimal region of mA3 containing the base editing inhibitory effect in the FANCF region. Figure 13A: Schematic showing co-transfection of a plasmid expressing sgFANCF with a plasmid expressing mA3rev-BE3-237, a plasmid expressing mA3rev-BE3-237-Del-255, a plasmid expressing mA3rev-BE3-237-Del-285, or a plasmid expressing mA3rev-BE3-237-Del-333. Figure 13B: The region spanning AA334 to AA429 of mA3 contains the base editing inhibitory effect. The dashed box represents the position of the C to T base edit in the sgFANCF target site. [Figure 14]Mapping of the minimal region of mA3 containing the base editing inhibitory effect in the SITE2 region. Figure 14A: Schematic diagram showing co-transfection of a plasmid expressing sgSITE2 with a plasmid expressing mA3rev-BE3-237, a plasmid expressing mA3rev-BE3-237-Del-255, a plasmid expressing mA3rev-BE3-237-Del-285, or a plasmid expressing mA3rev-BE3-237-Del-333. Figure 14B: The region spanning AA334 to AA429 of mA3 contains the base editing inhibitory effect. The dashed box represents the position of the C to T base edit at the sgSITE2 target site. [Figure 15] Mapping of the minimal region of mA3 containing the base editing inhibitory effect in the SITE4 region. Figure 15A: Schematic showing co-transfection of a plasmid expressing sgSITE4 with a plasmid expressing mA3rev-BE3-237, a plasmid expressing mA3rev-BE3-237-Del-255, a plasmid expressing mA3rev-BE3-237-Del-285, or a plasmid expressing mA3rev-BE3-237-Del-333. Figure 15B: The region spanning AA334 to AA429 of mA3 contains the base editing inhibitory effect. The dashed box represents the position of the C to T base edit at the sgSITE4 target site. [Figure 16] Schematic diagram showing the working process of BEsafe and BE3 or hA3A-BE3. Figure 16A: BEsafe induces a C to T base edit at the on-target site and avoids causing mutations in unassociated ssDNA regions. Figure 16B: BE3 or hA3A-BE3 induces a C to T base edit at the on-target site but causes a C to T mutation in unassociated ssDNA regions. [Figure 17]Comparison of hA3A-BE3 and BEsafe at the unrelated Sa-SITE31 ssDNA region and the TET1 on-target site. Figure 17A: Schematic diagram showing co-expression of Sa-sgSITE31 with SaD10A nickase, which triggers the formation of an ssDNA region at the Sa-sgSITE31 on-target site. Figure 17B: Schematic diagram showing co-transfection of a plasmid expressing Sa-sgSITE31 and a plasmid expressing SaD10A nickase with a plasmid expressing hA3A-BE3 and a plasmid expressing sgTET1, or a plasmid expressing BEsafe and a plasmid expressing MS2-sgTET1 and MCP-TEVc, or a plasmid expressing MCP-TEVc and a plasmid expressing MS2-sgTET1 and BEsafe. Figure 17C: Comparison of the frequency of unintended C to T mutations triggered by hA3A-BE3 and BEsafe at the unrelated Sa-SITE31 ssDNA region. The dashed box represents the position of the unintended base substitution at the Sa-sgSITE31 target site. Figure 17D: Comparison of base editing efficiencies of hA3A-BE3 and BEsafe at the TET1 site. The dashed box represents the position of the C to T base edit at the sgTET1 target site. [Figure 18]Comparison of hA3A-BE3 and BEsafe at the unrelated Sa-SITE32 ssDNA region and the RNF2 on-target site. Figure 18A: Schematic diagram showing co-expression of Sa-sgSITE32 with SaD10A nickase, which triggers the formation of an ssDNA region at the Sa-sgSITE32 on-target site. Figure 18B: Schematic diagram showing co-transfection of a plasmid expressing Sa-sgSITE32 and a plasmid expressing SaD10A nickase with a plasmid expressing hA3A-BE3 and a plasmid expressing sgRNF2, or a plasmid expressing BEsafe and a plasmid expressing MS2-sgRNF2 and MCP-TEVc, or a plasmid expressing MCP-TEVc and a plasmid expressing MS2-sgRNF2 and BEsafe. Figure 18C: Comparison of the frequency of unintended C to T mutations triggered by hA3A-BE3 and BEsafe at the unrelated Sa-SITE32 ssDNA region. The dashed box represents the position of the unintended base substitution at the Sa-sgSITE32 target site. Figure 18D: Comparison of base editing efficiencies of hA3A-BE3 and BEsafe at the RNF2 site. The dashed box represents the position of the C to T base edit at the sgRNF2 target site. [Figure 19]Comparison of hA3A-BE3 and BEsafe at the unrelated Sa-F1 ssDNA region and the SITE3 on-target site. Figure 19A: Schematic diagram showing co-expression of Sa-sgF1 with SaD10A nickase, which triggers the formation of an ssDNA region at the Sa-sgF1 on-target site. Figure 19B: Schematic diagram showing co-transfection of a plasmid expressing Sa-sgF1 and a plasmid expressing SaD10A nickase with a plasmid expressing hA3A-BE3 and a plasmid expressing sgSITE3, or a plasmid expressing BEsafe and a plasmid expressing MS2-sgSITE3 and MCP-TEVc, or a plasmid expressing MCP-TEVc and a plasmid expressing MS2-sgSITE3 and BEsafe. Figure 19C: Comparison of the frequency of unintended C to T mutations triggered by hA3A-BE3 and BEsafe at the unrelated Sa-F1 ssDNA region. The dashed box represents the position of the unintended base substitution at the Sa-sgF1 target site. Figure 19D: Comparison of base editing efficiencies of hA3A-BE3 and BEsafe at the SITE3 site. The dashed box represents the position of the C to T base edit at the sgSITE3 target site. [Figure 20]Identification of cytidine deaminase inhibitors. Figure 20a: Schematic diagrams show APOBEC family members with single or dual CDA domains (left) and paired base editors constructed with one or two CDAs of dual-domain APOBECs (right). Figure 20b: Editing frequencies induced by the indicated BEs at one representative genomic locus. Figure 20c: Statistical analysis of normalized editing frequencies. Induced by a single-CDA-containing BE was set to 100%. n = 78 from three independent experiments at the 26 editable cytosine sites shown in (b). Figure 20d: Schematic diagram showing conjugation of various cytidine deaminase inhibitors (CD1s) to the N-terminus of mA3CDA1-nSpCas9-BE. Figure 20e: Editing frequencies induced by the indicated BEs at one representative genomic locus. Figure 20f: Statistical analysis of normalized editing frequencies, those induced by BE without CDI set to 100%. n = 57 from three independent experiments at 19 editable cytosine sites shown in (e). (b), (e) Mean ± standard deviations were from three independent experiments. NT, non-transfected control. (c), (f) P values, one-tailed Student's t-test. Median and interquartile range (IQR) are shown. [Figure 21]Conjugation of mA3CDI reduced unintended base editing at sgRNA-independent OTss sites. Figure 21a: Schematic showing that BE3 induces C-to-T mutations, while CDI-conjugated iBE1 (CDI-conjugated iBE1) remains dormant at sgRNA-independent OTss sites. Figure 21b: Comparison of C-to-T editing frequencies induced by BE3 and iBE1 at ssDNA regions triggered by nSaCas9-generated SSBs. Figure 21c: Statistical analysis of normalized cumulative editing frequencies at the four ssDNA sites shown in (b). BE3-induced frequencies were set to 100%. n = 12 from three independent experiments. Figure 21d: Schematic showing that sgRNA-mediated CDI cleavage restores iBE editing activity at on-target sites. Figure 21e: Comparison of C-to-T editing frequencies induced by BE3 and iBE1 at on-target sites. Figure 21f: Statistical analysis of normalized cumulative editing frequencies at the four on-target sites shown in (e). BE3-induced was set to 100%. n=12 from three independent experiments. (c), (f) Mean ± standard deviation from three independent experiments. (d), (g) P values, one-tailed Student's t-test. Median and interquartile range (IQR) are shown. [Figure 22]neSpCas9 reduced unintended editing by iBE1 at OTsg sites. Figure 22a: Schematic showing that iBE1, but not iBE2, induces C-to-T editing at OTsg sites partially complementary to the sgRNA. Figure 22b: Comparison of C-to-T editing frequencies induced by iBE1 and iBEs with improved targeting specificity at the indicated OTsg sites. Figure 22c: Statistical analysis of normalized cumulative editing frequencies at OTsg sites for the two sgRNAs used in (b). That induced by iBE1 was set to 100%. n = 6 from three independent experiments. Figure 22d: Comparison of C-to-T editing frequencies induced by iBE1 and iBEs with improved targeting specificity at on-target sites. Figure 22e: Statistical analysis of normalized cumulative editing frequencies at the six on-target sites shown in (d). That induced by iBE1 was set to 100%. n = 18 from three independent experiments. (b), (d) Mean values ± standard deviations were obtained from three independent experiments. (c), (e) P values, one-tailed Student's t-test. Median and interquartile range (IQR) are shown. [Figure 23]Comparison of base editing induced by hA3A-BE3 and iBE2. Figure 23a: Comparison of C to T editing frequencies induced by hA3A-BE3 and iBE2 at representative OTss, OTsg, and on-target sites. Figures 23b-c: Statistical analysis of normalized cumulative editing frequencies at OTss, OTsg (b), and on-target sites (c) for the three sgRNAs used in (a). That induced by hA3A-BE3 was set to 100%. n=9 from three independent experiments. Figure 23d: Statistical analysis of the normalized ratio of on-target editing frequency to total editing frequency at OTss and OTsg sites for the three sgRNAs used in (a). That induced by hA3A-BE3 was set to 1. n=9 from three independent experiments. Figure 23e: Schematic diagram showing that iBE2 induces specific base editing at the on-target site but not at the OTss or OTsg site, whereas hA3A-BE3 induces base editing at both the on-target site and the OTss and OTsg sites. (a) Mean ± standard deviation from three independent experiments. (b-d) P values, one-tailed Student's t-test. Median and IQR are shown. [Figure 24] Schematic diagram showing the working process of isplitBE and conventional base editors. Figure 24A: isplitBE induces C to T base editing only at the on-target site, avoiding mutations at unrelated off-target ssDNA regions (OTss) or off-target sites with sequence similarity to the spacer region of the sgRNA (OTsg). Figure 24B: BE3 or hA3A-BE3 induces C to T base editing at the on-target site, but causes C to T mutations in the OTss and OTsg regions. [Figure 25] Schematic diagram showing different strategies for removing a cytidine deaminase inhibitor (mA3CDA2) at the on-target site. [Figure 26]C-to-T editing at the EMX1-ON, Sa-SITE31-OTss, and EMX1-OTsg sites induced by different combinations of nCas9(D10A), APOBEC cytidine deaminase, cytidine deaminase inhibitor (CDI), uracil-DNA glycosylase inhibitor (UGI), and TEV protease. Figure 26A: Schematic showing cotransfection of plasmids expressing Sa-sgSITE31 and SaD10A nicase with the indicated 10 pairs of plasmids expressing various base editors. Figure 26B: Comparison of editing efficiency at the EMX1-ON, Sa-SITE31-OTss, and EMX1-OTsg sites. isplitBE-rA1 (pair 9) induced substantial editing at the ON site but not at the OTss or OTsg sites. [Figure 27] C-to-T editing at the FANCF-ON, Sa-VEGFA-7-OTss, and FANCF-OTsg sites induced by different combinations of nCas9(D10A), APOBEC cytidine deaminase, cytidine deaminase inhibitor (CDI), uracil-DNA glycosylase inhibitor (UGI), and TEV protease. Figure 27A: Schematic showing cotransfection of plasmids expressing Sa-sgVEGFA-7 and SaD10A nickase with the indicated 10 pairs of plasmids expressing various base editors. Figure 27B: Comparison of editing efficiency at the FANCF-ON, Sa-VEGFA-7-OTss, and FANCF-OTsg sites. isplitBE-rA1 (pair 9) induced substantial editing at the ON site but not at the OTss or OTsg sites. [Figure 28]C-to-T editing at the V1B-ON, Sa-SITE42-OTss, and V1B-OTsg sites induced by different combinations of nCas9(D10A), APOBEC cytidine deaminase, cytidine deaminase inhibitor (CDI), uracil-DNA glycosylase inhibitor (UGI), and TEV protease. Figure 28A: Schematic showing cotransfection of plasmids expressing Sa-sgSITE42 and SaD10A nickase with the indicated 10 pairs of plasmids expressing various base editors. Figure 28B: Comparison of editing efficiency at the V1B-ON, Sa-SITE42-OTss, and V1B-OTsg sites. isplitBE-rA1 (pair 9) induced substantial editing at the ON site but not at the OTss or OTsg sites. [Figure 29] The effect of the distance between the helper sgRNA (hsgRNA) and sgRNA on base editing efficiency. Figure 29A: Schematic showing the distance between the hsgRNA and sgRNA at the DNTET1, EMX1, and FANCF sites. Figure 29B: Base editing frequencies induced by the indicated sgRNAs and hsgRNAs. Figure 29C: Summary of the effect of the distance between the hsgRNA and sgRNA. The distance range for the highest base editing efficiency is -91 to -34 bp from the PAM of the hsgRNA to the PAM of the sgRNA. [Figure 30] Effect of hsgRNA spacer length on base editing efficiency. Figure 30A: Schematic showing co-transfection of sgRNA and hsgRNA with different spacer lengths at the DNEMX1, FANCF, and V1A sites. Figure 30B: Base editing frequencies induced by the indicated sgRNA and hsgRNA at the hsgRNA and sgRNA target sites. Figure 30C: Statistical analysis of the effect of hsgRNA spacer length. Use of an hsgRNA with a 10-bp spacer significantly reduces editing efficiency at the hsgRNA target site but maintains editing efficiency at the sgRNA target site. [Figure 31] Comparison of editing efficiencies of isplitBE-rA1 and BE3. Editing frequencies induced by the indicated base editors at different target sites. [Figure 32] Comparison of genome-wide C to T mutations induced by isplitBE-rA1 and BE3. Figure 32A: mRNA expression levels in wild-type 293FT cells and APOBEC3 knockout 293FT cells (293FT-A3KO). Figure 32B: Schematic showing the procedure for determining genome-wide C to T mutations induced by base editors. Figure 32C: On-target editing efficiency (left) and the number of genome-wide C to T mutations induced by Cas9, BE3, hA3A-BE3-Y130F (Y130F), and isplitBE-rA1. [Figure 33] Comparison of transcriptome-wide C to U mutations induced by isplitBE-mA3, BE3, and hA3A-BE3-Y130F (Y130F). Figure 33A: Number of transcriptome-wide C to U mutations induced by Cas9, BE3, hA3A-BE3-Y130F (Y130F), and isplitBE-mA3. Figure 33B: Frequency of RNA C to U editing induced by Cas9, BE3, hA3A-BE3-Y130F (Y130F), and isplitBE-mA3. Figure 33C: Distribution of RNA C to U editing induced by BE3 replicate 1 and isplitBE-mA3 replicate 1. [Figure 34] Stop codons induced by isplitBE-mA3 in the human PCSK9 gene. Figure 34A: Schematic showing co-transfection of sgRNA and hsgRNA containing isplitBE-mA3 with nCas9. Figures 34B-34D: Editing efficiencies induced by isplitBE-mA3 at the indicated sites. [Figure 35] Inhibitory effect of mA3CDA2 on the editing efficiency of adenine base editor (ABE). Figure 35A: Schematic showing co-transfection of sgRNA with ABE fused or not to mA3CDA2. Figure 35B: Editing efficiency induced by ABE shown at RNF2 and FANCF sites. [Figure 36]Enhanced prime editing by engineering prime editing guide RNA (pegRNA). Figure 36A: Schematic showing RNA base pair changes to increase stem stability of enhanced pegRNA (epegRNA). Figure 36B: Schematic showing co-transfection of PE2, nicking sgRNA, and pegRNA or epegRNA-GC. Figures 36C-36D: Comparison of prime editing efficiency induced by pegRNA and epegRNA-GC. Figure 36E: Schematic showing RNA base pair changes to increase stem stability of enhanced pegRNA (epegRNA). Figure 36F: Schematic showing co-transfection of PE2, nicking sgRNA, and pegRNA or epegRNA-CG. Figure 36G: Comparison of prime editing efficiency induced by pegRNA and epegRNA-CG. [Figure 37] Prime editing system using PE containing different Cas9 proteins. Figure 37A: Schematic showing co-transfection of pegRNA, nicking sgRNA, and PE2-NG or xPE2. Figure 37B: Prime editing efficiency induced by PE2-NG and xPE2. [Figure 38] Split prime editing (split-PE) system. Figure 38A: Schematic diagram showing the working process of PE and split-PE system. Figure 38B: Schematic diagram showing co-transfection of PE and split-PE system. Figure 38C: Editing efficiency induced by PE and split-PE system at EMX1 site. [Figure 39] Figure 39A-C: Alignment of the mA3CDA2 core region with other cytidine deaminase domains. [Figure 40] Figure 40A-D: Alignment of hA3BCDA1 with other cytidine deaminase domains. DETAILED DESCRIPTION OF THE INVENTION
[0030] ·Definition The term "a" or "an" entity refers to the entity Note that the term "antibody" refers to one or more of the antibodies; for example, "an antibody" The term "antibody" is understood to refer to one or more antibodies. a)" (or "one (an)"), "one or more (one or more)", and " The terms "at least one" and "at least one" may be used interchangeably herein.
[0031] As used herein, the term "polypeptide" refers to a single "polypeptide" " and "polypeptides," and It is made up of monomers (amino acids) linearly linked by bonds (also known as cleavage bonds). The term "polypeptide" refers to any one or more amino acids. refers to multiple chains and not to a specific length of the product. Peptides, tripeptides, oligopeptides, "proteins," "amino acid chains," or 2 Any other term used to refer to one or more chains of one or more amino acids is "poly" The term "polypeptide" is included within the definition of "polypeptide" and the term "polypeptide" may be substituted for either of these terms. The term "polypeptide" can be used interchangeably with any of these terms. In addition, the amino acid sequence may be modified by, but not limited to, glycosylation, acetylation, phosphorylation, amidation, or alteration of an already existing amino acid. Derivatization with known protecting / blocking groups, proteolytic cleavage, or unnatural amino acids It is also intended to refer to the products of post-expression modification of polypeptides, including modification by Peptides may be derived from natural biological sources or produced by recombinant technology. It is not necessarily translated from a designated nucleic acid sequence, including chemical synthesis. The data may be generated in any manner, including
[0032] "Homology" or "identity" or "similarity" refers to the degree of similarity between two peptides or two nucleic acids. Refers to the sequence similarity between acid molecules. Homology is determined by comparing positions in each sequence. The sequences can be aligned for comparison purposes. If positions are occupied by the same base or amino acid, the molecules The degree of homology between sequences is determined by the number of identical or similar sequences shared by those sequences. is a function of the number of homologous positions. An "unrelated" or "non-homologous" sequence is a sequence shares less than 40% identity, but preferably less than 25% identity, with one of .
[0033] Polynucleotide or polynucleotide region (or polypeptide or polypeptide) peptide region) to another sequence at a specific percentage (e.g., 60%, 65%, 70%). %, 75%, 80%, 85%, 90%, 95%, 98% or 99%) of "sequence identity" If aligned, the percentage of bases (or amino acids) ) means that two sequences are the same when comparing. Percent identity or percent sequence identity can be calculated using software programs known in the art. Ram, e.g., Ausubel et al. (2007) Current Pro Use the ones described in the Guidelines in Molecular Biology Preferably, the default parameters are used for alignment. One of the alignment programs uses default parameters. It's BLAST.
[0034] The term "equivalent nucleic acid or polynucleotide" refers to the nucleotide sequence of a nucleic acid or its complement. a nucleic acid having a nucleotide sequence that has a degree of homology or sequence identity with the nucleotide sequence A homolog of a double-stranded nucleic acid is a nucleotide sequence that has some degree of homology with its complement. In one aspect, a homolog of a nucleic acid is intended to include a nucleic acid having the sequence Similarly, an "equivalent polypeptide" is a polypeptide that is capable of hybridizing to a reference polypeptide. A polypeptide having a degree of homology or sequence identity with the amino acid sequence of the reference polypeptide. In some embodiments, the sequence identity is at least about 70%, 75%, 80%, 85%, 90%, 95%, 98%, or 99%. In some embodiments, the equivalent percentage A polypeptide or polynucleotide is compared to a reference polypeptide or polynucleotide. and one, two, three, four, or five additions, deletions, substitutions, and combinations thereof. In some embodiments, the equivalent sequence has the activity (e.g., epitope binding) of the reference sequence. ) or structure (e.g., salt bridges).
[0035] The term "encode" as applied to a polynucleotide means that, in its natural state, or when engineered by methods well known to those skilled in the art, polypeptides and / or A polypeptide is a fragment of a gene that can be transcribed and / or translated to produce mRNA of that gene. The antisense strand refers to a polynucleotide that is said to "encode" such a peptide. The nucleic acid complement is a suitable nucleic acid from which the coding sequence can be deduced.
[0036] Nucleobase deaminase inhibitors to reduce random insertions and deletions use As shown in the experimental example and Figures 1-3, the currently commonly used base editor BE3 and hA3A-BE3, which contains a C to T mutation in an off-target single-stranded DNA region. This led to:
[0037] However, surprisingly, the expression of mouse APOBEC3 (mA3) in mA3-BE3 was not significantly different from that of the control. 3) (Figures 4B, 5B, 6B) generally demonstrated a significant increase in C at the target sites tested. It was found that mA3 does not induce editing of the two Ts (Fig. 4C, 5C, 6C). It contains cytidine deaminase (CDA) domains, i.e., CDA1 and CDA2 ( When the CDA2 domain was removed from full-length mA3, the resulting The base editor mA3CDA1-BE3 (Figures 4B, 5B, and 6B) is a These results suggest that the mA3-CDA2 domain is involved in the transcription of mAb-mediated transcription. This indicates that α-glucan is an inhibitor of base editing.
[0038] Surprisingly, the mA3-CDA2 domain also suppresses the base editing activity of mA3-CDA1. Not only can they inhibit the nucleotide sequence of the nucleotides in ... mA3-CDA2 was purified by three active BEs (mA3CDA1-BE3, BE3 and hA3A- When fused to the N-terminus of each of mA3rev-BE3 and mA3rev-BE3, the fusion protein mA3-CDA2-BE3 and mA3-CDA2-hA3A-BE3 (Figures 4B, 5B, 6B) clearly reduced base editing efficiency (Figures 4C, 5C, and 6C).
[0039] Furthermore, cleavage of mA3-CDA2 from the fusion protein restored base editing efficiency (Figure 1). 4C, 5C, 6C), indicating that the inhibition of mA3-CDA2 is related to its covalent binding to BE. It suggests.
[0040] Similar to mA3, human APOBEC3B (hA3B) also encodes two cytidine deaminases (CDA) domains CDA1 and CDA2 (Figures 7A, 8A, 9A). Incorporation of 3B into hA3B-BE3 (Figures 7B, 8B, and 9B) resulted in the activation of the three tested targets. Only relatively low levels of C to T editing were induced at the nucleotide site (Figs. 7C, 8C, 9 C). However, the hA3B-CDA1 domain was deleted. hA3B-CDA2-BE3 (Figures 7B, 8B, and 9B) induces higher C-to-T editing. These results suggest that hA3B-CDA1 plays a key role in the regulation of base editing. inhibitor, and that the inhibition of hA3B-CDA1 is related to its covalent binding to BE. This shows:
[0041] Using the sequences of mA3-CDA2 and hA3B-CDA1, we identified the proteins Identifying additional nucleobase deaminase inhibitors / domains in protein databases Table 1 shows significant sequence homology to the mA3-CDA2 core sequence. 44 proteins / domains are shown (Figure 39), and Table 2 shows the effective Forty-three proteins / domains with significant sequence homology are shown (Figure 40). All proteins and domains, and their variants and equivalents, are nucleotide bases. It is believed to have deaminase inhibitory activity.
[0042] Fusion proteins Based on these surprising and anticipated findings, base editing specificity and efficiency were improved. Fusion proteins are designed that can be used to generate base editors. The present disclosure provides a first fragment comprising a nucleobase deaminase or a catalytic domain thereof. a second fragment comprising a nucleobase deaminase inhibitor, and a protease cleavage site between said first fragment and said second fragment. Deliver quality.
[0043] Base editors incorporating such fusion proteins either have reduced editing ability or No off-target mutations occur, resulting in reduced or no off-target mutations. The cleavage site is cleaved, releasing the nucleobase deaminase inhibitor from the fusion protein at the target site. When the base is released, the base editor at the target site efficiently You will be able to edit it.
[0044] As used herein, the term "nucleobase deaminase" refers to a nucleobase deaminase that is capable of degrading cytidine, deoxycytidine, It catalyzes the hydrolytic deamination of nucleobases such as thiamine, adenosine, and deoxyadenosine. Non-limiting examples of nucleobase deaminases include cytidine deaminase, These include adenosine deaminase and adenosine deaminase.
[0045] "Cytidine deaminase" refers to the deoxycytidine deaminase of uridine, cytidine, and deoxycytidine. and refers to the enzyme that catalyzes the irreversible hydrolytic deamination of uridine to deoxyuridine. Cytidine deaminases maintain the cellular pyrimidine pool. Lee is a researcher at the University of Tokyo who is studying APOBEC (apolipoprotein B mRNA editing enzyme, catalytic polypeptide-like Members of this family are C to U editing enzymes. Family members have two domains and are classified as APOBEC-like proteins (APOBEC One domain of the ATP-like proteins is the catalytic domain and the other The domain is a pseudo-catalytic domain. More specifically, the catalytic domain is a zinc-dependent cytidine deoxynucleotide. Aminase domain, important for cytidine deamination by APOBEC-1 RNA editing requires homodimerization, and this complex interacts with RNA-binding proteins. This acts to form editosomes.
[0046] Non-limiting examples of APOBEC proteins include APOBEC1, APOBEC2, APOBEC3, APOBEC4, APOBEC5, APOBEC6, APOBEC7, APOBEC8, APOBEC9, APOBEC10, APOBEC11, APOBEC12, APOBEC13, APOBEC14, APOBEC15, APOBEC16, APOBEC17, APOBEC18, APOBEC19 ...9, APOBEC11, APOBEC12, APOBEC13, APOBEC14, APOBEC15, APOBEC16, APOBEC17, AP OBEC3A, APOBEC3B, APOBEC3C, APOBEC3D, APOBEC 3F, APOBEC3G, APOBEC3H, APOBEC4, and activation-induced (Citrine cytidine) deaminase (activation-induced (cytidine) deaminase (AID)).
[0047] Different variants of APOBEC proteins confer different editing properties to base editors For example, in the case of human APOBEC3A, certain mutants (e.g., W98 Y, Y130F, Y132D, W104A, D131Y and P134Y) are used to measure the editing efficiency. or in terms of editing window, it is superior to wild-type human APOBEC3A. Therefore, the term APOBEC and each of its family members refers to the corresponding wild-type APOB A certain level (e.g., 70%, 75%) of the EC protein or its catalytic domain , 80%, 85%, 90%, 95%, 98%, 99%) of sequence identity, and Variants and mutants that retain aminating activity are also included. The amino acid sequence may be derived from the addition, deletion and / or substitution of amino acids. In some embodiments, it is a conservative substitution.
[0048] Adenosine aminohydrolase, also known as ADA, Purine purine is an enzyme involved in purine metabolism (EC 3.5.4.4). It is required for the breakdown of adenosine and for the turnover of nucleic acids in tissues.
[0049] Non-limiting examples of adenosine deaminases include tRNA-specific adenosine deaminases. osine deaminase (TadA), adenosine deaminas e tRNA specific 1(ADAT1), adenosine deami nase tRNA specific 2(ADAT2), adenosine de aminase tRNA specific 3(ADAT3), adenosine deaminase RNA specific B1(ADARB1), adeno sine deaminase RNA specific B2(ADARB2), a denosine monophosphate deaminase 1(AMPD1) ), adenosine monophosphate deaminase 2(AM PD2), adenosine monophosphate deaminase 3 (AMPD3), adenosine deaminase (ADA), adenosi ne deaminase 2(ADA2), adenosine deaminase like(ADAL), adenosine deaminase domain c ontaining 1(ADAD1), adenosine deaminase d omain containing 2(ADAD2), adenosine deam inase RNA specific (ADAR) and adenosine dea It contains the enzyme ATP-dependent RNA specific B1 (ADARB1).
[0050] Some nucleobase deaminases have a single catalytic domain, while others have a single catalytic domain. They also contain other domains, such as an inhibitory domain as recently discovered by the authors. In some embodiments, the first fragment is a catalytic domain (e.g., mA3-C In some embodiments, the first flag The member comprises at least the catalytic core of the catalytic domain. As shown, when mA3-CDA1 is truncated at residues 196 / 197, The CDA1 domain still retained substantial editing efficiency (Figures 10C, 11C, 1 2C).
[0051] The present disclosure provides two nucleobase deaminases, which are inhibitory domains of the corresponding nucleobase deaminases. The enzyme inhibitors, mA3-CDA2 and hA3B-CDA1, were tested. Additional nucleobase deaminase inhibitors and inhibitors are available in the protein database. Harm domains were also identified (see Tables 1 and 2). Their biological equivalents (e.g., At least approximately 80%, 85%, 90%, 95%, 97%, 98%, 99%, 99.5% have sequence identity or have one, two or three amino acid additions / deletions / substitutions and have nucleobase deaminase inhibitor activity) may also be modified by conservative amino acid substitutions, etc. The nucleic acid deaminers can be prepared using any method known in the art. "Deaminase inhibitors" are proteins or enzymes that inhibit the deaminase activity of nucleic acid base deaminases. In some embodiments, the second fragment is an inhibitory It contains at least an inhibitory core of a protein / domain. For example, Thus, if mA3-CDA2 retains residues 334–429, CDA2 still As a result, it had an inhibitory effect on base editing (Figures 13B, 14B, and 15B).
[0052] In some embodiments, the fusion protein optionally contains, in the first fragment, a nucleic acid salt. clustered and regularly spaced adjacent to the base deaminase or its catalytic domain It further contains short palindromic repeat (CRISPR) associated (Cas) proteins.
[0053] "Cas proteins" or "clustered regularly spaced short palindromic structures" The term "CRISPR-associated repeat (Cas) protein" refers to the Streptococcus pyogenes (S CRISPR (class Streptococcus pyogenes) and other bacteria (Targeted regularly interspaced short palindromic repeats) associated with the adaptive immune system RNA-guided DNA endonuclease enzyme Cas proteins include the Cas9 protein , Cas12a (Cpf1) protein, Cas12b (formerly known as C2c1) The Cas13 protein and various engineered Examples of Cas proteins include SpCas9, FnCas9, and their corresponding Cas proteins. , St1Cas9, St3Cas9, NmCas9, SaCas9, AsCpf1, Lb Cpf1, FnCpf1, VQR SpCas9, EQR SpCas9, VRER S pCas9, SpCas9-NG, xSpCas9, RHA FnCas9, KKH SaCas9, NmeCas9, StCas9, CjCas9, AsCpf1, FnCp f1, SsCpf1, PcCpf1, BpCpf1, CmtCpf1, LiCpf1, P mCpf1, Pb3310Cpf1, Pb4417Cpf1, BsCpf1, EeCpf 1, BhCas12b, AkCas12b, EbCas12b, LsCas12b, Rf Cas13d, LwaCas13a, PspCas13b, PguCas13b, Ran Cas13b and those shown in Table A below.
[0054] Table A. Examples of Cas proteins [Table 1]
[0055] The protease cleavage site between the first and second fragments allows for the selection of any protease. The protease cleavage site (peptide) may be any known protease cleavage site (peptide) for a protease. Non-limiting examples of proteases include TEV protease, TuMV protease, PPV protease, protease, PVY protease, ZIKV protease and WNV protease The protein sequences of exemplary proteases and their corresponding cleavage sites are shown in Table B. provide.
[0056] Table B. Example sequences [Table 2-1] [Table 2-2]
[0057] In some embodiments, the protease cleavage site is a self-cleaving peptide, such as a 2A peptide. The "2A peptide" is a peptide that is produced by the "cleavage" of a polypeptide during translation in eukaryotic cells. The virus is a 18-22 amino acid long oligopeptide that transmits the virus. Refers to a specific region of the viral genome, and different viral 2A generally refer to the region of the virus from which they are derived. The first 2A virus discovered was F2A (foot-and-mouth disease virus). ot-and-mouth disease virus), and then E2A (horse equine rhinitis A virus), P2A (porcine rhinitis virus) porcine teschovirus-1 2A , and T2A (thosea asigna virus 2A (thosea asigna virus A virus 2A) peptide has also been identified. Some non-limiting examples of 2A peptides include: They are provided in columns 26 to 28.
[0058] In some embodiments, the protease cleavage site is a cleavage site for a TEV protease (e.g., For example, SEQ ID NO: 5. In some embodiments, the fusion protein comprises a TEV protein. The present invention further comprises a third fragment comprising a lyase or a fragment thereof. In this form, the TEV protease fragment in the fusion protein is inactive, i.e. That is, it cannot cleave the TEV cleavage site by itself. In the presence of the remaining part of the enzyme, this fragment is able to carry out the cleavage. As further described below, such configurations provide additional control and flexibility in base editing capabilities. The TEV fragment provides the TEV N-terminal domain (e.g., SEQ ID NO: 3). Or it can be the TEV C-terminal domain (eg, SEQ ID NO: 4).
[0059] Various configurations of the fragments can be made. Non-limiting examples include N-terminal to C Distally comprising: (1) a first fragment (e.g., catalytic domain)-protease cleavage site-second fragment a fragment of (e.g., inhibitory domain); (2) a first fragment (e.g., a catalytic domain and a Cas protein)—a tease cleavage site - second fragment (e.g., inhibitory domain); (3) The first fragment (e.g., catalytic domain, Cas protein, and TEV) N-terminal domain) - protease cleavage site (e.g., TEV cleavage site) - second flag ment (e.g., inhibitory domains); (4) a second fragment (e.g., an inhibitory domain)—a protease cleavage site (e.g., a first fragment (e.g., a catalytic domain, a Cas protein) and the TEV N-terminal domain); and (5) a second fragment (e.g., an inhibitory domain)—a protease cleavage site (e.g., a first fragment (e.g., a TEV cleavage site) - a second fragment (e.g., a Cas protein, catalytic domain) , and the TEV C-terminal domain).
[0060] In some embodiments, a first nucleobase deaminase (e.g., a cytidine deaminase a first fragment comprising a nucleotide sequence encoding ... a second fragment comprising an inhibitory domain of said first nucleobase deaminase; is different from the second nucleobase deaminase. In one embodiment, each of the first and second nucleobase deaminases is human and Mouse APOBEC3B (A3B), APOBEC3C (A3C), and APOBEC3D ( A3D), APOBEC3F(A3F), APOBEC3G(A3G), APOBEC3H (A3H), APOBEC1(A1), APOBEC3(A3), APOBEC2(A2 ), APOBEC4 (A4) and AICDA (AID).
[0061] The fusion protein may contain other fragments (e.g., uracil DNA glycosylase inhibitors) It may contain a nucleotide sequence (UGI) and a nuclear localization sequence (NLS).
[0062] Prepared from Bacillus subtilis bacteriophage PBS1 A possible "uracil glycosylase inhibitor" (UGI) is E. coli uracil-DNase. A glycosylase (UDG) and small proteins that inhibit UDG from other species (9 UDG inhibition is achieved by a stoichiometry of UDG:UGI=1:1 (s UGI is caused by reversible protein binding of UDG. A non-limiting example of a UGI is Bacillus fungus. In some embodiments, the phage AR9 (YP_009283008.1) The UGI comprises the amino acid sequence of SEQ ID NO: 25 or is at least partially similar to SEQ ID NO: 25. also have 70%, 75%, 80%, 85%, 90% or 95% sequence identity, It retains glycosylase inhibitory activity.
[0063] The fusion protein, in some embodiments, includes one or more nuclear localization sequences (NLS). It can be seen.
[0064] "Nuclear localization signals or sequences" (NLS) are used to identify proteins that are imported into the cell nucleus via nuclear transport. A signal is an amino acid sequence that is tagged onto a protein for transcription. Typically, this signal is One or more short sequences of positively charged lysines or arginines exposed on the protein surface It is possible that different nuclear-localized proteins share the same NLS. It has the opposite function to the nuclear export signal (NES), which targets proteins to leave the nucleus. A non-limiting example of an NLS is the internal SV40 nuclear localization sequence (NLS). SV40 nuclear localization sequence (iNLS) )
[0065] In some embodiments, a peptide linker is optionally provided between each flag in the fusion protein. In some embodiments, the peptide linker is provided between 1 and 100 amino acids. having amino acid residues (or 3-20, 4-15, but not limited to these) In some embodiments, at least 10%, 20%, or more of the amino acid residues of the peptide linker %, 30%, 40%, 50%, 60%, 70%, 80% or 90% of the The amino acid residue is selected from the group consisting of arginine, cysteine, and serine.
[0066] For any fusion protein of the present disclosure, a biological equivalent thereof is also provided. In some embodiments, a biological equivalent is at least about 70%, 75% or more of a fusion protein that is biocompatible with a reference fusion protein. , 80%, 85%, 90%, 95%, 98%, or 99% sequence identity. Alternatively, a biological equivalent retains the desired activity of the reference fusion protein. In this embodiment, a bioequivalent may be one, two, three, four, five or more amino acids. The amino acid sequence may be derived by including amino acid additions, deletions, substitutions, or combinations thereof. In some embodiments, the substitutions are conservative amino acid substitutions.
[0067] A "conservative amino acid substitution" is one in which an amino acid residue is replaced with an amino acid residue having a similar side chain. Families of amino acid residues having similar side chains have been defined in the art. These include basic side chains (e.g., lysine, arginine, histidine), acidic side chains (e.g., aspartic acid, glutamic acid), uncharged polar side chains (e.g., glycine, aspartic acid, glutamic acid), paragine, glutamine, serine, threonine, tyrosine, cysteine), non-polar side chains (e.g. For example, alanine, valine, leucine, isoleucine, proline, phenylalanine, methyl threonine, tryptophan), beta-branched side chains (e.g., threonine, valine, isoleucine amino acids) and aromatic side chains (e.g., tyrosine, phenylalanine, tryptophan, histidine, Therefore, the non-essential amino acid residues in immunoglobulin polypeptides are , preferably with another amino acid residue from the same side chain family. In these sequences, the strings of amino acids differ in the order and / or composition of side chain family members. The nucleotide sequence may be replaced with a structurally similar string.
[0068] Non-limiting examples of conservative amino acid substitutions are shown in the table below. In the table, a similarity score of 0 or greater indicates , indicates a conservative substitution between two amino acids.
[0069] Table C. Amino acid similarity matrix [Table 3]
[0070] Table D. Conservative Amino Acid Substitutions [Table 4]
[0071] On-target activation of fusion proteins The present disclosure also provides a method for detecting nucleobase deaminases or both their catalytic domains and inhibitors. and methods for activating a fusion protein of the present disclosure when its activity is desired. This technique is shown in Figure 16.
[0072] In an exemplary configuration, the fusion protein (A) comprises (a) a nucleobase deaminase (e.g., cytidine deaminase) or its catalytic domain, optionally clustered and regularly CRISPR-associated (Cas) proteins, and and a first fragment comprising a first TEV protease fragment; (b) a nucleobase derivative; a second fragment comprising an aminase inhibitor; and (c) a second fragment comprising the first fragment. The fragment contains a TEV protease cleavage site between the first fragment and the second fragment. In embodiments, the first TEV protease fragment alone does not inhibit the TEV protease. The ase cleavage site cannot be cleaved.
[0073] Fusion proteins can be used in vitro or in vivo to induce gene editing in cells. When performing aggregation, two additional molecules can be introduced. In one example, one molecule (B) The single guideline further incorporates a tag sequence that can be recognized by an RNA recognition peptide. Alternatively, the sgRNA may target the target site. crRNA and CRISPR RNA (crRNA) alone or transactivating The tag sequence and the CRISPR RNA (tracrRNA) can be used in combination. Examples of the corresponding RNA recognition peptides include MS2 / MS2 coat protein (MCP) , PP7 / PP7 coat protein (PCP), and boxB / boxB coat protein The sequences of these proteins are shown in Table B. Molecules (B) are RNA It may be provided as a DNA sequence encoding the molecule.
[0074] Other additional molecules (C) may, in some embodiments, be RNA recognition peptides (e.g., M It contains a second TEV protease fragment bound to the CP, PCP, and N22p. The first TEV fragment and the second TEV fragment are, in some embodiments, When present together, the TEV protease site can be cleaved.
[0075] Such coexistence is due to the interaction of the tag sequence with the RNA recognition protein, molecule (C) can be triggered by binding to molecule (B). On the other hand, fusion protein (A) and Both molecules (B) will be present at the target genomic locus for gene editing. Therefore, molecule (B) inhibits the TEV protease of both fusion protein (A) and molecule (C). The fragments are combined, which activates the TEV protease and forms the fusion protein. This results in the removal of nucleobase deaminase inhibitors from the target and activation of base editors. Such activation occurs only at the target genomic site, and off-target single-stranded DNA fragments are not detected. It is easy to understand that base editing does not occur in NA regions. Therefore, base editing is This does not occur in single-stranded DNA regions where A does not bind (as shown in Figures 17-19).
[0076] "Guide RNA" is a non-coding RNA that binds to a complementary target DNA sequence. Guide RNA is a non-coding short RNA sequence. First, the Cas enzyme binds to the gRNA, which then binds to a specific site on the DNA. Cas then guides the endonuclease to the target DNA strand, where it activates the endonuclease. Single-guide RNAs, often simply called "guide RNAs," exert cleavage activity. "RNA" refers to a synthetic or recombinant RNA consisting of both crRNA and tracrRNA as a single construct. refers to the expressed single guide RNA (sgRNA). The tracrRNA portion is Ca The crRNA portion is responsible for the endonuclease activity and binds to the target-specific DNA region. Therefore, transactivating RNA (tracrRNA, or scaffold The nucleotide sequence (nucleotide region) and crRNA are two important components, connected by a tetraloop. This results in the formation of the sgRNA.
[0077] The scaffold of the guide RNA itself has a stem-loop structure, and the endonuclease A typical scaffold has the structure shown in Figure 36A (top). It has a structure comprising, from the 5' end to the 3' end, (A) a repeat region, (b) a tetraloop, (c) an anti-repeat that is at least partially complementary to the repeat region; (d) a stem-loop. (e) linker, (f) stem loop 2, and (g) stem loop 3. The nucleotide sequences are generally conserved, but the sequences of stem-loop 1 and stem-loop 3 are The loops can have different sequences. More importantly, the tetraloops and stem loops The loop of 2 can be completely replaced with a much longer sequence. A sequence such as P7, box B) can be inserted here to allow recognition by the corresponding recognition peptide. An example of a scaffold sequence is shown below.
[0078] Table E. Example sgRNA scaffold sequences [Table 5]
[0079] With reference to these exemplary scaffold sequences, the fragment at positions 1-12 ( For example, GUUUUAGAGCUA, SEQ ID NO: 197; GUUUGAGAGCUA, SEQ ID NO: Number 198) represents the repeat region, which is located at positions 17-30 (e.g., UAGCAAG UUAAAAU, SEQ ID NO: 199) and approximately 8 to 12 base pairs. The GAAA loop (SEQ ID NO: 200) between them is a tetraloop. As shown in number 17, this entire loop can be replaced with the MS2 sequence. Loop 1 includes approximately positions 31-39 and is a small loop (e.g., but not limited to, UA , AU, AA, or UU). Stem loop 1 generally consists of 3-4 amino acids in the stem. Positions 48 to 61 (e.g., AACUUGAAAAAGUG, SEQ ID NO: Stem-loop 2, which contains 201), generally involves four base pairs in the stem and a completely substituted nucleotide. The remaining positions 62-76 (e.g., GCACCGAGUCGGUGC, SEQ ID NO: 202;GCACCGAUUCGGUGC; SEQ ID NO: 203) constitutes stem loop 3, which generally consists of four base pairs in the stem. The small loop (U and G in this example) can be any nucleotide.
[0080] Thus, the scaffold sequence can be represented as follows: [ka] Here, N represents any base, and X1 and X2 represent any nucleotides having a length of 2 to 50 bases. The terms "guide RNA" and "single guide RNA" refer to RNA sequences. Additional sequences such as MS2, PP7 and boxB inserted into one or more loops in A It includes those that include.
[0081] Nucleobase deaminases, catalytic domains, nucleobase deaminase inhibitors, and C Various embodiments and examples of AS proteins are provided in the present disclosure. For example, The nucleobase deaminase can be a cytidine deaminase and an adenosine deaminase. Non-limiting examples of cytidine deaminases include APOBEC1, APOBEC2, APOBEC3, APOBEC4, APOBEC5, APOBEC6, APOBEC7, APOBEC8, APOBEC9, APOBEC10, APOBEC11, APOBEC12, APOBEC13, APOBEC14, APOBEC15, APOBEC16, APOBEC17, APOBEC18, APOBEC19 ...9, APOBEC11, APOBEC12, APOBEC13, APOBEC14, APOBEC15, APOBEC16, APOBEC OBEC3A, APOBEC3B, APOBEC3C, APOBEC3D, APOBEC 3F, APOBEC3G, APOBEC3H, APOBEC4, and activation-induced (Citrine cytidine) deaminase (activation-induced (cytidine) deaminase).
[0082] Non-limiting examples of adenosine deaminases include tRNA-specific adenosine deaminases. osine deaminase (TadA), adenosine deaminas e tRNA specific 1(ADAT1), adenosine deami nase tRNA specific 2(ADAT2), adenosine de aminase tRNA specific 3(ADAT3), adenosine deaminase RNA specific B1(ADARB1), adeno sine deaminase RNA specific B2(ADARB2), a denosine monophosphate deaminase 1(AMPD1) ), adenosine monophosphate deaminase 2(AM PD2), adenosine monophosphate deaminase 3 (AMPD3), adenosine deaminase (ADA), adenosi ne deaminase 2(ADA2), adenosine deaminase like(ADAL), adenosine deaminase domain c ontaining 1(ADAD1), adenosine deaminase d omain containing 2(ADAD2), adenosine deam inase RNA specific (ADAR) and adenosine dea It contains the enzyme ATP-dependent RNA specific B1 (ADARB1).
[0083] Examples of Cas proteins include SpCas9, FnCas9, St1Cas9, and St3C. as9, NmCas9, SaCas9, AsCpf1, LbCpf1, FnCpf1, V QR SpCas9, EQR SpCas9, VRER SpCas9, RHA FnC as9 and KKH SaCas9, as well as those provided in Table A.
[0084] The fusion protein may contain other fragments (e.g., uracil DNA glycosylase inhibitors) The nucleic acid sequence may include a nucleic acid sequence encoding a nucleic acid fragment (GUI) and a nuclear localization sequence (NLS), each of which is described herein. It is discussed in the book.
[0085] The base editors and base editing methods described in this disclosure can be used to identify genes in the genomes of various eukaryotic organisms. It can be applied to perform highly specific and efficient base editing in DNA.
[0086] The present disclosure provides compositions and methods. Such compositions comprise an effective amount of a fusion protein. In some embodiments, the composition comprises a target D Such compositions further comprise a guide RNA having the desired complementarity to the NA. It can be used for base editing in samples.
[0087] The fusion proteins and compositions can be used for base editing. In one embodiment, A method for editing a target polynucleotide, comprising: a guide having at least partial sequence complementarity to the target polynucleotide; contacting said target polynucleotide with a target RNA, said editing deamination of cytosine (C) in the target polynucleotide. can be.
[0088] In one embodiment, a method for editing cytosines on a nucleic acid sequence in a sample is provided. In some embodiments, the method comprises administering to a subject a fusion protein of the present disclosure or a method comprising administering to a subject a fusion protein of the present disclosure. Some embodiments involve contacting a sample with a polynucleotide encoding the protein. In this embodiment, an appropriate guide RNA is further added. Design of the guide RNA is easily accomplished by those skilled in the art. is available for.
[0089] Contact between the fusion protein (and guide RNA) and the target polynucleotide , in vitro, especially in a cell culture If the contacting is ex vivo or in vivo, the fusion protein may be used in clinical trials. In vivo contact may include, but is not limited to, The administration may be to a living subject, such as a human, animal, yeast, plant, bacteria, virus, etc.
[0090] Configuration of guided and split base editors Various configurations of the constructs were generated using the inducible and split base editor (isplitBE) The design was tested for implementation (Fig. 24). Among the tested configurations (Fig. 25), Example 3 Pair 9 showed excellent editing efficiency and minimized off-target editing (significantly improved specificity). Pair 9 uses a dual-sgRNA system, with the helper sgR hsgRNA is used to target sites adjacent to the main target site. Such dual targeting improves specificity (Figures 32-33).
[0091] In construct pair 9 (Figures 25-28), the nucleobase deaminase inhibitor binds both sg Only when RNA binds to the target sequence is it released, and the nucleobase deaminase is released This ensures that no editing occurs at the target site. They can be generated, for example, from two separate constructs (Figures 26A and 34A).
[0092] The first molecule can include only a Cas protein, and the Cas protein is generally The second molecule has a size suitable for packaging into an AAV, which is a suitable vehicle. , among others, nucleobase deaminases (e.g., APOBEC), nucleobase deaminases inhibitors (e.g., mA3-CDA2), and RNA recognition peptides (e.g., MCP ) containing a protease cleavage site (e.g., a TEV site) that allows the nucleobase deaminase and nucleic acid Inserted between the base deaminase inhibitor and the nucleotide sequence, allowing for proper timing and location. Optionally, the second molecule is UG Further includes I.
[0093] The third molecule is a proteasome fused to a different RNA recognition peptide (e.g., N22p). The fourth molecule is a fusion between the first portion and an inactive portion of a fusion protein (e.g., TEVc). In combination, they perform a protease activity to release nucleobase deaminase inhibitors from a second molecule. It is a standalone TEVn that can remove bitterness. do.
[0094] The fifth molecule contains an RNA recognition site recognizable by an RNA recognition peptide in the second molecule. The sixth molecule is a helper sgRNA containing the R in the third molecule (e.g., MS2). A typical RNA recognition site (e.g., box B) that can be recognized by an RNA recognition peptide. It is an sgRNA.
[0095] Both the hsgRNA and sgRNA are expressed at the correct target site in the genome (or RNA). Each binds to the target site and recruits a Cas protein to the target site. The binding between MS2 and MCP also recruits a second molecule, and the sgRNA binds to boxB. The binding between N22p and the third molecule is then recruited. EVc is in contact with the TEV site. Standalone TEVn ) is present throughout the cell, so it may be present here as well. This allows TEVc to become active. The nucleobase deaminase inhibitor is cleaved from the nucleobase deaminase in the second molecule. This ensures that the nucleobase deaminase is activated.
[0096] Furthermore, the optimal distance between the hsgRNA binding site and the normal sgRNA binding site is 34 It was found that the gene is ~91 bp (PAM to PAM) and that the hsgRNA is located upstream. It was.
[0097] Furthermore, intentional editing at the target site for a typical sgRNA involves hsg Proper binding of both RNA and regular sgRNA is required, but for hsgRNA Editing at the target site is undesirable. If the spacer length (the spacer is the target-complementary region) is 8 to 15 bases, such A suitable hsgRNA still provides dual recognition to ensure binding specificity. found that this was sufficient to significantly reduce editing at the hsgRNA target site. will be done.
[0098] Thus, according to one embodiment of the present disclosure, a nucleobase deaminase or its catalytic domain is a nucleic acid, a nucleic acid base deaminase inhibitor, a first RNA recognition peptide, and the nucleic acid A base deaminase or its catalytic domain and said nucleobase deaminase inhibitor A fusion protein is provided, which includes a first fragment containing a TEV protease cleavage site between the first and second fragments. It is served.
[0099] In some embodiments, the fusion protein comprises the TEV protease cleavage site. TEV protease fragments that cannot be cleaved alone and a second RNA recognition In some embodiments, the fusion further comprises a second fragment comprising a recognition peptide. The protein has a self-cleavage site between the first fragment and the second fragment. Further includes:
[0100] In some embodiments, the fusion protein comprises a second TEV protease fragment. a third fragment containing the first TEV protease fragment, The TEV protease fragment is cleaved in the presence of the second TEV protease fragment. In some embodiments, the fusion protein can cleave the second a second self-cleavage site between the first fragment and the third fragment, Upon cleavage of the second self-cleavage site, the fusion protein is fused to an RNA recognition peptide. This releases the second TEV protease fragment.
[0101] Also, in one embodiment, a sequence A target single guide RNA containing a first spacer with complementarity to a second PAM a second spacer having sequence complementarity to a second nucleic acid sequence adjacent to the site; Per single guide RNA, a set of clustered regularly spaced short palindromic repeats CRISPR-associated (Cas) proteins, and nucleobase deaminases, Dual guide RNA systems are provided.
[0102] In some embodiments, the second PAM site is located 150 or more bases from the second PAM site. Within, or 140, 130, 120, 110, 100, 95, 94, 93, 92, 91 , 90, 89, 88, 87, 86, 85, 84, 83, 82, 81, 80, 75 or 7 In some embodiments, the second PAM site is located within 0 bases of the first PAM. or at least 10 bases, or at least 15, 20, 25, 30, 31, 32, 3 3, 34, 35, 36, 37, 38, 39, 40, 45, 50, 55, or 60 bases In some embodiments, the second PAM site is upstream of the first PAM site. In some embodiments, the second PAM site is downstream of the first PAM site. In some embodiments, the distance may be, but is not limited to, 20-100, 25-95, 3 0~95, 34~95, 34~91, 34~90, 35~90, 40~90, 40~84 , 45 to 85, or 50 to 80 bases.
[0103] In some embodiments, the second (helper) spacer is 8 to 15 bases in length. In some embodiments, the second spacer is 8 to 14, 8 to 13, or 8 to 1 2, 8-11, 8-10, 9-15, 9-14, 9-13, 9-12, 9-11, 9-1 0, 10-15, 10-14, 10-13, 10-12, 10-11, 11-15, 11 ~14, 11~13, 11~12, 12~15, 12~14, 12~13, or 13~ In contrast, the first spacer is at least 16, 17, 18, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, It is 8 or 19 bases in length.
[0104] Cas protein and nucleobase deaminase are delivered in separate delivery vehicles (e.g., AAV ) to allow for packaging of various "split" base editing systems. Also described herein.
[0105] In some embodiments, a method for generating a premature stop codon in a PCSK9 gene Conventional sgRNAs and sgRNAs can mediate efficient editing and may have clinical benefits. Based on the findings herein, pairs of hsgRNAs and hsgRs are provided. A suitable target site for NA is selected to convert a non-stop codon to a stop codon. For example, when editing C to T / U, the non-stop codons are CAG, CAA, or CG. It could be A.
[0106] Examples of such target sites are shown in Table 4. The sequences in Table 4 indicate the target locations. However, it is readily understood that the actual sgRNA and hsgRNA sequences are used for A does not need to bind to the entire sequence. Indeed, for example, for hsgRNA, As mentioned above, binding of 8 to 15 nucleotides may be sufficient. The spacer sequence on A can be complementary to any of the subsequences shown in Table 4, or The same applies to the sgRNA, which is preferred. The length of the spacer is not limited, but is 18 to 24 nucleotides.
[0107] In one embodiment, a helper guide RNA / guide for editing a human PCSK9 nucleic acid sequence is provided. A pair of guide RNAs is provided, wherein the guide RNA is a first site on the PCSK9 nucleic acid. It specifically targets the target gene and converts non-stop codons into stop codons by base editing. the helper guide RNA is a PCSK9 nucleic acid 20 to 100 bases away from the first site In some embodiments, the second site is specifically targeted to: From the first site, approximately 20-100, 25-95, 30-95, 34-95, 34-91, 3 4-90, 35-90, 40-90, 40-84, 45-85, or 50-80 bases apart It is being done.
[0108] In some embodiments, the hsgRNA has a spacer that is 8-15 bases in length. In some embodiments, the spacer is 8 to 14, 8 to 13, 8 to 12, 8 to 11, 8 ~10, 9~15, 9~14, 9~13, 9~12, 9~11, 9~10, 10~15, 10-14, 10-13, 10-12, 10-11, 11-15, 11-14, 11-1 3, 11–12, 12–15, 12–14, 12–13, or 13–15 bases in length In some embodiments, the sgRNA comprises at least 16, 17, 18, or 19 amino acids. It has a base-long spacer.
[0109] The spacer sequence of the sgRNA / hsgRNA can be easily designed. For example, For each target site shown in Table 4, the spacer is a complementary sequence of the desired length (i.e., (i.e., complementary to any subsequence of SEQ ID NOs: 166 to 180 or 181 to 195) Specific examples of binding site pairs include, but are not limited to, SEQ ID NO: 166 and and 181; SEQ ID NOs: 167 and 182; SEQ ID NOs: 168 and 183; SEQ ID NO: 169 and 184; SEQ ID NOs: 170 and 185; SEQ ID NOs: 171 and 186; SEQ ID NO: 1 72 and 187; SEQ ID NOs: 173 and 188; SEQ ID NOs: 174 and 189; SEQ ID NOs: SEQ ID NOs: 175 and 190; SEQ ID NOs: 176 and 191; SEQ ID NOs: 177 and 192; Sequence numbers 178 and 193; sequence numbers 179 and 194; and sequence numbers 180 and and 195 are included.
[0110] Exemplary sgRNA / hsgRNA sequences were also designed and tested. See Table 3. Furthermore, polynucleotide sequences encoding helper guide RNAs and guide RNAs is also provided.
[0111] Using such sgRNA / hsgRNA sequence pairs, we investigated the expression of the PSCK9 gene in cells. In some embodiments, the method comprises inactivating the cells of the present invention. Disclosed are pairs of helper guide RNAs and guide RNAs, clustered and regularly spaced. CRISPR-associated (Cas) proteins, and the nucleus Each of these elements is described in this disclosure. is further explained.
[0112] Enhanced Prime Editing In some embodiments, an improved prime editing system is also provided. The specific prime-editing guide RNA (pegRNA) molecules provided herein are These pegRNAs have one additional nucleotide sequence compared to conventional guide RNAs. The scaffold contains 100 base pairs (see Figures 36A and 36E). The standard scaffold (SEQ ID NO: 31) was used in the template to develop the improved scaffold. The fold may have the sequence of any of SEQ ID NOs: 32-43.
[0113] As mentioned above, a typical guide RNA scaffold consists of a 5'-end to a 3'-end , (a) repeat region, (b) tetraloop, (c) at least partially in the repeat region complementary antirepeat, (d) stem-loop 1, (e) linker, (f) stem-loop 2, and (g) a stem-loop 3. In other words, the scaffold , containing four stem loops. The third stem loop (5 The nucleotide sequence (counting from ' to 3') contains four base pairings in the conventional design. In the original design, this stem loop has five base pairs.
[0114] In one embodiment, in the 5' to 3' direction, a first stem loop portion, a second stem loop portion, a scaffold comprising a first portion, a third stem-loop portion, and a fourth stem-loop portion wherein the third stem loop contains five base pairs therein. A guide RNA is provided.
[0115] The sequence of the scaffold can be represented as follows: [ka] Here, N represents any base, and X1 and X2 each represent 2 to 50 bases (or 2 to 40, 3 to 50, or 4 to 50, or 5 to 50, or 6 to 60, or 7 to 80, or 8 to 90, or 9 to 100, or 10 ... Any nucleotide sequence of length (40, 4-40, 4-30, 2-30, 4-20 bases) Thus, in some embodiments, the base pairing occurs at position 31 of SEQ ID NO: 31. Thus, it includes one between positions 45 and 55. In some embodiments, the scaffold The group has at least 75%, 80%, 85%, 90%, 95%, 96%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 29 6%, 97%, 98% or 99% sequence identity and Includes involution.
[0116] Thus, in one embodiment, the stem-loop structure or scaffold / guide RNA A base pairing is introduced between the bases at positions 45 and 55, as long as functionality is maintained, and optionally, Addition, deletion, substitution, or any combination of 1, 2, 3, 4, or 5 bases is possible. By enabling the synthesis of guide RNAs containing scaffolds derived from SEQ ID NO: 31, In some embodiments, the scaffold comprises a sequence selected from the group consisting of SEQ ID NOs: 32-43. In some embodiments, the guide RNA comprises a sequence selected from the group consisting of: 100 nucleotides, or 105, 110, 115, 120, 125, 130, 140 or 150 nucleotides in length. In some embodiments, the guide RNA is a spacer (e.g., 8-25 nucleotides), a reverse transcriptase template, and / or further comprises a primer binding site.
[0117] In some embodiments, improved prime editor proteins are also provided. In an embodiment, the Prime Editor performs testing to optimize the Prime Editor's performance. In one embodiment, the Cas protein and the reverse transcriptase are linked via a linker. In one embodiment, the prime editor comprises the amino acid sequence of SEQ ID NO: 44. The prime editor comprises the amino acid sequence of SEQ ID NO: 45. Both of these prime editors It has been tested and shown to exhibit excellent editing efficiency and specificity.
[0118] Various "split" prime editing systems are also described herein, including Therefore, the Cas protein and reverse transcriptase can be packaged in separate delivery vehicles (e.g., AAV). You can do this.
[0119] Using a split-prime editing system, gene editing can be performed in cells at the target site. Also provided are methods for performing the method. In some embodiments, the method comprises administering to the cell a clonal antibody. Clustered regularly interspaced short palindromic repeats (CRISPR)-related (C as) a first viral particle encapsulating a first construct encoding a protein, and A second virus encapsulating a second construct encoding a reverse transcriptase fused to an A recognition peptide. In some embodiments, the second construct comprises introducing a R The nucleic acid sequence further encodes a guide RNA that includes an RNA recognition site to which the RNA recognition peptide binds.
[0120] In some embodiments, the second construct comprises an R In some embodiments, the C further encodes a guide RNA comprising a guide RNA recognition site. The as proteins are SpCas9, FnCas9, St1Cas9, St3Cas9, and N mCas9, SaCas9, AsCpf1, LbCpf1, FnCpf1, VQR Sp Cas9, EQR SpCas9, VRER SpCas9, SpCas9-NG, xS pCas9, RHA FnCas9, KKH SaCas9, NmeCas9, StCa s9, CjCas9, AsCpf1, FnCpf1, SsCpf1, PcCpf1, Bp Cpf1, CmtCpf1, LiCpf1, PmCpf1, Pb3310Cpf1, Pb 4417Cpf1, BsCpf1, EeCpf1, BhCas12b, AkCas12b , EbCas12b, LsCas12b, RfCas13d, LwaCas13a, Ps pCas13b, PguCas13b, and RanCas13b In some embodiments, the Cas protein is SpCas9-NG or x It is SpCas9.
[0121] Non-limiting examples of reverse transcriptases include human immunodeficiency virus (HIV) reverse transcriptase, Moroni reverse transcriptase, - Murine leukemia virus (MMLV) reverse transcriptase and avian myeloblastosis virus (AMV) ) reverse transcriptase. [Example]
[0122] Example 1: Fusion base editor with reduced off-target editing activity Single guide RNAs (sgRNAs) and base editors (BEs) mentioned in the Examples Unless otherwise noted, the sgRNA for SpCas9 (e.g., sgRNA for SaCas9 (S Current base editing systems are based on the a-sgRNA, which converts C to T in ssDNA regions. To test whether mutations could be induced into SaD10A nickase and Sa-sg RNA was used to create a DNA single-strand break (SSB), which triggered end retraction. SaD10A, Sa- can generate ssDNA regions (Figs. 1A, 2A, and 3A). sgRNAs (Sa-sgSITE31, Sa-sgSITE42, and Sa-sgF1) The two publicly available BEs (i.e., BE3 and hA3A-BE3) or an empty BE The vectors were co-transfected with SaD10A (Figures 1B, 2B, and 3B) and then transfected with SaD10A. Mutagenesis around the targeted ssDNA region was determined using three test sites (Sa-SITE). 31, Sa-SITE42 and Sa-F1), BE3 or hA3A-BE3 Expression of the vector induced a C-to-T mutation, whereas expression of the empty vector did not (Fig. 1C, 2C and 3C). These results support the current understanding of the catalytically active cytidine deaminase. Base editors actually induce unintended mutations in unrelated ssDNA regions (Figs. 1, 2 and 3).
[0123] To inhibit the activity of cytidine deaminase at non-relevant sites, e.g., ssDNA regions Therefore, the present inventors proposed fusing a base editor with a base editing inhibitor. Mouse APOBEC3 (mA3) contains two cytidine deaminase (CDA) domains (CD A1 and CDA2 (Figures 4A, 5A, 6A), and the full-length mA3 in mA3-BE3. The use of (Figures 4B, 5B, and 6B) resulted in C-to-T editing at the three tested target sites. However, the mA3-BE3-mA3CDA2 gene was not induced by the mA3-BE3-mA3CDA2 gene (Fig. 4C, 5C, 6C). mA3CDA1-BE3 (Fig. 4B, 5B, 6B) generated by deleting induced substantial C-to-T editing (Figures 4C, 5C, and 6C). These results suggest that 3CDA2 is a natural inhibitor of base editing. CDA2 was cloned into three active BEs (i.e., mA3CDA1-BE3, BE3, and hA3 mA3rev-BE3, mA3CDA2-BE3, and mA3CDA2-hA3A-BE3 was generated (Figures 4B, 5B, and 6B). Addition of 3CDA2 to the N-terminus clearly reduced base editing efficiency (Figures 4C, 5C, 6 C).
[0124] Next, we investigated whether truncation of mA3CDA2 could restore base editing efficiency. in rev-BE3, mA3CDA2-BE3 and mA3CDA2-hA3A‐BE3. Insertion of the 2A self-cleaving peptide between mA3CDA2 and the rest of the BE, resulting in mA3re v-2A-BE3, mA3CDA2-2A-BE3 and mA3CDA2-2A-hA3 A-BE3 was generated (Figures 4B, 5B, and 6B). BE3, mA3CDA2-2A-BE3 and mA3CDA2-2A-hA3A-BE3 Base editing efficiency was restored (Figures 4C, 5C, and 6C), indicating that inhibition of mA3CDA2 leads to BE. Furthermore, the protein database shows that mA3C A search for domains similar to the DA2 core sequence revealed at least 44 proteins. It was revealed that they have similar domains (Table 1).
[0125] Human APOBEC3B (hA3B) also contains two cytidine deaminase (CDA) domains. (CDA1 and CDA2, Figures 7A, 8A, 9A) in hA3B-BE3 The use of full-length hA3B (Figures 7B, 8B, and 9B) compared across the three tested target sites. However, it induced only relatively low levels of C to T editing (Figs. 7C, 8C, and 9C). hA3BCDA2-BE3 was generated by deleting hA3BCDA1. hA3BCDA2-BE3 (Figures 7B, 8B, 9B) induces higher C to T editing. Furthermore, the nucleotide sequence between hA3BCDA1 and hA3BCDA2 was also expressed (Fig. 7C, 8C, 9C). The 2A self-cleaving peptide was inserted to generate hA3B-2A-BE3 (Fig. 7B, 8 B, 9B), hA3B-2A-BE3 has a higher C to T editing efficiency than hA3B-BE3. These results suggest that hA3BCDA1 induces base editing. Another inhibitor, hA3BCDA1, whose inhibition depends on its covalent binding to BE, In addition, the protein database shows that there is a domain similar to hA3BCDA1. A search revealed that at least 43 proteins share similar domains. It became clear (Table 2).
[0126] Next, we planned to develop a novel BE using mA3. Dividing mA3 between 7 and AA208 yields two BEs (mA3rev-BE 3 and mA3rev-2A-BE3), and then divide mA3CDA2. We determined whether the highest editing efficiency could be maintained by using mA3C (Figures 10A, 11A, and 12A). DA1 ends at amino acid (AA) 154, and mA3CDA2 begins at AA238. , mA3CDA2 AA196 / AA197, AA215 / AA216, AA229 / A Divided by A230 and AA237 / AA238, mA3rev-BE3-196, m A3rev-2A-BE3-196, mA3rev-BE3-215, mA3rev-2 A-BE3-215, mA3rev-BE3-229, mA3rev-2A-BE3-2 29, producing mA3rev-BE3-237 and mA3rev-2A-BE3-237 (Figures 10B, 11B, and 12B). AA207 / AA208 and AA215 / AA2 Although splitting mA3 at 16 maintains the highest editing efficiency, the results show that AA196 / AA Split sites spanning 197 to AA237 / AA238 generally maintain substantial editing efficiency We also showed that this is the case (Figures 10C, 11C, and 12C).
[0127] Furthermore, we attempted to determine the minimal region of mA3 that has the inhibitory effect on base editing. Various N-terminal deletions of mA3CDA2 in rev-BE-237 were performed to obtain mA3 rev-BE-237-Del-255, mA3rev-BE-237-Del-285 and mA3rev-BE-237-Del-333. 237-Del-255, mA3rev-BE-237-Del-285 and mA3r ev-BE‐237-Del-333 act as base editing inhibitors and mA 3 AA256-AA429, AA286-AA429 and AA334-AA429 parts mA3 (Figures 13A, 14A, and 15A). Compared with A3rev-BE-237, mA3rev-BE-237-Del-255, mA3rev-BE-237-Del-285, mA3rev-BE-237-Del- 333 showed similar editing efficiency (Figures 13B, 14B, and 15B). This indicates that the AA334-AA429 portion of 3 still has an inhibitory effect on base editing. there was.
[0128] Developing a base editor that does not cause C to T mutations in unrelated ssDNA regions To achieve this, the 2A self-cleavage site of mA3rev-2A-BE3 was converted to the cleavage site of TEV protease. The N-terminal end of the TEV protease was then replaced with another TEV cleavage site. minutes (TEVn) [Gray et al, 2010, Cell, doi:10.101 6 / j.cell.2010.07.014] to the C-terminus of mA3rev-2A-BE3 The newly developed BE was named BEsafe. 2 loop into the sgRNA to generate MS2-sgRNA [Ma et al., 20 16, Nature Biotechnology, doi:10.1038 / nbt 3526], and then the C-terminal portion of the TEV protease (TEVc) was ligated to the MS2 loop. BEsafe was fused to the MS2 coat protein (MCP), which can bind to BEsafe (Figure 16A). When MS2-sgRNA and MCP-TEVc were co-expressed, they were fused in BEsafe TEVn and TEV in MCP-TEVc that can be recruited by MS2-sgRNA c associates with the TEV site, restoring protease activity at the on-target site. Cleavage at the N- and C-termini of BEsafe results in the mA3CDA2 and TEVn The resulting mA3CDA1-BE3 efficiently binds to the on-target site. In contrast, the cytidine residues in mA3CDA1 can induce cytogenetic editing (Figure 16A). Deaminase activity is inhibited by mA3CDA2, so that in unrelated ssDNA regions However, BEsafe does not induce C to T mutations (Figure 16B).
[0129] Next, BEsafe and hA3 at the on-target site and non-associated ssDNA regions The performance of hA3A-BE3 was compared with that of A-BE3 (Figures 17, 18, and 19). and together with sgRNA expression plasmid, or together with BEsafe expression plasmid and MS2- together with a plasmid expressing sgRNA and MCP-TEVc, or MCP-TEV c Expression plasmid together with the plasmid expressing MS2-sgRNA and BEsafe , Sa-sgRN, which can trigger ssDNA formation at the Sa-sgRNA target site Plasmids expressing SaD10A and SaD10A (Figures 17A, 18A, 19A) were cotransfected into The unrelated ssDNA region (Sa-sgRNA) was transfected (Figures 17B, 18B, and 19B). C to T mutation frequency at on-target sites (not related to those of SpCas9) (Figures 17C, 18C, 19C) and hA3A-BE3 and BE, both derived from SpCas. Base editing efficiency at the safe sgRNA on-target site (Figures 17D, 18D, 1 9D). BEsafe was used to identify unrelated ssDNA regions (Sa-sgRNA on-target). hA3A-BE3 does not cause a C to T mutation at the nucleotide site (the nucleotide site), but hA3A-BE3 causes a clear mutation. It was revealed that sgRNA induces mutations (Figures 17C, 18C, and 19C). BEsafe induced base editing at the target site comparable to that of hA3A-BE3 However, expression of both MS2-sgRNA and BEsafe from one plasmid is not possible with a single plasmid. This resulted in higher base editing efficiency than expression of BEsafe alone from the plasmid (Figure 17 D, 18D, 19D).
[0130] The base editors and base editing methods described in this invention can be used to sequence and / or clone the genomes of various eukaryotic organisms. It can be applied to perform highly specific and efficient base editing in mammals.
[0131] Avoid C to T mutations in unrelated ssDNA regions and A base editing system was established for the first time to induce efficient base editing at the target site. The BEsafe base editing system and associated methods disclosed in the present invention are currently Cytidine deaminase in the BE induces unintended mutations in unrelated ssDNA regions. This allows for highly specific base editing that cannot be achieved with currently available BEs due to the possibility of cross-reaction. Importantly, the high specificity of the BEsafe base editing system The isomerism and efficiency are important for potential clinical trials, especially in gene therapy involving the restoration of disease-associated mutations. This will promote saturation.
[0132] Table 1: mA3CDA2 core sequence-related domains [Table 6-1] [Table 6-2] [Table 6-3]
[0133] Table 2: hA3BCDA1-related domains [Table 7-1] [Table 7-2] [Table 7-3]
[0134] Example 2: Further evaluation of inhibitor-conjugated base editors This example demonstrates how the APOBEC portion of the base editor (BE) can be turned off in an sgRNA-independent manner. Efficiency for demonstrating direct mutation induction at target single-stranded DNA (OTss) sites We developed a novel method to identify a series of A nucleotides containing two cytidine deaminase (CDA) domains. By testing APOBEC proteins, we were able to identify specific dual-domain APOBECs. Catalytically inactive CDA domains act as cytidine deaminase inhibitors (CDIs) This finding and the concept of split-TEV protease were utilized to By doing so, an inducible base editor (iBE) was developed using sgRNA-guided cleavage of CDI. This uses a TEV cleavage site to link nSpCas9-BE and CDI. At sgRNA-independent OTss sites, iBE1 remains dormant due to the covalently bound CDI. On the other hand, at the on-target site, iBE1 was involved in the sgRNA-guided TE of CDI. Activated by V cleavage, it resulted in efficient base editing. iBE2 reduces unintended OTsg mutations by using Cas9 nickase Its minimal off-target effects and uncompromised on-target efficacy have been further developed. The obtained editing efficiency indicates that the editing specificity of iBE is significantly higher than that of previously reported BE. Therefore, the iBE system described in this example is superior to current base editing systems. A new layer of regulation for stem specificity n) and ensure its applicability to off-target mutations.
[0135] ·method Cell culture and transfection HEK293FT cells (ATCC) were cultured in DMEM (10566, Gibco / Thermo Fisher Scientific, 1999). rmo Fisher Scientific)+10%FBS(16000-044, Gibco / Thermo Fisher Scientific) and The plasma was tested periodically to rule out contamination.
[0136] For base editing in genomic DNA, HEK293FT cells were cultured in a 24-well plate. 1.1 x 10 per well 5 and 250 μl serum-free Opti-MEM. 250 μl of serum-free Opti-MEM was transfected with 5.35 μl of LIP. OFECTAMINE LTX (Life, Invitrogen), 2.14 μl L IPOFECTAMINE plus (Life, Invitrogen), 1 μg p CMV-BE3 (or hA3B-BE3, hA3BCDA2-nSpCas9-BE, hA3D-BE3, hA3DCDA2-nSpCas9-BE, hA3F-BE3, hA 3FCDA2-nSpCas9-BE, hA3G-BE3, hA3GCDA2-nSpC as9-BE, mA3-BE3, mA3CDA1-nSpCas9-BE, mA3CDA 2-mA3CDA1-nSpCas9-BE, hA3FCDA1-mA3CDA1-nS pCas9-BE, hA3BCDA1-mA3CDA1-nSpCas9-BE, mA3 CDA2-rA1-nSpCas9-BE, hA3FCDA1-rA1-nSpCas9 -BE, hA3BCDA1-rA1-nSpCas9-BE, hA3A-BE3, mA3 CDA2-hA3A-nSpCas9-BE, hA3FCDA1-hA3A-nSpCa s9-BE, hA3BCDA1-hA3A-nSpCas9-BE, mA3CDA2F1 -mA3CDA1-nSpCas9-BE, mA3CDA2F2-mA3CDA1-nS pCas9-BE, mA3CDA2F3-mA3CDA1-nSpCas9-BE, mA 3CDAI-T2A-mA3CDA1-nSpCas9-BE, EGFP-mA3CDA 1-nSpCas9-BE, EGFP-T2A-mA3CDA1-nSpCas9-BE , mA3CDAI-T2A-rA1-nSpCas9-BE, EGFP-rA1-nSp [[ID=2——6]]Cas9-BE, EGFP-T2A-rA1-nSpCas9-BE, mA3CDAI- T2A-hA3A-nSpCas9-BE, EGFP-hA3A-nSpCas9-BE , EGFP-T2A-hA3A-nSpCas9-BE, pCMV-dSpCas9, i[[ID=——1]] BE1, iBE2, mA3CDAI-TS-mA3CDA1-nSpCas9HF1-B E-NTEV or mA3CDAI-TS-mA3CDA1-nHypaSpCas9 -BE-NTEV) expression vector, containing 0.64 μg of sgRNA expression vector, 0. After 24 hours, the cells were cultured in the presence or absence of 5 μg of Sa-sg-SaD10A expression vector. uromycin (ant-pr-1, InvivoGen) at a final concentration of 4 μg / ml After another 48 hours, the cells were cultured in Quick Extract™ DNA Extraction Solution (QE090 Genomic DNA was extracted from the cells using a 500-well plate (Epicentre).
[0137] DNA library preparation and sequencing The target genomic sequence is determined by the primers flanking the sgRNA target site under test. The set includes the high-fidelity DNA polymerase PrimeSTAR HS (Clonetec h). The indexed DNA library was PCR amplified using TruSeq. ChIP samples were prepared using a ChIP sample preparation kit (Illumina) with minor modifications. Briefly, PCR products amplified from genomic DNA regions were used as Covert The fragmented DNA was then purified using TruSeq Ch PCR amplification was performed by using an IP sample preparation kit (Illumina). Qubit High-Sensitivity DNA Kit (Invitrogen ) and then quantified by CAS-MPG Partner Institute for Co Computational Biology Omics Core (Shanghai, China) Illumina Hiseq X10 (2x150) or NextSeq 500 (2 × 150) for deep sequencing by using different tags PCR products with α, β and β were pooled together. The raw read quality was Assessed by stQC. For paired-end sequencing, only R1 reads were used. Adapter sequences and double-ended reads with Phred quality scores less than 30 were used. The sequences were trimmed. The trimmed reads were then run through the BWA-MEM algorithm. The sequences were mapped to the target sequence using the BWA program (v0.7.17). After piled up using tools (v1.9), the base substitutions were further Further calculations were carried out.
[0138] Calculation of base substitutions Base substitutions were determined by examining sg that were mapped with at least 1000 independent reads. Selected at each position of the RNA target site, obvious base substitutions are generated in the target base editing site. The base substitution frequency was calculated by dividing the base substitution reads by the total reads. For each sgRNA, the C to T salt for indels was calculated. The base substitution rate is calculated by dividing the total frequency of C to T base substitutions at all editing sites by the sgR A 50-bp region around the target site (8 nucleotides upstream of the target site) indel frequency (from up to 19 nucleotides downstream to the PAM site) was calculated by dividing by
[0139] ·result Naturally occurring cytidine deaminase or in vitro evolved adenosine deaminase Cytosine or adenine base editors (CBEs) that fuse enzymes with CRISPR-Cas9 BE or ABE) are targeted C to T or adenine to guanine (A to BE was developed to induce the conversion of catalytically insensitive Cas9 to β-actin (G) with high efficiency. (dCas9) protein or Cas9 nickase (nCas9) to The unintentional base substitutions partially interact with the sgRNA, inducing its binding to genomic DNA. In this scenario, the high activity of BE The use of precise Cas9 can reduce these OTsg mutations. APOBEC can induce unexpected C-to-T mutations in single-stranded DNA (ssDNA) regions Therefore, the APOBEC portion of the BE may directly trigger unexpected mutations at the OTss site. In other words, off-target mutations induced by BE may be due to the genomic DNA fragmentation of sgRNA. may also occur at the OTss site independent of guidance; however, quantitative OTss mutations were not revealed due to the lack of a reproducible detection method.
[0140] This example shows the effect of Staphylococcus aureus (S. aureus) and Streptococcus pyogenes (S. pyogenes) on the bacterial resistance of S. By co-expressing the Cas9 orthologue (CESSCO) of the BE gene, We established an efficient method to quantitatively evaluate induced OTss mutations. Expression of the aCas9 / Sa-sgRNA pair results in single-stranded DNA at specific genomic loci. Breaks (SSBs) were generated to form genomic ssDNA regions in a programmable manner At the same time, BE3 was co-expressed in the absence of sgRNA (hereafter, sgRNA is referred to as Sp-sgRNA). A) was used to visualize the area around the SSB introduced by nSaCas9 / Sa-sgRNA. In the ssDNA region generated by the sgRNA, a C to T base substitution occurs independently of the sgRNA. We investigated whether nSaCas9 / Sa-sgRNA could be induced by E3 alone. After deep sequencing of the targeted genomic region, C at the OTss site was Mutation to T was induced by rat APOBEC1 (rA1)-containing BE3 in the absence of sgRNA. It was clearly shown that OTss mutations are induced by dSpCas9 but not by dSpCas9. It was confirmed that this is triggered by the APOBEC portion of the BE in an sgRNA-independent manner. It was.
[0141] Next, this example demonstrates the identification of APOBEC family members suitable for highly specific BE construction. We sought to reduce OTss mutations by utilizing a commonly used The majority of BEs identified are single-domain APOBEs, such as rA1 in BE3. C, but not dual-domain APOBEC. In APOBECs with two CDA domains, one is catalytically active while the other is non-catalytic. Because OTs are catalytically inactive and play a role in regulating cytidine deamination activity, This may be suitable for constructing highly specific BEs with reduced β-effects. To test this possibility, five dual-domain APOBECs (i.e., human APOBEC 3B (hA3B), human APOBEC3D (hA3D), human APOBEC3F (hA3 F), human APOBEC3G (hA3G) and mouse APOBEC3 (mA3) 10 with either one catalytically active CDA domain or two CDA domains Pairs of BEs (Figure 20a) were constructed to compare C to T editing efficiencies.
[0142] As shown in Fig. 20b, c, certain APOBECs containing two CDA domains The BE constructed with (hA3B, hA3F, and mA3) contains only the active CDA domain. The pair with BE induced significantly lower editing efficiency than the pair with BE. Catalytically active dual-domain APOBECs (i.e., hA3B, hA3F, and mA3) Inactive CDA domains exert inhibitory functions on their corresponding active CDA domains. This indicates that...
[0143] To investigate whether the inhibitory function is general, we investigated the effects of mA3, hA3F, or hA3B. The catalytically inactive CDA domain was cloned into mA3CDA1-nSpCas9-BE (Figure 20d ) and two other commonly used BEs (i.e., BE3 and hA3A-BE3 ) were individually covalently linked to the N-terminus of each of these catalytically inactive CDA domains. It showed broad-spectrum inhibitory activity against all BEs tested. Among them, CDA2 of mA3 (mA3CDA2) showed the strongest inhibitory effect (Figure 20 e, f). Detailed mapping analysis revealed that residues 282–355 of mA3CDA2 are the full-length m It was further revealed that A3CDA2 showed the same inhibitory effect. The results suggest that the catalytically inactive domains of certain dual-domain APOBECs are indeed synaptonemal. It has been shown that it exhibits a general inhibitory effect on cysteine deaminase activity, and therefore , they were defined as cytidine deaminase inhibitors (CDIs).
[0144] Next, cleavage of mA3CDI (mA3CDA2) from the covalently bound BE results in the base pairing. The self-cleaving peptide (T2A) was used to test whether it could restore the cell adhesion ability. To achieve this, mA3CDI and mA3CDA1-nSpCas9-BE were ligated. After self-cleavage of mA3CDI-T2A-mA3CDA1-nSpCas9-BE, the editing efficiency of mA3CDI-T2A-mA3CDA1-nSpCas9-BE was The rate was determined by EGFP-mA3CDA1-nSpCas9-BE or EGFP-T2A-mA mA3CDA1-nSpCas9-BE and the non-cleaving mA3CD1 fusion protein The self-cleavage of mA3CDI from BE3 and hA3A-BE3 was approximately 10 times that of the combined BE. The cuts have also improved their editing efficiency, although to different degrees.
[0145] These results demonstrate the development of an iBE system for precise base editing with fewer OTss mutations. iBE1 served as an important proof-of-concept for the TEV protease cleavage site ( TS) to identify three important modules, namely mA3CDI, mA3CDA1- Ligation of nSpCas9-BE and the N-terminal half of TEV protease (NTEV) Theoretically, the iB E1 remains dormant when bound to the OTss site by its APOBEC moiety In particular, NTEV itself is inactive, but the C-terminal half (CTEV) Only when the ATP-binding domain is recruited does the functional TEV protease form. iBE1, guided by its CRISPR-Cas portion, induces functional TEV promoters. at the on-target site where CDI is cleaved by sgRNA-guided assembly of TEase Efficient base editing can be performed in this region (Figure 21d).
[0146] After expression in cells, iBE1 expressed the sgRNA-independent OTss region as expected. remained dormant (Fig. 21b) and exhibited much lower (~20%) levels of C compared to BE3. A mutation from α to T was induced (Fig. 21c). The MCP-fused CTEV was recruited by the MS2-fused sgRNA (Figure 21d). This results in the removal of mA3CDI from iBE1, allowing efficient base editing. BE3- and iBE1-induced on-target expression across multiple genomic loci Comparison of the efficiency of iBE1-mediated on-target editing (Figure 21e) showed that iBE1 achieved similar levels of on-target editing as BE3. It was shown that this induces base editing (Figure 21f, approximately 80% of BE3). For example, through manipulation of CDI, we have demonstrated that OTss mutations can be suppressed while enhancing efficacy at the on-target site. We demonstrate that we have developed an iBE system that catalyzes efficient base editing.
[0147] Cas9 non-intentionally targets OTsg sites with partial sequence complementarity to the sgRNA. Since unmodified nSpCas9 in iBE1 is known to induce phenotypic editing, Engineered version with improved targeting specificity We aimed to further reduce OTsg mutations by replacing it with OTsg (Figure 22a). Three engineered versions of nSpCas9 (i.e., neSpCas9, nS pCas9HF1 and nHypaSpCas9) were tested to confirm their targeting characteristics. Using one of the isomerically improved Cas9 proteins significantly reduced OTsg mutations. On the other hand, the use of neSpCas9 was found to be effective in reducing the number of While the other two did not impair on-target editing efficiency, In this scenario, nSpCas9 was replaced by neSpCas9. and set it to build iBE2.
[0148] As an early development of BE, the editing efficiency of BE3 is limited under certain conditions. Additional BEs with improved performance (e.g., AncBE4max or hA3A-BE3) were subsequently developed. hA3A-BE3 is a highly active BE in various conditions, The performance of iBE2 was compared with that of hA3A-BE3 in terms of editing efficiency and specificity. The average on-target editing frequency of iBE2 was approximately 50% of that of hA3A-BE3 (Fig. 23a). (Fig. 23a, c), but was induced by iBE2 at the OTss and OTsg sites. The C to T mutation was close to background level, whereas hA3A-BE3 Substantial mutations were induced at these off-target sites (Fig. 23a, b). The average editing specificity of BE2 was approximately 40-fold higher than that of hA3A-BE3 (Figure 23d). .
[0149] In this example, we first demonstrate an efficient method for quantitatively assessing sgRNA-independent OTss mutations. We developed a method (CESSCO) to detect the normal APOBEC-nCas9 backbone. We confirmed that BE indeed induces OTss mutations in an sgRNA-independent manner (Figure 2 1a, 21b, 23a, 23b). Consistent with our findings, recent whole-genome sequencing Studies have also shown that BE3 likely functions in an sgRNA-independent manner in mice and rice plants. Importantly, we have demonstrated that CDI induces substantial off-target mutations. The iBE was developed using the OTss site for covalent binding of CDI. It remains dormant at the target site but is activated by sgRNA-mediated cleavage of CDI at the on-target site. iBE can be activated by efficient on-target editing (Fig. 21a, d). They found that significantly lower levels of unintended mutations were induced in the sgRNA-independent ssDNA region. (Fig. 21b, c, e, f).
[0150] By replacing nSpCas9 with enSpCas9, which has improved specificity, Highly specific iBE2 was developed to further reduce unintended editing at Tsg sites The iBE system was developed using BEs and BEs with different Cas moieties (Figs. 22 and 23e). Compatible with and built on engineered BEs with improved performance It does not change the properties of the BE (e.g., the edit window). Because the family has many members, it is possible that other CDIs will be identified in the future. This will further enrich the repertoire of CDI-conjugated iBE systems. Both editing accuracy and efficiency are important for base editors, especially in their therapeutic applications. The iBE system developed here is essential for the current base editing system. A new layer of regulation for specificity This ensures its coverage against off-target mutations.
[0151] Example 3: Testing different configurations of guided and split base editors In this example, an induced and split base editor Molecules for implementing the lit base editor (isplitBE) system Several different configurations of were tested.
[0152] Compared with the conventional BE shown in FIG. 24B, the operation process of isplitBE is shown in FIG. In the illustrated isplitBE system, the nCas9-D10A construct is expressed as AA A typical AAV vehicle has a capacity of 4.7 kb. The nCas9 construct is approximately 4.7 kb in length. Another AVV vehicle encodes It can package nucleic acids (total length approximately 4.4 kb) that contain: (a) MCP, UGI, Fusion protein containing APOBEC, a TEV recognition site (TEV site), and mA3CDA2 (b) fusion protein with TEVc and N22p; (c) standalone TEVn (sta (d) helper sgRNA with MS2 tag (hsgR NA), and (e) another sgRNA with a boxB tag.
[0153] At the target site (ON, bottom left branch), the hsgRNA and sgRNA are It binds to two adjacent sites on the target DNA and binds to the fusion protein containing MCP and N22p. Proteins bind to the MS2 and boxB tags of hsgRNA and sgRNA, respectively. The proximity of the TEVc (in the presence of free TEVn) and TEV sites allows c / TEVn cleaves the TEV site, removing mA3CDA2 from APOBEC. Without the mA3CDA2 protein, APOBECs are unable to carry out the desired editing very efficiently. This can be done.
[0154] Binds to a nonspecific binding site (OTss, central lower branch) or only one of the guide RNAs At potential off-target sites, the TEVc / TEVn complex contains the hTEV site. The APOBECs are not recruited to the fusion protein containing the ATP-binding domain and therefore cannot be activated. In contrast, in conventional BE systems (Figure 24B), APOBECs are already active and single-chain triggering C to T editing whenever recruited to a nucleotide sequence can be done.
[0155] Ten different configurations (pairs 1-10) were prepared and tested as shown in Figure 25. For example, as shown in Figure 26A, pair 1 contains two constructs, the first of which is Contains rA1 fused to nCas9-D10A (spD10A) with UGI and NLS. Pair 2 was a clone of pair 1, and the second contained an sgRNA targeting EMX1. Pair 3 is similar, but with rA1 replaced by hA3A. The mutant hA3A(Y130F) was used.
[0156] In pair 4, the rA1 and nCas9 proteins were placed in different constructs. was further fused to the MCP protein, which recognizes the MS2 tag on the helper sgRNA In pair 5, mA3CDA2 binds to rA1 via the TEV recognition site (filled box). In pair 6, the TEV protein was fused to rA1- via the self-cleavage site 2A. The self-cleavage of 2A was further fused to the mA3CDA2 fusion protein, resulting in the release of TE from the fusion protein. Emits V.
[0157] Pair 7 fuses TEV to the N22p protein, which recognizes the boxB tag on the sgRNA. Pair 8 differs from pair 6 in that it contains two TEV proteins, TEVn and TEVp. In pair 9, only TEVc was split into two TEVs, separated by a 2A self-cleavage site. The RNA tag was fused to N22p, and TEVn did not contain any RNA tag-binding protein. In A10, the helper sgRNA targeted GFP rather than a nearby site.
[0158] The construct in Figure 26A is designed for C to T editing at the target site EMX1-ON. The off-target activity at the Sa-SITE31-OTss and EMX1-OTsg sites was calculated. The results are shown in Figure 26B. induced substantial editing at ON sites but not at OTss or OTsg sites. There wasn't.
[0159] Similarly, all of these configurations were FANCF-ON, Sa-VEGFA-7-OTs s and FANCF-OTsg sites (see schematic in Figure 27A). , FANCF-ON, Sa-VEGFA-7-OTss and FANCF-OTsg sites This shows a comparison of editing efficiency for different base editors in ispli. tBE-rA1 (pair 9) induced substantial editing at ON sites but not at OTss or O sites. No editing was induced at the Tsg site.
[0160] Further investigations were performed at the V1B-ON, Sa-SITE42-OTss and V1B-OTsg sites. Tests were performed (see schematic in Figure 28A). Again, as shown in Figure 28B, ispl itBE-rA1 (pair 9) induced substantial editing at ON sites but not at OTss or O sites. No editing was induced at the Tsg site.
[0161] (Example 4) Adjusting parameters of the isplitBE system Of the 10 configurations tested, Pair 9 performed best in terms of editing specificity. uses two sgRNAs: a helper sgRNA (hsgRNA) and a regular sgRNA. Dual use of sgRNAs requires that both target sites are in close proximity to each other. Therefore, further increasing specificity is required.
[0162] The first assay in this example evaluated the optimal distance between two target sites. A schematic diagram is shown in Figure 29A. hsgR at the DNTET1, EMX1 and FANCF sites The distance between the NA and the sgRNA is shown. The frequency of base edits induced by hsgRNAs is shown. The effect of distance between A and sgRNA is shown. Based on the overview, the best base editing efficiency The optimal distance range for obtaining the desired signal is -9 mm from the PAM of the hsgRNA to the PAM of the sgRNA. 1 to -34bp.
[0163] The second assay assessed the effect of hsgRNA spacer length on base editing efficiency and accuracy. The effects of the following were examined: Co-transfection of sgRNA and hsgRNA with different spacer lengths Figure 30B provides a schematic diagram showing the targeting portions of hsgRNA and sgRNA. Frequencies of base edits induced by the indicated sgRNAs and hsgRNAs at the positions Statistical analysis in Figure 30C shows the effect of hsgRNA spacer length. As shown, the use of hsgRNAs with a 10-nt spacer significantly reduces the length of the hsgRNA Although the editing efficiency at the target site was significantly reduced, editing at the sgRNA target site was Therefore, the 9-15 nt spacer in the helper sgRNA sequence , while minimizing editing at the hsgRNA target site. This may be a good range to ensure efficient editing in the region.
[0164] Example 5: Genome-wide and transcriptome-wide evaluation The overall efficiency of the isplitBE system was compared with the conventional BE3. The results are shown in Figure 31. Editing frequencies induced by the indicated base editors at different target sites are shown. Even with the significant improvement in specificity of isplitBE, there is a clear sacrifice in efficiency. Not at all.
[0165] Normal cells produce background levels of C due to their endogenous APOBEC3 activity. To more accurately measure off-target C to T mutations, , the APOBEC3 knockout 293FT cell line (293FT-A3KO) was used. Figure 32A shows wild-type 293FT cells and APOBEC3 knockout 293FT cells. Figure 32B shows the mRNA expression levels of the base editor-induced A schematic diagram showing the procedure for determining genome-wide C-to-T mutations is presented. The experimental results are shown in Figure 32C (Cas9, BE3, hA3A-BE3-Y130F ( On-target editing efficiency induced by isplitBE-rA1 and isplitBE-rA1 The rate (left) and genome-wide number of C to T mutations. Both BE3 and Y130F are Although isplitBE-rA1 had higher off-target editing than isplitBE-rA1, The editing rate is close to background (Cas9 only).
[0166] Next, in this example, split BE-mA3, BE3 and hA3A-BE3- Transcriptome-wide C to U induced by Y130F (Y130F) The mutations in Cas9, BE3, hA3A-BE3-Y130F (Y130F), and and isplitBE-mA3-induced transcriptome-wide C The number of mutations to T(U) is shown in Figure 33A. Figure 33B shows the results of the Cas9, BE3, and hA mutations. Induced by 3A-BE3-Y130F(Y130F) and isplitBE-mA3 Figure 33C shows the C to U editing frequency of BE3 replicate 1 and BE3 replicate 2. C to U editing of RNA induced by isplitBE-mA3 replicate 1 Again, isplitBE exhibits much lower C to U editing than BE3. Led.
[0167] Example 6: PCSK9 knockout Proprotein convertase subtilisin / kexin type 9 vertase subtilisin / kexin type 9 (PCSK9)) PCSK9 is an enzyme encoded by the human PCSK9 gene on chromosome 1. is the ninth member of the proprotein convertase family that activates other proteins PCSK9 is synthesized initially because part of the peptide chain blocks its activity. When the enzyme is activated, it is inactive; proprotein convertases remove the moiety and activate the enzyme. The PCSK9 gene is one of 27 genetic loci associated with an increased risk of coronary artery disease. include.
[0168] PCSK9 is ubiquitously expressed in many tissues and cell types. binds to receptors on low-density lipoprotein particles (LDL), which are typically 3,000-6,000 lipid molecules (including cholesterol) per particle in the extracellular fluid LDL receptors (LDLRs) on the membranes of liver and other cells bind to LDL particles and transport them. initiates the uptake of LDL particles from the extracellular fluid into the cells, reducing the concentration of LDL particles When PCSK9 is blocked, more of it is used to remove LDL particles from the extracellular fluid. The LDLR is recycled and becomes present on the surface of the cell. Blocking K9 can reduce blood LDL particle concentrations.
[0169] In this example, a stop codon was introduced by base editing using this technology. Approaches to inactivate PCSK9 were tested. The sgRNA / hsgRNA used were: The sequences are shown in Table 3 and the target sites on PCSK9 are shown in Table 4.
[0170] The number of stop codons generated by base editing was measured in the human PCSK9 gene. Figure 34A shows the sgRNA and hsplitBE-mA3 and nCas9 gene expression. A schematic diagram showing co-transfection of sgRNAs is presented. The editing efficiency induced by plitBE-mA3 is shown in Figures 34B-D. demonstrates the high efficiency and specificity of this method.
[0171] Table 3: Regular sgRNA and hsgRNA scaffolds in the PCSK9 gene and target area [Table 8-1] [Table 8-2]
[0172] Table 4: Target sites in the PCSK9 gene [Table 9]
[0173] (Example 7) Applicability of isplitBE design in adenine base editors This example illustrates the use of guided and split base editors in other types of base editors. ta(induced and split base editor(isplitBE )) Confirm the applicability of the design. The inhibitor used was mA3CDA2, and The editor was adenine base editor (ABE).
[0174] ABE fused to sgRNA and mA3CDA2 (i.e., not as a control) A schematic diagram showing the co-transfection of RNF2 and FA is shown in Figure 35A. The editing efficiency induced by the indicated ABE at the NCF site is shown in Figure 35B. When 3CDA2 was connected to ABE, the editing efficiency was reduced compared to ABE alone. When A2 is cleaved by 2A, the editing efficiency of ABE is restored, and ispl The effectiveness of the itBE approach has been proven.
[0175] Example 8: Enhanced Prime Editing Conventional base editors are limited to base transitions. There are no base transversions, insertions or deletions. Cas9 nickase is conjugated to reverse transcriptase (RTase) to enhance the A primer editing system using primer editor (PE) has been proposed. The system detects all types of base substitutions, small indels and It is possible to write genomes with almost any intentional modification, including combinations of these. However, the overall efficiency and specificity of the PE system remains limited. It is being done.
[0176] In the first assay, in this example, a primer-edited guide RNA (pegRNP) was used. A) A new design for the guide RNA was tested. Conventionally, each guide RNA contains a scaffold. Commonly used scaffold sequences are: [ka] Another example is [ka] The more common consensus sequence is [ka] where N represents any base, and X1 and X2 are any bases having a length of 2 to 50 bases. The nucleotide sequence is shown.
[0177] The scaffold has a secondary structure (Figure 36A, SEQ ID NO: 30) due to the complementary sequence within it. A typical sgRN used in base editors is expected to form a A is approximately 96 nt long and contains a spacer of approximately 20 nt long that binds to the target site. In pegRNA, the reverse transcription template and primer binding site are scaffolded. Surprisingly, in the present specification, the original The scaffold is not stable enough in the context of pegRNA It was discovered that...
[0178] Thus, positions 48 (e.g., A in SEQ ID NO: 30) and 61 (e.g., A new scaffold was prepared that forms a new pairing between the nucleoside and the nucleoside. In the example shown in 36E, the novel scaffold instead contains G and C or C and and G (SEQ ID NOs: 36, 37). This mutant scaffold and additional mutants Examples of scaffolds are shown in Table 5 below.
[0179] Table 5. Guide RNA scaffold sequences [Table 10]
[0180] Conventional pegRNA and newly designed enhanced pegRNA (epegRNA Constructs for testing ) were prepared as shown in Figures 36B and 36F for PE2. The test results are shown in Figures 36C to 36D and 36G. Comparison of prime editing efficiency induced by epegRNA and epegRNA. Stem stability was significantly improved. The epegRNA showed much higher editing efficiency overall than pegRNA.
[0181] Similarly, according to the schematic diagram in Figure 37A, PE2-NG (SEQ ID NO: 132) or xPE2 Co-transfection of pegRNA, nicking sgRNA with (SEQ ID NO: 133) The results are shown in Figure 37B. NG can recognize relaxed NG PAMs with engineered Cas9 (e.g., Nishimasu et al. , Science 361, pp. 1259-62 (2018). xPE2 is Relaxed NG, GAA and GAT PAM (relaxed NG, GAA and and GAT PAMs) (See, for example, Hu et al., Nature 556, pp. 57-63 (2018)) PE2-NG (SEQ ID NO: 44), xPE2 (SEQ ID NO: 45), SpCas9-NG (SEQ ID NO: 46) The sequences of xSpCas9 (sequence number 46), and xSpCas9 (sequence number 47) are shown in Table 6 below.
[0182] Table 6. Cas and PE sequences [Table 11-1] [Table 11-2] [Table 11-3]
[0183] The complete prime editor is a much larger construct than an AAV vehicle can accommodate. Therefore, a split PE system was designed and tested. The original PE system is shown in the left panel of Figure 38A, and the newly designed The designed split PE system is shown in the right panel, where nickase and RT RTase is packaged into different AAV particles. RTase is an RNA-binding protein. The pegRNA is fused to MCP and contains the binding site MS2. The RTase is recruited by pegRNA via MS2-MCP binding, It can be contacted with a lysozyme.
[0184] An exemplary co-transfection system is shown in Figure 38B, where The test results are shown in Figure 38C.
[0185] The present disclosure does not limit the scope of the specific embodiments described, which are intended as single illustrations of individual aspects of the disclosure. The scope should not be limited by the embodiment, and any composition or The methods are within the scope of the present disclosure. Those skilled in the art will recognize that various modifications and variations can be made in the methods and compositions. It will be apparent that the present disclosure is intended to cover all aspects of the present invention as defined by the appended claims and their equivalents. It is intended to cover the modifications and variations of this disclosure insofar as they come within their scope.
[0186] All publications and patent applications mentioned herein are hereby incorporated by reference in their entirety. The reference is to the same extent as if the application were specifically and individually indicated to be incorporated by reference. and is incorporated herein by reference.
Claims
1. a first fragment comprising a nucleobase deaminase or a catalytic domain thereof; a second fragment comprising a nucleobase deaminase inhibitor, and a protease cleavage site between the first fragment and the second fragment A fusion protein comprising:
2. The fusion protein of claim 1, wherein the nucleobase deaminase is adenosine deaminase. Protein.
3. The adenosine deaminase is a tRNA-specific adenosine deaminase. deaminase (TadA), adenosine deaminase tRNA specific 1 (ADAT1), adenosine deaminase t RNA specific 2 (ADAT2), adenosine deaminas e tRNA specific 3 (ADAT3), adenosine deami nase RNA specific B1 (ADARB1), adenosine d eaminase RNA specific B2 (ADARB2), adenosi ne monophosphate deaminase 1 (AMPD1), aden osine monophosphate deaminase 2 (AMPD2), a denosine monophosphate deaminase 3 (AMPD3 ), adenosine deaminase (ADA), adenosine dea minase 2 (ADA2), adenosine deaminase like ( ADAL), adenosine deaminase domain contain ing 1 (ADAD1), adenosine deaminase domain containing 2 (ADAD2), adenosine deaminase RNA specific (ADAR) and adenosine deaminase RNA specific B1 (ADARB1), 3. The fusion protein according to claim 2.
4. The fusion protein of claim 1, wherein the nucleobase deaminase is a cytidine deaminase. Plagiarism.
5. The cytidine deaminase is APOBEC3B (A3B), APOBEC3C (A3C ), APOBEC3D (A3D), APOBEC3F (A3F), APOBEC3G (A3 G), APOBEC3H (A3H), APOBEC1 (A1), APOBEC3 (A3) , APOBEC2 (A2), APOBEC4 (A4), and AICDA (AID) The fusion protein of claim 4 selected from the group:
6. The cytidine deaminase is a human or mouse cytidine deaminase. Item 5. The fusion protein according to item 4.
7. The catalytic domain is mouse A3 cytidine deaminase domain 1 (CDA1) or 7. The fusion of claim 6, which is human A3B cytidine deaminase domain 2 (CDA2). protein.
8. The nucleobase deaminase inhibitor is an inhibitory domain of a nucleobase deaminase. The fusion protein according to any one of claims 1 to 7.
9. The nucleobase deaminase inhibitor is an inhibitory domain of cytidine deaminase. The fusion protein of claim 8.
10. The nucleobase deaminase inhibitor is selected from SEQ ID NOs: 1-2 and 48-135. or an amino acid sequence selected from SEQ ID NOs: 1-2 and 48-135 an amino acid sequence having at least 85% sequence identity to any of the amino acid sequences of The fusion protein of claim 9.
11. The nucleobase deaminase inhibitor has the amino acid sequence of SEQ ID NO: 1, The amino acid sequence of claim 2, comprising amino acid residues AA128 to AA223, or the amino acid sequence of SEQ ID NO:
2.
9. A fusion protein according to claim 9.
12. The first fragment is composed of clustered regularly spaced short palindromic repeats. Any of claims 1 to 11, further comprising a CRISPR-associated (Cas) protein. The fusion protein according to any one of claims 1 to 10.
13. The Cas protein is SpCas9, FnCas9, St1Cas9, St3Ca s9, NmCas9, SaCas9, AsCpf1, LbCpf1, FnCpf1, VQ R SpCas9, EQR SpCas9, VRER SpCas9, SpCas9-N G, xSpCas9, RHA FnCas9, KKH SaCas9, NmeCas9, StCas9, CjCas9, AsCpf1, FnCpf1, SsCpf1, PcCpf 1, BpCpf1, CmtCpf1, LiCpf1, PmCpf1, Pb3310Cpf 1, Pb4417Cpf1, BsCpf1, EeCpf1, BhCas12b, AkCa s12b, EbCas12b, LsCas12b, RfCas13d, LwaCas13 a, PspCas13b, PguCas13b and RanCas13b The fusion protein of claim 12, wherein the fusion protein is selected from the group consisting of:
14. The protease cleavage site is selected from the group consisting of TuMV protease, PPV protease, and PVY protease. protease, ZIKV protease and WNV protease The fusion protein according to any one of claims 1 to 13, which is a protease cleavage site of a protease. Synthetic protein.
15. The method according to any one of claims 1 to 13, wherein the protease cleavage site is an autocleavage site. The fusion protein described above.
16. 14. The method according to claim 1, wherein the protease cleavage site is a TEV protease cleavage site. A fusion protein according to any one of claims 1 to 4.
17. Further comprising a third fragment comprising TEV protease or a fragment thereof; The fusion protein of claim 16.
18. The third fragment alone is not capable of cleaving the TEV protease cleavage site.
18. The fusion protein of claim 17, comprising a TEV protease fragment that is incapable of 。
19. Nucleobase deaminase or its catalytic domain, clustered and regularly spaced CRISPR-associated (Cas) protein and the first TE a first fragment comprising a V protease fragment; a second fragment comprising a nucleobase deaminase inhibitor; and a TEV protease cleavage site between the first fragment and the second fragment; Rank and Including, The first TEV protease fragment alone does not contain the TEV protease cleavage site. A fusion protein that cannot be cleaved at this position.
20. 20. The fusion protein of claim 19, further comprising a uracil glycosylase inhibitor (UGI). Synthetic protein.
21. the nucleobase deaminase inhibitor, the TEV protease cleavage site, the nucleic acid a base deaminase or catalytic domain thereof, the Cas protein, and the first T 20. The method of claim 19, wherein the EV protease fragments are arranged from the N-terminus to the C-terminus. Fusion proteins.
22. The first TEV protease fragment is a fragment of the N-terminal domain of the TEV protease.
20. The fusion protein of claim 19, wherein the C-terminal domain is the C-terminal domain (SEQ ID NO: 3) or the C-terminal domain (SEQ ID NO: 4). Synthetic protein.
23. 22. The TEV protease cleavage site has the amino acid sequence of SEQ ID NO:
5. The fusion protein according to claim 1.
24. The nucleic acid base deaminase is APOBEC3B (A3B), APOBEC3C (A3C ), APOBEC3D (A3D), APOBEC3F (A3F), APOBEC3G (A3 G), APOBEC3H (A3H), APOBEC1 (A1), APOBEC3 (A3) , APOBEC2 (A2), APOBEC4 (A4), and AICDA (AID) 20. The fusion protein of claim 19 selected from the group:
25. The nucleobase deaminase inhibitor is selected from SEQ ID NOs: 1-2 and 48-135. or an amino acid sequence selected from SEQ ID NOs: 1-2 and 48-135 an amino acid sequence having at least 85% sequence identity to any of the amino acid sequences of 20. The fusion protein of claim 19.
26. The Cas protein is SpCas9, FnCas9, St1Cas9, St3Ca s9, NmCas9, SaCas9, AsCpf1, LbCpf1, FnCpf1, VQ R SpCas9, EQR SpCas9, VRER SpCas9, SpCas9-N G, xSpCas9, RHA FnCas9, KKH SaCas9, NmeCas9, StCas9, CjCas9, AsCpf1, FnCpf1, SsCpf1, PcCpf 1, BpCpf1, CmtCpf1, LiCpf1, PmCpf1, Pb3310Cpf 1, Pb4417Cpf1, BsCpf1, EeCpf1, BhCas12b, AkCa s12b, EbCas12b, LsCas12b, RfCas13d, LwaCas13 a, PspCas13b, PguCas13b and RanCas13b 20. The fusion protein of claim 19, wherein the fusion protein is selected from the group consisting of:
27. 1. A method for performing gene editing in a cell at a target site, comprising administering to the cell: (a) a fusion protein according to any one of claims 19 to 26; (b) a guide RNA targeting the target site or the target site and a crRNA that targets tracrRNA and further includes a tag sequence; (c) a second TE linked to an RNA recognition peptide capable of binding to the tag sequence; V protease fragment and 20. A method comprising:
28. One or more of the molecules may be introduced into the cell by a polynucleotide encoding the molecule.
28. The method of claim 27, wherein
29. The first TEV protease fragment and the second TEV protease fragment The fragment cleaves the TEV protease cleavage site upon interaction. The method of claim 27,
30. The second TEV protease fragment is fused to the RNA recognition peptide.
28. The method of claim 27.
31. Any one of claims 27 to 30, wherein the tag sequence comprises the MS2 sequence (SEQ ID NO: 16). The method described below.
32. The RNA recognition peptide comprises MS2 coat protein (MCP, SEQ ID NO: 22).
32. The method of claim 31 .
33. The tag sequence comprises a PP7 sequence (SEQ ID NO: 18), and the RNA recognition peptide is PP7 coat protein (PCP, SEQ ID NO: 23), or The tag sequence comprises a boxB sequence (SEQ ID NO: 20), and the RNA recognition peptide comprises a boxB sequence (SEQ ID NO: 20). Any of claims 27 to 30, comprising the xB coat protein (N22p, SEQ ID NO: 24).
2. The method according to claim 1.
34. 28. The method of claim 27, wherein the cell is within a living organism.
35. 35. The method of claim 34, wherein the cell is a human cell in a human patient.
36. (a) a fusion protein according to any one of claims 19 to 26, and (b) a second TEV protein bound to an RNA recognition peptide capable of binding to an RNA sequence; Protease fragment 1. A kit or package for performing gene editing, comprising:
37. a first nucleobase deaminase or a first fragment comprising a catalytic domain thereof; and Beauty a second fragment comprising an inhibitory domain of a second nucleobase deaminase; Including, The first nucleobase deaminase may be the same as or different from the second nucleobase deaminase. A fusion protein.
38. The first and second nucleobase deaminases are human and mouse APOB, respectively. EC3B (A3B), APOBEC3C (A3C), APOBEC3D (A3D), APO BEC3F (A3F), APOBEC3G (A3G), APOBEC3H (A3H), AP OBEC1 (A1), APOBEC3 (A3), APOBEC2 (A2), APOBEC 4 (A4) and AICDA (AID). Fusion proteins.
39. nucleobase deaminase or catalytic domain thereof; nucleobase deaminase inhibitors, a first RNA recognition peptide, and The nucleobase deaminase or its catalytic domain and the nucleobase deaminase inhibitor TEV protease cleavage site between A fusion protein comprising a first fragment comprising:
40. A TEV protease that cannot cleave the TEV protease cleavage site by itself Zefragment, and Second RNA recognition peptide 40. The fusion protein of claim 39, further comprising a second fragment comprising:
41. further comprising a self-cleavage site between the first fragment and the second fragment.
41. The fusion protein of claim 40.
42. and a third fragment comprising a second TEV protease fragment, The first TEV protease fragment is a TEV protease fragment having a length greater than or equal to the length of the second TEV protease fragment.
42. The fusion protein of claim 41, wherein the fusion protein is capable of cleaving the TEV protease site in the presence of Synthetic protein.
43. A second self-cleavage site is added between the second fragment and the third fragment.
43. The fusion protein of claim 42, comprising, upon cleavage of the second self-cleavage site, The fusion protein comprises the second TEV protein that is not fused to an RNA recognition peptide. ase fragments, fusion proteins.
44. A first sequence having sequence complementarity to the target nucleic acid sequence adjacent to the first PAM site. a target single guide RNA comprising a pacer; a second spacer having sequence complementarity to a second nucleic acid sequence adjacent to the second PAM site; a helper single guide RNA comprising a Ser; Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-related (Cas) protein; and Nucleobase deaminase and Including, The second PAM site is 34 to 91 bases away from the first PAM site. Alguide RNA system.
45. The dual guide of claim 44, wherein the second spacer is 7 to 23 bases in length. DoRNA system.
46. The dual guide of claim 44, wherein the second spacer is 8 to 15 bases in length. DoRNA system.
47. The dual guide of claim 44, wherein the second spacer is 9 to 12 bases in length. DoRNA system.
48. In the 5' to 3' direction, a first stem-loop portion, a second stem-loop portion, a third stem-loop portion, and The fourth stem-loop portion A guide RNA comprising a scaffold comprising: A guide RNA, wherein the third stem loop contains five base pairs therein.
49. The base pairing includes one between positions 45 and 55 according to the positions of SEQ ID NO:
31. The guide RNA of claim 48.
50. The scaffold comprises a sequence selected from the group consisting of SEQ ID NOs: 32-43.
50. A guide RNA according to claim 48 or 49.
51. One or more spacers for recognizing the target sequence, a reverse transcriptase template, or or a primer binding site.
52. 1. A method for performing gene editing in a cell at a target site, comprising administering to the cell: Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) a first viral particle encapsulating a first construct encoding a Cas protein; and Beauty A second construct encoding a reverse transcriptase fused to an RNA recognition peptide is encapsulated. 2 virus particles 20. A method comprising:
53. The second construct is a guide construct comprising an RNA recognition site to which the RNA recognition peptide binds.
53. The method of claim 52, further encoding RNA.
54. The Cas protein is SpCas9, FnCas9, St1Cas9, St3Ca s9, NmCas9, SaCas9, AsCpf1, LbCpf1, FnCpf1, VQ R SpCas9, EQR SpCas9, VRER SpCas9, SpCas9-N G, xSpCas9, RHA FnCas9, KKH SaCas9, NmeCas9, StCas9, CjCas9, AsCpf1, FnCpf1, SsCpf1, PcCpf 1, BpCpf1, CmtCpf1, LiCpf1, PmCpf1, Pb3310Cpf 1, Pb4417Cpf1, BsCpf1, EeCpf1, BhCas12b, AkCa s12b, EbCas12b, LsCas12b, RfCas13d, LwaCas13 a, PspCas13b, PguCas13b and RanCas13b 53. The method of claim 52, wherein
55. The reverse transcriptase may be human immunodeficiency virus (HIV) reverse transcriptase, Moloney mouse leukemia virus (MLV) reverse transcriptase, or MMLV reverse transcriptase and avian myeloblastosis virus (AMV) reverse transcriptase 53. The method of claim 52, selected from the group consisting of:
56. Helper guide RNA / guide RNA pair for editing human PCSK9 nucleic acid sequence wherein the guide RNA specifically targets a first site on the PCSK9 nucleic acid. and allowing base editing to convert a non-stop codon into a stop codon, The guide RNA is a second site on the PCSK9 nucleic acid that is 20 to 100 bases away from the first site. A helper guide RNA / guide RNA pair that specifically targets two sites.
57. 57. The helper of claim 56, wherein the non-stop codon is CAG, CAA, or CGA. - guide RNA / guide RNA pair.
58. The helper guide RNA specifically binds to a sequence of 7 to 23 nucleotides in length.
58. A helper guide RNA / guide RNA pair according to claim 56 or 57.
59. The helper guide RNA specifically binds to a sequence of 8 to 15 nucleotides in length.
58. A helper guide RNA / guide RNA pair according to claim 56 or 57.
60. The helper guide RNA specifically binds to a sequence within the first region, and the guide R 60. The method of claim 58 or 59, wherein the NA specifically binds to a sequence within the second portion. Per-guide RNA / guide RNA pair.
61. 60. The method of claim 60, wherein the second site comprises the sequence of any one of SEQ ID NOs: 166 to 180. A helper guide RNA / guide RNA pair according to any one of claims 1 to 4.
62. The second portion and the first portion each include: SEQ ID NOs: 166 and 181; SEQ ID NOs: 167 and 182; SEQ ID NOs: 168 and 183; SEQ ID NOs: 169 and 184; SEQ ID NOs: 170 and 185; SEQ ID NOs: 171 and 186; SEQ ID NOs: 172 and 187; SEQ ID NOs: 173 and 188; SEQ ID NOs: 174 and 189; SEQ ID NOs: 175 and 190; SEQ ID NOs: 176 and 191; SEQ ID NOs: 177 and 192; SEQ ID NOs: 178 and 193; SEQ ID NOs: 179 and 194; or SEQ ID NOs: 180 and 195 61. The helper guide RNA / guide RNA pair of claim 60, comprising the sequence:
63. The helper guide RNA and guide RNA according to any one of claims 56 to 62 One or more polynucleotide sequences encoding the
64. A method for inactivating the PSCK9 gene in a cell, comprising:
2. A pair of helper guide RNA and guide RNA, a cluster CRISPR-associated (Cas) tag and a nucleic acid base deaminase.
Citation Information
Patent Citations
Site-specific DNA base editing using modified apobec enzymes
US20180170984A1
Directed editing of cellular RNA via nuclear delivery of crispr / CAS9
US20180334685A1
CAS variants for gene editing
WO2015089406A1
Gene editing of PCSK9
WO2018119354A1
Compositions and methods for treatment of proprotein convertase subtilisin / kexin type 9 (PCSK9)-related disorders
WO2018154380A1