Base editing enzyme
Engineered base editors with specific nucleic acid sequences and guide polynucleotides enhance the precision and efficiency of nucleobase conversions in nucleic acid editing, addressing off-target issues in CRISPR systems.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- METAGENOMI THERAPEUTICS INC
- Filing Date
- 2024-05-03
- Publication Date
- 2026-06-04
AI Technical Summary
Existing CRISPR-based gene editing systems face challenges in efficiently and specifically converting nucleobases, such as adenine to guanine, and cytosine to uracil, with high precision and minimal off-target effects.
Development of engineered base editors comprising specific nucleic acid sequences and guide polynucleotides that form complexes with endonucleases to target and modify nucleic acid sequences, including deaminases and uracil DNA glycosylase inhibitors, integrated into vectors like plasmids or viruses for delivery into various cell types.
Achieves precise and efficient conversion of nucleobases within target nucleic acid sequences with reduced off-target editing and indel formation, demonstrating high editing activity across diverse cell types.
Smart Images

Figure 2026518124000001_ABST
Abstract
Description
Background Art
[0001] Cross-reference This application claims the benefit and priority of U.S. Provisional Patent Application No. 63 / 499,912, filed May 3, 2023, U.S. Provisional Patent Application No. 63 / 519,790, filed Aug. 15, 2023, and U.S. Provisional Patent Application No. 63 / 611,049, filed Dec. 15, 2023, each of which is incorporated herein by reference in its entirety.
[0002] Cas enzymes, along with their associated clustered regularly interspaced short palindromic repeat (CRISPR) guide ribonucleic acids (RNAs), appear to be a widespread (about 45% of bacteria, about 84% of archaea) component of prokaryotic immune systems and serve to protect such microorganisms from non-self nucleic acids such as infectious viruses and plasmids by CRISPR-RNA-guided nucleic acid cleavage. Deoxyribonucleic acid (DNA) elements encoding CRISPR RNA elements can be relatively conserved in structure and length, but their CRISPR-associated (Cas) proteins are highly diverse and contain a wide variety of nucleic acid interaction domains. CRISPR DNA elements were observed as early as 1987, but the programmable endonuclease cleavage ability of CRISPR complexes has only recently been recognized and has led to the use of recombinant CRISPR systems in a variety of DNA manipulation and gene editing applications.
[0003] Sequence Listing This application includes a sequence listing submitted electronically in XML format, which is incorporated herein by reference in its entirety. The XML copy was created on May 3, 2024 and is 2,926,539 bytes in size.
Summary of the Invention
[0004] In this specification, in one embodiment, an engineered base editing system is described, comprising a base editor having a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% identity with any one of sequence numbers 1654-1703 and 2021-2023, wherein the sequence does not include any sequence selected from sequence numbers 1128-1160 and 1363-1415, and an engineered guide polynucleotide having a spacer sequence that forms a complex with the endonuclease of the base editor and hybridizes with a target nucleic acid sequence.
[0005] In some embodiments, the base editor includes a sequence having at least 95% sequence identity with any one of sequence numbers 1654-1703 and 2021-2023.
[0006] In some embodiments, the base editor includes a sequence having 100% sequence identity with any one of sequence numbers 1654-1703 and 2021-2023.
[0007] In this specification, in one embodiment, an engineered base editing system is described, comprising: a base editor encoded by a nucleic acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% identity with any one of sequence numbers 1727 to 1757; and an engineered guide polynucleotide comprising a spacer sequence that forms a complex with the endonuclease of the base editor and hybridizes with the target nucleic acid sequence.
[0008] In some embodiments, the base editor is encoded by a nucleic acid sequence having at least 90% identity with one of sequence numbers 1727 to 1757.
[0009] In some embodiments, the base editor is coded by a nucleic acid sequence that has 100% identity with any one of sequence numbers 1727 to 1757.
[0010] In some embodiments, the base editor includes a deaminase. In some embodiments, the deaminase is non-covalently bonded to an endonuclease. In some embodiments, the deaminase is covalently bonded to an endonuclease. In some embodiments, the deaminase is fused to an endonuclease. In some embodiments, the manipulated guide polynucleotide is a single guide nucleic acid. In some embodiments, the manipulated guide polynucleotide is a dual guide nucleic acid. In some embodiments, the manipulated guide polynucleotide is RNA.
[0011] In some embodiments, the endonuclease is non-covalently bonded to the manipulated guide polynucleotide. In some embodiments, the endonuclease is covalently bonded to the manipulated guide polynucleotide.
[0012] In some embodiments, the manipulated guide polynucleotide contains a sequence having at least 80% sequence identity with any one of the following sequences: 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, 1705-1710, and 1758-2019.
[0013] In some embodiments, the manipulated guide polynucleotide contains a sequence having at least 90% sequence identity with any one of the following sequences: 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, 1705-1710, and 1758-2019.
[0014] In some embodiments, the manipulated guide polynucleotide contains a sequence having 100% sequence identity with any one of the following sequence numbers: 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, 1705-1710, and 1758-2019.
[0015] In some embodiments, the manipulated guide polynucleotide includes a sequence having at least 80%, 85%, 90%, 95%, 98%, 99%, or 100% identity with any one of sequence numbers 1431-1454, 1479-1483, 1489-1490, 1705-1710, and 1758-2019.
[0016] In some embodiments, the manipulated guide polynucleotide contains a sequence having at least 80% sequence identity with one of sequence numbers 1431-1454, 1704, and 2010-2019.
[0017] In some embodiments, the manipulated guide polynucleotide comprises a sequence having at least 80% sequence identity with one of sequence numbers 1890-1976.
[0018] In some embodiments, the manipulated guide polynucleotide comprises a sequence having at least 80% sequence identity with one of the sequence numbers 1977–2009.
[0019] In some embodiments, the manipulated guide polynucleotide contains a sequence having at least 80% sequence identity with one of the sequence numbers 1479-1483 and 1758-1889.
[0020] In some embodiments, the base editor includes a nickase domain.
[0021] In some embodiments, the nickase includes a mutation from aspartic acid to alanine in residue 9 for SEQ ID NO: 70, residue 13 for SEQ ID NO: 71, 72, or 74, residue 12 for SEQ ID NO: 73, residue 17 for SEQ ID NO: 75, residue 23 for SEQ ID NO: 76, or residue 10 for SEQ ID NO: 597, or any combination thereof.
[0022] In some embodiments, the base editor further comprises a uracil DNA glycosylase inhibitor sequence. In some embodiments, the base editor further comprises a FAM72A sequence. In some embodiments, the FAM72A sequence has at least 80% identity with sequence number 1121.
[0023] In this specification, in some embodiments, nucleic acids encoding the manipulated base editing system described herein are described.
[0024] In this disclosure, in some embodiments, a vector comprising nucleic acids disclosed herein is described. In some embodiments, the vector is a plasmid, a minicircle, a CELiD, an adeno-associated virus (AAV) derived virion, a lentivirus, or an adenovirus.
[0025] In this disclosure, in some embodiments, cells comprising the engineered systems, nucleic acids, or vectors described herein are described. In some embodiments, the cells are eukaryotic cells. In some embodiments, the cells are mammalian cells. In some embodiments, the cells are immortalized cells. In some embodiments, the cells are insect cells. In some embodiments, the cells are yeast cells. In some embodiments, the cells are plant cells. In some embodiments, the cells are fungal cells. In some embodiments, the cells are prokaryotic cells. In some embodiments, the cells are A549, HEK-293, HEK-293T, BHK, CHO, HeLa, MRC5, Sf9, Cos-1, Cos-7, Vero, BSC1, BSC40, BMT10, WI38, HeLa, Saos, C2C12, L cells, HT1080, HepG2, Huh7, K562, primary cells, or derivatives thereof. In some embodiments, the cells are engineered cells. In some embodiments, the cells are stable cells.
[0026] In this specification, a method for modifying a target nucleic acid sequence is described, in some embodiments, which includes contacting the target nucleic acid sequence using an engineered base editing system described herein. In some embodiments, modifying the target nucleic acid sequence includes converting adenine to guanine within the target nucleic acid sequence. In some embodiments, modifying the target nucleic acid sequence includes converting cytosine to uracil within the target nucleic acid sequence. In some embodiments, the target nucleic acid sequence includes deoxyribonucleic acid (DNA). In some embodiments, the target nucleic acid sequence includes ribonucleic acid (RNA). In some embodiments, the target nucleic acid sequence includes genomic DNA, viral DNA, viral RNA, or bacterial DNA.
[0027] In some embodiments, the target nucleic acid sequence is modified in vitro. In some embodiments, the target nucleic acid sequence is modified in vivo. In some embodiments, the target nucleic acid sequence is modified ex vivo. In some embodiments, the target nucleic acid sequence is modified intracellularly. In some embodiments, the cell is a prokaryotic cell, a bacterial cell, a eukaryotic cell, a fungal cell, a plant cell, an animal cell, a mammalian cell, a rodent cell, a primate cell, a human cell, or a primary cell.
[0028] In some embodiments herein, a method of modifying a nucleic acid encoding ANGPTL3 is described, comprising contacting a nucleic acid sequence encoding ANGPTL3 with an engineered base editing system, wherein the aforementioned base editing system is a base editor comprising a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% identity to any one of SEQ ID NOs: 1654 - 1703 and 2021 - 2023, and the sequence does not contain any one of the sequences selected from SEQ ID NOs: 1128 - 1160 and 1363 - 1415, or a base editor encoded by a nucleic acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% identity to any one of SEQ ID NOs: 1727 - 1757, and an engineered guide polynucleotide that forms a complex with the endonuclease of the base editor and hybridizes to the target nucleic acid sequence. In some embodiments, the engineered guide polynucleotide comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 1479 - 1483 and 1758 - 1889.
[0029] In some embodiments described herein, a method of modifying a nucleic acid encoding APOA1 is provided that includes contacting a nucleic acid sequence encoding APOA1 with an engineered base editing system. The base editing system is a base editor that includes a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% identity to any one of SEQ ID NOs: 1654-1703 and 2021-2023, and the sequence does not include any one of the sequences selected from SEQ ID NOs: 1128-1160 and 1363-1415. Alternatively, the base editing system is a base editor encoded by a nucleic acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% identity to any one of SEQ ID NOs: 1727-1757, and includes an engineered guide polynucleotide that forms a complex with an endonuclease of the base editor and includes a spacer sequence that hybridizes to a target nucleic acid sequence. In some embodiments, the engineered guide polynucleotide includes a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 1431-1454, 1704, and 2010-2019.
[0030] In this specification, in one embodiment, a method for modifying a nucleic acid encoding BCL11A is described, comprising contacting the nucleic acid encoding BCL11A with a manipulated base editing system, wherein the base editing system is a base editor having at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% identity with any one of sequence numbers 1654-1703 and 2021-2023, the sequence being sequence numbers 1128-11 The invention comprises a base editor which does not contain any of the sequences selected from 60 and 1363-1415, or a base editor encoded by a nucleic acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% identity with any one of sequence numbers 1727-1757, and an engineered guide polynucleotide which includes a spacer sequence that forms a complex with the endonuclease of the base editor and hybridizes with the target nucleic acid sequence. In some embodiments, the engineered guide polynucleotide includes a sequence which has at least 80% sequence identity with any one of sequence numbers 1890-1976.
[0031] In this specification, in one embodiment, a method for modifying a nucleic acid encoding a PAH is described, comprising contacting the nucleic acid sequence encoding the PAH with an engineered base-editing system, wherein the base-editing system comprises a base editor having at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% identity with any one of sequence numbers 1654-1703 and 2021-2023, wherein the sequence comprises a base editor having at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% identity with any one of sequence numbers 1727-1757, and an engineered guide polynucleotide comprising a spacer sequence that complexes with the endonuclease of the base editor and hybridizes to the target nucleic acid sequence. In some embodiments, the manipulated guide polynucleotide comprises a sequence having at least 80% sequence identity with one of the sequence numbers 1977–2009.
[0032] Further aspects and advantages of the present disclosure will be readily apparent to those skilled in the art from the following detailed description, which shows and describes only exemplary embodiments of the present disclosure. As will be recognized, other different embodiments of the present disclosure are possible, and some of its details can be modified in various obvious ways without departing from the present disclosure. Accordingly, the drawings and specification should be considered illustrative and not restrictive. [Brief explanation of the drawing]
[0033] Novel features of this disclosure are specifically described in the appended claims. A better understanding of the features and advantages of this disclosure will be obtained by referring to the following detailed description illustrating exemplary embodiments in which the principles of this disclosure are utilized, and to the appended drawings (also referred to herein as “Figures” and “Diagrams”).
[0034] [Figure 1A] Figures 1A–1K show the dose-dependent A→G editing activity by ABE07 at 11 distinct target sites after transfection in primary mouse hepatocytes. The bar graph shows the average editing by ABE07 at adenines within the protospacer editing window accessible by the base editor. Target adenines are numbered relative to the distal PAM position. Values and error bars represent the mean ± sem of n=4 different biological replications. Various concentrations of transfected ABE07 mRNA are highlighted in the legend. [Figure 1B] Figures 1A–1K show the dose-dependent A→G editing activity by ABE07 at 11 distinct target sites after transfection in primary mouse hepatocytes. The bar graph shows the average editing by ABE07 at adenines within the protospacer editing window accessible by the base editor. Target adenines are numbered relative to the distal PAM position. Values and error bars represent the mean ± sem of n=4 different biological replications. Various concentrations of transfected ABE07 mRNA are highlighted in the legend. [Figure 1C] Figures 1A–1K show the dose-dependent A→G editing activity by ABE07 at 11 distinct target sites after transfection in primary mouse hepatocytes. The bar graph shows the average editing by ABE07 at adenines within the protospacer editing window accessible by the base editor. Target adenines are numbered relative to the distal PAM position. Values and error bars represent the mean ± sem of n=4 different biological replications. Various concentrations of transfected ABE07 mRNA are highlighted in the legend. [Figure 1D]Figures 1A–1K show the dose-dependent A→G editing activity by ABE07 at 11 distinct target sites after transfection in primary mouse hepatocytes. The bar graph shows the average editing by ABE07 at adenines within the protospacer editing window accessible by the base editor. Target adenines are numbered relative to the distal PAM position. Values and error bars represent the mean ± sem of n=4 different biological replications. Various concentrations of transfected ABE07 mRNA are highlighted in the legend. [Figure 1E] Figures 1A–1K show the dose-dependent A→G editing activity by ABE07 at 11 distinct target sites after transfection in primary mouse hepatocytes. The bar graph shows the average editing by ABE07 at adenines within the protospacer editing window accessible by the base editor. Target adenines are numbered relative to the distal PAM position. Values and error bars represent the mean ± sem of n=4 different biological replications. Various concentrations of transfected ABE07 mRNA are highlighted in the legend. [Figure 1F] Figures 1A–1K show the dose-dependent A→G editing activity by ABE07 at 11 distinct target sites after transfection in primary mouse hepatocytes. The bar graph shows the average editing by ABE07 at adenines within the protospacer editing window accessible by the base editor. Target adenines are numbered relative to the distal PAM position. Values and error bars represent the mean ± sem of n=4 different biological replications. Various concentrations of transfected ABE07 mRNA are highlighted in the legend. [Figure 1G]Figures 1A–1K show the dose-dependent A→G editing activity by ABE07 at 11 distinct target sites after transfection in primary mouse hepatocytes. The bar graph shows the average editing by ABE07 at adenines within the protospacer editing window accessible by the base editor. Target adenines are numbered relative to the distal PAM position. Values and error bars represent the mean ± sem of n=4 different biological replications. Various concentrations of transfected ABE07 mRNA are highlighted in the legend. [Figure 1H] Figures 1A–1K show the dose-dependent A→G editing activity by ABE07 at 11 distinct target sites after transfection in primary mouse hepatocytes. The bar graph shows the average editing by ABE07 at adenines within the protospacer editing window accessible by the base editor. Target adenines are numbered relative to the distal PAM position. Values and error bars represent the mean ± sem of n=4 different biological replications. Various concentrations of transfected ABE07 mRNA are highlighted in the legend. [Figure 1I] Figures 1A–1K show the dose-dependent A→G editing activity by ABE07 at 11 distinct target sites after transfection in primary mouse hepatocytes. The bar graph shows the average editing by ABE07 at adenines within the protospacer editing window accessible by the base editor. Target adenines are numbered relative to the distal PAM position. Values and error bars represent the mean ± sem of n=4 different biological replications. Various concentrations of transfected ABE07 mRNA are highlighted in the legend. [Figure 1J]Figures 1A–1K show the dose-dependent A→G editing activity by ABE07 at 11 distinct target sites after transfection in primary mouse hepatocytes. The bar graph shows the average editing by ABE07 at adenines within the protospacer editing window accessible by the base editor. Target adenines are numbered relative to the distal PAM position. Values and error bars represent the mean ± sem of n=4 different biological replications. Various concentrations of transfected ABE07 mRNA are highlighted in the legend. [Figure 1K] Figures 1A–1K show the dose-dependent A→G editing activity by ABE07 at 11 distinct target sites after transfection in primary mouse hepatocytes. The bar graph shows the average editing by ABE07 at adenines within the protospacer editing window accessible by the base editor. Target adenines are numbered relative to the distal PAM position. Values and error bars represent the mean ± sem of n=4 different biological replications. Various concentrations of transfected ABE07 mRNA are highlighted in the legend. [Figure 2] Figure 2 shows a schematic diagram highlighting the domain structures of homodimer and heterodimer ABEs. Two copies of the MG68-4 variant were inserted into the MG3-6 / 3-8 nuclease chassis, enabling PID exchange and diversifying the PAMs accessible by these ABE enzymes. [Figure 3A] Figures 3A and 3B show box plots comparing A→G editing by ABE variants in 31 distinct gRNAs targeting three different genes in Hepa1-6 cells. Figure 3A shows a box plot showing the average editing observed by individual ABE variants, and Figure 3B shows a box plot showing the maximum editing observed. The data corresponds to n=2 biologically independent replications for each target guide. Each independent data point represents a unique target gene locus. [Figure 3B]Figures 3A and 3B show box plots comparing A→G editing by ABE variants in 31 distinct gRNAs targeting three different genes in Hepa1-6 cells. Figure 3A shows a box plot showing the average editing observed by individual ABE variants, and Figure 3B shows a box plot showing the maximum editing observed. The data corresponds to n=2 biologically independent replications for each target guide. Each independent data point represents a unique target gene locus. [Figure 4A] Figures 4A–4B show box plots comparing C-to-G indiscriminate editing by ABE variants in 31 distinct gRNAs targeting three different genes in Hepa1–6 cells. Figure 4A shows a box plot showing the average editing observed by individual ABE variants, and Figure 4B shows a box plot showing the maximum editing observed. The data corresponds to n=2 biologically independent replications for each target guide. Each independent data point represents a unique target gene locus. [Figure 4B] Figures 4A–4B show box plots comparing C-to-G indiscriminate editing by ABE variants in 31 distinct gRNAs targeting three different genes in Hepa1–6 cells. Figure 4A shows a box plot showing the average editing observed by individual ABE variants, and Figure 4B shows a box plot showing the maximum editing observed. The data corresponds to n=2 biologically independent replications for each target guide. Each independent data point represents a unique target gene locus. [Figure 5] Figure 5 shows a box plot comparing undesirable indel formation by ABE variants in 31 distinct gRNAs targeting three different genes in Hepa1-6 cells. The box plot shows the average indel formation observed by each ABE variant, and the data corresponds to n=2 biologically independent replications for each target guide. Each independent data point represents a unique target gene locus. [Figure 6]Figure 6 shows a heatmap illustrating the breakdown of maximum A→G editing activity observed across highly edited gRNAs targeting three different genes in Hepa1-6 cells, broken down by guide and variant. Guides are ranked from left to right in decreasing order according to the observed A→G editing activity. ABE07-77 are highlighted along with guides selected for further in vivo investigation. [Figure 7] Figure 7 shows the in vivo A→G editing activity of the ABE07-77 variant across six different gRNA targets. The box plot shows the maximum editing observed by ABE07-77. Individual data points represent individual mice. Values and error bars represent the mean ± sem of n=4 different mice. [Figure 8] Figure 8 shows the in vivo indel formation activity of the ABE07-77 variant across six distinct gRNA targets. The bar graph shows indels formed by ABE07-77 and its corresponding parental nuclease, MG3-6 / 3-8. Each data point represents an individual mouse. Values and error bars represent the mean ± sem of n=4 different mice. [Figure 9] Figure 9 shows the in vivo A→G editing activity of the ABE07-77 variant across protospacers at the most edited target loci. The bar graph shows the average editing by ABE07-77 at adenines within the protospacer editing window accessible by the base editor. Target adenines are numbered relative to the distal PAM position. Individual data points represent individual mice. Values and error bars represent the mean ± sem of n=4 different mice. [Figure 10]Figure 10 shows a graph illustrating that the maximum C to T edits for 139-52-V2, 139-52-V13, 139-52-V14, 139-52-V17, 139-86v12, and 152-6v13 across all five guides were 26.2%, 58.2%, 22.7%, 16.8%, 12.6%, and 50.9%, respectively. Compared to the hyperactive positive control variant: 139-52-V2, 139-52-V13, 139-52-V14, 139-52-V17, 139-86v12, and 152-6v13 were edited at 159.2%, 102.1%, 353.7%, 137.8%, 76.3%, and 309.1% of the maximum positive control edit, respectively. [Figure 11] Figure 11 shows a bar graph indicating that the -1 nucleotide at the 5' position of the deaminated cytidine is important for binding and deamination reactions by cytidine deaminase. NGS data were used to measure the -1 nucleotide preference for each of the six engineered CDA variants. To do this, for each guide, the number of readings for the four most edited cytidine sites, edited by >1% (and their -1 nucleotide identity), were compiled into a table. This was done separately for all five guides, and then the preferences were averaged across all five guides. The resulting graph represents the relative -1 nucleotide preference for the type of cytidine that each CDA prefers to deaminate, across the five guides targeting the HEK293 engineered site. [Figure 12A] Figures 12A and 12B show graphs comparing targeted A→G editing by oligomeric ABE variants in 15 distinct gRNAs targeting two different genes in Hepa1-6 cells. The box plots show the observed mean editing (Figure 12A) and observed maximum editing (Figure 12B) by individual ABE variants, with data corresponding to n=2 biologically independent replications for each target guide. Each individual data point represents a unique target gene locus. [Figure 12B]Figures 12A and 12B show graphs comparing targeted A→G editing by oligomeric ABE variants in 15 distinct gRNAs targeting two different genes in Hepa1-6 cells. The box plots show the observed mean editing (Figure 12A) and observed maximum editing (Figure 12B) by individual ABE variants, with data corresponding to n=2 biologically independent replications for each target guide. Each individual data point represents a unique target gene locus. [Figure 13A] Figures 13A and 13B show graphs comparing indiscriminate C editing by oligomeric ABE variants in 15 distinct gRNAs targeting two different genes in Hepa1-6 cells. Box plots show the observed mean C editing (Figure 13A) and observed maximum C editing (Figure 13B) by individual ABE variants, with data corresponding to n=2 biologically independent replications for each target guide. Each individual data point represents a unique target gene locus. [Figure 13B] Figures 13A and 13B show graphs comparing indiscriminate C editing by oligomeric ABE variants in 15 distinct gRNAs targeting two different genes in Hepa1-6 cells. Box plots show the observed mean C editing (Figure 13A) and observed maximum C editing (Figure 13B) by individual ABE variants, with data corresponding to n=2 biologically independent replications for each target guide. Each individual data point represents a unique target gene locus. [Figure 14A] Figures 14A and 14B show graphs comparing targeted A→G editing by manipulated ABE variants in 15 distinct gRNAs targeting two different genes in Hepa1-6 cells. The box plots show the observed mean editing (Figure 14A) and observed maximum editing (Figure 14B) by individual ABE variants, with data corresponding to n=2 biologically independent replications for each target guide. Each individual data point represents a unique target gene locus. [Figure 14B]Figures 14A and 14B show graphs comparing targeted A→G editing by manipulated ABE variants in 15 distinct gRNAs targeting two different genes in Hepa1-6 cells. The box plots show the observed mean editing (Figure 14A) and observed maximum editing (Figure 14B) by individual ABE variants, with data corresponding to n=2 biologically independent replications for each target guide. Each individual data point represents a unique target gene locus. [Figure 15A] Figures 15A and 15B show graphs comparing indiscriminate C editing by manipulated ABE variants in 15 distinct gRNAs targeting two different genes in Hepa1-6 cells. Box plots show the observed mean C editing (Figure 15A) and observed maximum C editing (Figure 15B) by individual ABE variants, with data corresponding to n=2 biologically independent replications for each target guide. Each individual data point to a unique target locus. [Figure 15B] Figures 15A and 15B show graphs comparing indiscriminate C editing by manipulated ABE variants in 15 distinct gRNAs targeting two different genes in Hepa1-6 cells. Box plots show the observed mean C editing (Figure 15A) and observed maximum C editing (Figure 15B) by individual ABE variants, with data corresponding to n=2 biologically independent replications for each target guide. Each individual data point to a unique target locus. [Figure 16] Figure 16 shows a graph comparing undesirable indel formation by engineered ABE variants in 15 distinct gRNAs targeting two different genes in Hepa1-6 cells. The box plots show the average observed indel formation (indel percentage) for each ABE variant, with data corresponding to n=2 biologically independent replications for each target guide. Each individual data point represents a unique target gene locus. [Figure 17A]Figures 17A–17C show graphs illustrating the preferred sequence contexts for edited adenines. These graphs show the frequency distribution of bases surrounding all editable adenines within 15 distinct gRNAs tested in this study (Figure 17A), highly edited (over 30% editing activity) adenines by ABE15 (Figure 17B), and less edited (less than 30% editing activity) adenines by ABE15 (Figure 17C). [Figure 17B] Figures 17A–17C show graphs illustrating the preferred sequence contexts for edited adenines. These graphs show the frequency distribution of bases surrounding all editable adenines within 15 distinct gRNAs tested in this study (Figure 17A), highly edited (over 30% editing activity) adenines by ABE15 (Figure 17B), and less edited (less than 30% editing activity) adenines by ABE15 (Figure 17C). [Figure 17C] Figures 17A–17C show graphs illustrating the preferred sequence contexts for edited adenines. These graphs show the frequency distribution of bases surrounding all editable adenines within 15 distinct gRNAs tested in this study (Figure 17A), highly edited (over 30% editing activity) adenines by ABE15 (Figure 17B), and less edited (less than 30% editing activity) adenines by ABE15 (Figure 17C). [Figure 18A] Figures 18A and 18B show graphs comparing targeted A→G editing by manipulated ABE variants in 15 distinct gRNAs targeting two different genes in Hepa1-6 cells. The box plots show the observed mean editing (Figure 18A) and observed maximum editing (Figure 18B) by individual ABE variants, with data corresponding to n=2 biologically independent replications for each target guide. Each individual data point represents a unique target gene locus. [Figure 18B]Figures 18A and 18B show graphs comparing targeted A→G editing by manipulated ABE variants in 15 distinct gRNAs targeting two different genes in Hepa1-6 cells. The box plots show the observed mean editing (Figure 18A) and observed maximum editing (Figure 18B) by individual ABE variants, with data corresponding to n=2 biologically independent replications for each target guide. Each individual data point represents a unique target gene locus. [Figure 19A] Figures 19A and 19B show graphs comparing indiscriminate C-editing by engineered ABE variants in 15 distinct gRNAs targeting two different genes in Hepa1-6 cells. Box plots show the observed mean C-editing (Figure 19A) and observed maximum C-editing (Figure 19B) by individual ABE variants, with data corresponding to n=2 biologically independent replications for each target guide. Each individual data point represents a unique target gene locus. [Figure 19B] Figures 19A and 19B show graphs comparing indiscriminate C-editing by engineered ABE variants in 15 distinct gRNAs targeting two different genes in Hepa1-6 cells. Box plots show the observed mean C-editing (Figure 19A) and observed maximum C-editing (Figure 19B) by individual ABE variants, with data corresponding to n=2 biologically independent replications for each target guide. Each individual data point represents a unique target gene locus. [Figure 20A] Figures 20A and 20B show graphs illustrating undesirable indel formation for engineered ABE variants in 15 distinct gRNAs targeting two different genes in Hepa1-6 cells, and favorable sequence editing contexts for engineered ABE23. Figure 20A shows a box plot illustrating the mean of observed indel formation for each ABE variant, with data corresponding to n=2 biologically independent replications for each target guide. Each individual data point represents a unique target gene locus. [Figure 20B]Figures 20A and 20B show graphs illustrating undesirable indel formation of engineered ABE variants in 15 distinct gRNAs targeting two different genes in Hepa1-6 cells, and favorable sequence editing contexts for engineered ABE23. Figure 20B shows the frequency distribution of bases surrounding highly edited (over 30% editing activity) adenines by ABE23. [Figure 21A] Figures 21A and 21B show graphs comparing targeted A→G editing by engineered D109Q ABE variants and variants with additional proline mutations designed to mitigate sequence preference in 15 distinct gRNAs targeting two different genes in Hepa1-6 cells. Box plots show the observed mean edits (Figure 21A) and observed maximum edits (Figure 21B) by individual ABE variants, with data corresponding to n=2 biologically independent replications for each target guide. Each individual data point represents a unique target locus. D109Q mutations designed to suppress indiscriminate C deamination increased the mean and maximum A→G edits. [Figure 21B] Figures 21A and 21B show graphs comparing targeted A→G editing by engineered D109Q ABE variants and variants with additional proline mutations designed to mitigate sequence preference in 15 distinct gRNAs targeting two different genes in Hepa1-6 cells. Box plots show the observed mean edits (Figure 21A) and observed maximum edits (Figure 21B) by individual ABE variants, with data corresponding to n=2 biologically independent replications for each target guide. Each individual data point represents a unique target locus. D109Q mutations designed to suppress indiscriminate C deamination increased the mean and maximum A→G edits. [Figure 22A]Figures 22A and 22B show graphs comparing indiscriminate C editing by an engineered D109Q ABE variant and a variant with an additional proline mutation designed to mitigate sequence preference in 15 distinct gRNAs targeting two different genes in Hepa1-6 cells. Box plots show the observed mean C editing (Figure 22A) and observed maximum C editing (Figure 22B) by individual ABE variants, with data corresponding to n=2 biologically independent replications for each target guide. [Figure 22B] Figures 22A and 22B show graphs comparing indiscriminate C editing by an engineered D109Q ABE variant and a variant with an additional proline mutation designed to mitigate sequence preference in 15 distinct gRNAs targeting two different genes in Hepa1-6 cells. Box plots show the observed mean C editing (Figure 22A) and observed maximum C editing (Figure 22B) by individual ABE variants, with data corresponding to n=2 biologically independent replications for each target guide. [Figure 23] Figure 23 shows a graph comparing undesirable indel formation and editing context for 15 distinct gRNAs targeting two different genes in Hepa1-6 cells, specifically for an engineered D109Q variant and a variant with additional proline mutations designed to mitigate sequence preference. The box plot shows the mean of observed indel formation by each ABE variant, with data corresponding to n=2 biologically independent replications for each target guide. Each individual data point represents a unique target gene locus. [Figure 24A] Figures 24A and 24B show graphs comparing the most effective ABE-based targeted A→G editing in 15 distinct gRNAs targeting two different genes in Hepa1-6 cells. Box plots show the observed mean editing (Figure 24A) and observed maximum editing (Figure 24B) by individual ABE variants, with data corresponding to n=2 biologically independent replications for each target guide. Each individual data point represents a unique target gene locus. [Figure 24B] Figures 24A and 24B show graphs comparing the most effective ABE-based targeted A→G editing in 15 distinct gRNAs targeting two different genes in Hepa1-6 cells. Box plots show the observed mean editing (Figure 24A) and observed maximum editing (Figure 24B) by individual ABE variants, with data corresponding to n=2 biologically independent replications for each target guide. Each individual data point represents a unique target gene locus. [Figure 25A] Figures 25A and 25B show graphs comparing the most effective ABEs for indiscriminate C-editing in 15 distinct gRNAs targeting two different genes in Hepa1-6 cells. Box plots show the observed mean C-editing (Figure 25A) and observed maximum C-editing (Figure 25B) by individual ABE variants, with data corresponding to n=2 biologically independent replications for each target guide. Each individual data point represents a unique target gene locus. [Figure 25B] Figures 25A and 25B show graphs comparing the most effective ABEs for indiscriminate C-editing in 15 distinct gRNAs targeting two different genes in Hepa1-6 cells. Box plots show the observed mean C-editing (Figure 25A) and observed maximum C-editing (Figure 25B) by individual ABE variants, with data corresponding to n=2 biologically independent replications for each target guide. Each individual data point represents a unique target gene locus. [Figure 26] Figure 26 shows a graph comparing undesirable indel formation to the most effective ABE for 15 distinct gRNAs targeting two different genes in Hepa1-6 cells. The box plot shows the average observed indel formation for each ABE variant, with data corresponding to n=2 biologically independent replications for each target guide. Each individual data point represents a unique target gene locus. [Figure 27A]Figures 27A–27C show a small base editor fusion construct, including a schematic diagram of a small base editor showing the relative position of the deaminase domain to the nickase domain (Figure 27A), a structural alignment of the computationally predicted structure of MG3-6_3-8ABE (light gray) with the predicted structure of MG34-29 nuclease (dark gray) (Figure 27B), and an alternative diagram (Figure 27CB) in which MG3-6_3-8 nickase is omitted from the alignment shown in Figure 27B, revealing a potential loop in the predicted MG34-29 structure for fusing MG68-4 deaminase. [Figure 27B] Figures 27A–27C show a small base editor fusion construct, including a schematic diagram of a small base editor showing the relative position of the deaminase domain to the nickase domain (Figure 27A), a structural alignment of the computationally predicted structure of MG3-6_3-8ABE (light gray) with the predicted structure of MG34-29 nuclease (dark gray) (Figure 27B), and an alternative diagram (Figure 27CB) in which MG3-6_3-8 nickase is omitted from the alignment shown in Figure 27B, revealing a potential loop in the predicted MG34-29 structure for fusing MG68-4 deaminase. [Figure 27C] Figures 27A–27C show a small base editor fusion construct, including a schematic diagram of a small base editor showing the relative position of the deaminase domain to the nickase domain (Figure 27A), a structural alignment of the computationally predicted structure of MG3-6_3-8ABE (light gray) with the predicted structure of MG34-29 nuclease (dark gray) (Figure 27B), and an alternative diagram (Figure 27CB) in which MG3-6_3-8 nickase is omitted from the alignment shown in Figure 27B, revealing a potential loop in the predicted MG34-29 structure for fusing MG68-4 deaminase. [Figure 28A]Figures 28A–28D show graphs illustrating efficient A→G editing activity by small base editors in immortalized human K562 cells. The bar graphs show the maximum A→G editing observed across panels of multiple guides, with ABE33 (Figure 28A), a positive control for the MG34-29 ABE and hAPOA1 target guides, and with MG102-39 ABE (Figure 28C). The mean editing observed across highly edited target loci, AAVS1 E7, AAVS1 C7 (Figure 28B), and TRAC C11 (Figure 28D), is illustrated as bar graphs. Target adenines are numbered relative to their distal PAM position in each ABE variant. Values and error bars represent the mean ± sem of n=2 biologically independent replications for each target guide. [Figure 28B] Figures 28A–28D show graphs illustrating efficient A→G editing activity by small base editors in immortalized human K562 cells. The bar graphs show the maximum A→G editing observed across panels of multiple guides, with ABE33 (Figure 28A), a positive control for the MG34-29 ABE and hAPOA1 target guides, and with MG102-39 ABE (Figure 28C). The mean editing observed across highly edited target loci, AAVS1 E7, AAVS1 C7 (Figure 28B), and TRAC C11 (Figure 28D), is illustrated as bar graphs. Target adenines are numbered relative to their distal PAM position in each ABE variant. Values and error bars represent the mean ± sem of n=2 biologically independent replications for each target guide. [Figure 28C]Figures 28A–28D show graphs illustrating efficient A→G editing activity by small base editors in immortalized human K562 cells. The bar graphs show the maximum A→G editing observed across panels of multiple guides, with ABE33 (Figure 28A), a positive control for the MG34-29 ABE and hAPOA1 target guides, and with MG102-39 ABE (Figure 28C). The mean editing observed across highly edited target loci, AAVS1 E7, AAVS1 C7 (Figure 28B), and TRAC C11 (Figure 28D), is illustrated as bar graphs. Target adenines are numbered relative to their distal PAM position in each ABE variant. Values and error bars represent the mean ± sem of n=2 biologically independent replications for each target guide. [Figure 28D] Figures 28A–28D show graphs illustrating efficient A→G editing activity by small base editors in immortalized human K562 cells. The bar graphs show the maximum A→G editing observed across panels of multiple guides, with ABE33 (Figure 28A), a positive control for the MG34-29 ABE and hAPOA1 target guides, and with MG102-39 ABE (Figure 28C). The mean editing observed across highly edited target loci, AAVS1 E7, AAVS1 C7 (Figure 28B), and TRAC C11 (Figure 28D), is illustrated as bar graphs. Target adenines are numbered relative to their distal PAM position in each ABE variant. Values and error bars represent the mean ± sem of n=2 biologically independent replications for each target guide. [Figure 29A] Figures 29A–29C illustrate PAM interaction domain (PID) engineering used to develop a series of chimeric MG3-6 base editors with broad genome targeting capabilities. Figure 29A shows the proportion of genomic adenines in the hg38 human reference genome that can be targeted by SpCas9 ABE. [Figure 29B]Figures 29A–29C illustrate PAM interaction domain (PID) engineering used to develop a series of chimeric MG3-6 base editors with broad genome targeting capabilities. Figure 29B shows a schematic diagram of a PID-exchangeable chimeric nuclease platform, highlighting the diversity and range of PAMs accessible through PID exchange. [Figure 29C] Figures 29A–29C illustrate PAM interaction domain (PID) engineering used to develop a series of chimeric MG3-6 base editors with broad genome targeting capabilities. Figure 29C shows the proportion of genomic adenines in the hg38 human reference genome that can be targeted using an MG3-6 PID-exchangeable ABE platform. [Figure 30A] Figures 30A–30D show schematic diagrams of the experimental design for the high-throughput chimeric ABE test. Figure 30A shows the editing windows of two highly active MG3-6_3-8 ABEs. Each dot represents a unique editing event observed across more than 15 unique guides guided by ABE07 (SEQ ID NO: 1411) and ABE33 (SEQ ID NO: 1673). The Gaussian curve in the upper panel shows the estimated editing windows of the MG3-6_3-8 ABEs used to design the tiling guides used throughout this experiment. [Figure 30B] Figures 30A–30D show schematic diagrams of experimental designs for high-throughput chimeric ABE studies. Figure 30B shows a scheme demonstrating base editing of splice sites using ABE, which results in intron retention (splice donor disruption) or exon skipping (splice acceptor disruption), and can therefore be used for targeted gene knockout. [Figure 30C] Figures 30A–30D show schematic diagrams of the experimental design for the high-throughput chimeric ABE study. Plate configurations for (Figure 30C) high-throughput pool screen and (Figure 30D) deconvolution screen were used to identify highly active ABEs and derive combinations that resulted in efficient targeted knockout. [Figure 30D]Figures 30A–30D show schematic diagrams of the experimental design for the high-throughput chimeric ABE study. Plate configurations for (Figure 30C) high-throughput pool screen and (Figure 30D) deconvolution screen were used to identify highly active ABEs and derive combinations that resulted in efficient targeted knockout. [Figure 31A] Figures 31A and 31B show targeted knockout of hANGPTL3 using PID-exchanged chimeric ABEs. Figure 31A shows a graph comparing the theoretical splice site targeting potential of the reference ABE and the MG3-6_3-8 chimeric ABE. [Figure 31B] Figures 31A and 31B show targeted knockout of hANGPTL3 using PID-exchanged chimeric ABEs. Figure 31B shows the results of a pooled screening of compatible chimeric ABEs across the splice site of the hANGPTL3 gene tested in K562 cells. Heatmap values indicate the A→G editing rate observed at the splice site adenine by the corresponding chimeric ABE variant. [Figure 32A] Figures 32A–32D show targeted knockout of hANGPTL3 using PID-exchange chimeric ABEs. Bar graphs of A→G editing at splice site adenine used to deconvolve combinations of active guides and chimeric ABEs in hANGPTL3 exon 1 splice donor (Figure 32A), exon 4 splice acceptor and splice donor (Figure 32B), exon 5 splice donor (Figure 32C), and exon 7 splice acceptor (Figure 32D). [Figure 32B] Figures 32A–32D show targeted knockout of hANGPTL3 using PID-exchange chimeric ABEs. Bar graphs of A→G editing at splice site adenine used to deconvolve combinations of active guides and chimeric ABEs in hANGPTL3 exon 1 splice donor (Figure 32A), exon 4 splice acceptor and splice donor (Figure 32B), exon 5 splice donor (Figure 32C), and exon 7 splice acceptor (Figure 32D). [Figure 32C] Figures 32A–32D show targeted knockout of hANGPTL3 using PID-exchange chimeric ABEs. Bar graphs of A→G editing at splice site adenine used to deconvolve combinations of active guides and chimeric ABEs in hANGPTL3 exon 1 splice donor (Figure 32A), exon 4 splice acceptor and splice donor (Figure 32B), exon 5 splice donor (Figure 32C), and exon 7 splice acceptor (Figure 32D). [Figure 32D] Figures 32A–32D show targeted knockout of hANGPTL3 using PID-exchange chimeric ABEs. Bar graphs of A→G editing at splice site adenine used to deconvolve combinations of active guides and chimeric ABEs in hANGPTL3 exon 1 splice donor (Figure 32A), exon 4 splice acceptor and splice donor (Figure 32B), exon 5 splice donor (Figure 32C), and exon 7 splice acceptor (Figure 32D). [Figure 33A] Figures 33A to 33C show the target enhancer site disruption of the GATA1 binding site in hBCL11A using PID-exchanged chimeric ABEs. Figure 33A shows a graph comparing the theoretical targetability of GATA1 enhancer binding sites for the reference ABE and the MG3-6_3-8 chimeric ABE. [Figure 33B] Figures 33A–33C show targeted enhancer site disruption of the GATA1 binding site in hBCL11A using PID-exchanged chimeric ABEs. Figure 33B shows the results of a pooled screening of compatible chimeric ABEs across various enhancer factor binding sites of the BCL11A gene, tested in K562 cells. Heatmap values indicate the A→G editing rate observed at the GATA site adenine by the corresponding chimeric ABE variant. [Figure 33C]Figures 33A–33C show target enhancer site disruption of the GATA1 binding site in hBCL11A using PID-exchanged chimeric ABEs. Figure 33C shows a bar graph of A→G editing at GATA1 adenine used to deconvolve the combination of active guide and chimeric ABE at the DHS+58 enhancer binding site. [Figure 34A] Figures 34A to 34D show the target correction of hPAH SNVs using PID-exchanged chimeric ABEs. Figure 34A shows a table of the most common pathogenic SNVs reported in hPAHs and their responsiveness to sapropterin, quoted from Regier DS, Greene CL. Phenylalanine Hydroxylase Deficiency. [Figure 34B] Figures 34A–34D show targeted correction of hPAH SNVs using PID-exchanged chimeric ABEs. Figures 34B–34C show a pooled screening of compatible chimeric ABEs across SNV sites of hPAH genes tested in modified cell lines. Heatmap values indicate the percentage of A→G editing observed at the target adenine, as well as bystander A with the corresponding chimeric ABE variant for the following SNVs: 1222C>T(p.R408W)SNV (Figure 34B), 1066-11G>A (Figure 34C), and 1315+1G>A (Figure 34D). [Figure 34C] Figures 34A–34D show targeted correction of hPAH SNVs using PID-exchanged chimeric ABEs. Figures 34B–34C show a pooled screening of compatible chimeric ABEs across SNV sites of hPAH genes tested in modified cell lines. Heatmap values indicate the percentage of A→G editing observed at the target adenine, as well as bystander A with the corresponding chimeric ABE variant for the following SNVs: 1222C>T(p.R408W)SNV (Figure 34B), 1066-11G>A (Figure 34C), and 1315+1G>A (Figure 34D). [Figure 34D]Figures 34A–34D show targeted correction of hPAH SNVs using PID-exchanged chimeric ABEs. Figures 34B–34C show a pooled screening of compatible chimeric ABEs across SNV sites of hPAH genes tested in modified cell lines. Heatmap values indicate the percentage of A→G editing observed at the target adenine, as well as bystander A with the corresponding chimeric ABE variant for the following SNVs: 1222C>T(p.R408W)SNV (Figure 34B), 1066-11G>A (Figure 34C), and 1315+1G>A (Figure 34D).
[0035] A brief explanation of sequence listings The sequence listings submitted with this specification provide exemplary polynucleotide and polypeptide sequences for use in the methods, compositions, and systems described herein. The following is an exemplary description of some of these sequences.
[0036] Sequence IDs 1-47 show the full-length peptide sequences of MG66 deaminase suitable for the manipulated nucleic acid editing systems described herein.
[0037] Sequence IDs 48-49 show the full-length peptide sequences of MG67 deaminase suitable for the manipulated nucleic acid editing system described herein.
[0038] Sequence IDs 50-51 show the full-length peptide sequences of MG68 deaminase suitable for the engineered nucleic acid editing system described herein.
[0039] Sequence IDs 52-56 represent sequences of uracil DNA glycosylase inhibitors suitable for the manipulated nucleic acid editing systems described herein.
[0040] Sequence numbers 57-66 show the sequence of the reference deaminase.
[0041] Sequence ID 67 shows the sequence of the reference uracil DNA glycosylase inhibitor.
[0042] Sequence ID 68 shows the sequence of the adenine base editor.
[0043] Sequence ID 69 shows the sequence of the cytosine base editor.
[0044] Sequence IDs 70-78 represent the full-length peptide sequences of MG nickase suitable for the manipulated nucleic acid editing systems described herein.
[0045] Sequence IDs 79-87 represent the protospacers and PAMs used in the in vitro nickase assays described herein.
[0046] Sequence IDs 88-96 show the peptide sequences of single guide RNAs used in the in vitro niccasing assays described herein.
[0047] Sequence numbers 97-156 show the spacer sequence when targeting E. coli lacZ.
[0048] Sequence numbers 157-176 show the primer sequences used when performing site-directed mutagenesis.
[0049] Sequence numbers 177-178 show the primer sequences for lacZ sequencing.
[0050] Sequence numbers 179–342 show the primer sequences used during amplification.
[0051] Sequence numbers 343-345 show the primer sequences for lacZ sequencing.
[0052] Sequence numbers 346-359 show the primer sequences used during amplification.
[0053] Sequence IDs 360-368 represent protospacer adjacent motifs suitable for the manipulated nucleic acid editing systems described herein.
[0054] Sequence IDs 369-384 represent nuclear localization sequences (NLS) suitable for the manipulated nucleic acid editing systems described herein.
[0055] Sequence IDs 385-443 represent the full-length peptide sequences of MG68 deaminase suitable for the engineered nucleic acid editing systems described herein.
[0056] Sequence IDs 444-447 represent the full-length peptide sequences of MG121 deaminase suitable for the engineered nucleic acid editing system described herein.
[0057] Sequence IDs 448-475 represent the full-length peptide sequences of MG68 deaminase suitable for the engineered nucleic acid editing system described herein.
[0058] Sequence IDs 476 and 477 show the sequences of the adenine base editor.
[0059] Sequence numbers 478-482 show the cytosine base editor sequences.
[0060] Sequence IDs 483–487 represent plasmid sequences suitable for encoding the manipulated nucleic acid editing systems described herein.
[0061] Sequence IDs 488 and 489 show the sgRNA scaffold sequences for MG15-1 and MG34-1, respectively.
[0062] Sequence IDs 490–522 show the sequences of spacers used to target genomic loci in E. coli and HEK293T cells.
[0063] Sequence IDs 523–585 show the primer sequences used during amplification and Sanger sequencing.
[0064] Sequence numbers 584-585 show the primer sequences used during amplification.
[0065] Sequence ID 586 shows the sequence of the adenine base editor.
[0066] Sequence ID 587 shows the sequence of the cytosine base editor.
[0067] Sequence numbers 588-589 show the sequences of the adenine base editor.
[0068] Sequence IDs 590-593 represent the full-length peptide sequences of linkers suitable for the manipulated nucleic acid editing systems described herein.
[0069] Sequence ID 594 shows the sequence of cytosine deaminase.
[0070] Sequence ID 595 shows the sequence of adenosine deaminase.
[0071] Sequence ID 596 shows the sequence of an MG34 active effector suitable for the manipulated nucleic acid editing system described herein.
[0072] Sequence ID 597 shows the sequence of MG34 nickase suitable for the manipulated nucleic acid editing system described herein.
[0073] Sequence ID 598 shows the sequence of MG34 PAM.
[0074] Sequence IDs 599-638 represent the full-length peptide sequences of MG138 cytidine deaminase suitable for the engineered nucleic acid editing system described herein.
[0075] Sequence IDs 639-659 represent the full-length peptide sequences of MG139 cytidine deaminase suitable for the engineered nucleic acid editing system described herein.
[0076] Sequence IDs 660-662 represent the full-length peptide sequences of MG141 cytidine deaminase suitable for the engineered nucleic acid editing systems described herein.
[0077] Sequence IDs 663-664 represent the full-length peptide sequences of MG142 cytidine deaminase suitable for the engineered nucleic acid editing system described herein.
[0078] Sequence IDs 665-675 represent the full-length peptide sequences of MG93 cytidine deaminase suitable for the engineered nucleic acid editing system described herein.
[0079] Sequence numbers 676-678 show the sequences of the adenine base editor.
[0080] Sequence IDs 679-680 show the sgRNA scaffold sequences for MG34-1 and SpCas9.
[0081] Sequence IDs 681-689 represent spacer sequences used to target genomic loci in the guide RNA.
[0082] Sequence IDs 690–707 show the primer sequences used to amplify genomic targets of adenine base editors (ABEs) for next-generation sequencing (NGS) analysis.
[0083] Sequence ID 708 shows the sequence of a blast-cystic (BSD) resistant cassette.
[0084] Sequence IDs 709–719 represent spacer sequences used to target genomic loci in the guide RNA.
[0085] Sequence IDs 720–726 represent plasmid sequences suitable for encoding the manipulated nucleic acid editing systems described herein.
[0086] Sequence numbers 728-729 show the sequences of the adenine base editor.
[0087] Sequence IDs 730-736 represent spacer sequences used to target genomic loci in the guide RNA.
[0088] Sequence IDs 737-738 represent plasmid sequences suitable for encoding the manipulated nucleic acid editing systems described herein.
[0089] Sequence numbers 739-740 show the sequences of the cytidine base editor.
[0090] Sequence ID 741 shows a plasmid sequence suitable for encoding the A1CF gene.
[0091] Sequence ID 742 shows the RNA sequence used to test CDA for RNA activity.
[0092] Sequence ID 743 shows the sequence of a labeled primer for a poisoned primer extension assay used to test CDA for RNA activity.
[0093] Sequence IDs 744-827 represent the full-length peptide sequences of MG139 cytidine deaminase suitable for the engineered nucleic acid editing system described herein.
[0094] Sequence ID 828 shows the full-length peptide sequence of MG93 cytidine deaminase suitable for the engineered nucleic acid editing system described herein.
[0095] Sequence ID 829 shows the full-length peptide sequence of MG142 cytidine deaminase suitable for the engineered nucleic acid editing system described herein.
[0096] Sequence IDs 830-835 represent the full-length peptide sequences of MG152 cytidine deaminase suitable for the engineered nucleic acid editing system described herein.
[0097] Sequence numbers 836-860 show the sequences of the adenine base editor.
[0098] Sequence IDs 861-864 represent spacer sequences used to target genomic loci in the guide RNA.
[0099] Sequence IDs 865–872 show the primer sequences used to amplify genomic targets of adenine base editors (ABEs) for next-generation sequencing (NGS) analysis.
[0100] Sequence IDs 873–875 represent plasmid sequences suitable for encoding the manipulated nucleic acid editing systems described herein.
[0101] Sequence ID 876 shows the sgRNA scaffold sequence of MG34-1.
[0102] Sequence numbers 877-916 show the sequences of the cytosine base editor.
[0103] Sequence IDs 917-931 represent sgRNA sequences suitable for the manipulated nucleic acid editing systems described herein.
[0104] Sequence IDs 932–961 show the primer sequences used to amplify the genomic target of an adenine base editor (ABE) for next-generation sequencing (NGS) analysis.
[0105] Sequence ID 962 shows the manipulated site in a mammalian cell line with five PAMs compatible with Cas9 and MG3-6 editing.
[0106] Sequence IDs 963-967 represent sgRNA sequences suitable for the manipulated nucleic acid editing systems described herein.
[0107] Sequence numbers 968-969 show the cytosine base editor sequence.
[0108] Sequence ID 970 shows the full-length peptide sequence of MG139 cytidine deaminase suitable for the engineered nucleic acid editing system described herein.
[0109] Sequence IDs 971-977 represent the full-length peptide sequences of MG93 cytidine deaminase suitable for the manipulated nucleic acid editing system described herein.
[0110] Sequence IDs 978-981 represent the full-length peptide sequences of MG138 cytidine deaminase suitable for the engineered nucleic acid editing system described herein.
[0111] Sequence ID 982 shows the full-length peptide sequence of MG142 cytidine deaminase suitable for the manipulated nucleic acid editing system described herein.
[0112] Sequence IDs 983-1014 represent the full-length peptide sequences of MG128 deaminase suitable for the engineered nucleic acid editing systems described herein.
[0113] Sequence IDs 1015-1026 represent the full-length peptide sequences of MG129 deaminase suitable for the engineered nucleic acid editing system described herein.
[0114] Sequence IDs 1027-1031 represent the full-length peptide sequences of MG130 deaminase suitable for the engineered nucleic acid editing systems described herein.
[0115] Sequence IDs 1032-1040 represent the full-length peptide sequences of MG131 deaminase suitable for the engineered nucleic acid editing systems described herein.
[0116] Sequence IDs 1041-1043 represent the full-length peptide sequences of MG132 deaminase suitable for the engineered nucleic acid editing system described herein.
[0117] Sequence IDs 1044-1057 represent the full-length peptide sequences of MG133 deaminase suitable for the engineered nucleic acid editing system described herein.
[0118] Sequence IDs 1058-1061 represent the full-length peptide sequences of MG134 deaminase suitable for the engineered nucleic acid editing systems described herein.
[0119] Sequence IDs 1062-1069 represent the full-length peptide sequences of MG135 deaminase suitable for the engineered nucleic acid editing system described herein.
[0120] Sequence IDs 1070-1081 represent the full-length peptide sequences of MG136 deaminase suitable for the engineered nucleic acid editing systems described herein.
[0121] Sequence IDs 1082-1098 represent the full-length peptide sequences of MG137 deaminase suitable for the engineered nucleic acid editing systems described herein.
[0122] Sequence IDs 1099-1105 represent sgRNA sequences suitable for the manipulated nucleic acid editing systems described herein.
[0123] Sequence numbers 1106-1111 represent the sequence of MG35 PAM.
[0124] Sequence ID 1112 shows the DNA sequence of the gene encoding the ABE-MG35-1 adenine base editor.
[0125] Sequence ID 1113 shows the protein sequence of the ABE-MG35-1 adenine base editor.
[0126] Sequence ID 1114 shows the nucleotide sequence of a plasmid encoding a Cas9-based cytosine base editor (CBE).
[0127] Sequence ID 1115 shows the nucleotide sequence of the plasmid encoding Fam72a.
[0128] Sequence numbers 1116-1117 show the sequences of the Cas9-CBE target sites.
[0129] Sequence numbers 1118-1119 show the sequence of the NGS amplicon.
[0130] Sequence ID 1120 shows the full-length peptide sequence of MG35 nuclease.
[0131] Sequence ID 1121 shows the full-length peptide sequence of Fam72A.
[0132] Sequence IDs 1122-1127 show the full-length peptide sequences of the MG35 nuclease.
[0133] Sequence IDs 1128-1160 show the full-length peptide sequences of the MG3-6 / 3-8 adenine base editor.
[0134] Sequence IDs 1161-1186 show the full-length peptide sequences of the MG34-1 adenine base editor.
[0135] Sequence IDs 1187-1195 represent sgRNA sequences suitable for the manipulated nucleic acid editing systems described herein.
[0136] Sequence IDs 1196-1204 show spacer sequences used to target genomic loci in the guide RNA.
[0137] Sequence ID 1205 shows the nucleotide sequence of the plasmid encoding the MG3-6 / 3-8 adenine base editor.
[0138] Sequence ID 1206 shows the nucleotide sequence of a plasmid encoding an sgRNA suitable for the MG3-6 / 3-8 adenine base editor described herein.
[0139] Sequence ID 1207 shows the nucleotide sequence of the plasmid encoding the MG34-1 adenine base editor.
[0140] Sequence IDs 1208-1269 represent the full-length peptide sequences of MG93 deaminase suitable for the engineered nucleic acid editing systems described herein.
[0141] Sequence IDs 1270-1296 represent the full-length peptide sequences of MG139 deaminase suitable for the engineered nucleic acid editing system described herein.
[0142] Sequence IDs 1297-1311 represent the full-length peptide sequences of MG152 deaminase suitable for the engineered nucleic acid editing system described herein.
[0143] Sequence IDs 1312-1313 represent the full-length peptide sequences of MG138 deaminase suitable for the engineered nucleic acid editing systems described herein.
[0144] Sequence IDs 1314-1315 represent the full-length peptide sequences of MG139 deaminase suitable for the engineered nucleic acid editing system described herein.
[0145] Sequence IDs 1316-1319 show the nucleotide sequences of 5'-FAM-labeled ssDNA.
[0146] Sequence IDs 1320-1321 show the nucleotide sequences of Cy5.5-labeled ssDNA.
[0147] Sequence numbers 1322-1355 show the sequences of the cytidine base editor.
[0148] Sequence IDs 1356-1362 show the full-length peptide sequences of the MG34-1 adenine base editor.
[0149] Sequence IDs 1363-1415 show the full-length peptide sequences of the MG3-6 / 3-8 adenine base editor.
[0150] Sequence IDs 1416-1417 show nucleotide sequences of sgRNAs suitable for use with the MG34-1 adenine base editor described herein.
[0151] Sequence ID 1418 shows the nucleotide sequence of an sgRNA suitable for use with the MG3-6 / 3-8 adenine base editor described herein.
[0152] Sequence IDs 1419-1420 indicate DNA sequences of target sites suitable for targeting with the MG34-1 adenine base editor described herein.
[0153] Sequence ID 1421 shows a DNA sequence of a target site suitable for targeting with the MG3-6 / 3-8 adenine base editor described herein.
[0154] Sequence ID 1422 shows the nucleotide sequence of a plasmid suitable for expression of the MG34-1 adenine base editor described herein.
[0155] Sequence ID 1423 shows the nucleotide sequence of a plasmid suitable for expression of the MG3-6 / 3-8 adenine base editor described herein.
[0156] Sequence ID 1424 shows the full-length peptide sequence of the MG35-1 adenine base editor.
[0157] Sequence IDs 1425-1426 show the nucleotide sequences of plasmids suitable for expression of the MG35-1 adenine base editor and sgRNA described herein.
[0158] Sequence IDs 1427-1428 show nucleotide sequences of sgRNAs suitable for use with the MG35-1 adenine base editor described herein.
[0159] Sequence IDs 1429-1430 indicate DNA sequences of target sites suitable for targeting with the MG35-1 adenine base editor described herein.
[0160] Sequence IDs 1431–1454 show the nucleotide sequences of sgRNAs engineered to function with the MG3-6 / 3-8 adenine base editor to target APOA1.
[0161] Sequence IDs 1455-1478 show the DNA sequences of the APOA1 target site.
[0162] Sequence IDs 1479–1483 show the nucleotide sequences of sgRNAs engineered to function with the MG3-6 / 3-8 adenine base editor to target ANGPTL3.
[0163] Sequence IDs 1484-1488 show the DNA sequences of the ANGPTL3 target site.
[0164] Sequence IDs 1489–1490 show the nucleotide sequences of sgRNAs engineered to function with the MG3-6 / 3-8 adenine base editor in order to target TRAC.
[0165] Sequence numbers 1491-1492 show the DNA sequences of the TRAC region.
[0166] Sequence IDs 1493–1516 show the nucleotide sequences of NGS primers suitable for use when evaluating APOA1 base editing.
[0167] Sequence IDs 1517–1521 show the nucleotide sequences of NGS primers suitable for use when evaluating base editing of ANGPTL3.
[0168] Sequence IDs 1522-1523 show the nucleotide sequences of NGS primers suitable for use when evaluating TRAC base editing.
[0169] Sequence IDs 1524–1547 show the nucleotide sequences of NGS primers suitable for use when evaluating APOA1 base editing.
[0170] Sequence IDs 1548-1552 show the nucleotide sequences of NGS primers suitable for use when evaluating base editing of ANGPTL3.
[0171] Sequence IDs 1553-1554 show the nucleotide sequences of NGS primers suitable for use when evaluating TRAC base editing.
[0172] Sequence ID 1555 shows the nucleotide sequence of a plasmid suitable for use in mRNA production.
[0173] Sequence IDs 1556-1562 show the full-length peptide sequences of the MG131 adenine deaminase variant.
[0174] Sequence IDs 1563-1566 show the full-length peptide sequences of the MG134 adenine deaminase variant.
[0175] Sequence IDs 1567-1574 show the full-length peptide sequences of the MG135 adenine deaminase variant.
[0176] Sequence IDs 1575-1589 show the full-length peptide sequences of the MG137 adenine deaminase variant.
[0177] Sequence IDs 1590-1599 show the full-length peptide sequences of the MG68 adenine deaminase variant.
[0178] Sequence IDs 1600-1602 show the full-length peptide sequences of the MG132 adenine deaminase variant.
[0179] Sequence IDs 1603-1616 show the full-length peptide sequences of the MG133 adenine deaminase variant.
[0180] Sequence IDs 1617-1624 show the full-length peptide sequences of the MG136 adenine deaminase variant.
[0181] Sequence IDs 1625-1633 show the full-length peptide sequences of the MG129 adenine deaminase variant.
[0182] Sequence IDs 1634-1638 show the full-length peptide sequences of the MG130 adenine deaminase variant.
[0183] Sequence IDs 1639-1644 show the full-length peptide sequences of the MG34-1 adenine base editor.
[0184] Sequence IDs 1698-1703 show the full-length peptide sequences of the MG34 adenine base editor.
[0185] Sequence IDs 1645-1646 show nucleotide sequences of ssDNA substrates suitable for testing adenine deaminase activity in vitro.
[0186] Sequence IDs 1647-1653 represent the linker sequences of the deaminase system described herein.
[0187] Sequence numbers 1654-1658 and 1665-1694 show the full-length peptide sequences of the MG3-6 / 3-8 chimeric base editor.
[0188] Sequence IDs 1659 and 1661-1664 represent the full-length peptide sequences of MG139 cytidine deaminase suitable for the engineered nucleic acid editing systems described herein.
[0189] Sequence ID 1660 shows the full-length peptide sequence of MG152 cytidine deaminase suitable for the engineered nucleic acid editing system described herein.
[0190] Sequence IDs 1695-1697 show the full-length peptide sequences of the MG102 adenine base editor.
[0191] Sequence ID 1704 shows the full-length nucleotide sequence of the hApoA1_1 guide RNA.
[0192] Sequence IDs 1705-1710 show the full-length nucleotide sequences of MG34 effector chemosynthesized / modified sgRNAs.
[0193] Sequence IDs 1711-1719 show the full-length nucleotide sequences of the TRAC target site.
[0194] Sequence IDs 1720-1725 show the full-length nucleotide sequences of the AAVS1 target site.
[0195] Sequence ID 1726 shows the full-length nucleotide sequence of the hApoA1 target site.
[0196] Sequence IDs 1727-1728 and 1743 show the nucleotide sequences of the MG3-6 / 3-8 adenine base editor.
[0197] Sequence IDs 1729 and 1744 show the nucleotide sequences of the MG3-6 / 3-8 / 3-4 adenine base editor.
[0198] Sequence IDs 1730 and 1745 show the nucleotide sequences of the MG3-6 / 3-8 / 3-6 adenine base editor.
[0199] Sequence IDs 1731 and 1746 show the nucleotide sequences of the MG3-6 / 3-8 / 3-7 adenine base editor.
[0200] Sequence IDs 1732 and 1747 show the nucleotide sequences of the MG3-6 / 3-8 / 3-22 adenine base editor.
[0201] Sequence IDs 1733 and 1748 show the nucleotide sequences of the MG3-6 / 3-8 / 3-24 adenine base editor.
[0202] Sequence IDs 1734 and 1749 show the nucleotide sequences of the MG3-6 / 3-8 / 3-38 adenine base editor.
[0203] Sequence IDs 1735 and 1750 show the nucleotide sequences of the MG3-6 / 3-8 / 3-89 adenine base editor.
[0204] Sequence IDs 1736 and 1751 show the nucleotide sequences of the MG3-6 / 3-8 / 3-90 adenine base editor.
[0205] Sequence IDs 1737 and 1752 show the nucleotide sequences of the MG3-6 / 3-8 / 3-92 adenine base editor.
[0206] Sequence IDs 1738 and 1753 show the nucleotide sequences of the MG3-6 / 3-8 / 3-93 adenine base editor.
[0207] Sequence IDs 1739 and 1754 show the nucleotide sequences of the MG3-6 / 3-8 / 3-95 adenine base editor.
[0208] Sequence IDs 1740 and 1755 show the nucleotide sequences of the MG3-6 / 3-8 / 3-104 adenine base editor.
[0209] Sequence IDs 1741 and 1756 show the nucleotide sequences of the MG3-6 / 3-8 / 150-2 adenine base editor.
[0210] Sequence IDs 1742 and 1757 show the nucleotide sequences of the MG3-6 / 3-8 / 150-9 adenine base editor.
[0211] Sequence IDs 1758-1889 show the nucleotide sequences of the MG3-6 ABE hANGPTL3 guide.
[0212] Sequence IDs 1890–1976 show the nucleotide sequences of the MG3-6 ABE hBCL11A guide.
[0213] Sequence IDs 1977–2009 show the nucleotide sequences of the MG3-6 ABE hPAH guide.
[0214] Sequence IDs 2010-2019 show the nucleotide sequences of the MG3-6 ABE APOA1 guide.
[0215] Sequence ID 2020 shows the nucleotide sequence of the manipulated therapeutic sequence.
[0216] Sequence ID 2021 shows the protein sequence of the MG3-6 / 3-8 / 3-4 adenine base editor.
[0217] Sequence ID 2022 shows the protein sequence of the MG3-6 / 3-8 / 3-7 adenine base editor.
[0218] Sequence ID 2023 shows the protein sequence of the MG3-6 / 3-8 / 3-104 adenine base editor.
[0219] Sequence IDs 2024-2043 show the amino acid sequences of the nuclear localization signal (NLA). [Modes for carrying out the invention]
[0220] While various embodiments of the present disclosure are shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided merely as examples. Numerous variations, alterations, and substitutions can be conceived by those skilled in the art without departing from the present disclosure. It should be understood that various alternatives to the embodiments of the present disclosure described herein may be used.
[0221] The practices of some of the methods disclosed herein employ techniques from immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA, unless otherwise indicated. See, for example, Sambrook and Green, Molecular Cloning: A Laboratory Manual, 4th Edition (2012); the series Current Protocols in Molecular Biology (FMAusubel, et al. eds.); the series Methods In Enzymology (Academic Press, Inc.), PCR 2: A Practical Approach (MJ MacPherson, BD Hames and GRTaylor eds. (1995)), Harlow and Lane, eds. (1988) Antibodies: A Laboratory Manual, and Culture of Animal Cells: A Manual of Basic Technique and Specialized Applications, 6th Edition (RIFreshney, ed. (2010)).
[0222] As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context otherwise explicitly indicates. Furthermore, to the extent that the terms "including," "containing," "having," "having," or their variants are used in either the detailed description or the claims, such terms are intended to be inclusive in a manner similar to that of the term "including."
[0223] The terms “about” or “approximately” mean within an acceptable range of error for a particular value, as determined by those skilled in the art, which depends in part on how the value is measured or determined, i.e., the limits of the measuring system. For example, “about” may mean within one or more standard deviations according to the practice of the art. Alternatively, “about” may mean a range of up to 20%, up to 15%, up to 10%, up to 5%, or up to 1% of a given value.
[0224] As used in this disclosure, the term “nucleotide” refers to a base-sugar-phosphate combination. Nucleotides intended to be nucleotides include naturally occurring and synthetic nucleotides. A nucleotide is the monomeric unit of a nucleic acid sequence (e.g., deoxyribonucleic acid (DNA) and ribonucleic acid (RNA)). The term nucleotide includes ribonucleoside triphosphates such as adenosine triphosphate (ATP), uridine triphosphate (UTP), cytosine triphosphate (CTP), guanosine triphosphate (GTP), and deoxyribonucleoside triphosphates, e.g., dATP, dCTP, dITP, dUTP, dGTP, dTTP, or their derivatives. Such derivatives include, for example, [αS]dATP, 7-deaza-dGTP, and 7-deaza-dATP, as well as nucleotide derivatives that confer nuclease resistance to nucleic acid molecules containing them. As used in this disclosure, the term nucleotide also includes dideoxyribonucleoside triphosphate (ddNTP) and its derivatives. Exemplary examples of ddNTPs include, but are not limited to, ddATP, ddCTP, ddGTP, ddITP, and ddTTP. Nucleotides may be unlabeled or may be labeled in a detectable manner, such as by using optically detectable moieties (e.g., fluorophores) or moieties containing quantum dots. Examples of detectable labels include radioisotopes, fluorescent labels, chemiluminescent labels, bioluminescent labels, and enzymatic labels. Examples of fluorescent labels for nucleotides include, but are not limited to, fluorescein, 5-carboxyfluorescein (FAM), 2'7'-dimethoxy-4'5-dichloro-6-carboxyfluorescein (JOE), rhodamine, 6-carboxyrhodamine (R6G), N,N,N',N'-tetramethyl-6-carboxyrhodamine (TAMRA), 6-carboxy-X-rhodamine (ROX), 4-(4'-dimethylaminophenylazo)benzoic acid (DABCYL), Cascade Blue, Oregon Green, Texas Red, cyanine, and 5-(2'-aminoethyl)aminonaphthalene-1-sulfonic acid (EDANS).Specific examples of fluorescently labeled nucleotides include [R6G]dUTP, [TAMRA]dUTP, [R110]dCTP, [R6G]dCTP, [TAMRA]dCTP, [JOE]ddATP, [R6G]ddATP, [FAM]ddCTP, [R110]ddCTP, [TAMRA]ddGTP, [ROX]ddTTP, [dR6G]ddATP, [dR110]ddCTP, [dTAMRA]ddGTP, and [dROX]ddTTP, available from Perkin Elmer (Foster City, Calif), FluoroLink DeoxyNucleotides, FluoroLink Cy3-dCTP, FluoroLink Cy5-dCTP, FluoroLink Fluor X-dCTP, FluoroLink Cy3-dUTP, and FluoroLink Cy5-dUTP, fluorescein-15-dATP, fluorescein-12-dUTP, tetramethylrhodamine-6-dUTP, IR770-9-dATP, fluorescein-12-ddUTP, fluorescein-12-UTP, and fluorescein-15-2'-dATP, available from Boehringer (Mannheim, Indianapolis, India), and Molecular Chromosome-labeled nucleotides available from Probes (Eugene, Oreg) include BODIPY-FL-14-UTP, BODIPY-FL-4-UTP, BODIPY-TMR-14-UTP, BODIPY-TMR-14-dUTP, BODIPY-TR-14-UTP, BODIPY-TR-14-dUTP, Cascade Blue-7-UTP, Cascade Blue-7-dUTP, Fluorescein-12-UTP, Fluorescein-12-dUTP, Oregon Green 488-5-dUTP, Rhodamine Green-5-UTP, Rhodamine Green-5-dUTP, Tetramethylrhodamine-6-UTP, Tetramethylrhodamine-6-dUTP, Texas Red-5-UTP, Texas Red-5-dUTP, and Texas Red-12-dUTP. The term nucleotide encompasses chemically modified nucleotides. An exemplary chemically modified nucleotide is biotin-dNTP.Non-limiting examples of biotinylated dNTPs include biotin-dATP (e.g., bio-N6-ddATP, biotin-14-dATP), biotin-dCTP (e.g., biotin-11-dCTP, biotin-14-dCTP), and biotin-dUTP (e.g., biotin-11-dUTP, biotin-16-dUTP, biotin-20-dUTP).
[0225] The terms “polynucleotide,” “oligonucleotide,” and “nucleic acid” are used interchangeably to refer to polymeric forms of nucleotides of any length, whether single-stranded, double-stranded, or multi-stranded, that are either deoxyribonucleotides or ribonucleotides, or analogues thereof. Polynucleotides as intended include genes or fragments thereof. Exemplary polynucleotides include, but are not limited to, DNA, RNA, coding or non-coding regions of genes or gene fragments, multiple loci (a single locus) defined by binding analysis, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short hairpin RNA (shRNA), micro-RNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, cell-free polynucleotides including cell-free DNA (cfDNA) and cell-free RNA (cfRNA), nucleic acid probes, and primers. When T is referred to in the context of polynucleotides, T means U (uracil) in RNA and T (thymine) in DNA. Polynucleotides can be exogenous or endogenous to cells and / or present in a cell-free environment. The term polynucleotide encompasses modified polynucleotides (e.g., modified backbones, sugars, or nucleic acid bases). Where present, modifications to the nucleotide structure are conferred before or after polymer assembly. Non-limiting examples of modifications include 5-bromouracil, peptide nucleic acids, heteronucleotides, morpholino, locked nucleic acids, glycol nucleic acids, threose nucleic acids, dideoxynucleotides, cordycepin, 7-deaza-GTP, fluorophores (e.g., rhodamine or fluorescein conjugated to sugars), thiol-containing nucleotides, biotin-conjugated nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, queosin, and waiosin. The sequence of nucleotides can be interrupted by non-nucleotide components.
[0226] The terms “peptide,” “polypeptide,” and “protein” are used interchangeably herein to refer to polymers of at least two amino acid residues joined by peptide bonds. These terms do not imply a specific length of the polymer and are not intended to imply or distinguish whether peptides are produced using recombinant techniques, chemical or enzymatic synthesis, or naturally occurring. The terms apply to naturally occurring amino acid polymers as well as amino acid polymers containing at least one modified amino acid. In some cases, the polymer is interrupted by non-amino acids. The terms include amino acid chains of any length, including full-length proteins and proteins (e.g., domains) with or without secondary or tertiary structures. The terms also encompass amino acid polymers modified by any other operations, such as disulfide bond formation, glycosylation, lipid formation, acetylation, phosphorylation, oxidation, and conjugation with labeling components. As used in this disclosure, the terms “amino acid” and “multiple amino acids” refer to natural and non-natural amino acids, including but not limited to modified amino acids. Modified amino acids include amino acids that have been chemically modified to include a group or chemical moiety that is not naturally present on the amino acid. The term "amino acid" includes both D-amino acids and L-amino acids.
[0227] As used in this disclosure, “non-natural” means a nucleic acid or polypeptide sequence that does not exist in nature. Non-natural means a nucleic acid or polypeptide sequence that does not exist in nature, including modifications such as mutations, insertions, or deletions. The term non-natural encompasses fusion nucleic acids or polypeptides in which the non-natural sequence encodes or exhibits the activity of the nucleic acid or polypeptide sequence to which it is fused (e.g., enzyme activity, methyltransferase activity, acetyltransferase activity, kinase activity, ubiquitination activity, etc.). Non-natural nucleic acids or polypeptide sequences include those that are genetically engineered to ligate a natural nucleic acid or polypeptide sequence (or a variant thereof) to produce a chimeric nucleic acid or polypeptide sequence encoding a chimeric nucleic acid or polypeptide.
[0228] As used in this disclosure, “operably linked,” “operably linked,” “operably linked,” or their grammatical equivalents refer to the arrangement of gene elements, such as promoters, enhancers, polyadenylation sequences, etc., where the action (e.g., movement or activation) of a first gene element has some effect on a second gene element. The effect on the second gene element may, but does not have to be, the same type as the action of the first gene element. For example, if the movement of the first element causes the activation of the second element, then the two gene elements are operably linked. For example, if a regulatory element, which may include a promoter sequence and / or an enhancer sequence, helps initiate transcription of a coding sequence, then the regulatory element is operably linked to the coding region. Intervening residues may exist between the regulatory element and the coding region, as long as this functional relationship is maintained.
[0229] A “functional fragment” of a DNA or protein sequence refers to a fragment that possesses biological activity (either functional or structural) substantially similar to that of the full-length DNA or protein sequence. The biological activity of a DNA sequence includes its ability to influence expression in a manner attributable to the full-length sequence.
[0230] The terms “engineered,” “synthetic,” and “artificial” are used interchangeably herein to refer to objects modified by human intervention. For example, these terms refer to polynucleotides or polypeptides that do not exist in nature. Engineered peptides have, but do not require, low sequence identity to naturally occurring human proteins (e.g., less than 50%, less than 25%, less than 10%, less than 5%, less than 1%). For example, the VPR domain and VP64 domain are synthetic transactivation domains. Non-limiting examples include: nucleic acids modified by altering their sequence to a sequence that does not occur in nature; nucleic acids modified by ligating them to nucleic acids that are not naturally associated so that the ligated product possesses a function not present in the original nucleic acid; engineered nucleic acids synthesized in vitro using sequences that do not exist in nature; proteins modified by altering their amino acid sequence to a sequence that does not exist in nature; engineered proteins that acquire new functions or properties. An “engineered” system includes at least one engineered component.
[0231] The terms "tracrRNA" or "tracr sequence" refer to the transactivation of CRISPR RNA. TracrRNA interacts with CRISPR(cr)RNA to form the guide (g)RNA for the type II and subtype VB CRISPR-Cas systems. When tracrRNA is manipulated, it may have approximately 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 100% sequence identity and / or similarity to a wild-type exemplary tracrRNA sequence (e.g., tracrRNA derived from S. pyogenes, S. aureus). TracrRNA may refer to modified forms of tracrRNA, which may include nucleotide changes such as deletions, insertions, or substitutions, variants, mutations, or chimeric forms. The term tracrRNA encompasses nucleic acids that may be at least approximately 60% identical to a wild-type exemplary tracrRNA sequence (e.g., tracrRNA from S. pyogenes, S. aureus, etc.) over a stretch of at least six consecutive nucleotides. For example, a tracrRNA sequence may have at least approximately 60% identity, at least approximately 65% identity, at least approximately 70% identity, at least approximately 75% identity, at least approximately 80% identity, at least approximately 85% identity, at least approximately 90% identity, at least approximately 95% identity, at least approximately 98% identity, at least approximately 99% identity, or 100% identity over a stretch of at least six consecutive nucleotides to a wild-type exemplary tracrRNA sequence (e.g., tracrRNA from S. pyogenes, S. aureus, etc.). Type II tracrRNA sequences can be predicted on a genomic sequence by identifying regions that are complementary to a portion of the repetitive sequences in an adjacent CRISPR array.
[0232] As used herein, “guide nucleic acid” or “guide polynucleotide” refers to a nucleic acid that hybridizes to a target nucleic acid, thereby directing the associated nuclease to the target nucleic acid. Guide nucleic acids are, but are not limited to, RNA (guide RNA or gRNA), DNA, or a mixture of RNA and DNA. Guide nucleic acids may include crRNA or tracrRNA, or a combination of both. The term guide nucleic acid encompasses engineered guide nucleic acids and programmable guide nucleic acids that specifically bind to the target nucleic acid. A portion of the target nucleic acid may be complementary to a portion of the guide nucleic acid. A double-stranded target polynucleotide chain that is complementary to the guide nucleic acid and hybridizes with it is called the complementary chain. A double-stranded target polynucleotide chain that is complementary to the complementary chain and therefore not complementary to the guide nucleic acid is called the non-complementary chain. A guide nucleic acid having a polynucleotide chain is called a “single guide nucleic acid”. A guide nucleic acid having two polynucleotide chains is called a “double guide nucleic acid”. Unless otherwise specified, the term “guide nucleic acid” is inclusive and refers to both single guide nucleic acids and double guide nucleic acids. Guide nucleic acids may include segments referred to as “nucleic acid targeting segments,” “nucleic acid targeting sequences,” or “spacers.” Nucleic acid targeting segments may include subsegments referred to as “protein-binding segments,” “protein-binding sequences,” or “Cas protein-binding segments.”
[0233] In the context of two or more nucleic acid or polypeptide sequences, the terms “sequence identity” or “identity rate” generally refer to two (e.g., in paired alignments) or more (e.g., in alignments of multiple sequences) sequences that, when compared and aligned to achieve maximum correspondence across local or global comparison windows, as measured using sequence comparison algorithms, have identical amino acid residues or nucleotides, or identical to a certain percentage of each other. Suitable sequence comparison algorithms for polypeptide sequences include, for example, BLASTP, which uses a BLOSUM62 scoring matrix with parameters of word length (W) of 3, expected value (E) of 10, and gap costs of 11 presence and 1 extension, and uses conditional composition score matrix adjustment for polypeptide sequences longer than 30 residues; BLASTP, which uses a PAM30 scoring matrix with parameters of word length (W) of 2, expected value (E) of 1,000,000, and 30 residues (these are the default parameters for BLASTP in BLAST available at https: / / blast.ncbi.nlm.nih.gov), and sets gap costs of 9 for gap start and 1 for gap extension; CLUSTALW with parameters; the Smith-Waterman homology search algorithm with parameters of 2 match, -1 mismatch, and -1 gap; MUSCLE with default parameters; MAFFT with parameter retrieval of 2 and a maximum of 1,000 repeats; Novafold with default parameters; and HMMER hmmalign with default parameters.
[0234] As used herein, the term “RuvC_III domain” refers to the third discontinuous segment of the RuvC endonuclease domain (the RuvC nuclease domain consists of three discontinuous segments: RuvC_I, RuvC_II, and RuvC_III). The RuvC domain or its segments can generally be identified by alignment to a documented domain sequence, structural alignment to a protein with an annotated domain, or comparison with a hidden Markov model (HMM) constructed based on a documented domain sequence (e.g., Pfam HMM PF18541 of RuvC_III).
[0235] As used in this disclosure, the term “HNH domain” refers to an endonuclease domain having characteristic histidine and asparagine residues. HNH domains can generally be identified by alignment to a documented domain sequence, structural alignment to a protein having an annotated domain, or comparison to a hidden Markov model (HMM) constructed based on a documented domain sequence (e.g., Pfam HMM PF01844 of the HNH domain).
[0236] As used herein, the term “base editor” refers to an enzyme that catalyzes the conversion of one target base or base pair to another base or base pair (e.g., A:T to G:C, C:G to T:A) without requiring the generation and repair of double-strand breaks. An exemplary base editor is a deaminase. In some embodiments, the base editor comprises a deaminase and a nuclease lacking nuclease activity. In some embodiments, the base editor comprises a deaminase and a catalytically inactive nuclease. In some embodiments, the base editor comprises a fusion of a deaminase and a catalytically inactive nuclease.
[0237] As used herein, the term “deaminase” refers to a protein or enzyme that catalyzes a deamination reaction (e.g., a reaction that removes an amino group). Examples of deaminases include adenosine deaminases (e.g., engineered adenosine deaminases that deaminate adenosine in DNA) that catalyze the hydrolytic deamination of adenine or adenosine, and cytidine (or cytosine) deaminases that catalyze the hydrolytic deamination of cytidine (or cytosine) or deoxycytidine to convert them to uridine (or uracil) or deoxyuridine, respectively. A deaminase or deaminase domain may be a naturally occurring deaminase or deaminase domain from a biological source such as a human, chimpanzee, gorilla, monkey, cattle, dog, rat, mouse, or bacterium (e.g., Escherichia coli), a naturally occurring deaminase or deaminase domain, or a naturally occurring deaminase or deaminase domain.
[0238] In the context of two or more nucleic acid sequences or polypeptide sequences, the term "optimally aligned" generally refers to two (e.g., in a paired alignment) or more (e.g., in a multi-sequence alignment) sequences aligned to the maximum correspondence of amino acid residues or nucleotides, as determined, for example, by the alignment that produces the maximum or "optimized" identity score.
[0239] As used in this disclosure, the term “complex” refers to the joining of at least two components. Each of the two components may retain properties / activities it had before forming the complex, or may acquire properties as a result of forming the complex. Joining includes, but is not limited to, covalent bonds, non-covalent bonds (i.e., hydrogen bonds, ionic interactions, van der Waals interactions, and hydrophobic bonds), the use of linkers, fusion, or any other preferred method. The intended components of a complex include polynucleotides, polypeptides, or combinations thereof. For example, a complex may include an endonuclease and a guide polynucleotide.
[0240] Any variant of the enzymes described herein having one or more conserved amino acid substitutions is included in this disclosure. Such conserved substitutions can be made in the amino acid sequence of a polypeptide without disrupting the three-dimensional structure or function of the polypeptide. Conservative substitutions can be achieved by substituting amino acids having similar hydrophobicity, polarity, and R-chain length with each other. In addition, or alternatively, by comparing the aligned sequences of homologous proteins from different species, conserved substitutions can be identified by finding interspecies mutated amino acid residues (e.g., non-conserved residues) without altering the fundamental function of the encoded protein. Such conservatively substituted variants may include variants having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, and at least about 99% identity to any one of the endonuclease protein sequences described herein. In some embodiments, such conservatively substituted variants are functional variants. Such functional variants may include sequences with substitutions such that the activity of one or more important active site residues or guide RNA binding residues of the endonuclease is not disrupted.
[0241] Also, the present disclosure includes any variant of the enzymes described herein (e.g., an activity-reduced variant) having a substitution of one or more catalytic residues to reduce or eliminate the activity of the enzyme. In some embodiments, the activity-reduced variant as a protein described in the present disclosure includes disruptive substitutions of at least one, at least two, or all three catalytic residues. In some embodiments, any of the endonucleases described herein may include a nickase mutation. In some embodiments, any of the endonucleases described herein may include a RuvC domain lacking nuclease activity. In some embodiments, any of the endonucleases described herein may be configured to cleave one strand of double-stranded target deoxyribonucleic acid. In some embodiments, any of the endonucleases described herein may be configured to lack endonuclease activity or be catalytically ineffective.
[0242] Tables of conservative substitutions providing functionally similar amino acids are available from a variety of references (see, e.g., Creighton, Proteins: Structures and Molecular Properties (W H Freeman & Co.; 2nd edition (December 1993)). The following eight groups each contain amino acids that are conservative substitutions for one another. 1) Alanine (A), Glycine (G); 2) Aspartic acid (D), Glutamic acid (E); 3) Asparagine (N), Glutamine (Q); 4) Arginine (R), Lysine (K); 5) Isoleucine (I), Leucine (L), Methionine (M), Valine (V); 6) Phenylalanine (F), Tyrosine (Y), Tryptophan (W); 7) Serine (S), Threonine (T); and 8) Cysteine (C), Methionine (M).
[0243] overview The discovery of new CRISPR enzymes with inherent functionality and structure could further revolutionize deoxyribonucleic acid (DNA) editing technology, potentially improving speed, specificity, functionality, and ease of use. Compared to the predicted prevalence of clustered regularly interspaced short palindromic repeat (CRISPR) systems in microorganisms and the full diversity of microbial species, relatively few CRISPR enzymes have been functionally characterized in the literature. This is partly due to the fact that, under laboratory conditions, a vast number of microbial species cannot be easily cultured. Metagenomic sequencing from natural environmental niches representing numerous microbial species has the potential to dramatically increase the number of documented new CRISPR systems and hasten the discovery of new oligonucleotide editing functions. A recent example demonstrating the benefit of such an approach is the 2016 discovery of the CasX / CasY CRISPR system from metagenomic analysis of natural microbial communities.
[0244] The CRISPR system is an RNA-directed nuclease complex described as functioning as an adaptive immune system in microorganisms. In their natural context, the CRISPR system arises in a CRISPR (clustered, regularly spaced, short palindromic repeat) operon or locus, which generally consists of two parts: (i) an array of short repeat sequences (30-40 bp) separated by equally short spacer sequences encoding an RNA-based targeting element; and (ii) an ORF encoding a nuclease polypeptide directed by the RNA-based targeting element, alongside accessory proteins / enzymes. Efficient nuclease targeting of a particular target nucleic acid sequence generally requires both (i) complementary hybridization between the first 6-8 nucleic acids of the target (target seed) and a crRNA guide; and (ii) the presence of a protospacer adjacency motif (PAM) sequence within a defined neighborhood of the target seed (PAMs are typically sequences not commonly expressed in the host genome). Depending on the precise function and organization of the system, CRISPR systems are generally classified into two classes, five types, and 16 subtypes based on common functional characteristics and evolutionary similarities (see Figure 1).
[0245] Class 1 CRISPR systems have large, multi-subunit effector complexes and include types I, III, and IV.
[0246] The type I CRISPR system is considered to be of moderate complexity in terms of its components. In the type I CRISPR system, an array of RNA targeting elements is transcribed as a long precursor crRNA (precrRNA), which is processed with repeat elements to release a short mature crRNA that orients the nuclease complex to the nucleic acid target, when followed by a preferred short consensus sequence called a protospacer-adjacent motif (PAM). This processing occurs via the endoribonuclease subunit (Cas6) of a larger endonuclease complex called a cascade, which also includes the nuclease (Cas3) protein component of the crRNA-directed nuclease complex. Type I nucleases primarily function as DNA nucleases.
[0247] The type III CRISPR system can be characterized by the presence of a central nuclease known as Cas10, alongside a repeat-associated mysterious protein (RAMP) containing a Csm or Cmr protein subunit. Similar to the type I system, mature crRNA is processed from precrRNA using a Cas6-like enzyme. Unlike the type I and II systems, the type III system is thought to target and cleave DNA-RNA double strands (such as the DNA strand used as a template for RNA polymerase).
[0248] The type IV CRISPR system has an effector complex containing a highly reduced large subunit nuclease (csf1), two genes for RAMP proteins from the Cas5 (csf3) and Cas7 (csf2) groups, and, in some cases, a gene for a predicted smaller subunit. Such systems are commonly found in endogenous plasmids.
[0249] Class 2 CRISPR systems generally have a single polypeptide multi-domain nuclease effector and include types II, V, and VI.
[0250] Type II CRISPR systems are considered the simplest in terms of their components. In type II CRISPR systems, processing the CRISPR array into mature crRNA does not require the presence of a special endonuclease subunit, but rather a small transencoded crRNA (tracrRNA) having a region complementary to the array repeat sequence. The tracrRNA interacts with both its corresponding effector nuclease (e.g., Cas9) and the repeat sequence to form a precursor dsRNA structure, which is cleaved with endogenous RNAse III to produce a mature effector enzyme loaded with both tracrRNA and crRNA. Type II nucleases are known as DNA nucleases. Type II effectors generally exhibit a structure containing a RuvC-like endonuclease domain that fits into an RNase H fold, which has an unrelated HNH nuclease domain inserted into the fold of a RuvC-like nuclease domain. The RuvC-like domain is involved in cleaving target (e.g., crRNA-complementary) DNA strands, while the HNH domain is involved in cleaving displaced DNA strands.
[0251] Type V CRISPR systems are characterized by a nuclease effector structure (e.g., Cas12) similar to that of type II effectors, containing a RuvC-like domain. Like type II, most (but not all) type V CRISPR systems use tracrRNA to process precrRNA into mature crRNA; however, unlike type II systems which require RNAse III to cleave precrRNA into multiple crRNAs, type V systems can cleave precrRNA using the effector nuclease itself. Similar to type II CRISPR systems, type V CRISPR systems are also known as DNA nucleases. Unlike type II CRISPR systems, some type V enzymes (e.g., Cas12a) appear to possess robust single-strand nonspecific deoxyribonuclease activity, activated by the first crRNA-directed cleavage of a double-stranded target sequence.
[0252] Type VI CRIPSR systems possess RNA-inducible RNA endonucleases. Instead of a RuvC-like domain, a single polypeptide effector of a type VI system (e.g., Cas13) contains two HEPN ribonuclease domains. Unlike both type II and V systems, type VI systems may also not require tracrRNA to process precrRNA into crRNA. However, similar to type V systems, some type VI systems (e.g., C2C2) appear to possess robust single-strand nonspecific nuclease (ribonuclease) activity, which is activated by cleavage of the target RNA by the initial crRNA.
[0253] Due to its simpler structure, Class 2 CRISPR has been the most widely adopted for manipulation and development as a designer nuclease / genome editing application.
[0254] One of the initial adaptations of such a system for in vitro use is (i) recombinantly expressed, purified full-length Cas9 (e.g., class 2, type II Cas enzyme) isolated from S. pyogenes SF370, (ii) purified mature crRNA of approximately 42 nt carrying a 5' sequence of approximately 20 nt complementary to the target DNA sequence to be cleaved, followed by a 3' tracr binding sequence (total crRNA transcribed in vitro from a synthetic DNA template carrying a T7 promoter sequence), (iii) purified tracrRNA transcribed in vitro from a synthetic DNA template carrying a T7 promoter sequence, and (iv) Mg 2+ Subsequent improved, manipulated systems involved a crRNA of (ii) linked to the 5' end of (iii) by a linker (e.g., GAAA) to form a single fusion synthetic guide RNA (sgRNA) capable of inducing Cas9 to the target by itself.
[0255] Such manipulated systems can be adapted for use in mammalian cells by providing DNA vectors encoding (i) an ORF encoding codon-optimized Cas9 (e.g., class 2, type II Cas enzyme) under a suitable mammalian promoter having a C-terminal nuclear localization sequence (e.g., SV40 NLS) and a suitable polyadenylation signal (e.g., TK pA signal), and (ii) an ORF encoding sgRNA (having a 5′ sequence beginning with G, followed by a 20nt complementary targeting nucleic acid sequence, linker, and tracrRNA sequence bound to a 3′ tracr binding sequence) under a suitable polymerase III promoter (e.g., U6 promoter).
[0256] Base editing Base editing is the conversion of one target base or base pair to another target base or base pair (e.g., A:T to G:C, C:G to T:A) without requiring the generation and repair of double-strand breaks. Base editing can be achieved using DNA and RNA base editors that allow the introduction of point mutations at specific sites in either DNA or RNA. Generally, DNA base editors may involve the fusion of a catalytically inactive nuclease with a catalytically active base-modifying enzyme that acts on single-stranded DNA (ssDNA). RNA base editors may consist of similar RNA-specific enzymes. Base editing can increase the efficiency of gene modification while reducing off-target and random mutations in DNA.
[0257] DNA base editors are engineered ribonucleoprotein complexes that act as tools for single base substitutions in cells and organisms. They can be constructed by fusing an engineered base-modifying enzyme with a catalytically deficient CRISPR endonuclease variant that cannot cleave dsDNA, but can unfold dsDNA in a protospacer-adjacent motif (PAM) sequence-dependent manner, allowing the guide RNA to find its complementary target and indicate the ssDNA cleavage site. The guide RNA anneals to the complementary DNA, moves the ssDNA fragment, and directs the CRISPR "scissors" to the base-modification site. Cellular repair mechanisms use the information from the complementary edited template to repair the nicked, unedited strand.
[0258] To date, two types of DNA editors have been developed: cytosine base editors (CBEs) and adenine base editors (ABEs). However, recent findings indicate that off-target modifications exist in DNA, and that many of these off-target modifications can also be introduced into RNA by DNA base editors.
[0259] MG Base Editor In this specification, in some embodiments, an engineered system is described comprising an engineered guide polynucleotide comprising (a) a base editor, (b) an endonuclease lacking nuclease activity configured to bind to the base editor, and (c) a spacer sequence configured to form a complex with the endonuclease and to hybridize with a target nucleic acid sequence. In this specification, in some embodiments, an engineered base editing system is described comprising a base editor having at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% identity with any one of sequence numbers 1654-1703 and 2021-2023, wherein the sequence does not include any sequence selected from sequence numbers 1128-1160 and 1363-1415, and an engineered guide polynucleotide comprising a spacer sequence that forms a complex with the endonuclease of the base editor and hybridizes with a target nucleic acid sequence. In this specification, in one embodiment, an engineered base editing system is described, comprising a base editor having at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% identity with any one of sequence numbers 1654-1703 and 2021-2023, wherein the sequence does not include any sequence selected from sequence numbers 1128-1160 and 1363-1415, and an engineered guide polynucleotide comprising a spacer sequence that forms a complex with an endonuclease lacking the nuclease activity of the base editor and hybridizes with a target nucleic acid sequence.
[0260] In this specification, in one embodiment, an engineered base editing system is described, comprising an engineered base editing system comprising a base editor having a sequence having at least 85%, 90%, 95%, 98%, 99%, or 100% identity with any one of sequence numbers 1654-1703 and 2021-2023, and an engineered guide polynucleotide comprising a spacer sequence that forms a complex with the endonuclease of the base editor and hybridizes with a target nucleic acid sequence. In this specification, in one embodiment, an engineered base editing system is described, comprising an engineered base editing system comprising a base editor having a sequence having at least 85%, 90%, 95%, 98%, 99%, or 100% identity with any one of sequence numbers 1654-1703 and 2021-2023, and an engineered guide polynucleotide comprising a spacer sequence that forms a complex with the endonuclease lacking the nuclease activity of the base editor and hybridizes with a target nucleic acid sequence.
[0261] In some embodiments, the base editor includes sequences having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 99%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of the sequences 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703. In some embodiments, the base editor includes sequences that have at least about 70% identity with any one of the following sequences: SEQ ID NOs: 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703. In some embodiments, the base editor includes sequences that have at least about 75% identity with any one of the following sequences: SEQ ID NOs: 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703. In some embodiments, the base editor includes sequences that have at least about 80% identity with any one of the following sequences: SEQ ID NOs: 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703.In some embodiments, the base editor includes sequences that have at least about 85% identity with any one of the following sequences: SEQ ID NOs: 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703. In some embodiments, the base editor includes sequences that have at least about 90% identity with any one of the following sequences: SEQ ID NOs: 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703. In some embodiments, the base editor includes sequences that have at least about 95% identity with any one of the following sequences: SEQ ID NOs: 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703. In some embodiments, the base editor includes sequences that have at least about 96% identity with any one of the following sequences: SEQ ID NOs: 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703. In some embodiments, the base editor includes sequences that have at least about 97% identity with any one of the following sequences: SEQ ID NOs: 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703.In some embodiments, the base editor includes sequences that have at least about 98% identity with any one of the following sequences: SEQ ID NOs: 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703. In some embodiments, the base editor includes sequences that have at least about 99% identity with any one of the following sequences: SEQ ID NOs: 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703. In some embodiments, the base editor includes sequences that have 100% identity with any one of the following sequence numbers: 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703.
[0262] In some embodiments, the base editor includes a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of sequence numbers 1654-1703 and 2021-2023. In some embodiments, the base editor includes a sequence having at least about 70% identity with any one of sequence numbers 1654-1703 and 2021-2023. In some embodiments, the base editor includes a sequence having at least about 75% identity with any one of sequence numbers 1654-1703 and 2021-2023. In some embodiments, the base editor includes a sequence having at least about 80% identity with one of sequence numbers 1654-1703 and 2021-2023. In some embodiments, the base editor includes a sequence having at least about 85% identity with one of sequence numbers 1654-1703 and 2021-2023. In some embodiments, the base editor includes a sequence having at least about 86% identity with one of sequence numbers 1654-1703 and 2021-2023. In some embodiments, the base editor includes a sequence having at least about 87% identity with one of sequence numbers 1654-1703 and 2021-2023. In some embodiments, the base editor includes a sequence having at least about 88% identity with one of sequence numbers 1654-1703 and 2021-2023. In some embodiments, the base editor includes a sequence having at least about 89% identity with one of sequence numbers 1654-1703 and 2021-2023. In some embodiments, the base editor includes a sequence having at least about 90% identity with one of sequence numbers 1654-1703 and 2021-2023.In some embodiments, the base editor comprises a sequence having at least about 95% identity with any one of SEQ ID NOs: 1654-1703 and 2021-2023. In some embodiments, the base editor comprises a sequence having at least about 96% identity with any one of SEQ ID NOs: 1654-1703 and 2021-2023. In some embodiments, the base editor comprises a sequence having at least about 97% identity with any one of SEQ ID NOs: 1654-1703 and 2021-2023. In some embodiments, the base editor comprises a sequence having at least about 98% identity with any one of SEQ ID NOs: 1654-1703 and 2021-2023. In some embodiments, the base editor comprises a sequence having at least about 99% identity with any one of SEQ ID NOs: 1654-1703 and 2021-2023. In some embodiments, the base editor comprises a sequence having 100% identity with any one of SEQ ID NOs: 1654-1703 and 2021-2023.
[0263] As used herein, in certain embodiments, an engineered base editing system is described that comprises a base editor encoded by a nucleic acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% identity with any one of SEQ ID NOs: 1727-1757, an engineered guide polynucleotide that forms a complex with the endonuclease of the base editor and comprises a spacer sequence that hybridizes to a target nucleic acid sequence. As used herein, in certain embodiments, an engineered base editing system is described that comprises a base editor encoded by a nucleic acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% identity with any one of SEQ ID NOs: 1727-1757, an engineered guide polynucleotide that forms a complex with the endonuclease of the base editor lacking nuclease activity and comprises a spacer sequence that hybridizes to a target nucleic acid sequence.
[0264] In some embodiments, the base editor is encoded by a nucleic acid having a sequence having at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of sequence numbers 1727-1757. In some embodiments, the base editor is encoded by a nucleic acid sequence having at least about 70% identity with any one of sequence numbers 1727-1757. In some embodiments, the base editor includes a sequence having at least about 75% identity with any one of sequence numbers 1727-1757. In some embodiments, the bases are encoded by nucleic acids having a sequence that is at least about 80% identical to one of sequence numbers 1727-1757. In some embodiments, the base editor is encoded by nucleic acids having a sequence that is at least about 85% identical to one of sequence numbers 1727-1757. In some embodiments, the base editor is encoded by nucleic acids having a sequence that is at least about 90% identical to one of sequence numbers 1727-1757. In some embodiments, the base editor is encoded by nucleic acids having a sequence that is at least about 95% identical to one of sequence numbers 1727-1757. In some embodiments, the base editor is encoded by nucleic acids having a sequence that is at least about 96% identical to one of sequence numbers 1727-1757. In some embodiments, the base editor is encoded by nucleic acids having a sequence that is at least about 97% identical to one of sequence numbers 1727-1757. In some embodiments, the base editor is encoded by a nucleic acid having a sequence that is at least about 98% identical to one of sequence numbers 1727-1757.In some embodiments, the base editor is encoded by a nucleic acid having a sequence that is at least about 99% identical to one of sequence numbers 1727-1757. In some embodiments, the base editor is encoded by a nucleic acid having a sequence that is 100% identical to one of sequence numbers 1727-1757.
[0265] In some embodiments, the base editor includes a deaminase. In some embodiments, the deaminase is non-covalently bonded to the endonuclease. In some embodiments, the deaminase is covalently bonded to the endonuclease. In some embodiments, the deaminase is fused to the endonuclease.
[0266] In some embodiments, the base editor is an adenine deaminase. In some embodiments, the adenosine deaminase includes sequences having sequence identity of at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% with sequence numbers 50-51, 57, 385-443, 448-475, 595, 1356-1415, 1424, 1556-1644, 1654-1658, and 1665-1703. In some embodiments, the adenosine deaminase contains a sequence having at least about 70% identity with any one of sequence numbers 50-51, 57, 385-443, 448-475, 595, 1356-1415, 1424, 1556-1644, 1654-1658, and 1665-1703. In some embodiments, the adenosine deaminase contains a sequence having at least about 75% identity with any one of sequence numbers 50-51, 57, 385-443, 448-475, 595, 1356-1415, 1424, 1556-1644, 1654-1658, and 1665-1703. In some embodiments, the adenosine deaminase contains a sequence having at least approximately 80% identity with any one of sequence numbers 50-51, 57, 385-443, 448-475, 595, 1356-1415, 1424, 1556-1644, 1654-1658, and 1665-1703. In some embodiments, the adenosine deaminase contains a sequence having at least approximately 85% identity with any one of sequence numbers 50-51, 57, 385-443, 448-475, 595, 1356-1415, 1424, 1556-1644, 1654-1658, and 1665-1703.In some embodiments, the adenosine deaminase contains a sequence having at least about 90% identity with any one of sequence numbers 50-51, 57, 385-443, 448-475, 595, 1356-1415, 1424, 1556-1644, 1654-1658, and 1665-1703. In some embodiments, the adenosine deaminase contains a sequence having at least about 95% identity with any one of sequence numbers 50-51, 57, 385-443, 448-475, 595, 1356-1415, 1424, 1556-1644, 1654-1658, and 1665-1703. In some embodiments, the adenosine deaminase contains a sequence having at least approximately 96% identity with one of sequence numbers 50-51, 57, 385-443, 448-475, 595, 1356-1415, 1424, 1556-1644, 1654-1658, and 1665-1703. In some embodiments, the adenosine deaminase contains a sequence having at least approximately 97% identity with one of sequence numbers 50-51, 57, 385-443, 448-475, 595, 1356-1415, 1424, 1556-1644, 1654-1658, and 1665-1703. In some embodiments, the adenosine deaminase contains a sequence having at least about 98% identity with any one of sequence numbers 50-51, 57, 385-443, 448-475, 595, 1356-1415, 1424, 1556-1644, 1654-1658, and 1665-1703. In some embodiments, the adenosine deaminase contains a sequence having at least about 99% identity with any one of sequence numbers 50-51, 57, 385-443, 448-475, 595, 1356-1415, 1424, 1556-1644, 1654-1658, and 1665-1703. In some embodiments, the adenosine deaminase contains a sequence that is 100% identical to any one of sequence numbers 50-51, 57, 385-443, 448-475, 595, 1356-1415, 1424, 1556-1644, 1654-1658, and 1665-1703.
[0267] In some embodiments, the base editor is a cytosine deaminase. In some embodiments, the cytosine deaminase contains a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NOs: 1-49, 58-66, 444-447, 594, 744-835, 970-982, and 1659-1664. In some embodiments, the cytosine deaminase contains a sequence having at least about 70% identity with one of sequence numbers 1-49, 58-66, 444-447, 594, 744-835, 970-982, and 1659-1664. In some embodiments, the cytosine deaminase contains a sequence having at least about 75% identity with one of sequence numbers 1-49, 58-66, 444-447, 594, 744-835, 970-982, and 1659-1664. In some embodiments, the cytosine deaminase contains a sequence having at least about 80% identity with one of sequence numbers 1-49, 58-66, 444-447, 594, 744-835, 970-982, and 1659-1664. In some embodiments, the cytosine deaminase contains a sequence having at least about 85% identity with one of sequence numbers 1-49, 58-66, 444-447, 594, 744-835, 970-982, and 1659-1664. In some embodiments, the cytosine deaminase contains a sequence having at least about 90% identity with one of sequence numbers 1-49, 58-66, 444-447, 594, 744-835, 970-982, and 1659-1664.In some embodiments, the cytosine deaminase contains a sequence having at least about 95% identity with one of sequence numbers 1-49, 58-66, 444-447, 594, 744-835, 970-982, and 1659-1664. In some embodiments, the cytosine deaminase contains a sequence having at least about 96% identity with one of sequence numbers 1-49, 58-66, 444-447, 594, 744-835, 970-982, and 1659-1664. In some embodiments, the cytosine deaminase contains a sequence having at least about 97% identity with one of sequence numbers 1-49, 58-66, 444-447, 594, 744-835, 970-982, and 1659-1664. In some embodiments, the cytosine deaminase contains a sequence having at least approximately 98% identity with one of sequence numbers 1-49, 58-66, 444-447, 594, 744-835, 970-982, and 1659-1664. In some embodiments, the cytosine deaminase contains a sequence having at least approximately 99% identity with one of sequence numbers 1-49, 58-66, 444-447, 594, 744-835, 970-982, and 1659-1664. In some embodiments, the cytosine deaminase contains a sequence having 100% identity with one of sequence numbers 1-49, 58-66, 444-447, 594, 744-835, 970-982, and 1659-1664.
[0268] In some embodiments, the base editor includes one or more modifications. In some embodiments, when optimally aligned, the base editor includes at least one substitution of residues T2, D7, E10, M13, W24, G32, K38, G45, G51, A63, E66, R75, C91, G93, H97, A107, E108, D109, P110, H124, A126, H129, F150, or S165 for sequence number 50, or any combination thereof. In some embodiments, when optimally aligned, the substitution includes W24G, G51V, E108D, P110H, F150P, D7G, E10G, or H129N for sequence number 50, or any combination thereof. In some embodiments, the substitutions, when optimally aligned, are T2X1, D7X1, E10X1, M13X4, W24X1, G32X1, K38X2, G45X2, G51X5, A63X7, E66X5, E66X2, R75H, C91R, G93X6, H97X6, H97X5, A107X5, E108X2, D109N, P110H, H The set includes 124X6, A126X2, H129R, H129N, F150P, F150S, S165X5, or any combination thereof, where X1 is A or G; X2 is D or E; X3 is N or Q; X4 is R or K; X5 is I, L, M, or V; X6 is F, Y, or W; and X7 is S or T.
[0269] Endonuclease In this specification, in some embodiments, an endonucleases lacking nuclease activity are described. In some embodiments, the endonucleases include a RuvC domain and an HNH domain. In some embodiments, the RuvC domain lacks nuclease activity. In some embodiments, the endonucleases include a nickas mutation. In some embodiments, the endonucleases are derived from uncultured microorganisms. In some embodiments, the endonucleases are class 2, type II endonucleases. In some embodiments, the endonucleases are configured to cleave one strand of a target nucleic acid (e.g., DNA).
[0270] In some embodiments, the endonuclease is not Cas9 endonuclease, Cas14 endonuclease, Cas12a endonuclease, Cas12b endonuclease, Cas12c endonuclease, Cas12d endonuclease, Cas12e endonuclease, Cas13a endonuclease, Cas13b endonuclease, Cas13c endonuclease, or Cas13d endonuclease. In some embodiments, the endonuclease has less than 80% identity with Cas9 endonuclease.
[0271] In some embodiments, the endonuclease contains a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NOs. In some embodiments, the endonuclease includes a sequence having at least about 75% identity with one of sequence numbers 70-78, 596, 597, 1120, and 1122-1127. In some embodiments, the endonuclease includes a sequence having at least about 80% identity with one of sequence numbers 70-78, 596, 597, 1120, and 1122-1127. In some embodiments, the endonuclease includes a sequence having at least about 85% identity with one of sequence numbers 70-78, 596, 597, 1120, and 1122-1127. In some embodiments, the endonuclease includes a sequence having at least about 90% identity with one of sequence numbers 70-78, 596, 597, 1120, and 1122-1127. In some embodiments, the endonuclease includes a sequence having at least about 95% identity with one of sequence numbers 70-78, 596, 597, 1120, and 1122-1127. In some embodiments, the endonuclease includes a sequence having at least about 96% identity with one of sequence numbers 70-78, 596, 597, 1120, and 1122-1127.In some embodiments, the endonuclease includes a sequence having at least about 97% identity with one of sequence numbers 70-78, 596, 597, 1120, and 1122-1127. In some embodiments, the endonuclease includes a sequence having at least about 98% identity with one of sequence numbers 70-78, 596, 597, 1120, and 1122-1127. In some embodiments, the endonuclease includes a sequence having at least about 99% identity with one of sequence numbers 70-78, 596, 597, 1120, and 1122-1127. In some embodiments, the endonuclease includes a sequence having 100% identity with one of sequence numbers 70-78, 596, 597, 1120, and 1122-1127.
[0272] In some embodiments, the endonuclease is configured to bind to a protospacer adjacent motif (PAM) sequence containing one of sequence numbers 360-368 and 598.
[0273] Guide polynucleotides In some embodiments, the manipulated systems disclosed herein include manipulated guide polynucleotides, such as guide ribonucleic acid (gRNA), single gRNA, or dual guide RNA.
[0274] In some embodiments, the engineered guide polynucleotide (e.g., engineered guide RNA) is configured to form a complex with the engineered endonuclease. In some embodiments, the engineered guide polynucleotide includes a spacer sequence. In some embodiments, the spacer sequence is configured to hybridize to a target nucleic acid sequence. In some embodiments, the endonuclease is configured to bind to a protospacer adjacent motif (PAM) sequence.
[0275] In some embodiments, the guide polynucleotide is at least about 20%, at least about 25%, at least about 30%, at least about 35%, and at least one of the sequence numbers 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, 1705-1710, and 1758-2019. Includes sequences encoded by sequences having approximately 40%, at least approximately 45%, at least approximately 50%, at least approximately 55%, at least approximately 60%, at least approximately 65%, at least approximately 70%, at least approximately 75%, at least approximately 80%, at least approximately 85%, at least approximately 90%, at least approximately 91%, at least approximately 92%, at least approximately 93%, at least approximately 94%, at least approximately 95%, at least approximately 96%, at least approximately 97%, at least approximately 98%, and at least approximately 99% identity. In some embodiments, the guide polynucleotide is encoded by a sequence having at least approximately 80% identity with any one of the following sequence numbers: 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, 1705-1710, and 1758-2019. In some embodiments, the guide polynucleotide is encoded by a sequence having at least approximately 85% identity with any one of the following sequence numbers: 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, 1705-1710, and 1758-2019.In some embodiments, the guide polynucleotide is encoded by a sequence having at least approximately 90% identity with any one of the following sequence numbers: 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, 1705-1710, and 1758-2019. In some embodiments, the guide polynucleotide is encoded by a sequence having at least approximately 95% identity with one of the following sequence numbers: 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, 1705-1710, and 1758-2019. In some embodiments, the guide polynucleotide is encoded by a sequence having at least approximately 96% identity with any one of the following sequence numbers: 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, 1705-1710, and 1758-2019. In some embodiments, the guide polynucleotide is encoded by a sequence having at least approximately 97% identity with any one of the following sequence numbers: 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, 1705-1710, and 1758-2019.In some embodiments, the guide polynucleotide is encoded by a sequence having at least approximately 98% identity with any one of the following sequence numbers: 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, 1705-1710, and 1758-2019. In some embodiments, the guide polynucleotide is encoded by a sequence having at least approximately 99% identity with any one of the following sequence numbers: 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, 1705-1710, and 1758-2019. In some embodiments, the guide polynucleotide is encoded by a sequence having 100% identity with any one of the following sequence numbers: 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, 1705-1710, and 1758-2019.
[0276] In some embodiments, the guide polynucleotide is a sequence complementary to any one of the following sequence numbers: 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, 1705-1710, and 1758-2019, or sequence numbers 88-96, 488- Hybridize or target sequences that have at least 90%, 95%, 97%, 98%, or 99% sequence identity with any one of the following sequences: 489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, 1705-1710, and 1758-2019. In some embodiments, the guide polynucleotides hybridize to or target sequences that are complementary to sequences having at least approximately 80% identity with any one of the following sequences: SEQ ID NOs: 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, 1705-1710, and 1758-2019. In some embodiments, the guide polynucleotides hybridize to or target sequences that are complementary to sequences having at least approximately 85% identity with any one of the following sequences: SEQ ID NOs: 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, 1705-1710, and 1758-2019.In some embodiments, the guide polynucleotides hybridize to or target sequences that are complementary to sequences having at least about 90% identity with any one of the following sequences: SEQ ID NOs: 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, 1705-1710, and 1758-2019. In some embodiments, the guide polynucleotides hybridize to or target sequences that are complementary to sequences having at least approximately 95% identity with any one of the following sequences: SEQ ID NOs: 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, 1705-1710, and 1758-2019. In some embodiments, the guide polynucleotides hybridize to or target sequences that are complementary to sequences having at least approximately 96% identity with any one of the following sequences: SEQ ID NOs: 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, 1705-1710, and 1758-2019. In some embodiments, the guide polynucleotides hybridize to or target sequences that are complementary to sequences having at least approximately 97% identity with any one of the following sequences: SEQ ID NOs: 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, 1705-1710, and 1758-2019.In some embodiments, the guide polynucleotides hybridize to or target sequences that are complementary to sequences having at least approximately 98% identity with any one of the following sequences: SEQ ID NOs: 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, 1705-1710, and 1758-2019. In some embodiments, the guide polynucleotides hybridize to or target sequences that are complementary to sequences having at least approximately 99% identity with any one of the following sequences: SEQ ID NOs: 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, 1705-1710, and 1758-2019. In some embodiments, the guide polynucleotides hybridize to or target sequences that are complementary to sequences having 100% identity with any one of the following sequences: SEQ ID NOs: 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, 1705-1710, and 1758-2019.
[0277] In some embodiments, the manipulated guide polynucleotide includes a sequence having at least 80%, 85%, 90%, 95%, 98%, 99%, or 100% identity with any one of sequence numbers 1431-1454, 1479-1483, 1489-1490, 1705-1710, and 1758-2019.
[0278] In some embodiments, the manipulated guide polynucleotide includes a sequence having at least 80% sequence identity with one of SEQ ID NOs: 1431-1454, 1704, and 2010-2019. In some embodiments, the manipulated guide polynucleotide includes a sequence having at least 85% sequence identity with one of SEQ ID NOs: 1431-1454, 1704, and 2010-2019. In some embodiments, the manipulated guide polynucleotide includes a sequence having at least 90% sequence identity with one of SEQ ID NOs: 1431-1454, 1704, and 2010-2019. In some embodiments, the manipulated guide polynucleotide includes a sequence having at least 95% sequence identity with one of SEQ ID NOs: 1431-1454, 1704, and 2010-2019. In some embodiments, the manipulated guide polynucleotide contains a sequence having 100% sequence identity with one of sequence numbers 1431-1454, 1704, and 2010-2019.
[0279] In some embodiments, the manipulated guide polynucleotide includes a sequence having at least 80% sequence identity with one of the sequence numbers 1890-1976. In some embodiments, the manipulated guide polynucleotide includes a sequence having at least 85% sequence identity with one of the sequence numbers 1890-1976. In some embodiments, the manipulated guide polynucleotide includes a sequence having at least 90% sequence identity with one of the sequence numbers 1890-1976. In some embodiments, the manipulated guide polynucleotide includes a sequence having at least 95% sequence identity with one of the sequence numbers 1890-1976. In some embodiments, the manipulated guide polynucleotide includes a sequence having 100% sequence identity with one of the sequence numbers 1890-1976.
[0280] In some embodiments, the manipulated guide polynucleotide includes a sequence having at least 80% sequence identity with one of SEQ ID NOs. 1977-2009. In some embodiments, the manipulated guide polynucleotide includes a sequence having at least 85% sequence identity with one of SEQ ID NOs. 1977-2009. In some embodiments, the manipulated guide polynucleotide includes a sequence having at least 90% sequence identity with one of SEQ ID NOs. 1977-2009. In some embodiments, the manipulated guide polynucleotide includes a sequence having at least 95% sequence identity with one of SEQ ID NOs. 1977-2009. In some embodiments, the manipulated guide polynucleotide includes a sequence having 100% sequence identity with one of SEQ ID NOs. 1977-2009.
[0281] In some embodiments, the manipulated guide polynucleotide includes a sequence having at least 80% sequence identity with one of the sequence numbers 1479-1483 and 1758-1889. In some embodiments, the manipulated guide polynucleotide includes a sequence having at least 85% sequence identity with one of the sequence numbers 1479-1483 and 1758-1889. In some embodiments, the manipulated guide polynucleotide includes a sequence having at least 90% sequence identity with one of the sequence numbers 1479-1483 and 1758-1889. In some embodiments, the manipulated guide polynucleotide includes a sequence having at least 95% sequence identity with one of the sequence numbers 1479-1483 and 1758-1889. In some embodiments, the manipulated guide polynucleotide includes a sequence having 100% sequence identity with one of the sequence numbers 1479-1483 and 1758-1889.
[0282] In some embodiments, the guide polynucleotide includes a sequence complementary to a eukaryotic, fungal, plant, mammalian, or human genome polynucleotide sequence. In some embodiments, the guide polynucleotide includes a sequence complementary to a eukaryotic genome polynucleotide sequence. In some embodiments, the guide polynucleotide includes a sequence complementary to a fungal genome polynucleotide sequence. In some embodiments, the guide polynucleotide includes a sequence complementary to a plant genome polynucleotide sequence. In some embodiments, the guide polynucleotide includes a sequence complementary to a mammalian genome polynucleotide sequence. In some embodiments, the guide polynucleotide includes a sequence complementary to a human genome polynucleotide sequence.
[0283] In some embodiments, the guide polynucleotide is 30 to 250 nucleotides long. In some embodiments, the guide polynucleotide is 42 to 44 nucleotides long. In some embodiments, the guide polynucleotide is 42 nucleotides long. In some embodiments, the guide polynucleotide is 43 nucleotides long. In some embodiments, the guide polynucleotide is 44 nucleotides long. In some embodiments, the guide polynucleotide is 85 to 245 nucleotides long. In some embodiments, the guide polynucleotide is more than 90 nucleotides long. In some embodiments, the guide polynucleotide is less than 245 nucleotides long. In some embodiments, the guide RNA is 30, 40, 50, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, 220, 240, or more than 240 nucleotides long. In the embodiment where guide RNA is added, the number of RNA molecules is approximately 30-40, 30-50, 30-60, 30-70, 30-80, 30-90, 30-100, 30-120, 30-140, 30-160, 30-180, 30-200, 30-220, 30-240, 50-60, 50-70, 50-80, 50-90, 50-100, and 5 The number of nucleotides is approximately 0 to 120, 50 to 140, 50 to 160, 50 to 180, 50 to 200, 50 to 220, 50 to 240, 100 to 120, 100 to 140, 100 to 160, 100 to 180, 100 to 200, 100 to 220, 100 to 240, 160 to 180, 160 to 200, 160 to 220, or approximately 160 to 240.
[0284] In some embodiments, the manipulated guide polynucleotide comprises a synthetic or modified nucleotide. In some embodiments, the manipulated guide polynucleotide comprises one or more internucleoside linkers modified from natural phosphodiesters. In some embodiments, all or their continuous nucleotide sequences are modified. For example, in some embodiments, the internucleoside linkages include sulfur (S), such as phosphorothioate internucleoside linkages.
[0285] In some embodiments, the manipulated guide polynucleotide comprises modifications to the ribose sugar or nucleic acid base. In some embodiments, the manipulated guide polynucleotide comprises one or more nucleosides containing a modified sugar moiety, where the modified sugar moiety is a modification of the sugar moiety compared to the ribose sugar moiety found in deoxyribose nucleic acids (DNA) and RNA. In some embodiments, the modification is within the ribose ring structure. Exemplary modifications include, but are not limited to, substitution with a hexose ring (HNA), a bicyclic ring having a biradical bridge between the C2 and C4 carbons on the ribose ring (e.g., locked nucleic acid (LNA)), or an unbound ribose ring typically lacking a bond between the C2 and C3 carbons (e.g., UNA). In some embodiments, the sugar-modified nucleoside comprises a bicyclohexose nucleic acid or a tricyclic nucleic acid. In some embodiments, the modified nucleoside comprises a nucleoside in which the sugar moiety is substituted with a non-sugar moiety, such as a peptide nucleic acid (PNA) or morpholino nucleic acid.
[0286] In some embodiments, the manipulated guide polynucleotide contains one or more modified sugars. In some embodiments, the sugar modification includes modifications by altering substituents on the ribose ring to non-hydrogen groups or to naturally occurring 2'-OH groups in DNA and RNA nucleosides. In some embodiments, substituents are introduced at the 2', 3', 4', or 5' positions, or combinations thereof. In some embodiments, the nucleoside having a modified sugar moiety includes 2'-modified nucleosides, e.g., 2'-substituted nucleosides. In some embodiments, the 2'-sugar-modified nucleosides are nucleosides having substituents other than -H or -OH at the 2' position (2'-substituted nucleosides), or include 2'-linked biradicals and include 2'-substituted nucleosides and LNA (2'-4' biradical-bridged) nucleosides. Examples of 2'-substituted nucleosides include, but are not limited to, 2'-O-alkyl-RNA, 2'-O-methyl-RNA, 2'-alkoxy-RNA, 2'-O-methoxyethyl-RNA (MOE), 2'-amino-DNA, 2'-fluoro-RNA, and 2'-F-ANA nucleosides. In some embodiments, the modification in the ribose group involves a modification at the 2' position of the ribose group. In some embodiments, the modification at the 2' position of the ribose group is selected from the group consisting of 2'-O-methyl, 2'-fluoro, 2'-deoxy, and 2'-O-(2-methoxyethyl).
[0287] In some embodiments, the manipulated guide polynucleotide contains one or more modified sugars. In some embodiments, the manipulated guide polynucleotide contains only modified sugars. In some embodiments, the manipulated guide polynucleotide contains more than 10%, 25%, 50%, 75%, or 90% modified sugars. In some embodiments, the modified sugar is a bicyclic sugar. In some embodiments, the modified sugar contains a 2'-O-methoxyethyl group. In some embodiments, the manipulated guide polynucleotide contains both internucleoside linker modification and nucleoside modification.
[0288] In some embodiments, the manipulated guide polynucleotide includes a hairpin containing at least eight base-pair ribonucleotides. In some embodiments, the manipulated guide polynucleotide includes a hairpin containing at least nine base-pair ribonucleotides. In some embodiments, the manipulated guide polynucleotide includes a hairpin containing at least ten base-pair ribonucleotides. In some embodiments, the manipulated guide polynucleotide includes a hairpin containing at least eleven base-pair ribonucleotides. In some embodiments, the manipulated guide polynucleotide includes a hairpin containing at least twelve base-pair ribonucleotides.
[0289] In some embodiments, the manipulated guide polynucleotide includes a DNA targeting segment. In some embodiments, the DNA targeting segment includes a nucleotide sequence complementary to the target sequence. In some embodiments, the target sequence is located in the target DNA molecule. In some embodiments, the manipulated guide polynucleotide includes a protein-binding segment. In some embodiments, the protein-binding segment includes two complementary stretches of nucleotides. In some embodiments, the two complementary stretches of nucleotides hybridize to form a double-stranded RNA (dsRNA) double helix. In some embodiments, the two complementary stretches of nucleotides are covalently linked to each other by an intervening nucleotide.
[0290] Base editing system In this specification, in some embodiments, an engineered system is described comprising an engineered guide polynucleotide comprising (a) a base editor, (b) an endonuclease lacking nuclease activity configured to bind to the base editor, and (c) a spacer sequence configured to form a complex with the endonuclease and to hybridize with a target nucleic acid sequence. When T is referred to in polynucleotides, T means U (uracil) in RNA and T (thymine) in DNA.
[0291] In this specification, in one embodiment, an engineered base editing system is described, comprising a base editor having a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% identity with any one of sequence numbers 1654-1703 and 2021-2023, wherein the sequence does not include any sequence selected from sequence numbers 1128-1160 and 1363-1415, and an engineered guide polynucleotide having a spacer sequence that forms a complex with the endonuclease of the base editor and hybridizes with a target nucleic acid sequence. In this specification, in one embodiment, an engineered base editing system is described, comprising a base editor having at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% identity with any one of sequence numbers 1654-1703 and 2021-2023, wherein the sequence does not include any sequence selected from sequence numbers 1128-1160 and 1363-1415, and an engineered guide polynucleotide comprising a spacer sequence that forms a complex with an endonuclease lacking the nuclease activity of the base editor and hybridizes with a target nucleic acid sequence.
[0292] In this specification, in one embodiment, an engineered base editing system is described, comprising an engineered base editing system comprising a base editor having a sequence having at least 85%, 90%, 95%, 98%, 99%, or 100% identity with any one of sequence numbers 1654-1703 and 2021-2023, and an engineered guide polynucleotide comprising a spacer sequence that forms a complex with the endonuclease of the base editor and hybridizes with a target nucleic acid sequence. In this specification, in one embodiment, an engineered base editing system is described, comprising an engineered base editing system comprising a base editor having a sequence having at least 85%, 90%, 95%, 98%, 99%, or 100% identity with any one of sequence numbers 1654-1703 and 2021-2023, and an engineered guide polynucleotide comprising a spacer sequence that forms a complex with the endonuclease lacking the nuclease activity of the base editor and hybridizes with a target nucleic acid sequence.
[0293] In some embodiments, the manipulated system comprises a) a base editor containing a sequence having at least about 70% identity with any one of sequence numbers 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703; b) an endonuclease; and c) a manipulated guide polynucleotide. In some embodiments, the manipulated system comprises a) a base editor containing a sequence having at least about 75% identity with any one of SEQ ID NOs: 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703; b) an endonuclease; and c) a manipulated guide polynucleotide. In some embodiments, the manipulated system comprises a) a base editor containing a sequence having at least about 80% identity with any one of sequence numbers 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703; b) an endonuclease; and c) a manipulated guide polynucleotide. In some embodiments, the manipulated system comprises a) a base editor containing a sequence having at least about 85% identity with any one of SEQ ID NOs: 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703; b) an endonuclease; and c) a manipulated guide polynucleotide.In some embodiments, the manipulated system comprises a) a base editor containing a sequence having at least about 90% identity with any one of sequence numbers 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703; b) an endonuclease; and c) a manipulated guide polynucleotide. In some embodiments, the manipulated system comprises a) a base editor having at least about 95% identity with any one of sequence numbers 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703; b) an endonuclease; and c) a manipulated guide polynucleotide. In some embodiments, the manipulated system comprises a) a base editor containing a sequence having at least about 96% identity with any one of SEQ ID NOs: 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703; b) an endonuclease; and c) a manipulated guide polynucleotide. In some embodiments, the manipulated system comprises a) a base editor containing a sequence having at least about 97% identity with any one of SEQ ID NOs: 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703; b) an endonuclease; and c) a manipulated guide polynucleotide.In some embodiments, the manipulated system comprises a) a base editor containing a sequence having at least about 98% identity with any one of SEQ ID NOs: 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703; b) an endonuclease; and c) a manipulated guide polynucleotide. In some embodiments, the manipulated system comprises a) a base editor containing a sequence having at least about 99% identity with any one of SEQ ID NOs: 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703; b) an endonuclease; and c) a manipulated guide polynucleotide. In some embodiments, the manipulated system comprises a) a base editor having 100% identity with any one of SEQ ID NOs: 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703; b) an endonuclease; and c) a manipulated guide polynucleotide.
[0294] In some embodiments, the manipulated system includes a) a base editor having at least about 70% identity with any one of sequence numbers 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703; b) an endonuclease configured to bind to the base editor; and c) a manipulated guide polynucleotide having a spacer sequence configured to form a complex with the endonuclease and to hybridize with a target nucleic acid sequence. In some embodiments, the manipulated system includes a) a base editor having at least about 75% identity with any one of sequence numbers 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703; b) an endonuclease configured to bind to the base editor; and c) a manipulated guide polynucleotide having a spacer sequence configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence. In some embodiments, the manipulated system includes a) a base editor having at least about 80% identity with any one of sequence numbers 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703; b) an endonuclease configured to bind to the base editor; and c) a manipulated guide polynucleotide having a spacer sequence configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence.In some embodiments, the manipulated system includes a) a base editor having at least about 85% identity with any one of sequence numbers 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703; b) an endonuclease configured to bind to the base editor; and c) a manipulated guide polynucleotide having a spacer sequence configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence. In some embodiments, the manipulated system includes a) a base editor having at least about 90% identity with any one of sequence numbers 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703; b) an endonuclease configured to bind to the base editor; and c) a manipulated guide polynucleotide having a spacer sequence configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence. In some embodiments, the manipulated system includes a) a base editor having at least about 95% identity with any one of sequence numbers 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703; b) an endonuclease configured to bind to the base editor; and c) a manipulated guide polynucleotide having a spacer sequence configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence.In some embodiments, the manipulated system includes a) a base editor having at least about 96% identity with any one of sequence numbers 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703; b) an endonuclease configured to bind to the base editor; and c) a manipulated guide polynucleotide having a spacer sequence configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence. In some embodiments, the manipulated system includes a) a base editor having at least about 97% identity with any one of sequence numbers 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703; b) an endonuclease configured to bind to the base editor; and c) a manipulated guide polynucleotide having a spacer sequence configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence. In some embodiments, the manipulated system includes a) a base editor having at least about 98% identity with any one of sequence numbers 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703; b) an endonuclease configured to bind to the base editor; and c) a manipulated guide polynucleotide having a spacer sequence configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence.In some embodiments, the manipulated system includes a) a base editor having at least about 99% identity with any one of sequence numbers 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703; b) an endonuclease configured to bind to the base editor; and c) a manipulated guide polynucleotide having a spacer sequence configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence. In some embodiments, the manipulated system includes an manipulated guide polynucleotide comprising: a) a base editor having 100% identity with any one of sequence numbers 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703; b) an endonuclease configured to bind to the base editor; and c) a spacer sequence configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence.
[0295] In some embodiments, the operated system includes: a) a base editor containing a sequence having at least about 70% identity with any one of sequence numbers 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703; b) an endonuclease configured to bind to the base editor; and c) a complex that forms with the endonuclease. The device comprises an engineered guide polynucleotide comprising a spacer sequence configured to hybridize to a target nucleic acid sequence, wherein the engineered guide polynucleotide contains a sequence having at least approximately 70% identity with any one of the following sequences: SEQ ID NOs: 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, and 1704-1710. In some embodiments, the operated system includes: a) a base editor containing a sequence having at least about 75% identity with any one of sequence numbers 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703; b) an endonuclease configured to bind to the base editor; and c) a complex that forms with the endonuclease. The device comprises an engineered guide polynucleotide comprising a spacer sequence configured to hybridize to a target nucleic acid sequence, wherein the engineered guide polynucleotide contains a sequence having at least approximately 75% identity with one of the following sequences: SEQ ID NOs: 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, and 1704-1710.In some embodiments, the operated system includes: a) a base editor containing a sequence having at least about 80% identity with any one of sequence numbers 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703; b) an endonuclease configured to bind to the base editor; and c) a complex that forms with the endonuclease. The device comprises an engineered guide polynucleotide which includes a spacer sequence configured to hybridize to a target nucleic acid sequence, the engineered guide polynucleotide which includes a sequence having at least approximately 80% identity with one of the following sequences: SEQ ID NOs: 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, and 1704-1710. In some embodiments, the operated system includes: a) a base editor containing a sequence having at least approximately 85% identity with any one of sequence numbers 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703; b) an endonuclease configured to bind to the base editor; and c) a complex that forms with the endonuclease. The device comprises an engineered guide polynucleotide comprising a spacer sequence configured to hybridize to a target nucleic acid sequence, wherein the engineered guide polynucleotide contains a sequence having at least approximately 85% identity with one of the following sequences: SEQ ID NOs: 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, and 1704-1710.In some embodiments, the operated system includes: a) a base editor containing a sequence having at least about 90% identity with any one of sequence numbers 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703; b) an endonuclease configured to bind to the base editor; and c) a complex that forms with the endonuclease. The device comprises an engineered guide polynucleotide which includes a spacer sequence configured to hybridize to a target nucleic acid sequence, the engineered guide polynucleotide which includes a sequence having at least approximately 90% identity with one of the following sequences: SEQ ID NOs: 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, and 1704-1710. In some embodiments, the operated system includes: a) a base editor containing a sequence having at least about 95% identity with any one of sequence numbers 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703; b) an endonuclease configured to bind to the base editor; and c) a complex that forms with the endonuclease. The device comprises an engineered guide polynucleotide comprising a spacer sequence configured to hybridize to a target nucleic acid sequence, wherein the engineered guide polynucleotide contains a sequence having at least approximately 95% identity with one of the following sequences: SEQ ID NOs: 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, and 1704-1710.In some embodiments, the operated system includes: a) a base editor containing a sequence having at least approximately 96% identity with any one of sequence numbers 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703; b) an endonuclease configured to bind to the base editor; and c) a complex that forms with the endonuclease. The device comprises an engineered guide polynucleotide comprising a spacer sequence configured to hybridize to a target nucleic acid sequence, wherein the engineered guide polynucleotide contains a sequence having at least approximately 96% identity with one of the following sequences: SEQ ID NOs: 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, and 1704-1710. In some embodiments, the operated system includes: a) a base editor containing a sequence having at least about 97% identity with any one of sequence numbers 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703; b) an endonuclease configured to bind to the base editor; and c) a complex that forms with the endonuclease. The device comprises an engineered guide polynucleotide comprising a spacer sequence configured to hybridize to a target nucleic acid sequence, wherein the engineered guide polynucleotide contains a sequence having at least approximately 97% identity with one of the following sequences: SEQ ID NOs: 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, and 1704-1710.In some embodiments, the operated system includes: a) a base editor containing a sequence having at least approximately 98% identity with any one of sequence numbers 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703; b) an endonuclease configured to bind to the base editor; and c) a system that forms a complex with the endonuclease. The device comprises an engineered guide polynucleotide comprising a spacer sequence configured to hybridize to a target nucleic acid sequence, wherein the engineered guide polynucleotide comprises a sequence having at least approximately 98% identity with one of the following sequences: SEQ ID NOs: 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, and 1704-1710. In some embodiments, the operated system includes: a) a base editor containing a sequence having at least approximately 99% identity with any one of sequence numbers 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703; b) an endonuclease configured to bind to the base editor; and c) a complex that forms with the endonuclease. The device comprises an engineered guide polynucleotide comprising a spacer sequence configured to hybridize to a target nucleic acid sequence, wherein the engineered guide polynucleotide contains a sequence having at least approximately 99% identity with one of the following sequences: SEQ ID NOs: 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, and 1704-1710.In some embodiments, the manipulated system comprises a) a base editor having 100% identity with any one of sequence numbers 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703; b) an endonuclease configured to bind to the base editor; and c) a manipulated guide polynucleotide comprising a spacer sequence configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence. The manipulated guide polynucleotide has 100% identity with any one of the following sequence numbers: 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, and 1704-1710. In some embodiments, the guide polynucleotide is sequence number A sequence complementary to any one of the following: 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, and 1704-1710, or sequence numbers 88-96, 488-489, 679-680, 876 Hybridize or target sequences that have at least 90%, 95%, 97%, 98%, or 99% sequence identity with any one of the following sequences: 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, and 1704-1710.
[0296] In some embodiments, the operated system includes: a) a base editor containing a sequence having at least about 70% identity with any one of sequence numbers 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703; and b) a sequence configured to bind to the base editor, containing a sequence having at least about 70% identity with any one of sequence numbers 70-78, 596, 597, 1120, and 1122-1127. The invention comprises an endonuclease, and c) an engineered guide polynucleotide comprising a spacer sequence configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence, wherein the engineered guide polynucleotide comprises a sequence having at least approximately 70% identity with any one of sequence numbers 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, and 1704-1710.In some embodiments, the operated system includes: a) a base editor containing a sequence having at least about 75% identity with any one of sequence numbers 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703; and b) a sequence configured to bind to the base editor, containing a sequence having at least about 75% identity with any one of sequence numbers 70-78, 596, 597, 1120, and 1122-1127. The invention comprises an endonuclease, and c) an engineered guide polynucleotide comprising a spacer sequence configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence, wherein the engineered guide polynucleotide comprises a sequence having at least approximately 75% identity with any one of sequence numbers 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, and 1704-1710.In some embodiments, the operated system includes: a) a base editor containing a sequence having at least approximately 80% identity with any one of sequence numbers 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703; and b) a sequence configured to bind to the base editor, containing at least approximately 80% identity with any one of sequence numbers 70-78, 596, 597, 1120, and 1122-1127. The invention comprises an endonuclease, and c) an engineered guide polynucleotide comprising a spacer sequence configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence, wherein the engineered guide polynucleotide comprises a sequence having at least approximately 80% identity with any one of sequence numbers 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, and 1704-1710.In some embodiments, the operated system includes: a) a base editor containing a sequence having at least approximately 85% identity with any one of sequence numbers 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703; and b) a sequence configured to bind to the base editor, containing at least approximately 85% identity with any one of sequence numbers 70-78, 596, 597, 1120, or 1122-1127. The invention comprises an endonuclease, and c) an engineered guide polynucleotide comprising a spacer sequence configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence, wherein the engineered guide polynucleotide comprises a sequence having at least approximately 85% identity with any one of sequence numbers 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, and 1704-1710.In some embodiments, the operated system includes: a) a base editor containing a sequence having at least about 90% identity with any one of sequence numbers 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703; and b) a sequence configured to bind to the base editor, containing a sequence having at least about 90% identity with any one of sequence numbers 70-78, 596, 597, 1120, and 1122-1127. The invention comprises an endonuclease, and c) an engineered guide polynucleotide comprising a spacer sequence configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence, wherein the engineered guide polynucleotide comprises a sequence having at least approximately 90% identity with any one of sequence numbers 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, and 1704-1710.In some embodiments, the operated system includes: a) a base editor containing a sequence having at least about 95% identity with any one of sequence numbers 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703; and b) a sequence configured to bind to the base editor, containing a sequence having at least about 95% identity with any one of sequence numbers 70-78, 596, 597, 1120, and 1122-1127. The invention comprises an endonuclease, and c) an engineered guide polynucleotide comprising a spacer sequence configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence, wherein the engineered guide polynucleotide comprises a sequence having at least approximately 95% identity with any one of sequence numbers 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, and 1704-1710.In some embodiments, the operated system includes: a) a base editor containing a sequence having at least about 96% identity with any one of sequence numbers 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703; and b) a sequence configured to bind to the base editor, containing a sequence having at least about 96% identity with any one of sequence numbers 70-78, 596, 597, 1120, and 1122-1127. The invention comprises an endonuclease, and c) an engineered guide polynucleotide comprising a spacer sequence configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence, wherein the engineered guide polynucleotide comprises a sequence having at least approximately 96% identity with any one of sequence numbers 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, and 1704-1710.In some embodiments, the operated system includes: a) a base editor containing a sequence having at least approximately 97% identity with any one of sequence numbers 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703; and b) a sequence configured to bind to the base editor, containing at least approximately 97% identity with any one of sequence numbers 70-78, 596, 597, 1120, and 1122-1127. The invention comprises an endonuclease, and c) an engineered guide polynucleotide comprising a spacer sequence configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence, wherein the engineered guide polynucleotide comprises a sequence having at least approximately 97% identity with any one of sequence numbers 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, and 1704-1710.In some embodiments, the operated system includes: a) a base editor containing a sequence having at least approximately 98% identity with any one of sequence numbers 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703; and b) a sequence configured to bind to the base editor, containing a sequence having at least approximately 98% identity with any one of sequence numbers 70-78, 596, 597, 1120, and 1122-1127. The invention comprises an endonuclease, and c) an engineered guide polynucleotide comprising a spacer sequence configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence, wherein the engineered guide polynucleotide comprises a sequence having at least approximately 98% identity with any one of sequence numbers 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, and 1704-1710. In some embodiments, the operated system is a) at least one of sequence numbers 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703. a) a base editor containing a sequence having approximately 99% identity, b) an endonuclease configured to bind to the base editor and containing a sequence having at least approximately 99% identity with one of sequence numbers 70-78, 596, 597, 1120, and 1122-1127, and c) a spacer sequence configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence. The manipulated guide polynucleotides include a sequence having at least approximately 99% identity with one of the following sequence numbers: 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, and 1704-1710. In some embodiments, the operated system includes: a) a base editor having 100% identity with any one of sequence numbers 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703; and b) an end sequence configured to bind to the base editor and having 100% identity with any one of sequence numbers 70-78, 596, 597, 1120, and 1122-1127. The invention comprises a nuclease, and a modified guide polynucleotide comprising a spacer sequence configured to form a complex with a nuclease and a endonuclease, and to hybridize to a target nucleic acid sequence, wherein the modified guide polynucleotide has 100% identity with any one of sequence numbers 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, and 1704-1710.In some embodiments, the guide polynucleotide is a sequence complementary to any one of the following sequence numbers: 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, and 1704-1710, or sequence numbers 88-96, 488-4 Hybridize or target sequences that have at least 90%, 95%, 97%, 98%, or 99% sequence identity with any one of the following sequences: 89, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, and 1704-1710.
[0297] In some embodiments, the endonuclease comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to SEQ ID NO: 70, the guide RNA structure comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to at least one of SEQ ID NOs: 88, and the endonuclease is configured to bind to a PAM comprising SEQ ID NO: 360.
[0298] In some embodiments, the endonuclease comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to SEQ ID NO: 71, the guide RNA structure comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to at least one of SEQ ID NO: 89, and the endonuclease is configured to bind to a PAM comprising SEQ ID NO: 361.
[0299] In some embodiments, the endonuclease comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to SEQ ID NO: 73, the guide RNA structure comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to at least one of SEQ ID NO: 91, and the endonuclease is configured to bind to a PAM comprising SEQ ID NO: 363.
[0300] In some embodiments, the endonuclease comprises a sequence or variant thereof that is at least 70%, at least 80%, or at least 90% identical to SEQ ID NO: 75, the guide RNA structure comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to at least one of SEQ ID NO: 93, and the endonuclease is configured to bind to PAM containing SEQ ID NO: 365.
[0301] In some embodiments, the endonuclease comprises a sequence or variant thereof that is at least 70%, at least 80%, or at least 90% identical to SEQ ID NO: 76, the guide RNA structure comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to at least one of SEQ ID NO: 94, and the endonuclease is configured to bind to PAM containing SEQ ID NO: 366.
[0302] In some embodiments, the endonuclease comprises a sequence or variant thereof that is at least 70%, at least 80%, or at least 90% identical to SEQ ID NO: 77, the guide RNA structure comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to at least one of SEQ ID NO: 95, and the endonuclease is configured to bind to PAM containing SEQ ID NO: 367.
[0303] In some embodiments, the endonuclease comprises a sequence or variant thereof that is at least 70%, at least 80%, or at least 90% identical to SEQ ID NO: 78, the guide RNA structure comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to at least one of SEQ ID NOs: 96, and the endonuclease is configured to bind to PAM containing SEQ ID NO: 368.
[0304] In some embodiments, the base editor includes adenine deaminase. In some embodiments, the adenine deaminase includes sequence number 57. In some embodiments, the base editor includes cytosine deaminase. In some embodiments, the cytosine deaminase includes sequence number 58.
[0305] In some embodiments, the endonuclease or base editor includes one or more modifications to the nickase domain. In some embodiments, the nickase domain includes a mutation from aspartic acid to alanine in residue 9 for SEQ ID NO: 70, residue 13 for SEQ ID NO: 71, 72, or 74, residue 12 for SEQ ID NO: 73, residue 17 for SEQ ID NO: 75, residue 23 for SEQ ID NO: 76, or residue 10 for SEQ ID NO: 597, or any combination thereof. In some embodiments, the endonuclease or base editor, when optimally aligned, includes the substitution of 109N for SEQ ID NO: 386 and at least one other substitution including any one of 24R, 37L, 49A, 52L, 83S, 85F, 107V, 110S, 112R, 120N, 123N, 124Y, 147C, 148Y, 148R, 150Y, 156V, 157F, 158N, 166I, or 129N, or any combination thereof. Endonucleases or base editors are used for W90A, W90F, W90H, W90Y, Y120F, Y120H, Y121F, Y121H, Y121Q, Y121A, Y121D, Y121W, H122Y, H122F, H122I, H122A, H122W, H122D, Y121T, R33A, R34A, R3 4K, H122A, R33A, R34A, R52A, N57G, H122A, E123A, E123Q, W127F, W127H, W127Q, W127A, W127D, R39A, K40A, H128A, N63G, R58A, H121F, H121Y, H121Q, H121A, H121D, H121W, R33 The substitution of at least one wild-type amino acid with a non-wild-type amino acid, which includes one of the following: A, K34A, H122A, H121A, R52A, P26R, P26A, N27R, N27A, W44A, W45A, K49G, S50G, R51G, R121A, I122A, N123A, Y88F, Y120F, P22R, P22A, K23A, K41R, K41A, E54A, E54A, E55A, K30A, K30R, M32A, M32K, Y117A, K118A, I119A, I119H, R120A, R121A, P46A, P46R, N29A, R27A, or N50G, or any combination thereof.In some embodiments, the nickase includes a mutation from aspartic acid to alanine in residue 9 for SEQ ID NO: 70, residue 13 for SEQ ID NO: 71, 72, or 74, residue 12 for SEQ ID NO: 73, residue 17 for SEQ ID NO: 75, residue 23 for SEQ ID NO: 76, or residue 10 for SEQ ID NO: 597, or any combination thereof.
[0306] In some embodiments, the operated system further comprises a uracil DNA glycosylase inhibitor. In some embodiments, the uracil DNA glycosylase inhibitor comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NOs. In some embodiments, the uracil DNA glycosylase inhibitor comprises a sequence having at least 70% identity with any one of SEQ ID NOs. In some embodiments, the uracil DNA glycosylase inhibitor comprises a sequence having at least 75% identity with any one of SEQ ID NOs. In some embodiments, the uracil DNA glycosylase inhibitor includes a sequence having at least 80% identity with one of SEQ ID NOs. 52-56 or 67. In some embodiments, the uracil DNA glycosylase inhibitor includes a sequence having at least 85% identity with one of SEQ ID NOs. 52-56 or 67. In some embodiments, the uracil DNA glycosylase inhibitor includes a sequence having at least 90% identity with one of SEQ ID NOs. 52-56 or 67. In some embodiments, the uracil DNA glycosylase inhibitor includes a sequence having at least 95% identity with one of SEQ ID NOs. 52-56 or 67. In some embodiments, the uracil DNA glycosylase inhibitor includes a sequence having at least 96% identity with one of SEQ ID NOs. 52-56 or 67. In some embodiments, the uracil DNA glycosylase inhibitor includes a sequence having at least 97% identity with one of SEQ ID NOs. 52-56 or 67.In some embodiments, the uracil DNA glycosylase inhibitor includes a sequence having at least 98% identity with one of SEQ ID NOs. 52-56 or 67. In some embodiments, the uracil DNA glycosylase inhibitor includes a sequence having at least 99% identity with one of SEQ ID NOs. 52-56 or 67. In some embodiments, the uracil DNA glycosylase inhibitor includes a sequence having 100% identity with one of SEQ ID NOs. 52-56 or 67.
[0307] In some embodiments, the base editor is non-covalently bonded to the endonuclease. In some embodiments, the base editor is covalently bonded to the endonuclease. In some embodiments, the base editor is fused to the endonuclease at the N-terminus or C-terminus. In some embodiments, the base editor is fused to the endonuclease.
[0308] In some embodiments, the endonuclease is covalently bonded to the base editor or covalently bonded to the base editor via a linker. In some embodiments, the linker includes a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SGGSSGGSSGSETPGTSESATPESSGGSSGGS, SGSETPGTSESATPESA, GSGGS, SGSETPGTSESATPES, SGGSS, or GAAA. In some embodiments, the linker includes a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with one of sequence numbers 1647-1653.
[0309] [Table 1]
[0310] In some embodiments, the system is Mg 2+ This further includes the sources of supply.
[0311] In some embodiments, the endonuclease includes one or more nuclear localization sequences (NLS) near the N-terminus or C-terminus of the endonuclease. In some embodiments, the base editor includes one or more nuclear localization sequences (NLS) near the N-terminus or C-terminus of the endonuclease. The NLS may include any or a combination thereof of the sequences in Table 2 below.
[0312] In some embodiments, the NLS includes a sequence that has at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with one of the sequences from sequence numbers 369-384 and 2024-2053. In some embodiments, the NLS includes a sequence having at least approximately 85% identity with one of sequence numbers 369-384 and 2024-2053. In some embodiments, the NLS includes a sequence having at least approximately 90% identity with one of sequence numbers 369-384 and 2024-2053. In some embodiments, the NLS includes a sequence having at least approximately 91% identity with one of sequence numbers 369-384 and 2024-2053. In some embodiments, the NLS includes a sequence having at least approximately 92% identity with one of sequence numbers 369-384 and 2024-2053. In some embodiments, the NLS includes a sequence having at least approximately 93% identity with one of sequence numbers 369-384 and 2024-2053. In some embodiments, the NLS includes a sequence having at least approximately 94% identity with one of sequence numbers 369-384 and 2024-2053. In some embodiments, the NLS includes a sequence having at least about 95% identity with one of sequence numbers 369-384 and 2024-2053.In some embodiments, the NLS includes a sequence having at least approximately 96% identity with one of sequence numbers 369-384 and 2024-2053. In some embodiments, the NLS includes a sequence having at least approximately 97% identity with one of sequence numbers 369-384 and 2024-2053. In some embodiments, the NLS includes a sequence having at least approximately 98% identity with one of sequence numbers 369-384 and 2024-2053. In some embodiments, the NLS includes a sequence having at least approximately 99% identity with one of sequence numbers 369-384 and 2024-2053. In some embodiments, the NLS includes a sequence having 100% identity with one of sequence numbers 369-384 and 2024-2053.
[0313] [Table 2-1]
[0314] [Table 2-2]
[0315] [Table 2-3]
[0316] cell In certain embodiments, cells comprising the system described herein are described herein.
[0317] In some embodiments, the cells may be eukaryotic cells (e.g., plant cells, animal cells, protist cells, or fungal cells), mammalian cells (Chinese hamster ovary (CHO) cells, baby hamster kidney (BHK), human fetal kidney (HEK), mouse myeloma (NS0), or human retinal cells), immortalized cells (e.g., HeLa cells, COS cells, HEK-293T cells, MDCK cells, 3T3 cells, PC12 cells, Huh7 cells, HepG2 cells, K562 cells, N2a cells, or SY5Y cells), insect cells (e.g., Spodoptera frugiperda cells, Trichoplusia ni cells, Drosophila melanogaster cells, S2 cells, or Heliothis virescens cells), or yeast cells (e.g., Saccharomyces). These include cerevisiae cells, Cryptococcus cells, or Candida cells, plant cells (e.g., parenchymal cells, plaque cells, or plaque-walled cells), fungal cells (e.g., Saccharomyces cerevisiae cells, Cryptococcus cells, or Candida cells), or prokaryotic cells (e.g., E. coli cells, Streptococcus bacterial cells, Streptomyces soil bacterial cells, or archaeal cells). In some embodiments, the cells are eukaryotic cells. In some embodiments, the cells are mammalian cells. In some embodiments, the cells are immortalized cells. In some embodiments, the cells are insect cells. In some embodiments, the cells are yeast cells. In some embodiments, the cells are plant cells. In some embodiments, the cells are fungal cells. In some embodiments, the cells are prokaryotic cells.
[0318] In some embodiments, the cells are A549, HEK-293, HEK-293T, BHK, CHO, HeLa, MRC5, Sf9, Cos-1, Cos-7, Vero, BSC1, BSC40, BMT10, WI38, HeLa, Saos, C2C12, L cells, HT1080, HepG2, Huh7, K562, primary cells, or derivatives thereof.
[0319] In some embodiments, this disclosure provides cells (e.g., host cells) containing the vectors described herein. In some embodiments, the cells express the manipulated systems or components described herein. In some embodiments, the cells are human cells. In some embodiments, the cells are genome-edited ex vivo. In some embodiments, the cells are genome-edited in vivo.
[0320] In some embodiments of this specification, a host cell is described that includes a heterologous endonuclease and a heterologous base editor having at least 75% sequence identity with any one of SEQ ID NOs: 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703. In some embodiments, the aforementioned heterogeneous base editor includes sequences having at least 75%, at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 97%, at least about 98%, at least about 98%, at least about 99%, or 100% sequence identity with any one of sequence numbers 1-51, 57-66, 385-443, 444-475, 594-595, 599-675, 744-835, 970-1098, 1128-1186, 1208-1315, 1356-1415, 1424, 1556-1644, and 1654-1703.
[0321] In some embodiments, the host cell is a bacterial cell. In some embodiments, the bacterial cell is Bifidobacterium longum, Bifidobacterium lactis, Bifidobacterium animalis, Bifidobacterium breve, Bifidobacterium infantis, Bifidobacterium adolescentis, Lactobacillus acidophilus, Lactobacillus casei, Lactobacillus paracasei, Lactobacillus salivarius, Lactobacillus reuteri, Lactobacillus rhamnosus, Lactobacillus johnsonii, Lactobacillus plantarum, Lactobacillus fermentum, Lactococcus lactis, Streptococcus thermophilus, Lactococcus lactis, Lactococcus diacetylactis, Lactococcus cremoris, Lactobacillus bulgaricus, Lactobacillus helveticus, Lactobacillus The host cell is delbrueckii, or Escherichia coli. In some embodiments, the host cell is an E. coli cell. In some embodiments, the E. coli cell is the λDE3 lysogen, or BL21(DE3) strain. In some embodiments, the E. coli cell has the ompT lon genotype.
[0322] In some embodiments, the cells are located within the cochlea. In some embodiments, the cells are located within the embryo. In some embodiments, the embryo is a two-cell stage embryo. In some embodiments, the embryo is a mouse embryo.
[0323] Lipid nanoparticles The lipid nanoparticles described herein may be four-component lipid nanoparticles. Such nanoparticles can be configured for the delivery of RNA or other nucleic acids (e.g., synthetic RNA, mRNA, or mRNA synthesized in vitro) and can generally be formulated as described in WO2012135805A2. Such nanoparticles can generally include (a) cationic lipids (e.g., 98N12-5 (TETA5-LAP), DLin DMA, DLin-K-DMA (2,2-dilinoleyl-4-dimethylaminomethyl-[1,3]-dioxolane), DLin-KC2-DMA, DLin-MC3-DMA, or C12-200), (b) neutral lipids (e.g., DSPC or DOPE), (c) sterols (e.g., cholesterol or cholesterol analogs), and (d) PEG-modified lipids (e.g., PEG-DMG).
[0324] Cationic lipids referred to herein as "C12-200" are disclosed in Love et al., Proc Natl Acad Sci USA. 2010 107:1864-1869 and Liu and Huang, Molecular Therapy. 2010 669-670. Cationic lipid formulations may include particles containing three, four, or more components in addition to polynucleotides, primary constructs, or RNA (e.g., mRNA). For example, formulations containing a particular cationic lipid include, but are not limited to, 98N12-5, which may contain 42% lipidoid, 48% cholesterol, and 10% PEG (alkyl chain length of C14 or greater). Another example of a formulation containing a specific lipidoid is C12-200, which may contain 50% cationic lipids, 10% disteloylphosphatidylcholine, 38.5% cholesterol, and 1.5% PEG-DMG.
[0325] In some embodiments, the lipid nanoparticles are formulated as described in U.S. Patent No. 10709779B2. In some embodiments, the cationic lipid nanoparticles comprise a cationic lipid, a PEG-modified lipid, a sterol, and a non-cationic lipid. In some embodiments, the cationic lipid is selected from the group consisting of 98N12-5 (TETA5-LAP), DLin DMA, DLin-K-DMA (2,2-dilinoleyl-4-dimethylaminomethyl-[1,3]-dioxolane), DLin-KC2-DMA, DLin-MC3-DMA, and C12-200. In some embodiments, the cationic lipid nanoparticles have a molar ratio of about 20-60% cationic lipid, about 5-25% non-cationic lipid, about 25-55% sterol, and about 0.5-15% PEG-modified lipid. In some embodiments, the cationic lipid nanoparticles contain a molar ratio of about 50% cationic lipid, about 1.5% PEG-modified lipid, about 38.5% cholesterol, and about 10% non-cationic lipid. In some embodiments, the cationic lipid nanoparticles contain a molar ratio of about 55% cationic lipid, about 2.5% PEG-modified lipid, about 32.5% cholesterol, and about 10% non-cationic lipid. In some embodiments, the cationic lipid is an ionic cationic lipid, the non-cationic lipid is a neutral lipid, and the sterol is cholesterol. In some embodiments, the cationic lipid nanoparticles have a molar ratio of 50:38.5:10:1.5 cationic lipid:cholesterol:PEG2000-DMG:DSPC or DMG:DOPE. In some embodiments, the lipid nanoparticles described herein may contain cholesterol, 1,2-dioleoyl-sn-glycero-3-phosphoethanolamine (DOPE), 1,1'-((2-(4-(2-((2-((bis(2-hydroxydodecyl)amino)ethyl)(2-hydroxydodecyl)amino)ethyl)piperazine-1-yl)ethyl)azandiyl)bis(dodecane-2-ol)(C12-200), and DMG-PEG-2000 in a molar ratio of 47.5:16:35:1.5.
[0326] Delivery and vector In this specification, in some embodiments, nucleic acids encoding the manipulated system described herein are disclosed, comprising a base editor, an endonuclease, and a manipulated guide polynucleotide or its components (e.g., a base editor, an endonuclease, or a manipulated guide polynucleotide).
[0327] In some embodiments, the nucleic acid encoding the manipulated system or its components is DNA, such as linear DNA, plasmid DNA, or minicircle DNA. In some embodiments, the nucleic acid encoding the manipulated system is RNA, such as mRNA.
[0328] In some embodiments, nucleic acids encoding the manipulated system or its components are delivered by nucleic acid-based vectors. In some embodiments, the nucleic acid-based vectors are plasmids (e.g., circular DNA molecules that can autonomously replicate inside a cell), cosmids (e.g., pWE or sCos vectors), artificial chromosomes, human artificial chromosomes (HACs), yeast artificial chromosomes (YACs), bacterial artificial chromosomes (BACs), P1-derived artificial chromosomes (PACs), phagemids, phage derivatives, bacmids, or viruses. In some embodiments, the nucleic acid-based vectors are pSF-CMV-NEO-NH2-PPT-3XFLAG, pSF-CMV-NEO-COOH-3XFLAG, pSF-CMV-PURO-NH2-GST-TEV, pSF-OXB20-COOH-TEV-FLAG(R)-6His, pCEP4 pDEST27, pSF-CMV-Ub-KrYFP, pSF-CMV-FMDV-daGFP, pEF1a-mCherry-N1 vector, pEF1a-tdTomato vector, pSF-CMV-FMDV-Hygro, pSF-CMV-PGK-Puro, pMCP-tag(m), pSF-CMV-PURO-NH2-CMYC, pSF-OXB20-BetaGal, pSF-OXB20-Fluc, pSF-OXB20, pSF-Tac, and pRI 101-AN. The selection is made from a list consisting of DNA, pCambia2301, pTYB21, pKLAC2, pAc5.1 / V5-His A, and pDEST8.
[0329] In some embodiments, the nucleic acid-based vector includes a promoter. In some embodiments, an open reading frame is operably linked to the promoter. In some embodiments, the promoter is selected from the group consisting of mini-promoters, inducible promoters, constitutive promoters, and their derivatives. In some embodiments, the promoter is selected from the group consisting of CMV, CBA, EF1a, CAG, PGK, TRE, U6, UAS, T7, Sp6, lac, araBad, trp, Ptac, p5, p19, p40, synapsin, CaMKII, GRK1, and their derivatives. In some embodiments, the promoter is the U6 promoter. In some embodiments, the promoter is the CAG promoter.
[0330] In some embodiments, the open reading frame includes the T7 promoter sequence, T7-lac promoter sequence, lac promoter sequence, tac promoter sequence, trc promoter sequence, ParaBAD promoter sequence, PrhaBAD promoter sequence, T5 promoter sequence, cspA promoter sequence, and araP BAD It is operably coupled to a promoter, a strong leftward promoter (pL promoter) from a phage lambda, or any combination thereof.
[0331] In some embodiments, the open reading frame includes a sequence encoding an affinity tag in-frame to the sequence encoding the aforementioned base editor. In some embodiments, the affinity tag is an immobilized metal affinity chromatography (IMAC) tag. In some embodiments, the IMAC tag is a polyhistidine tag. In some embodiments, the affinity tag is a myc tag, a human influenza hemagglutinin (HA) tag, a maltose-binding protein (MBP) tag, a glutathione S-transferase (GST) tag, a streptavidin tag, a FLAG tag, or any combination thereof.
[0332] In some embodiments, the affinity tag is in-frame concatenated to the aforementioned sequence encoding the aforementioned base editor via a linker sequence encoding a protease cleavage site. In some embodiments, the protease cleavage site is a tobacco etch virus (TEV) protease cleavage site, a PreScission® protease (PSP) cleavage site, a thrombin cleavage site, a factor Xa cleavage site, an enterokinase cleavage site, or any combination thereof.
[0333] In some embodiments, the open reading frame is codon-optimized for expression in the aforementioned host cell. In some embodiments, the open reading frame is provided on a vector. In some embodiments, the open reading frame is incorporated into the genome of the aforementioned host cell.
[0334] In some embodiments, the nucleic acid-based vector is a virus. In some embodiments, the virus is an alphavirus, parvovirus, adenovirus, AAV, baculovirus, dengue virus, lentivirus, herpesvirus, poxvirus, anellovirus, bocavirus, vacciniavirus, or retrovirus. In some embodiments, the virus is an alphavirus. In some embodiments, the virus is a parvovirus. In some embodiments, the virus is an adenovirus. In some embodiments, the virus is an AAV. In some embodiments, the virus is a baculovirus. In some embodiments, the virus is a dengue virus. In some embodiments, the virus is a lentivirus. In some embodiments, the virus is a herpesvirus. In some embodiments, the virus is a poxvirus. In some embodiments, the virus is an anellovirus. In some embodiments, the virus is a bocavirus. In some embodiments, the virus is a vacciniavirus. In some embodiments, the virus is a retrovirus.
[0335] In some embodiments, the AAV is AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAV14, AAV15, AAV16, AAV-rh8, AAV-rh1 0, AAV-rh20, AAV-rh39, AAV-rh74, AAV-rhM4-1, AAV-hu37, AAV-Anc80, AAV-Anc80L65, AAV-7m8, AAV-PHP-B, AAV-PHP-EB, AAV-2.5, AAV-2tYF, A The following are examples of the following: AV-3B, AAV-LK03, AAV-HSC1, AAV-HSC2, AAV-HSC3, AAV-HSC4, AAV-HSC5, AAV-HSC6, AAV-HSC7, AAV-HSC8, AAV-HSC9, AAV-HSC10, AAV-HSC11, AAV-HSC12, AAV-HSC13, AAV-HSC14, AAV-HSC15, AAV-TT, AAV-DJ / 8, AAV-Myo, AAV-NP40, AAV-NP59, AAV-NP22, AAV-NP66, AAV-HSC16, or derivatives thereof. In some embodiments, the herpesvirus is HSV type 1, HSV-2, VZV, EBV, CMV, HHV-6, HHV-7, or HHV-8.
[0336] In some embodiments, the virus is AAV1 or a derivative thereof. In some embodiments, the virus is AAV2 or a derivative thereof. In some embodiments, the virus is AAV3 or a derivative thereof. In some embodiments, the virus is AAV4 or a derivative thereof. In some embodiments, the virus is AAV5 or a derivative thereof. In some embodiments, the virus is AAV6 or a derivative thereof. In some embodiments, the virus is AAV7 or a derivative thereof. In some embodiments, the virus is AAV8 or a derivative thereof. In some embodiments, the virus is AAV9 or a derivative thereof. In some embodiments, the virus is AAV10 or a derivative thereof. In some embodiments, the virus is AAV11 or a derivative thereof. In some embodiments, the virus is AAV12 or a derivative thereof. In some embodiments, the virus is AAV13 or a derivative thereof. In some embodiments, the virus is AAV14 or a derivative thereof. In some embodiments, the virus is AAV15 or a derivative thereof. In some embodiments, the virus is AAV16 or a derivative thereof. In some embodiments, the virus is AAV-rh8 or a derivative thereof. In some embodiments, the virus is AAV-rh10 or a derivative thereof. In some embodiments, the virus is AAV-rh20 or a derivative thereof. In some embodiments, the virus is AAV-rh39 or a derivative thereof. In some embodiments, the virus is AAV-rh74 or a derivative thereof. In some embodiments, the virus is AAV-rhM4-1 or a derivative thereof. In some embodiments, the virus is AAV-hu37 or a derivative thereof. In some embodiments, the virus is AAV-Anc80 or a derivative thereof. In some embodiments, the virus is AAV-Anc80L65 or a derivative thereof. In some embodiments, the virus is AAV-7m8 or a derivative thereof. In some embodiments, the virus is AAV-PHP-B or a derivative thereof.In some embodiments, the virus is AAV-PHP-EB or a derivative thereof. In some embodiments, the virus is AAV-2.5 or a derivative thereof. In some embodiments, the virus is AAV-2tYF or a derivative thereof. In some embodiments, the virus is AAV-3B or a derivative thereof. In some embodiments, the virus is AAV-LK03 or a derivative thereof. In some embodiments, the virus is AAV-HSC1 or a derivative thereof. In some embodiments, the virus is AAV-HSC2 or a derivative thereof. In some embodiments, the virus is AAV-HSC3 or a derivative thereof. In some embodiments, the virus is AAV-HSC4 or a derivative thereof. In some embodiments, the virus is AAV-HSC5 or a derivative thereof. In some embodiments, the virus is AAV-HSC6 or a derivative thereof. In some embodiments, the virus is AAV-HSC7 or a derivative thereof. In some embodiments, the virus is AAV-HSC8 or a derivative thereof. In some embodiments, the virus is AAV-HSC9 or a derivative thereof. In some embodiments, the virus is AAV-HSC10 or a derivative thereof. In some embodiments, the virus is AAV-HSC11 or a derivative thereof. In some embodiments, the virus is AAV-HSC12 or a derivative thereof. In some embodiments, the virus is AAV-HSC13 or a derivative thereof. In some embodiments, the virus is AAV-HSC14 or a derivative thereof. In some embodiments, the virus is AAV-HSC15 or a derivative thereof. In some embodiments, the virus is AAV-TT or a derivative thereof. In some embodiments, the virus is AAV-DJ / 8 or a derivative thereof. In some embodiments, the virus is AAV-Myo or a derivative thereof. In some embodiments, the virus is AAV-NP40 or a derivative thereof. In some embodiments, the virus is AAV-NP59 or a derivative thereof. In some embodiments, the virus is AAV-NP22 or a derivative thereof.In some embodiments, the virus is AAV-NP66 or a derivative thereof. In some embodiments, the virus is AAV-HSC16 or a derivative thereof.
[0337] In some embodiments, the virus is HSV-1 or a derivative thereof. In some embodiments, the virus is HSV-2 or a derivative thereof. In some embodiments, the virus is VZV or a derivative thereof. In some embodiments, the virus is EBV or a derivative thereof. In some embodiments, the virus is CMV or a derivative thereof. In some embodiments, the virus is HHV-6 or a derivative thereof. In some embodiments, the virus is HHV-7 or a derivative thereof. In some embodiments, the virus is HHV-8 or a derivative thereof.
[0338] In some embodiments, the nucleic acid encoding the engineered system, endonuclease, or engineered guide polynucleotide is delivered by a non-nucleic acid-based delivery system (e.g., a non-viral delivery system). In some embodiments, the non-viral delivery system is a liposome. In some embodiments, the nucleic acid is related to lipids. In some embodiments, the lipid-bound nucleic acid is encapsulated within the aqueous interior of a liposome, dispersed within the lipid bilayer of a liposome, attached to a liposome via binding molecules that bind to both the liposome and the nucleic acid, confined within a liposome, complexed with a liposome, dispersed in a lipid-containing solution, mixed with a lipid, combined with a lipid, contained as a suspension in a lipid, contained in or complexed with a micelle, or otherwise bound to a lipid. In some embodiments, the nucleic acid is contained in lipid nanoparticles (LNPs).
[0339] In some embodiments, the engineered system, endonuclease, or engineered guide polynucleotide is introduced into the cell in any preferred manner, either stably or transiently. In some embodiments, the engineered system, endonuclease, or engineered guide polynucleotide is transfected into the cell. In some embodiments, the cell is transduced or transfected with a nucleic acid construct encoding the engineered system, endonuclease, or engineered guide polynucleotide. For example, the cell is transduced (e.g., with a virus encoding the engineered system, endonuclease, or engineered guide polynucleotide) or transfected (e.g., with a plasmid encoding the engineered system, endonuclease, or engineered guide polynucleotide) with a nucleic acid encoding the engineered system, endonuclease, or engineered guide polynucleotide, or with a translated engineered system or endonuclease. In some embodiments, the transduction is stably or transiently transduced. In some embodiments, cells expressing the engineered system, endonuclease, or engineered guide polynucleotide are transfected with one or more gRNA molecules. In some embodiments, plasmids expressing the engineered system, endonuclease, or engineered guide polynucleotide are introduced into cells by electroporation, transient (e.g., lipofection) and stable genome integration (e.g., piggyback), as well as viral transduction (e.g., lentivirus or AAV), or other methods known to those skilled in the art. In some embodiments, the engineered system, endonuclease, or engineered guide polynucleotide is introduced into cells as one or more polypeptides. In some embodiments, delivery is achieved by the use of RNP complexes. For example, methods for delivering polypeptides and / or RNPs into cells by electroporation or cell compression are known in the art.
[0340] Exemplary methods for nucleic acid delivery include lipofection, nucleofection, electroporation, stable genomic integration (e.g., piggyback), microinjection, biolistek, virosomes, liposomes, immunoliposomes, polycationic or lipid nucleic acid conjugates, naked DNA, artificial virions, and drug-enhanced uptake of DNA. Lipofection is described, for example, in U.S. Patents 5,049,386, 4,946,787, and 4,897,355, and lipofection reagents are commercially available (e.g., Transfectam®, Lipofectin®, and SF Cell Line 4D-Nucleofector X Kit® (Lonza)). Cationic and neutral lipids suitable for efficient receptor recognition lipofection of polynucleotides include those of WO91 / 17424 and WO91 / 16024. In some embodiments, delivery is to cells (e.g., in vitro or ex vivo administration) or target tissue (e.g., in vivo administration). In some embodiments, the nucleic acid is contained in liposomes or nanoparticles that specifically target host cells.
[0341] Additional methods for delivering nucleic acids to cells are known to those skilled in the art. See, for example, U.S. Patent Application Publication No. 2003 / 0087817.
[0342] How to use In this specification, in some embodiments, a method for modifying a target nucleic acid using a base editor or manipulated system described herein, a method for disrupting a gene locus using a base editor or manipulated system described herein, or a method for manufacturing a base editor or manipulated system described herein is described herein.
[0343] In some embodiments, the manipulated system comprises an adenine deaminase base editor, where the nucleotide is adenine, and modifying the target nucleic acid locus involves converting adenine to guanine. In some embodiments, the manipulated system comprises a cytidine deaminase base editor and a uracil DNA glycosylase inhibitor, where the nucleotide is cytosine, and modifying the target nucleic acid locus involves converting cytosine to uracil.
[0344] In some embodiments, the method is used to introduce modifications into the genome of a cell. In some embodiments, the target nucleic acid is modified in vitro. In some embodiments, the target nucleic acid sequence is modified in vivo. In some embodiments, the target nucleic acid sequence is modified ex vivo.
[0345] In some embodiments, the target nucleic acid includes genomic DNA, viral DNA, or bacterial DNA. In some embodiments, the target nucleic acid is intracellular. In some embodiments, the cell is a prokaryotic cell, bacterial cell, eukaryotic cell, fungal cell, plant cell, animal cell, mammalian cell, rodent cell, primate cell, or human cell. In some embodiments, the cell is intracellular.
[0346] In some embodiments, the target nucleic acid includes DNA. In some embodiments, the DNA includes a first strand containing a sequence complementary to the sequence of the manipulated guide polynucleotide, and a second strand containing the PAM. In some embodiments, the PAM is directly adjacent to the 3' end of the sequence complementary to the sequence of the manipulated guide polynucleotide. In some embodiments, the PAM contains a sequence selected from the group consisting of SEQ ID NOs. 360-368 or 598.
[0347] In some embodiments, the disclosure provides methods for modifying a target nucleic acid (e.g., gene) locus. In some embodiments, the method includes delivering the engineered system described herein to the target nucleic acid locus. In some embodiments, the endonuclease is configured to form a complex with an engineered guide polynucleotide. In some embodiments, the complex is configured to modify the target nucleic acid locus upon binding of the complex to the target nucleic acid locus.
[0348] In some embodiments, delivery of the manipulated system to a target nucleic acid locus includes delivery of a nucleic acid or vector described herein. In some embodiments, delivery of the manipulated system to a target nucleic acid locus includes delivery of a nucleic acid comprising a base editor and an open reading frame encoding an endonuclease. In some embodiments, the nucleic acid comprises a promoter. In some embodiments, the open reading frame encoding the base editor and the endonuclease is operably linked to the promoter.
[0349] In some embodiments, delivery of the engineered system to the target nucleic acid locus includes delivery of capped mRNA containing a base editor and an open reading frame encoding an endonuclease. In some embodiments, delivery of the engineered system to the target nucleic acid locus includes delivery of a translated polypeptide. In some embodiments, delivery of the engineered system to the target nucleic acid locus includes delivery of deoxyribonucleic acid (DNA) encoding an engineered guide RNA operably ligated to a ribonucleic acid (RNA) pol III promoter.
[0350] In some embodiments, the target gene is TRAC. In some embodiments, the gRNA contains a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with SEQ ID NO: 1489 or SEQ ID NO: 1490. In some embodiments, the gRNA contains a sequence having at least about 70% identity with SEQ ID NO: 1489 or SEQ ID NO: 1490. In some embodiments, the gRNA contains a sequence having at least about 80% identity with SEQ ID NO: 1489 or SEQ ID NO: 1490. In some embodiments, the gRNA contains a sequence having at least about 85% identity with SEQ ID NO: 1489 or SEQ ID NO: 1490. In some embodiments, the gRNA contains a sequence having at least about 90% identity with SEQ ID NO: 1489 or SEQ ID NO: 1490. In some embodiments, the gRNA contains a sequence having at least about 91% identity with SEQ ID NO: 1489 or SEQ ID NO: 1490. In some embodiments, the gRNA contains a sequence having at least about 92% identity with SEQ ID NO: 1489 or SEQ ID NO: 1490. In some embodiments, the gRNA contains a sequence having at least about 93% identity with SEQ ID NO: 1489 or SEQ ID NO: 1490. In some embodiments, the gRNA contains a sequence having at least about 94% identity with SEQ ID NO: 1489 or SEQ ID NO: 1490. In some embodiments, the gRNA contains a sequence having at least about 95% identity with SEQ ID NO: 1489 or SEQ ID NO: 1490. In some embodiments, the gRNA contains a sequence having at least about 96% identity with SEQ ID NO: 1489 or SEQ ID NO: 1490.In some embodiments, the gRNA contains a sequence having at least about 97% identity with SEQ ID NO: 1489 or SEQ ID NO: 1490. In some embodiments, the gRNA contains a sequence having at least about 98% identity with SEQ ID NO: 1489 or SEQ ID NO: 1490. In some embodiments, the gRNA contains a sequence having at least about 99% identity with SEQ ID NO: 1489 or SEQ ID NO: 1490. In some embodiments, the gRNA contains a sequence having 100% identity with SEQ ID NO: 1489 or SEQ ID NO: 1490.
[0351] In some embodiments, the gRNA hybridizes with a TRAC sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with one of the SEQ ID NOs. In some embodiments, the gRNA hybridizes with a TRAC sequence having at least about 75% identity with one of sequence numbers 1491-1492 and 1711-1719. In some embodiments, the gRNA hybridizes with a TRAC sequence having at least about 80% identity with one of sequence numbers 1491-1492 and 1711-1719. In some embodiments, the gRNA hybridizes with a TRAC sequence having at least about 85% identity with one of sequence numbers 1491-1492 and 1711-1719. In some embodiments, the gRNA hybridizes with a TRAC sequence having at least about 90% identity with one of sequence numbers 1491-1492 and 1711-1719. In some embodiments, the gRNA hybridizes with a TRAC sequence having at least approximately 91% identity with one of sequence numbers 1491-1492 and 1711-1719. In some embodiments, the gRNA hybridizes with a TRAC sequence having at least approximately 92% identity with one of sequence numbers 1491-1492 and 1711-1719.In some embodiments, the gRNA hybridizes with a TRAC sequence having at least approximately 93% identity with one of sequence numbers 1491-1492 and 1711-1719. In some embodiments, the gRNA hybridizes with a TRAC sequence having at least approximately 94% identity with one of sequence numbers 1491-1492 and 1711-1719. In some embodiments, the gRNA hybridizes with a TRAC sequence having at least approximately 95% identity with one of sequence numbers 1491-1492 and 1711-1719. In some embodiments, the gRNA hybridizes with a TRAC sequence having at least approximately 96% identity with one of sequence numbers 1491-1492 and 1711-1719. In some embodiments, the gRNA hybridizes with a TRAC sequence having at least approximately 97% identity with one of sequence numbers 1491-1492 and 1711-1719. In some embodiments, the gRNA hybridizes with a TRAC sequence having at least approximately 98% identity with one of sequence numbers 1491-1492 and 1711-1719. In some embodiments, the gRNA hybridizes with a TRAC sequence having at least approximately 99% identity with one of sequence numbers 1491-1492 and 1711-1719. In some embodiments, the gRNA hybridizes with a TRAC sequence having 100% identity with one of sequence numbers 1491-1492 and 1711-1719.
[0352] In some embodiments, the target gene is AAVS1. In some embodiments, the gRNA contains sequences that have at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with sequence numbers 1705-1710. In some embodiments, the gRNA contains sequences that have at least about 70% identity with sequence numbers 1705-1710. In some embodiments, the gRNA contains sequences that are at least about 80% identical to sequence numbers 1705-1710. In some embodiments, the gRNA contains sequences that are at least about 85% identical to sequence numbers 1705-1710. In some embodiments, the gRNA contains sequences that are at least about 90% identical to sequence numbers 1705-1710. In some embodiments, the gRNA contains sequences that are at least about 91% identical to sequence numbers 1705-1710. In some embodiments, the gRNA contains sequences that are at least about 92% identical to sequence numbers 1705-1710. In some embodiments, the gRNA contains sequences that are at least about 93% identical to sequence numbers 1705-1710. In some embodiments, the gRNA contains sequences that are at least about 94% identical to sequence numbers 1705-1710. In some embodiments, the gRNA contains sequences that are at least about 95% identical to sequence numbers 1705-1710. In some embodiments, the gRNA contains sequences that are at least about 96% identical to sequence numbers 1705-1710. In some embodiments, the gRNA contains sequences that are at least about 97% identical to sequence numbers 1705-1710.In some embodiments, the gRNA contains sequences that are at least about 98% identical to sequence numbers 1705-1710. In some embodiments, the gRNA contains sequences that are at least about 99% identical to sequence numbers 1705-1710. In some embodiments, the gRNA contains sequences that are 100% identical to sequence numbers 1705-1710.
[0353] In some embodiments, the gRNA hybridizes with an AAVS1 sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with any one of the SEQ ID NOs. In some embodiments, the gRNA hybridizes with an AAVS1 sequence having at least about 75% identity with one of sequence numbers 1720-1725. In some embodiments, the gRNA hybridizes with an AAVS1 sequence having at least about 80% identity with one of sequence numbers 1720-1725. In some embodiments, the gRNA hybridizes with an AAVS1 sequence having at least about 85% identity with one of sequence numbers 1720-1725. In some embodiments, the gRNA hybridizes with an AAVS1 sequence having at least about 90% identity with one of sequence numbers 1720-1725. In some embodiments, the gRNA hybridizes with an AAVS1 sequence having at least about 91% identity with one of sequence numbers 1720-1725. In some embodiments, the gRNA hybridizes with an AAVS1 sequence having at least about 92% identity with one of sequence numbers 1720-1725. In some embodiments, the gRNA hybridizes with an AAVS1 sequence that has at least approximately 93% identity with one of sequence numbers 1720-1725.In some embodiments, the gRNA hybridizes with an AAVS1 sequence having at least about 94% identity with one of sequence numbers 1720-1725. In some embodiments, the gRNA hybridizes with an AAVS1 sequence having at least about 95% identity with one of sequence numbers 1720-1725. In some embodiments, the gRNA hybridizes with an AAVS1 sequence having at least about 96% identity with one of sequence numbers 1720-1725. In some embodiments, the gRNA hybridizes with an AAVS1 sequence having at least about 97% identity with one of sequence numbers 1720-1725. In some embodiments, the gRNA hybridizes with an AAVS1 sequence having at least about 98% identity with one of sequence numbers 1720-1725. In some embodiments, the gRNA hybridizes with an AAVS1 sequence having at least about 99% identity with one of sequence numbers 1720-1725. In some embodiments, the gRNA hybridizes with an AAVS1 sequence that has 100% identity with any one of sequence numbers 1720-1725.
[0354] In some embodiments, the target gene is hApoA1. In some embodiments, the gRNA contains a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with SEQ ID NO: 1704. In some embodiments, the gRNA contains a sequence having at least about 70% identity with SEQ ID NO: 1704. In some embodiments, the gRNA contains a sequence having at least about 75% identity with SEQ ID NO: 1704. In some embodiments, the gRNA contains a sequence having at least about 80% identity with SEQ ID NO: 1704. In some embodiments, the gRNA contains a sequence having at least about 85% identity with SEQ ID NO: 1704. In some embodiments, the gRNA contains a sequence having at least about 90% identity with SEQ ID NO: 1704. In some embodiments, the gRNA contains a sequence having at least about 91% identity with SEQ ID NO: 1704. In some embodiments, the gRNA contains a sequence having at least about 92% identity with SEQ ID NO: 1704. In some embodiments, the gRNA contains a sequence having at least about 93% identity with SEQ ID NO: 1704. In some embodiments, the gRNA contains a sequence having at least about 94% identity with SEQ ID NO: 1704. In some embodiments, the gRNA contains a sequence having at least about 95% identity with SEQ ID NO: 1704. In some embodiments, the gRNA contains a sequence having at least about 96% identity with SEQ ID NO: 1704. In some embodiments, the gRNA contains a sequence having at least about 97% identity with SEQ ID NO: 1704. In some embodiments, the gRNA contains a sequence having at least about 98% identity with SEQ ID NO: 1704.In some embodiments, the gRNA contains a sequence that is at least about 99% identical to sequence number 1704. In some embodiments, the gRNA contains a sequence that is 100% identical to sequence number 1704.
[0355] In some embodiments, the gRNA hybridizes with an hApoA1 sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with any one of sequence numbers 1726. In some embodiments, the gRNA hybridizes with an hApoA1 sequence having at least about 70% identity with any one of sequence numbers 1726. In some embodiments, the gRNA hybridizes with an hApoA1 sequence having at least about 75% identity with one of the sequence numbers 1726. In some embodiments, the gRNA hybridizes with an hApoA1 sequence having at least about 80% identity with one of the sequence numbers 1726. In some embodiments, the gRNA hybridizes with an hApoA1 sequence having at least about 85% identity with one of the sequence numbers 1726. In some embodiments, the gRNA hybridizes with an hApoA1 sequence having at least about 90% identity with one of the sequence numbers 1726. In some embodiments, the gRNA hybridizes with an hApoA1 sequence having at least about 91% identity with one of the sequence numbers 1726. In some embodiments, the gRNA hybridizes with an hApoA1 sequence having at least about 92% identity with one of the sequence numbers 1726. In some embodiments, the gRNA hybridizes with an hApoA1 sequence having at least about 93% identity with one of the sequence numbers 1726. In some embodiments, the gRNA hybridizes with an hApoA1 sequence having at least about 94% identity with one of the sequence numbers 1726.In some embodiments, the gRNA hybridizes with an hApoA1 sequence having at least about 95% identity with any one of the sequence numbers 1726. In some embodiments, the gRNA hybridizes with an hApoA1 sequence having at least about 96% identity with any one of the sequence numbers 1726. In some embodiments, the gRNA hybridizes with an hApoA1 sequence having at least about 97% identity with any one of the sequence numbers 1726. In some embodiments, the gRNA hybridizes with an hApoA1 sequence having at least about 98% identity with any one of the sequence numbers 1726. In some embodiments, the gRNA hybridizes with an hApoA1 sequence having at least about 99% identity with any one of the sequence numbers 1726. In some embodiments, the gRNA hybridizes with an hApoA1 sequence having 100% identity with any one of the sequence numbers 1726.
[0356] In this specification, in one embodiment, a method for modifying a nucleic acid encoding ANGPTL3 is described, comprising contacting the nucleic acid encoding ANGPTL3 with a manipulated base editing system, wherein the base editing system is a base editor having at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% identity with any one of sequence numbers 1654-1703 and 2021-2023, the sequence being sequence numbers 1128-1 The invention comprises a base editor which does not contain any of the sequences selected from 160 and 1363-1415, or a base editor encoded by a nucleic acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% identity with any one of sequence numbers 1727-1757, and an engineered guide polynucleotide which includes a spacer sequence that forms a complex with the endonuclease of the base editor and hybridizes with the target nucleic acid sequence. In some embodiments, the engineered guide polynucleotide includes a sequence which has at least 80% sequence identity with any one of sequence numbers 1479-1483 and 1758-1889.
[0357] In this specification, in one embodiment, a method for modifying a nucleic acid encoding APOA1 is described, comprising contacting the nucleic acid encoding APOA1 with a manipulated base editing system, wherein the base editing system is a base editor having at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% identity with any one of sequence numbers 1654-1703 and 2021-2023, the sequence being sequence numbers 1128-116 The invention comprises a base editor which does not contain any of the sequences selected from 0 and 1363-1415, or a base editor encoded by a nucleic acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% identity with any one of sequence numbers 1727-1757, and an engineered guide polynucleotide which includes a spacer sequence that forms a complex with the endonuclease of the base editor and hybridizes with the target nucleic acid sequence. In some embodiments, the engineered guide polynucleotide includes a sequence which has at least 80% sequence identity with any one of sequence numbers 1431-1454, 1704, and 2010-2019.
[0358] In this specification, in one embodiment, a method for modifying a nucleic acid encoding BCL11A is described, comprising contacting the nucleic acid encoding BCL11A with a manipulated base editing system, wherein the base editing system is a base editor having at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% identity with any one of sequence numbers 1654-1703 and 2021-2023, the sequence being sequence numbers 1128-11 The invention comprises a base editor which does not contain any of the sequences selected from 60 and 1363-1415, or a base editor encoded by a nucleic acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% identity with any one of sequence numbers 1727-1757, and an engineered guide polynucleotide which includes a spacer sequence that forms a complex with the endonuclease of the base editor and hybridizes with the target nucleic acid sequence. In some embodiments, the engineered guide polynucleotide includes a sequence which has at least 80% sequence identity with any one of sequence numbers 1890-1976.
[0359] In some embodiments, the endonuclease induces single-strand or double-strand breaks at or near the target locus. In some embodiments, the endonuclease alternately induces single-strand breaks within or 5′ to the aforementioned target locus. In some embodiments, the endonuclease does not induce breaks at or near the target locus.
[0360] In some embodiments, the disclosure provides methods for manufacturing or producing a base editor. In some embodiments, the method comprises culturing cells. In some embodiments, the method for producing a base editor comprises culturing the host cells described herein in a suitable growth medium. In some embodiments, the method further comprises inducing the expression of the aforementioned base editor by adding an additional chemical agent or an increased amount of a nutrient. In some embodiments, the chemical agent is isopropyl β-D-1-thiogalactopyranoside (IPTG). In some embodiments, the nutrient is lactose. In some embodiments, the method further comprises isolating the aforementioned host cells after the aforementioned culture and lysing the aforementioned host cells to produce a protein extract. In some embodiments, the method further comprises subjecting the aforementioned protein extract to IMAC, or ion affinity chromatography. In some embodiments, the method further comprises cleaving the aforementioned IMAC affinity tag by contacting the aforementioned base editor with a protease corresponding to the aforementioned protease cleavage site. In some embodiments, the method further comprises performing subtractive IMAC affinity chromatography to remove the aforementioned affinity tag from the composition containing the aforementioned base editor.
[0361] The systems of this disclosure can be used for a variety of applications, such as nucleic acid editing (e.g., gene editing) and binding to nucleic acid molecules (e.g., sequence-specific binding). Such systems can be used, for example, to address (e.g., remove or replace) genetically inherited mutations that may cause disease in a subject; to inactivate genes to confirm gene function in cells; as a diagnostic tool to detect disease-causing genetic elements (e.g., via cleavage of reverse-transcribed viral RNA or amplified DNA sequences encoding disease-causing mutations); as an inactivating enzyme combined with a probe to target and detect specific nucleotide sequences (e.g., sequences encoding antibiotic resistance in bacteria); to inactivate viruses or prevent them from infecting host cells by targeting viral genomes; to manipulate organisms to produce valuable small, large, or secondary metabolites by adding genes or modifying metabolic pathways; to establish gene-driven elements for evolutionary selection; and as a biosensor to detect cellular perturbations by foreign small molecules and nucleotides.
[0362] kit In some embodiments, the Disclosure provides a kit comprising one or more nucleic acid constructs encoding various components of the manipulated system described herein, for example, a nucleotide sequence encoding a component of the manipulated editing system capable of modifying a target DNA sequence. In some embodiments, the nucleotide sequence includes a heterologous promoter that drives the expression of the manipulated system component.
[0363] In some embodiments, any of the operated editing systems disclosed herein are incorporated into a pharmaceutical, diagnostic, or research kit to facilitate its use in therapeutic, diagnostic, or research applications. The kit may include one or more containers containing any of the vectors disclosed herein, and instructions for use.
[0364] The kit may be designed to facilitate researchers' use of the methods described herein and may take various forms. Each of the components of the kit may be provided in liquid form (e.g., in solution) or solid form (e.g., dry powder), where applicable. In certain cases, some of the components may be configured to be (e.g., in an active form) or otherwise processable by the addition of a suitable solvent or other kind (e.g., water or cell culture medium), which may or may not be provided with the kit. Where used herein, “instructions” defines the components of the instructions and / or promotional materials and may typically be accompanied by written instructions on or associated with the packaging of the herein. Instructions may also include any oral or electronic instructions provided in any format that clearly indicates to the user that the instructions relate to the kit, such as audiovisual (e.g., videotape, DVD, etc.), the internet, and / or web-based communications. In some embodiments, written instructions may be in the form prescribed by a government agency that regulates the manufacture, use, or sale of a pharmaceutical or biological product, and such instructions may also reflect approval by an agency for manufacture, use, or sale for animal administration. [Examples]
[0365] The following embodiments are given for the purpose of illustrating various embodiments of the Disclosure and are not intended to limit the Disclosure in any way. These embodiments, along with the methods described herein, represent preferred embodiments at present and are illustrative and are not intended to limit the scope of the Disclosure. Modifications and other uses therein, which are included within the spirit of the Disclosure as defined by the claims, will be recalled by those skilled in the art.
[0366] Example 1 - The base editor ABE07 targets the APOA1 and ANGPTL3 genes in primary mouse hepatocytes at various doses. Hypercholesterolemia is a metabolic disorder characterized by elevated plasma levels of low-density lipoprotein (LDL), which can lead to atherosclerosis, heart attack, and stroke. The proteins that regulate plasma lipoprotein levels are primarily expressed in hepatocytes of the liver and are encoded by the APOA1 and ANGPTL3 genes. Knockdown of these genes may enable the treatment of human lipoprotein metabolic disorders such as hypercholesterolemia, and this can be achieved by precisely introducing mutations into their coding sequences through base editing.
[0367] In this example, eleven guides targeting the mouse APOA1 gene and one guide targeting the ANGPTL3 gene (Table 3) were co-transfected with ABE07 mRNA (SEQ ID NO: 1411) via lipofection in primary mouse hepatocytes. A range of mRNA doses was administered, and conversion from A to G was induced in the eleven guides across four mRNA dose ranges (Figures 1A-1K).
[0368] Primary cell culture, transfection, next-generation sequencing, and base editing activity analysis
[0369] Primary mouse hepatocytes were seeded at 100,000 viable cells per well in collagen I precoated plates. After incubation overnight for 24 hours at 37°C and 5% CO2 in hepatocyte basal medium, ABE07 mRNA (0.21, 0.42, 0.625, 1.25 μg) was co-transfected with mRNA:guide RNA in a 1:20 molar ratio via lipofection. Genomic DNA was collected from cells three days after transfection. PCR primers suitable for use in NGS-based DNA sequencing were generated and optimized and used to amplify the individual target sequences of each guide RNA. The amplicons were sequenced and analyzed to measure the conversion rate from A to G at the relevant spacer position for each guide.
[0370] [Table 3-1]
[0371] [Table 3-2]
[0372] result
[0373] NGS analysis of genomic DNA isolated from primary mouse hepatocytes three days after transfection revealed the base editing activity of eleven APOA1 guides (A5, A8, C5, D7, D11, E1, F4, and F12) and one ANGPTL3 guide (C12). A maximum A-to-G conversion of 70.3% was achieved at spacer position A11 with the lowest mRNA dose of 0.21 μg in ABE07 and the transfected APOA1 guide F12 (Figure 1A-1K). The second lowest mRNA dose of 0.42 μg also resulted in a maximum A-to-G conversion of 64.4% in APOA1 guide F12 and at spacer position A11. The two highest mRNA doses, 0.625 and 1.25 μg, resulted in up to 55.0% and 35.9% A-to-G conversion, respectively, at spacer position A5 in APOA1 guide F4.
[0374] Maximum A-to-G conversion was frequently observed towards the center of the mRNA dose range within a given spacer position, resulting in an inverted U-shaped dose-effect curve. This shape is well demonstrated at spacer position A7 in the APOA1 G6 guided results graph (Figure 1A-1K), where mRNA doses of 0.21, 0.42, 0.625, and 1.25 μg yielded A-to-G conversions of 47.4%, 54.6%, 54.4%, and 35.9%, respectively. Capturing the range of base editing activity within spacer positions with relatively small deviations between data points at a given dose suggests a well-characterized, robust base editing tool ready for downstream optimization and in vivo delivery.
[0375] Example 2 - Manipulation and optimization of ABE variants by extensive guided screening in Hepa1-6 cells The DV83S, T112R, H129N, and A155R mutations, along with the D109N mutation, were inserted into the 3-68_DIV30_HT / HM nickases chassis (where 3-68, DIV30, and HT / HM represent MG3-6 / 3-8 nickases, domain-integrated version 30, and heterodimer / homodimer, respectively) to generate novel dimeric ABE variants (SEQ ID NOs. 1654-1658). In the homodimer configuration, two copies of the mutant were fused with a linker, while in the heterodimer, the wild-type deaminase was fused to the beneficial mutant via the linker (Figure 2). Further details of these variants are summarized in Table 4.
[0376] In this example, these ABE variants were screened across 31 pre-characterized loci in the Hepa1-6 cell line, along with the homodimeric D109N variant, ABE07 (SEQ ID NO: 1411), and the heterodimeric D109N variant, ABE07-74 (SEQ ID NO: 1654). (Table 5; SEQ ID NOs: 1455-1478 and 1484-1488).
[0377] mRNA production, cell culture, transfection, next-generation sequencing, and base editing analysis for ABE screening
[0378] mRNA corresponding to the ABE variants listed in Table 4 was produced. After mRNA production, these ABE variants were nucleofected into Hepa1-6 cells along with chemically synthesized sgRNAs targeting the gene loci listed in Table 5. The amplicons were sequenced and analyzed to measure gene editing.
[0379] [Table 4]
[0380] result
[0381] Editing across all spacers and all ABE variants averaged 11.19%, with the ABE07 homodimer construct outperforming its heterodimer homolog ABE07-74 (Figure 3A, Table 5). However, the incorporation of additional mutations in ABE07-75, ABE07-76, ABE07-77, and ABE07-78 enabled these heterodimer variants to function comparably to the less evolved homodimer ABE07 (Figure 3A, Table 6). Specifically, comparing the largest A-to-G base conversions observed across different genomic targets, ABE07-77 (36.13% across all 31 guides tested) and ABE07-78 (36.66% across all 31 guides tested) showed comparable editing efficiency to the ABE07 variant (36.88% across all 31 guides tested; Figure 3B, Table 6). Furthermore, these variants performed slightly better than ABE07 when comparing their average editing efficiency across the entire guide panel: ABE07 (10.36% across all 31 guides tested) versus ABE07-77 (13.07% across all 31 guides tested) and ABE07-78 (14.22% across all 31 guides tested; Figure 3A, Table 6).
[0382] [Table 5-1]
[0383] [Table 5-2]
[0384] [Table 5-3]
[0385] [Table 6-1]
[0386] [Table 6-2]
[0387] By comparing the most effective heterodimer ABE variants, we determined that ABE07-78 introduced C-to-G editing across the different guides tested. This indiscriminate editing by the ABE07-78 variant ranged from an average of 5% to a maximum of 27.78% at one of the sites (Figures 4A-4B). Furthermore, ABE07-78 also introduced indels (insertions and deletions) at a much higher rate (up to 7% at specific sites) than other ABE variants (Figure 5).
[0388] Therefore, based on its high on-target A-to-G editing efficiency (Figures 3A-3B) and favorable off-target editing activity (Figures 4A-4B and 5), the ABE07-77 variant with the D109N+T112R+A155R mutation combination in the heterodimer structure showed the highest levels of efficiency and specificity compared to other tested variants.
[0389] Example 3 - In vivo gene editing in mouse liver using ABE07-77 delivered by systemic administration of lipid nanoparticles ABE07-77 was subjected to further in vivo testing by lipid nanoparticle delivery of sgRNAs targeting highly edited gene loci identified in protein-coding mRNA and Hepa1-6 screening experiments and primary mouse hepatocytes (Table 7). To understand any undesirable indel formation by ABE07-77, its parent nuclease, MG3-6 / 3-8, was also included in this study as a control. Furthermore, dose-escalation experiments were performed to assess the dose-dependence of ABE07-77's editing activity. Finally, numerous chemical modifications of the native RNA structure were incorporated into these sgRNAs (Table 7). These chemical modifications were selected based on their ability to improve the in vitro stability of the sgRNAs when incubated in mammalian cell-derived extracts without adversely affecting editing activity.
[0390] mRNA preparation
[0391] mRNA encoding ABE07-77 was generated by in vitro transcription of a linearized plasmid template using T7 RNA polymerase, nucleotides, and enzymes. The DNA sequence transcribed to RNA contained the following elements in 5' to 3' order: T7 RNA polymerase promoter, 5' untranslated region (5'UTR), nuclear localization signal, short linker, ABE07-77 coding sequence, short linker, nuclear localization signal, 3' untranslated region, and a poly-A tail of approximately 100 nucleotides.
[0392] The protein sequence encoded within the synthetic mRNA encoded in this ABE07-77 cassette contained the following elements from 5' to 3': a nuclear localization signal derived from SV40, a five-amino acid linker (GGGGS), the ABE07-77 protein coding sequence with the start methionine codon removed, a three-amino acid linker (SGG), and a nuclear localization signal derived from nucleoplasmin. The DNA sequence of the protein coding region of this cassette was modified using a commercially available algorithm to reflect codon usage in humans. A poly(A) tail of approximately 100 nucleotides was encoded in the plasmid used for in vitro transcription, and the mRNA was co-transcribed and capped. Uridine in the mRNA was replaced with N1-methylpseudridine. A similar procedure was performed to prepare mRNA encoding the MG3-6 / 3-8 nuclease.
[0393] Preparation of lipid nanoparticles
[0394] The lipid nanoparticle (LNP) formulation used to deliver ABE07-77 mRNA and guide RNA is based on LNP formulations described in the literature, including Kauffman et al. (Nano Lett. 2015, 15, 11, 7300-7306). A lipid working mix was prepared by dissolving four lipid components in ethanol and mixing them in appropriate molar ratios. mRNA and guide RNA were mixed in a 1:1 mass ratio before formulation. An RNA working stock was prepared by diluting the RNA in 100 mM sodium acetate (pH 4.0). The lipid working stock and RNA working stock were mixed in a microfluidic apparatus at a flow ratio of 1:3 and a flow rate of 12 mL / min. The LNPs were dialyzed against phosphate-buffered saline (PBS) for 2 hours and then concentrated until a reduced volume was achieved. The concentration of RNA in the LNP formulation was measured using Ribogreen reagent. The diameter and polydispersity (PDI) of the LNPs were determined by dynamic light scattering. Typical LNP diameters ranged from 65 nm to 120 nm, and PDI values ranged from 0.05 to 0.20.
[0395] Mouse administration and collection
[0396] mRNA and sgRNA LNPs were mixed in a 1:1 mass ratio and intravenously injected into 7-week-old C57Bl6 wild-type mice via the tail vein (0.1 mL per mouse) at a total RNA dose of either 1.5 mg or 1 mg of RNA per kg of body weight (Table 7). Seven days after administration, all mice in each group were sacrificed. The left lobe of the liver was collected and rapidly frozen.
[0397] [Table 7-1]
[0398] [Table 7-2]
[0399] Genomic DNA preparation and editing analysis using next-generation sequencing (NGS).
[0400] The left lobe of the liver (100 mg) was homogenized in digestion buffer. Genomic DNA was purified from the resulting homogenate and quantified by measuring the absorbance at 260 nm. Genomic DNA purified from mice injected with PBS buffer alone was used as a control. Regions of the APOA1 and ANGPTL3 genes targeted by each specific sgRNA were PCR-amplified for a total of 29 cycles using DNA polymerase and gene-specific primers with adapters complementary to the barcoded primers used for next-generation sequencing (NGS). The product of this first PCR reaction was PCR-amplified for a total of 10 cycles using barcoded primers for NGS. The resulting products were subjected to NGS, and the results were processed to generate sequencing readout percentages containing insertions or deletions (indels) at the target sites in the APOA1 and ANGPTL3 genes.
[0401] result
[0402] NGS analysis of mouse liver genomic DNA showed up to 35% A-to-G conversion in mice treated with LNPs containing ABE07-77 mRNA and mApoa1 A4 gRNA (Figure 7). Furthermore, target sites mApoa1 A5 and mApoa1 C5 showed average A-to-G editing activity exceeding 5%. Indel formation by ABE07-77 remained low at all sites, particularly in comparison to MG3-6 / 3-8 nucleases, which showed up to 80% indel formation activity at some target sites (Figure 8). The ABE07-77 base editor maintained its relatively large editing window in vivo because multiple adenines within the protospacer were edited (Figure 9). Except for slight differences between treatment groups, ABE07-77 demonstrated robust and reproducible editing results.
[0403] Example 4 - Mammalian editing activity of manipulated CDA as CBE To test the activity of manipulated CDA variants, we devised a manipulated cell line with five consecutive PAMs compatible with MG3-6 and Cas9. This cell line allows gRNA tiling to test editing efficiency and find the preference for -1nt of candidate CDAs.
[0404] To test the engineered CDA, CDA was cloned in a plasmid backbone containing MG3-6 and MG uracilglycosylase inhibitors. CDA was cloned at the N-terminus (SEQ ID NOs: 1659-1664). Once the cloning of variant CDA was confirmed, they were transiently transfected into engineered HEK293T cells using lipofectamine 2000 with the appropriate guide. In the gRNA tiling experiment described above, a total of six manipulated variants were tested: 139-52-V2 (SEQ ID NO: 1271), 139-52-V13 (SEQ ID NO: 1282), 139-52-V14 (SEQ ID NO: 1283), 139-52-V17 (SEQ ID NO: 1314), 139-86v12 (SEQ ID NO: 1296), and 152-6v13 (SEQ ID NO: 1309). Of the six manipulated CDAs, all showed editing activity higher than 9% (Figure 10). Editing activity was compared to a positive control (a known high-activity CDA). When normalized for each experimental condition compared to A:A0A2K5RDN7, five of the six candidates were observed to show higher activity than the A0A2K5RDN7 hyperactivity-positive control (Figure 10). 139-52-V2, 139-52-V13, 139-52-V14, 139-52-V17, 139-86v12, and 152-6v13 were edited to 159.2%, 102.1%, 353.7%, 137.8%, 76.3%, and 309.1% of the maximum A0A2K5RDN7 edit, respectively.
[0405] To characterize -1nt preference, six manipulated target candidates were selected (139-52-V2, 139-52-V13, 139-52-V14, 139-52-V17, 139-86v12, and 152-6v13). -1nt mammalian cell preference was calculated by selecting the top four modified cytosines per guide RNA and calculating the ratio per -1 position. In this analysis, only cytosines with >1% editing were considered. The average ratio for all five guides was plotted. In vitro -1nt preference was plotted by calculating the sum of cleavage rates per -1nt preference (cleavage rate measures deamination rate) and then calculating the ratio per -1 nucleotide. Mammalian cell and in vitro -1nt preference is shown in Figure 11. Notably, different CDA families tend to have different -1nt preferences, and their preferences tend to be conserved among proteins belonging to the same family. For example, candidates have different -1nt preferences: 152-6WT (SEQ ID NO: 1322) prefers T at the -1 position, while 139-52 (WT, SEQ ID NO: 1325; and the manipulated variant) has a strong preference for C at the -1 position. Having tighter nt preferences improves off-target activity, so it is preferable to have candidates with strong -1nt preferences. Candidates with different and strong -1nt preferences allow targeting different loci without the risk of high off-target activity. Candidates are identified by purine preference: 139-86, which is predominantly G and / or A.
[0406] All of these factors allowed us to characterize engineered CDA candidates that act as CBEs in mammalian cells. Six variants derived from three native active CDAs (MG139-52 (SEQ ID NO: 1325), MG139-86 (SEQ ID NO: 810), and MG152-6 (SEQ ID NO: 1322)) were tested, each possessing different -1nt preferences and varying levels of activity.
[0407] Example 5 - Optimization of ABE structure In the above examples, either the homodimer (ABE07) or heterodimer (ABE07-77, hereafter referred to as ABE15) structure (SEQ ID NOs. 1411 and 1657) was utilized. These dimer states were selected to mimic the natural oligomerization state of the deaminase domain. To reduce the size of the base editor and to understand the function of each deaminase subunit in relation to the base editor, several novel ABE variants (Table 9) were constructed, which were either monomeric (ABE01, ABE02, ABE55), homodimer (ABE54, ABE58), or heterodimer variants (ABE51, ABE52, ABE57). Furthermore, in these dimeric variants, either the N-terminal or C-terminal deaminase domain was selectively inactivated by disrupting the catalytic site with the E60A mutation (ABE51, ABE52, ABE54, and ABE57) to evaluate which of the two deaminase copies is important for ABE's DNA editing ability. These variants were tested using the protocol described below on a panel of pre-characterized target sites across the mouse APOA1 and ANGPTL3 genes (Table 8).
[0408] Production of guide RNA and mRNA
[0409] Table 8 lists the guide sequences used to screen for the activity of novel ABE variants. These were selected based on the editing results observed for ABE07-77. The first and last three bases of the guide sequences were modified with 2'-O-methyl and phosphorothioate groups. mRNA corresponding to the ABE variants listed in Tables 9-12 was produced.
[0410] [Table 8-1]
[0411] [Table 8-2]
[0412] [Table 8-3]
[0413] Mouse cell culture, transfection, and data analysis
[0414] Hepa1-6 cells were obtained from ATCC. Hepa1-6 cells were nucleofected with 500 ng of mRNA and 150 pmol of chemically synthesized sgRNA using an electroporator. Each nucleofect reaction yielded 100,000 cells. Immediately after transfection, cells were cultured for three days at 37°C in 5% CO2 in DMEM supplemented with 10% FBS and 1XMEM containing non-essential amino acids. Genomic DNA was collected. Target sequences were amplified with NGS primers. Amplicons approximately 250 bp in length were checked by gel, sequenced by NGS, and computationally analyzed to measure the gene editing results.
[0415] result
[0416] MG68-4 deaminase (SEQ ID NO: 386) is a putative tRNA adenosine deaminase that functions naturally as an obligate dimer, with the dimer construct being superior to its monomeric form (Figures 12A and 12B). When comparing D109N variants, the homodimer ABE07 (SEQ ID NO: 1411) was superior to the monomer ABE01 (SEQ ID NO: 1410) with the MG68-4(D109N) variant incorporated into the RuvC-III domain, as well as to ABE02 (SEQ ID NO: 1665) with the MG68-4(D109N) variant incorporated into the REC domain. Adding T112R and A155R mutations to the D109N mutation appears to enhance monomer activity, and a similar trend is also observed for triplet mutations, with the homodimer ABE58 (SEQ ID NO: 1666) being superior to the monomer ABE55 (SEQ ID NO: 1667, Figures 12A and 12B).
[0417] Comparing selectively inactivated ABE variants ABE51, ABE52, ABE54, and ABE57 (SEQ ID NOs. 1668–1671), it was observed that only the C-terminal deaminase subunit was actually involved in genomic DNA editing, with the N-terminal subunit merely assisting the activity of the C-terminal subunit. This was demonstrated by ABE51, which has a C-terminal inactivated deaminase domain that functions comparably to a completely inactivated ABE54, clearly indicating that the N-terminal subunit is not precisely positioned to access exposed DNA. In contrast, the N-terminal inactivated variants ABE52 and ABE57 showed higher editing activity than monomeric and sometimes homodimeric ABEs (Figures 12A–12B, and 13A–13B).
[0418] In conclusion, only the C-terminal deaminase domain is involved in DNA editing, and it may be possible to compensate for the lack of the heterodimer deaminase domain by introducing beneficial mutations into the monomeric variant.
[0419] [Table 9]
[0420] Example 6 - Operation of ABE with Enhanced On-Target Editing Next, the editing activity of the heterodimer ABE15 variant (SEQ ID NO: 1657) was improved by incorporating an additional mutation into the C-terminal deaminase subunit. Based on combined mutagenesis screening, ABE variants ABE32–ABE41 (Table 10, SEQ ID NOs: 1672–1681) were generated and screened across the guides listed in Table 8 using the protocol described in Example 5.
[0421] result
[0422] Addition of the L85F mutation to the existing D109N+T112R+A155R mutation in ABE33 (SEQ ID NO: 1665) was found to result in the most significant increase in both mean and maximum A:T to G:C editing activity. Specifically, ABE33 was observed to show a mean A:T to G:C editing of 32.97% across 15 guides tested, which is 1.5 times higher than ABE15, which has an average A:T to G:C editing activity of approximately 22.8% (Figures 14A-14B). Furthermore, ABE33 showed lower indiscriminate C deamination compared to ABE15 (Figures 15A-15B), but also lower indel formation activity (Figure 16).
[0423] Unlike the addition of L85F, the incorporation of A143W (ABE34), A155E (ABE35), E10Y (ABE37), and A126D (ABE40) was determined to result in a reduction of A:T to G:C editing. The addition of V83S (ABE32) significantly increased the undesirable indiscriminate C deamination and indel formation activity of ABE without increasing A:T to G:C editing activity (Figures 14A-14B, 15A-15B, and 16).
[0424] Overall, it was found that the L85F mutation can be successfully added to the ABE15 construct to enhance the editing properties of the resulting ABE33 variant.
[0425] [Table 10]
[0426] Example 7 - Operation of ABE with a relaxed editing context While analyzing the editing activity of ABE15 across 15 guide panels, even the mutated MG68-4 variant showed the same native -U enzyme present in the native substrates of other tRNA adenosine deaminase enzymes. A CG - Imitating motifs - N AIt was found that it has an inherent preference for C-motifs. In particular, the guides tested (Table 8) showed a large number of target adenines (A) adjacent to guanine (Gs) at both +1 and -1 positions (Figure 17A), but ABE15 tended to preferentially edit A adjacent at +1C (Figure 17B) at a much faster rate than A adjacent at +1G (Figure 17C).
[0427] To suppress this sequence preference and broaden the editing context of ABE, we designed a series of proline variants that allow flexibility in the terminal alpha-helix of the deaminase domain (Table 11, SEQ ID NOs. 1682-1685), and screened them using the protocol described in Example 5, according to the guides listed in Table 8.
[0428] result
[0429] Of the proline variants tested, the simultaneous addition of the R153P and R154P mutations in addition to the ABE15 mutation, i.e., ABE23 (SEQ ID NO: 1682), was observed to result in an overall increase in both mean and maximum A:T to G:C editing activity without significant changes in C deamination (Figures 19A-19B) and indel formation rate (Figure 20A) (Figures 18A-18B). Specifically, ABE23 showed a mean A:T to G:C editing of 30.8% compared to 21.4% with ABE15, which was a similar increase to the mean maximum observed A:T to G:C editing of ABE23 of 82.1%, compared to 73.5% across all 15 guides tested in this study. Overall, this increase in A:T to G:C editing activity in ABE23 is evident in the relaxation of the strict +1C context compared to ABE15 (Figure 20B).
[0430] [Table 11]
[0431] Example 8 - Operation of ABE with low indifferentiation Low levels of cytidine editing were observed in almost all ABE variants tested and described in Examples 2 and 5-7 above. This indifferentiation was shown to be reduced by the introduction of the D109Q mutation.
[0432] By employing a similar strategy, the undesirable C-deamination activity of ABE was suppressed in several of the best-performing variants (Table 12) by changing the underlying D109N mutation to Q, and these were screened across the guides listed in Table 8 using the protocol described in Example 5.
[0433] Specifically, the D109Q mutation was incorporated into the top variants ABE12, ABE13, ABE14, ABE15, and ABE16 (SEQ ID NOs. 1654-1658) to obtain ABE18, ABE19, ABE20, ABE21, and ABE22 (SEQ ID NOs. 1686-1690), respectively. These mutations were also included in the proline variants tested in Example 7 to obtain ABE28, ABE29, ABE30, and ABE31 (SEQ ID NOs. 1691-1694).
[0434] result
[0435] The addition of the D109Q mutation was determined to have various effects on ABE activity. In ABE18 and ABE12, ABE19 and ABE13, and ABE21 and ABE15, replacing D109N with D109Q had no significant effect compared to their on-target A:T to G:C editing activity. However, this mutation swap resulted in a slight increase in on-target editing activity in ABE22 and ABE29, and a significant decrease in activity in ABE28, ABE30, and ABE31 (Figures 21A-21B). The D109Q mutation had little to no effect on the indiscriminate C deamination (Figures 22A-22B) and indel formation rate (Figure 23) of these ABE variants.
[0436] Overall, the operations and screening efforts described herein identified key mutations in the deaminase domain of ABE, resulting in higher levels of on-target A:T to G:C editing and a reduced rate of indiscriminate C deamination. In particular, ABE23 and ABE33 emerged as important novel ABE variants with higher on-target editing activity (Figures 24A-24B), lower undesirable C deamination rates (Figures 25A-25B), and lower undesirable indel activity (Figure 26) compared to their parent ABE15.
[0437] [Table 12-1]
[0438] [Table 12-2]
[0439] Example 9 - A novel, small base editor exhibits high activity in human cells. Small base editor construct
[0440] The manipulated MG68-4 adenine deaminase was incorporated into the C-terminus, N-terminus, or inside nickase MG102-39 or nickase MG34-29 and fused to generate adenine base editors ABE75-ABE77 and ABE87-88 (sequence numbers 1695-1699), respectively (Figure 27A and Table 8).
[0441] [Table 13]
[0442] Design of guide RNA and production of mRNA
[0443] The small ABE sequences were codon-optimized for human expression, and each was cloned into an expression vector containing a T7 promoter and capping start sequence, 5'UTR and 3'UTR, and a poly-A tail. The coding sequences contained an N-terminal SV40 nuclear localization signal and a C-terminal nucleoplasmin nuclear localization signal. The expression vectors were midiprep, linearized with SpeI, washed, and used for in vitro transcription using Hi-T7 polymerase. The in vitro transcription reaction included N1-methylpseudridine instead of uridine and capping reagent was added. The resulting mRNA was washed, and the size and purity of the product were checked by spectrophotometric and gel electrophoresis, and diluted to 250 ng / μL in sterile water for use in nucleofection. ABE33 (SEQ ID NO: 1672) was used as a positive control using the hAPOA1 target guide (Table 14).
[0444] The base-editing efficiency of these small ABEs was evaluated across target guides selected based on previously observed nuclease activity (Table 14). The first and last three bases of these guides were modified with 2'-O-methyl and phosphorothioate groups.
[0445] Mammalian cell culture, transfection, and data analysis
[0446] K562 cells (ATCC #CCL-243) were passaged 1-2 times in IMDM + L-alanine-L-glutamine dipeptide medium and 10% FBS prior to nucleofection. On the day of nucleofection, cells were harvested, counted, washed in 1XPBS, and resuspended in nucleofection buffer according to the manufacturer's instructions. 120,000 cells were distributed per well, and 500 ng of mRNA and 200 pmol of sgRNA were nucleofected using a nucleofection kit. In some experiments, the amount of guide added varied from 100 to 400 pmol. Cells were added to the harvest medium and grown for 72 hours before genomic DNA harvesting. The obtained gDNA was diluted 1:3 and used as a template for NGS PCR. Target sequences were amplified with NGS primers (IDT). An amplicon approximately 250 bp in length was checked by gel electrophoresis, sequenced using an NGS sequencing machine, and computationally analyzed to measure the gene editing results.
[0447] result
[0448] High A→G editing was observed for ABEs derived from MG34-29 across the panel of guides tested (Figure 28A). ABE87 (Figure 27C), containing MG68-4 (D109N, T112R, A155R) deaminase incorporated into MG34-29 nickase, showed the highest editing performance among all the small ABEs tested, with mean editing efficiencies of 74.34% at the AAVS1 E7 locus and 32.41% at the AAVS1 C7 locus (Figure 28B). Incorporation structure preference has also been observed in other nickases, such as MG3-6_3-8 ABE.
[0449] For the small ABE (ABE75) derived from MG102-39, the observed editing efficiency was low, but a similar preference for the embedded structure was observed (Figure 28C). ABE75 outperformed its C-terminal and N-terminal variants, exhibiting the best performance of 7.10% to 1.60% at the TRAC C11 locus (Figure 28D).
[0450] [Table 14-1]
[0451] [Table 14-2]
[0452] Example 10 - PAM interaction domain exchange increases ABE targetability Adenine base editors (ABEs) can install precise A:T→G:C edits in a programmable manner, allowing them to correct more than half of known pathogenic single nucleotide polymorphisms and efficiently knock out proteins by disrupting splice sites. However, due to their programmable nature, the usefulness of ABEs is highly dependent on the presence of appropriate PAMs downstream of the desired target adenine. In fact, SpCas9-based ABEs (such as ABE7.10) can only access 18% of all adenines present in the human reference genome due to the lack of NGGs near target A (Figure 29A).
[0453] Here, this problem is mitigated by generating a set of ABEs that collectively provide broad PAM compatibility by leveraging a PID-exchangeable MG3-6 nuclease chassis and a highly efficient adenine deaminase repository (Figure 29B). This group of chimeric ABEs can theoretically target 95% of all adenines in the human genome (Figure 29C). We demonstrate that these chimeric ABEs can be used for disruption and correction of targeted splice sites in several different therapeutically relevant contexts, achieving highly efficient base editing with several different ABEs and multiple guides on each target locus. Due to their diverse PAM compatibility, this group of chimeric ABEs may offer lifelong, permanent cures for a wide range of genetic disorders.
[0454] method
[0455] Design of guide RNA and production of mRNA
[0456] ABE was codon-optimized for human expression and subsequently cloned into an expression vector containing a CleanCap T7 promoter, 5' and 3' UTRs, and a poly-A tail. The coding sequence contained an N-terminal SV40 nuclear localization signal and a C-terminal nucleoplasmin nuclear localization signal. The expression vector was midiprep, and the mRNA template was amplified from the plasmid using Q5 High-Fidelity 2X Master Mix, washed with HighPrep PCR, and used for in vitro transcription using Hi-T7. The in vitro transcription reaction included N1-methylpseudridine instead of uridine and CleanCap reagent was added. The resulting mRNA was washed with Rneasy, and the size and purity of the product were checked by NanoDrop and Tapestation, and diluted to 250 ng / μL in sterile water for use in nucleofection.
[0457] [Table 15-1]
[0458] [Table 15-2]
[0459] The MG3-6 ABE was determined to function most efficiently as an adenine base located in the editing window spanning bases 3–13 of the protospacer (Figure 30A). Using this editing window, a guide was designed that tiled across the splice site of hANGPTL3 (SEQ ID NOs. 1758–1889), the GATA1 binding site of hBCL11A (SEQ ID NOs. 1890–1976), and the most frequently occurring SNVs in hPAH (SEQ ID NOs. 1977–2009) (Figure 30B). The first and last three bases on the guide were modified with a 2'-O methyl group and a phosphorothioate group (IDT).
[0460] High-throughput pooling and arrayed mammalian cell screening
[0461] K562 cells were obtained from ATCC. On the day of nucleofection, cells were harvested, counted, washed in 1XPBS, and resuspended in SF buffer in a 4D electroporator according to the manufacturer's instructions. To maximize the number of chimeric ABEs at each locus, a high-throughput pooled screening method was developed. A typical plate layout for this approach is shown in Figure 30C. In each well, 15 pmol of ten guides tiling a specific splice site were pooled together with 500 ng of mRNA related to a unique chimeric ABE (Figure 30B). This pooled approach was used to identify the chimeric ABE exhibiting the highest editing efficiency at each splice site. A secondary deconvolution screen was performed to determine the active guide and chimeric ABE combinations at each splice site. This was done by testing each active chimeric ABE individually by "unpooling" the test guides from the most edited guide pool (Figure 30D). 150 pmol of each test guide was added to 500 ng of chimeric ABE mRNA. Each nucleofection reaction involved 120,000 cells per guide. Immediately after transfection, cells were cultured for three days at 37°C and 5% CO2 in GlutaMAX supplemented with IMDM + 10% FBS. Genomic DNA was extracted from the cells using Quick Extract. Target sequences were amplified with NGS primers. Amplicons approximately 250 bp in length were checked on an Agilent TapeStation D1000 gel, sequenced using a MiSeq machine, and analyzed with CRISPResso2 to measure gene editing outcomes.
[0462] For some experiments, the K562 cell line was engineered with “disease” state nucleotides using a lentiviral vector encoding a combination of therapeutically relevant targets (approximately 4.5 kb, SEQ ID NO: 2020) at a low MOI. This 4.5 kb sequence contained the PAH gene with the relevant SNVs 1222C>T(p.R408W), 1066-11G>A, and 1315+1G>A, and was approximately 250 nt on both sides of the therapeutically relevant target nucleotide. The mutations corresponding to the target gene were then modified to the disease state, allowing for correction of the therapeutically relevant edit. To amplify the engineered target target more than the endogenous target, a 20 nt non-natural sequence substituted approximately 125 nt of the endogenous gene sequence at the 5' and 3' ends of the target target, providing a primer binding site resulting in a 250 nt amplicon for NGS treatment. Transduced cells were selected with puromycin 3–10 days after transduction. These stable cell lines are referred to here as engineered K562 cells. The same nucleofection approach was used to determine the deamination events at the therapeutically relevant sites described above.
[0463] result
[0464] Over 100 PID-exchanged MG3-6-based chimeras were constructed, each recognizing a unique PAM sequence. Fifteen of these chimeric nucleases were selected by evaluating the PAM targeting ability of each chimeric nuclease across the human genome and engineered into nickases for ABE construct design. This was done by inactivating their RuvC nuclease domains and inserting highly efficient adenine deaminases (mutant MG68-4[D109N, T112R, R153P, R154P, A155R] or mutant MG68-4[L85F / D109N / T112R / A155R]) to generate a series of chimeric ABEs (Table 15, SEQ ID NOs: 1727-1757).
[0465] Similar to observations made when comparing targetability across the entire genome (Figures 29A-29C), a series of PID-exchangeable MG3-6 chimeric ABEs theoretically provide more guides for disrupting splice sites (Figure 31A), as well as access to enhancer-binding sites (Figure 33A) and single nucleotide variants (Figure 34A). While SpCas9 ABEs with their NGG PAMs have limited targetability across putative targets of interest, MG3-6 ABEs can target these cumulatively with multiple possible guides. This is also true when comparing them to other well-characterized ABEs with relaxed PAM requirements, such as SaCas9, SaKKHCas9, CjCas9, Nme2Cas9, and SauriCas9. Overall, MG3-6 chimeric ABEs offer more options for guides targeting genes of interest.
[0466] hANGPTL3
[0467] ANGPTL3 is an endogenous inhibitor of lipoprotein lipase (LPL), primarily expressed in the liver and associated with cholesterol metabolism. Knockdown of ANGPTL3 has been shown to increase LPL levels, a key enzyme involved in triglyceride hydrolysis. Therefore, ANGPTL3 knockout is considered a potential permanent treatment for coronary artery disease.
[0468] In a pooled screening approach, we screened the splice sites of all seven exons of hANGPTL3 in K562 cells. We identified several chimeric ABEs that disrupted the splice sites of this gene (Figure 31B), and the deconvolution of the pooled wells showed high base editing overall (Figures 32A-32D). For example, ABE100 (SEQ ID NO: 1743), which translates to amino acid sequence ABE33 (SEQ ID NO: 1673), edited the exon 1 splice donor (Figure 32A), while ABE101, ABE103, and ABE112 (SEQ ID NOs: 1744, 1746, and 1755), which translate to amino acid sequences of SEQ ID NOs: 2021-2023, were active at the exon 4 splice acceptor and splice donor, exon 5 splice donor, and exon 7 splice acceptor (Figures 32B-32D). During deconvolution, ABE100 was observed to destroy the exon 1 splice donor site with an average editing efficiency of 45.9% (Figure 32A). ABE112 (SEQ ID NO: 1755) showed the best editing performance in ANGPTL3, with 21.9% editing in the exon 4 splice acceptor site, 61.5% editing in the exon 4 splice donor site, and 25% editing in the exon 7 splice acceptor site.
[0469] hBCL11A
[0470] To demonstrate the applicability of a chimeric ABE platform for targeted gene silencing by enhancer site disruption, we selected the BCL11A gene, known to play a crucial role in β-thalassemia, as our target. BCL11A possesses three well-characterized human BCL11A composite enhancer DHS sites (i.e., DHS+55, DHS+58, and DHS+62). These enhancer sites have consensus GATA motifs that bind to GATA1 and TAL1 enhancers, upregulating BCL11A expression. Mutations in these GATA sites may suppress BCL11A levels, which subsequently lead to increased γ-globin levels and rescue the pathogenesis of β-thalassemia. We designed all possible sgRNAs that could target these three GATA sites from both the forward and reverse strands and screened them together with chimeric ABEs using the aforementioned high-throughput pool-guided screening approach (Figure 33B). Significant editing (16.5%) was observed in guide pool editing of the DHS+55 site in ABE101 (SEQ ID NOs: 1744 and 2021) based on MG3-6_3-8_3-4, PAM NNRMWW, which correlates with theoretical predictions of the targetability of this ABE (Figure 33A). Deconvolution of the pooled guides identified two guides that edit both adenines at the GATA1 site, with mean A→G editing efficiencies of 44.81% and 24.02% across both A's (Figure 33C). DHS+55 is a critical enhancer binding site in the BCL11A gene. Disrupting this site with ABE to increase γ-globin levels is presumed to be a safer alternative to Cas9-mediated indels.
[0471] hPAH
[0472] Phenylalanine hydroxylase (PAH) deficiency results in intolerance to the dietary intake of the essential amino acid phenylalanine, causing a variety of disorders. The risk of adverse outcomes varies depending on the severity of the PAH deficiency. More than 500 mutations have been reported in the coding and intercalating sequences of the PAH gene (Regier and Greene, 2000). The best current treatment for PAH deficiency is the oral drug sapropterin, which acts as a cofactor for the PAH protein and can improve the activity of several variants of PAH. However, it is not effective against the most commonly reported SNVs in the PAH gene: 1222C>T(p.R408W), 1066-11G>A, and 1315+1G>A (Figure 34A).
[0473] We hypothesized that the chimeric ABEs disclosed herein could provide permanent and effective substitutes for sapropterin by directly correcting the causative SNVs. To test chimeric ABEs on these SNVs, we generated novel cell lines carrying these various mutations and performed a pooled ABE screening pipeline on these engineered K562 cells. The chimeric ABEs were tested with 10 guides tiled across three of the SNVs within this gene.
[0474] Chimeric ABEs did not show significant editing at the 1066-11G>A or 1315+1G>A sites (Figures 34C and 34D), however, ABE101 (SEQ ID NOs. 1744 and 2021) corrected 1222C>T(p.R408W)SNVs with a pooled efficiency of 15% in high-throughput pool screening (Figure 34B). In 1222C>T(p.R408W)SNVs, the possibility of bystander editing results due to adjacency editing was observed, similar to the editing window of a base editor.
[0475] The data presented in this embodiment highlight the strength of the PID-exchangeable chimeric ABE toolbox disclosed herein, particularly when used in conjunction with high-throughput guided screening. The series of novel ABEs boast greater targetability across therapeutically relevant genes and theoretical sites in the human genome when compared with well-characterized SpCas9 or SaCas9 ABEs. This is supported by the findings herein, as many chimeric ABEs demonstrated efficient editing with multiple guides within the same gene. Due to their broad targetability and high efficiency, the chimeric ABEs disclosed herein have the potential to introduce precise base editing to many therapeutically relevant gene loci that were typically inaccessible with previously reported base editors.
[0476] References
[0477] Regier DS,Greene CL.Phenylalanine Hydroxylase Deficiency.2000 Jan 10 [Updated 2017 Jan 5].In:Adam MP,Feldman J,Mirzaa GM,et al.,Editors.GeneReviews(R)[Internet].Seattle(WA):University of Washington,Seattle;1993-2024.
[0478] Equal parts This disclosure may be implemented in other specific forms without departing from their spirit or essential features. Therefore, the embodiments described above should be considered illustrative in all respects and not limiting the disclosures described herein. Accordingly, the scope of this disclosure is indicated by the appended claims rather than the foregoing description, and all modifications that fall within the meaning and scope of equivalence of the claims are intended to be encompassed therein.
Claims
1. A manipulated base editing system, (a) A base editor comprising a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% identity with any one of sequence numbers 1654-1703 and 2021-2023, wherein the sequence does not comprise any sequence selected from sequence numbers 1128-1160 and 1363-1415, (b) An engineered base editing system comprising an engineered guide polynucleotide, which includes an engineered guide polynucleotide that forms a complex with the endonuclease of the base editor and hybridizes with a target nucleic acid sequence.
2. The manipulated base editing system according to claim 1, wherein the base editor includes a sequence having at least 95% sequence identity with any one of sequence numbers 1654-1703 and 2021-2023.
3. The manipulated base editing system according to claim 1, wherein the base editor includes a sequence having 100% sequence identity with any one of sequence numbers 1654-1703 and 2021-2023.
4. A manipulated base editing system, (a) A base editor encoded by a nucleic acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% identity with any one of sequence numbers 1727 to 1757, (b) An engineered base editing system comprising an engineered guide polynucleotide, which includes an engineered guide polynucleotide that forms a complex with the endonuclease of the base editor and hybridizes with a target nucleic acid sequence.
5. The manipulated base editing system according to claim 4, wherein the base editor is encoded by a nucleic acid sequence having at least 90% identity with any one of sequence numbers 1727 to 1757.
6. The manipulated base editing system according to claim 4, wherein the base editor is encoded by a nucleic acid sequence having 100% identity with any one of sequence numbers 1727 to 1757.
7. The operated base editing system according to any one of claims 1 to 6, wherein the base editor comprises a deaminase.
8. The manipulated base editing system according to claim 7, wherein the deaminase is non-covalently bonded to the endonuclease.
9. The manipulated base editing system according to claim 7, wherein the deaminase is covalently bound to the endonuclease.
10. The manipulated base editing system according to claim 7, wherein the deaminase is fused to the endonuclease.
11. The manipulated base editing system according to any one of claims 1 to 10, wherein the manipulated guide polynucleotide is a single guide nucleic acid.
12. The manipulated base editing system according to any one of claims 1 to 11, wherein the manipulated guide polynucleotide is a dual guide nucleic acid.
13. The manipulated base editing system according to any one of claims 1 to 12, wherein the manipulated guide polynucleotide is RNA.
14. The manipulated base editing system according to any one of claims 1 to 13, wherein the endonuclease is non-covalently bonded to the manipulated guide polynucleotide.
15. The manipulated base editing system according to any one of claims 1 to 13, wherein the endonuclease is covalently bonded to the manipulated guide polynucleotide.
16. The manipulated base editing system according to any one of claims 1 to 15, wherein the manipulated guide polynucleotide includes a sequence having at least 80% sequence identity with any one of sequence numbers 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, 1705-1710, and 1758-2019.
17. The manipulated base editing system according to any one of claims 1 to 16, wherein the manipulated guide polynucleotide comprises a sequence having at least 90% sequence identity with any one of sequence numbers 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, 1705-1710, and 1758-2019.
18. The manipulated base editing system according to any one of claims 1 to 17, wherein the manipulated guide polynucleotide includes a sequence having 100% sequence identity with any one of SEQ ID NOs: 88-96, 488-489, 679-680, 876, 917-931, 963-967, 1099-1105, 1187-1195, 1416-1418, 1427-1428, 1431-1454, 1479-1483, 1489-1490, 1705-1710, and 1758-2019.
19. The manipulated base editing system according to any one of claims 1 to 15, wherein the manipulated guide polynucleotide comprises a sequence having at least 80%, 85%, 90%, 95%, 98%, 99%, or 100% identity with any one of SEQ ID NOs: 1431-1454, 1479-1483, 1489-1490, 1705-1710, and 1758-2019.
20. The manipulated base editing system according to claim 19, wherein the manipulated guide polynucleotide comprises a sequence having at least 80% sequence identity with any one of SEQ ID NOs: 1431-1454, 1704, and 2010-2019.
21. The manipulated base editing system according to claim 19, wherein the manipulated guide polynucleotide comprises a sequence having at least 80% sequence identity with any one of sequence numbers 1890 to 1976.
22. The manipulated base editing system according to claim 19, wherein the manipulated guide polynucleotide comprises a sequence having at least 80% sequence identity with any one of sequence numbers 1977 to 2009.
23. The manipulated base editing system according to claim 19, wherein the manipulated guide polynucleotide comprises a sequence having at least 80% sequence identity with any one of sequence numbers 1479-1483 and 1758-1889.
24. The operated base editing system according to any one of claims 1 to 23, wherein the base editor includes a nickase domain.
25. The manipulated base editing system according to claim 24, wherein the nickase includes a mutation from aspartic acid to alanine in residue 9 of SEQ ID NO: 70, residue 13 of SEQ ID NO: 71, 72, or 74, residue 12 of SEQ ID NO: 73, residue 17 of SEQ ID NO: 75, residue 23 of SEQ ID NO: 76, or residue 10 of SEQ ID NO: 597, or any combination thereof.
26. The manipulated base editing system according to any one of claims 1 to 25, wherein the base editor further comprises a uracil DNA glycosylase inhibitor sequence.
27. The operated base editing system according to any one of claims 1 to 26, wherein the base editor further comprises the FAM72A sequence.
28. The manipulated base editing system according to claim 27, wherein the FAM72A sequence has at least 80% identity with sequence number 1121.
29. A nucleic acid encoding an operated base editing system according to any one of claims 1 to 28.
30. A vector comprising the nucleic acid described in claim 29.
31. The vector according to claim 30, wherein the vector is a plasmid, a minicircle, CELiD, an adeno-associated virus (AAV) derived virion, a lentivirus, or an adenovirus.
32. A cell comprising an engineered base editing system according to any one of claims 1 to 28, a nucleic acid according to claim 29, or a vector according to any one of claims 30 to 31.
33. The cell according to claim 32, wherein the cell is a eukaryotic cell.
34. The cell according to claim 32, wherein the cell is a mammalian cell.
35. The cell according to claim 32, wherein the cell is an immortalized cell.
36. The cell according to claim 32, wherein the cell is an insect cell.
37. The cell according to claim 32, wherein the cell is a yeast cell.
38. The cell according to claim 32, wherein the cell is a plant cell.
39. The cell according to claim 32, wherein the cell is a fungal cell.
40. The cell according to claim 32, wherein the cell is a prokaryotic cell.
41. The cell according to claim 32, wherein the cell is A549, HEK-293, HEK-293T, BHK, CHO, HeLa, MRC5, Sf9, Cos-1, Cos-7, Vero, BSC1, BSC40, BMT10, WI38, HeLa, Saos, C2C12, L cell, HT1080, HepG2, Huh7, K562, primary cell, or a derivative thereof.
42. The cell according to claim 32, wherein the cell is an engineered cell.
43. The cell according to claim 32, wherein the cell is a stable cell.
44. A method for modifying a target nucleic acid sequence, comprising contacting the target nucleic acid sequence using an operated base editing system according to any one of claims 1 to 29.
45. The method according to claim 44, wherein modifying the target nucleic acid sequence includes converting adenine to guanine within the target nucleic acid sequence.
46. The method according to claim 44, wherein modifying the target nucleic acid sequence includes converting cytosine to uracil within the target nucleic acid sequence.
47. The method according to any one of claims 44 to 45, wherein the target nucleic acid sequence comprises deoxyribonucleic acid (DNA).
48. The method according to any one of claims 44 to 47, wherein the target nucleic acid sequence includes ribonucleic acid (RNA).
49. The method according to any one of claims 44 to 48, wherein the target nucleic acid sequence includes genomic DNA, viral DNA, viral RNA, or bacterial DNA.
50. The method according to any one of claims 44 to 49, wherein the target nucleic acid sequence is modified in vitro.
51. The method according to any one of claims 44 to 49, wherein the target nucleic acid sequence is modified in vivo.
52. The method according to any one of claims 44 to 49, wherein the target nucleic acid sequence is modified ex vivo.
53. The method according to any one of claims 44 to 52, wherein the target nucleic acid sequence is modified in a cell.
54. The method according to claim 53, wherein the cell is a prokaryotic cell, bacterial cell, eukaryotic cell, fungal cell, plant cell, animal cell, mammalian cell, rodent cell, primate cell, human cell, or primary cell.
55. A method for modifying a nucleic acid encoding ANGPTL3, comprising contacting a nucleic acid sequence encoding ANGPTL3 with a manipulated base editing system, wherein the base editing system is a) A base editor comprising a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% identity with any one of sequence numbers 1654-1703 and 2021-2023, wherein the sequence does not include any of the sequences selected from sequence numbers 1128-1160 and 1363-1415, or a base editor encoded by a nucleic acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% identity with any one of sequence numbers 1727-1757, b) A method comprising an engineered guide polynucleotide comprising a spacer sequence that forms a complex with the endonuclease of the base editor and hybridizes with a target nucleic acid sequence.
56. The method according to claim 55, wherein the manipulated guide polynucleotide comprises a sequence having at least 80% sequence identity with any one of SEQ ID NOs: 1479-1483 and 1758-1889.
57. A method for modifying a nucleic acid encoding APOA1, comprising contacting a nucleic acid sequence encoding APOA1 with a manipulated base editing system, wherein the base editing system is a) A base editor comprising a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% identity with any one of sequence numbers 1654-1703 and 2021-2023, wherein the sequence does not include any of the sequences selected from sequence numbers 1128-1160 and 1363-1415, or a base editor encoded by a nucleic acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% identity with any one of sequence numbers 1727-1757, b) A method comprising an engineered guide polynucleotide comprising a spacer sequence that forms a complex with the endonuclease of the base editor and hybridizes with a target nucleic acid sequence.
58. The method according to claim 57, wherein the manipulated guide polynucleotide comprises a sequence having at least 80% sequence identity with any one of SEQ ID NOs: 1431-1454, 1704, and 2010-2019.
59. A method for modifying a nucleic acid encoding BCL11A, comprising contacting a nucleic acid sequence encoding BCL11A with a manipulated base editing system, wherein the base editing system is a) A base editor comprising a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% identity with any one of sequence numbers 1654-1703 and 2021-2023, wherein the sequence does not include any of the sequences selected from sequence numbers 1128-1160 and 1363-1415, or a base editor encoded by a nucleic acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% identity with any one of sequence numbers 1727-1757, b) A method comprising an engineered guide polynucleotide comprising a spacer sequence that forms a complex with the endonuclease of the base editor and hybridizes with a target nucleic acid sequence.
60. The method according to claim 59, wherein the manipulated guide polynucleotide comprises a sequence having at least 80% sequence identity with any one of sequence numbers 1890 to 1976.
61. A method for modifying a nucleic acid encoding a PAH, comprising contacting a nucleic acid sequence encoding a PAH with a manipulated base editing system, wherein the base editing system is a) A base editor comprising a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% identity with any one of sequence numbers 1654-1703 and 2021-2023, wherein the sequence does not include any of the sequences selected from sequence numbers 1128-1160 and 1363-1415, or a base editor encoded by a nucleic acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% identity with any one of sequence numbers 1727-1757, b) A method comprising an engineered guide polynucleotide comprising a spacer sequence that forms a complex with the endonuclease of the base editor and hybridizes with a target nucleic acid sequence.
62. The method according to claim 61, wherein the manipulated guide polynucleotide comprises a sequence having at least 80% sequence identity with any one of sequence numbers 1977 to 2009.