Engineered V-type RNA programmable endonucleases and uses thereof
By engineering type V endonucleases, particularly by introducing non-aspartic amino acids at the D504 or D501 positions, and by improving the oligonucleotide binding domain (OBD), the problem of insufficient gene editing efficiency in eukaryotic systems of existing CRISPR-Cas systems has been solved, achieving highly efficient gene editing without the need for tracrRNA.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-04
- Publication Date
- 2026-05-01
AI Technical Summary
Existing CRISPR-Cas systems have shortcomings in gene editing efficiency, especially the poor performance of type V endonucleases in eukaryotic genome editing, and some systems require the assistance of tracrRNA, which increases complexity.
By engineering type V endonucleases, particularly by introducing non-aspartic amino acids, such as arginine, at the D504 or D501 position, and by improving the oligonucleotide binding domain (OBD), engineered type V endonucleases have been developed, simplifying system requirements and improving gene editing efficiency.
It significantly improves the efficiency and precision of gene editing without the need for tracrRNA, is applicable to genome editing in eukaryotic cells, and provides a simpler and more efficient gene editing tool.
Smart Images

Figure CN121969744A_ABST
Abstract
Description
Engineered V-type RNA programmable endonucleases and their applications
[0001] 1. Cross-Reference to Related Patent Applications This application claims the benefit of priority to U.S. Provisional Application No. 63 / 588,636, filed October 6, 2023, the contents of which are incorporated herein by reference in their entirety.
[0002] 2. Sequence List This application includes a sequence list, which has been submitted electronically in XML format and is incorporated herein by reference in its entirety. The XML sequence list was created on September 26, 2024, is named BRT-008WO_SL.xml, and has a size of 451,059 bytes. Background Technology
[0003] 3. Background Technology Clusters of regularly spaced short palindromic repeats (CRISPR) and CRISPR-associated (Cas) genes, collectively known as the CRISPR-Cas or CRISPR / Cas system, are currently considered to provide bacteria and archaea with immunity to bacteriophage infection. The CRISPR-Cas system for adaptive immunity in prokaryotes is a highly diverse set of protein effector non-coding elements and locus architectures, some of which have been engineered and adapted to generate important biotechnologies.
[0004] The components of a host defense system include one or more effector proteins capable of modifying DNA or RNA, and RNA guide elements responsible for targeting the activity of these proteins to specific sequences on phage DNA or RNA. Guide RNAs consist of CRISPR RNA (crRNA) and may require additional trans-acting RNA (tracrRNA) to manipulate target nucleic acids via the effector proteins. crRNA consists of segments called “positive repeat sequences” and segments called “spacers sequences.” The “positive repeat sequences” are responsible for binding the crRNA to the effector proteins, and the “spacers sequences” are complementary to the desired nucleic acid target sequence. CRISPR systems, broadly speaking, can be divided into two categories: Category 1 systems consist of multiple effector proteins that collectively form a complex around the crRNA, and Category 2 systems consist of a single effector protein that, in conjunction with the crRNA guide, targets DNA or RNA substrates. The single-subunit effector compositions of Category 2 systems offer a simpler set of components for engineering and application, and are therefore a significant source of programmable effectors to date. Therefore, the discovery, engineering, and optimization of new Class 2 systems can lead to powerful programmable technologies with broad applications in genome engineering and other fields.
[0005] The RNA-guided DNA targeting principle of CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats)-Cas (CRISPR-associated proteins) for genome editing has been widely used in recent years. Five types of CRISPR-Cas systems have been documented (Type I, II, IIb, Ill, V, and VI). The most widely used CRISPR-Cas system for genome editing is the Type II system. The main advantage offered by the bacterial Type II CRISPR-Cas system lies in its minimal requirement for programmable DNA interference: the Cas9 endonuclease guided by a customizable dual-RNA structure. As demonstrated in the primitive type II system of *Streptococcus pyogenes*, trans-activated CRISPR RNA (tracrRNA) binds to an immutable repeat sequence of precursor CRISPR RNA (pre-crRNA) to form a dual RNA, which is essential for RNase III co-maturation of crRNA in the presence of Cas9 and for Cas9 cleavage of invading DNA. As demonstrated in *Streptococcus pyogenes*, Cas9, guided by the duplex formed between the mature activated tracrRNA and the target crRNA, introduces site-specific double-stranded DNA (dsDNA) breaks in invading homologous DNA. Cas9 is a multidomain enzyme that uses an HNH nuclease domain to cleave the target strand (defined as complementary to the spacer region sequence of crRNA) and a RuvC-like domain to cleave the non-target strand.
[0006] In addition to type II CRISPR Cas9 nucleases, many different type V endonucleases have been documented, such as Cas12a, Cas12b, Cas12e, Cas12f, Cas13a, and Cas13b (Koonin et al., Curr Opin Microbiol. 2017 Jun; 37: 67–78, and Makarova et al., Nat Rev Microbiol. 2020 Feb; 18(2):67–83.). Some of these systems do not require tracr RNA (Cas 12a, Cas 13a, Cas 13b), however, Cas 12b nucleases generally do require tracr RNA (Koonin et al., Curr Opin Microbiol. 2017 Jun; 37: 67–78).
[0007] WO2022258753A1 discloses novel type V endonuclease peptides, designated B-GEn.1, B-GEn.1.2, and B-GEn.2, which possess particularly advantageous characteristics for eukaryotic genome editing compared to other endonucleases. This disclosure provides engineered type V endonuclease peptides with improved gene editing efficiency. Summary of the Invention
[0008] 4. Summary of the Invention This disclosure relates to engineered type V endonucleases with improved gene editing efficiency. Unbound by theory, engineered type V endonucleases are thought to possess improved gene editing efficiency through better target interactions via their oligonucleotide-binding domains (OBDs) (e.g., OBD-II).
[0009] This disclosure provides engineered type V endonucleases that contain an amino acid other than aspartic acid at position D504 of the type V endonuclease corresponding to SEQ ID NO: 1 (B-GEn.1) or position D501 of the type V endonuclease corresponding to SEQ ID NO: 2 (B-GEn.1.2) or SEQ ID NO: 3 (B-GEn.2). In some embodiments, the amino acid at position D504 of the type V endonuclease corresponding to SEQ ID NO: 1 (B-GEn.1) or position D501 of the type V endonuclease corresponding to SEQ ID NO: 2 (B-GEn.1.2) or SEQ ID NO: 3 (B-GEn.2) is arginine.
[0010] This disclosure also provides engineered type V endonucleases comprising an OBD (e.g., OBD-II) containing a target interaction sequence motif GX1X2X3X4NX5X6X7DX8 (SEQ ID NO: 204), wherein each of X1 to X8 is any amino acid. In an exemplary embodiment, the target interaction sequence motif is any one of SEQ ID NO: 201, SEQ ID NO: 202, and SEQ ID NO: 3.
[0011] Exemplary engineered type V endonucleases include the B-GEn polypeptide. The engineered type V endonucleases and B-GEn polypeptides include those described in section 6.2 (and optionally include the nuclear localization signal described in section 6.3 and / or the adapter sequence described in section 6.4) and those described in embodiments numbered 1 to 52.
[0012] This disclosure further provides engineered type V endonuclease systems, such as an engineered B-Gen type V endonuclease system, comprising an engineered B-GEn polypeptide and a suitable guide RNA and / or nucleic acid encoding them. Exemplary engineered type V endonuclease systems are disclosed in section 6.5, and exemplary guide RNAs are disclosed in section 6.6. In some embodiments, the engineered type V endonuclease system is a ribonucleoprotein (RNP) complex comprising an engineered type V endonuclease or a B-GEn polypeptide and a guide RNA. Ribonucleoprotein complexes are described in section 6.7 and numbered embodiments 72 through 77.
[0013] This disclosure also provides nucleic acids encoding engineered type V endonucleases and B-GEn peptides, such as expression vectors for engineered type V endonucleases and B-GEn peptides, and recombinant cells engineered to express engineered type V endonucleases and B-GEn peptides. Exemplary nucleic acids are disclosed in section 6.8 and numbered embodiments 53 to 62, exemplary recombinant cells and their use for producing engineered type V endonucleases and B-GEn peptides are described in section 6.10 and numbered embodiments 63 to 69, and exemplary vectors are disclosed in section 6.9.
[0014] In some respects, this document provides a method for targeting, editing, modifying, or manipulating target DNA at one or more sites in cells or in vitro. The method typically requires introducing an engineered type V endonuclease or B-GEn peptide system into a cellular or in vitro environment under conditions suitable for engineered type V endonucleases or B-GEn peptides to create one or more nicks or cuts or base editing in the target DNA, wherein the engineered type V endonuclease or B-GEn peptide, in its processed or unprocessed form, is guided to the target DNA via guide RNA. As used herein, the term "type V endonuclease or B-GEn peptide system" refers to any combination of nucleic acid and peptide components that can be delivered to cells such that an RNP containing a type V endonuclease or B-GEn peptide is operatively constructed in the cell, thereby enabling editing. Therefore, a type V endonuclease or B-GEn polypeptide system may include any combination of the following: (a) (i) a type V endonuclease or B-GEn polypeptide and / or (ii) one or more nucleic acids containing a nucleotide sequence encoding a type V endonuclease or B-GEn polypeptide; and (b) (i) a guide RNA and / or (ii) a nucleic acid containing a nucleotide sequence encoding a guide RNA.
[0015] In some embodiments, the RNP of this disclosure (comprising an engineered type V endonuclease or B-GEn peptide and a guide RNA) is used to edit the genome of a cell. In some embodiments, methods for genomic DNA editing using RNPs include nuclear transfection of target cells containing genomic DNA with the RNP and exposing the target cells to conditions suitable for gene editing, such as culturing the target cells under conditions appropriate for genome editing by engineered type V endonucleases or B-GEn peptides.
[0016] In some embodiments, one or more viruses (e.g., one or more adeno-associated viruses (AAVs)) are used to deliver engineered type V endonucleases or B-GEn peptide systems (e.g., one or more viruses containing one or more nucleic acids encoding engineered type V endonucleases or B-GEn peptides and one or more nucleic acids encoding guide RNA) into cells, enabling their genome to be edited. In some embodiments, methods for genomic DNA editing using one or more viruses include contacting target cells containing genomic DNA with one or more viruses and exposing the target cells to conditions conducive to gene editing, such as culturing the target cells under conditions suitable for expressing engineered type V endonucleases or B-GEn peptides and guide RNA, and performing genome editing with engineered type V endonucleases or B-GEn peptides guided by guide RNA.
[0017] In some embodiments, lipid nanoparticles (LNPs) are used to deliver engineered type V endonucleases or B-GEn peptide systems (e.g., LNPs comprising one or more nucleic acids encoding engineered type V endonucleases or B-GEn peptides and guide RNA or nucleic acids encoding guide RNA) to cells, enabling their genome to be edited. In some embodiments, methods for genomic DNA editing using one or more viruses include contacting target cells containing genomic DNA with lipid nanoparticles and exposing the target cells to conditions conducive to gene editing, such as culturing the target cells under conditions suitable for expressing engineered type V endonucleases or B-GEn peptides and optionally guide RNA, and genome editing by engineered type V endonucleases or B-GEn peptides guided by guide RNA.
[0018] Exemplary methods for editing cellular genomes using engineered type V endonucleases and B-GEn peptides are described in section 6.13 and numbered embodiments 78 to 97.
[0019] Editing of the cell genome can be performed in vitro (e.g., in cell cultures), ex vivo, or in vivo (e.g., by administering RNP, AAV, or LNP to a subject for gene therapy purposes).
[0020] Exemplary cells containing engineered type V endonucleases and B-GEn peptides and nucleic acids (e.g., cells in contact with the RNP, AAV, or LNP disclosed herein) are described in section 6.11 and numbered embodiments 98 to 110.
[0021] The following describes in more detail the additional features, advantages, and applications of the engineered type V endonuclease and B-GEn peptide of this disclosure.
[0022] 5. Figure 1 illustrates the amino acid sequence alignments of three B-GEn enzymes: B-GEN 1 (SEQ ID NO: 1), B-GEN 1.2 (SEQ ID NO: 2), and B-GEn.2 (SEQ ID NO: 3), as well as an exemplary common sequence (SEQ ID NO: 205). Residues of each B-GEn enzyme that differ from the other two B-GEn enzymes are shown within black boxes, while matching amino acids are unlabeled.
[0023] Figures 2A-2C show a phylogenetic assessment of endonucleases derived from various species in the order Bacillales. Figure 2A is a phylogenetic tree showing genera derived from the genus *Bacillus*, adapted from Suzuki, 2018, Appl Microbiol Biotechnol. 102:10425-10437. Figure 2B shows the phylogenetic origin of several endonucleases. Figure 2C is a phylogenetic tree of endonucleases derived from genera in the order Bacillales, based on NCBI blast output of the B-GEn.2 OBD II and RuvC I domain sequences, using the tree-building algorithm in Geneious Prime software.
[0024] Figures 3A-3C show the enzyme structures of AacC2c1 and B-GEn.2. Figure 3A is the crystal structure of AacC2c1 (5U33) visualized using DNAStar. Figure 3B shows the predicted structure of B-GEn.2 generated using AlphaFold2 and visualized using DNAStar. Figure 3C shows a comparison between the crystal structure of AacC2c1 (5U33) and the predicted structure of B-GEn.2.
[0025] Figure 4 shows the amino acid sequence alignment of B-GEn.1 and B-GEn.2 with 20 different Cas nucleases from the Bacillus order (created using the MUSCLE alignment algorithm in Geneious Prime software). The domains labeled on the Bth C2C1 sequence are as follows: bold amino acids represent oligonucleotide-binding domains (OBD domains) (OBD-I and OBD-II), single consecutive underlined amino acids represent recognition (REC) domains (REC1-I, REC1-II, and REC2), italicized amino acids represent PAM interaction (PI) domains, and double consecutive underlined amino acids represent RuvC domains (RuvC-I, RuvC-II, and RuvC-III) that together form the nuclease domain. The plus sign above the common sequence indicates the amino acid corresponding to the D501 residue of B-GEn.2. The asterisk above the residues marks the catalytic residues at the active site within the RuvC domain.
[0026] Figure 5 is a Coomassie staining image of an SDS-PAGE gel, showing the results of single-step purification of B-GEn.2D501R based on heparin sulfate using the method described in Section 8.1.3.
[0027] Figure 6 shows the B2M target site cleavage efficiency of B-GEn.2D501R relative to B-GEn.2WT and Cpf1 in in vitro plasmid cleavage assays with different nuclease:linearized plasmid ratios.
[0028] Figure 7 is a bar graph showing the results of two evaluations of B2M gene editing in iPSCs using 50 pmol RNPs containing B-GEn.2WT, B-GEn.2D501R, or Cpf1. The percentage of insertion / deletion formation on the Y-axis represents the percentage of gene editing performed via RNPs formed with the specified enzyme.
[0029] Figure 8 shows the superimposed structure of AacC2c1 (labeled 5U31 on the image) and B-GEn.2 generated using AlphaFold2. The aligned amino acid backbones of AacC2c1 (black residues) and B-GEn.2 (white residues) adjacent to the P-1 nucleotide of the target DNA are shown, illustrating the superimposed target-interaction chains of the endonucleases. The target-interaction chains of both endonucleases are located near the PAM recognition region of each enzyme and contain amino acids corresponding to R507 (side chain showing hydrogen bond to one of the DNA phosphate groups) of AacC2c1 and D501 (gray, distant from the target DNA chain) of B-GEn.2. The amino acid sequence of the target interaction chain of B-GEn.1 with D504R substitution corresponds to SEQ ID NO: 201, the amino acid sequence of the target interaction chain of B-GEn.1.2 with D501R substitution corresponds to SEQ ID NO: 202, and the amino acid sequence of the target interaction chain of B-GEn.2 with D501R substitution corresponds to SEQ ID NO: 203. The common amino acid sequence of the target interaction chains in the endonucleases shown in Figure 4, and the presence of the amino acid arginine at the D501 position corresponding to SEQ ID NO: 3, is provided as SEQ ID NO: 204.
[0030] 6. Detailed Description 6.1. Definitions Unless otherwise defined herein, scientific and technical terms used in connection with this disclosure shall have the meanings commonly understood by one of ordinary skill in the art. Exemplary methods and materials are described below, but similar or equivalent methods and materials may also be used in the practice or testing of this disclosure. In case of conflict, this specification (including definitions) shall prevail. Generally, the nomenclature and techniques used in connection with cell and tissue culture, molecular biology, immunology, microbiology, genetics, analytical chemistry, synthetic organic chemistry, medical and medicinal chemistry, and protein and nucleic acid chemistry and hybridization as described herein are those well known and commonly used in the art. Enzymatic reactions and purification techniques are performed according to the manufacturer's instructions, as is commonly done in the art or as described herein. Furthermore, unless the context requires otherwise, singular terms shall include plural terms, and plural terms shall include singular terms. Throughout this specification and the embodiments, the words “have” and “compare”, or variations such as “has,” “having,” “comprises,” or “comprising”, will be understood to imply inclusion of the whole or group thereof, but not to exclude any other whole or group thereof. All publications and other references mentioned herein are incorporated herein by reference in their entirety. Although numerous references are cited herein, such citations do not constitute an admission that any of these references constitutes part of common general knowledge in the art.
[0031] B-GEn polypeptide: As used herein, the term "B-GEn polypeptide" refers to a polypeptide containing a nuclease domain having an amino acid sequence associated with or derived from Brevibacillus type V endonucleases, such as B-GEn.1 (SEQ ID NO: 1), B-GEn.1.2 (SEQ ID NO: 2), or B-GEn.2 (SEQ ID NO: 3). B-GEn.1 (SEQ ID NO: 1) has a nuclease domain containing RuvC I, RuvC II, and RuvC III subdomains corresponding to SEQ ID NO: 8-10. B-GEn.1.2 (SEQ ID NO: 2) has a nuclease domain containing RuvC I, RuvC II, and RuvC III subdomains corresponding to SEQ ID NO: 11-13. B-GEn.2 (SEQ ID NO: 3) has a nuclease domain comprising the RuvC I, RuvC II, and RuvC III subdomains corresponding to SEQ ID NO: 11-13. The term "B-GEn polypeptide" encompasses a polypeptide comprising an amino acid sequence having at least 40% sequence identity with the RuvC I, RuvC II, and RuvC III domains (alone or together) of any one of B-GEn 1, B-GEn 1.2, and B-GEn.2. "B-GEn polypeptide" also encompasses variants of any one of B-GEn.1 (SEQ ID NO: 1), B-GEn.1.2 (SEQ ID NO: 2), and B-GEn.2 (SEQ ID NO: 3), such as variants comprising an amino acid sequence having at least 50% sequence identity with SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3 and / or (2) variants comprising an amino acid sequence differing from SEQ ID NO: 4, SEQ ID NO: 5, or SEQ ID NO: 6 by up to 25 amino acids. In some embodiments, the B-GEn polypeptide has nuclease activity.The term “B-GEn polypeptide” encompasses engineered fusion polypeptides comprising the amino acid sequences of B-GEn.1 (SEQ ID NO: 1), B-GEn.1.2 (SEQ ID NO: 2), B-GEn.2 (SEQ ID NO: 3), or any variant thereof described in Section 6.2, for example, amino acid sequences having the following characteristics: (1) having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of the nuclease domains or full lengths of SEQ ID Nos: 1 to 3, and / or (2) having the same sequence identity as SEQ ID Nos: 1 to 3. NO: 1 to 3 differs from each other by up to 25, 20, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, or 5 amino acids, and also includes additional sequences (e.g., one or more nuclear localization and / or linker sequences described in Sections 6.3 and / or 6.4). The term “B-GEn polypeptide” covers its variants in which the amino acid at position D504 of B-GEn.1 (SEQ ID NO: 1), D501 of B-GEn.1.2 (SEQ ID NO: 2), or D501 of B-GEn.2 (SEQ ID NO: 3) is not aspartic acid, but is, for example, arginine.
[0032] Binding: As used herein, the term “binding” (e.g., regarding the RNA-binding domain of a polypeptide) refers to a non-covalent interaction between macromolecules (e.g., between proteins and nucleic acids). When in a non-covalent interaction state, macromolecules are referred to as “associated,” “interacting,” or “bound” (e.g., when molecule X is said to interact with molecule Y, this means that molecule X binds to molecule Y in a non-covalent manner). Not all components of a binding interaction require sequence specificity (e.g., contact with phosphate residues in the DNA backbone), but certain parts of a binding interaction can be sequence specific. Binding interactions are generally characterized by a dissociation constant (Kd) less than 10. -6 M, less than 10 -7 M, less than 10 -8 M, less than 10 -9 M, less than 10 -10 M, less than 10 -11 M, less than 10 -12 M, less than 10 -13 M, less than 10 -14M or less than 10 -15 M. “Affinity” refers to the strength of the binding; increased binding affinity is associated with lower Kd.
[0033] Cell therapy: As used herein, the term "cell therapy" refers to a therapy that administers cellular material to a patient. The cellular material can be intact, living cells. For example, T cells capable of fighting cancer cells via cell-mediated immunity can be administered during immunotherapy. Cell therapy is also known as cellular therapy or cytotherapy.
[0034] Coding sequence: As used herein, the term "coding sequence" or "coding nucleic acid" refers to a sequence within a nucleic acid (RNA or DNA) molecule that encodes a protein or RNA molecule. The coding sequence may further include start and stop signals operatively linked to regulatory elements, including promoters and polyadenylation signals capable of directing expression in human or mammalian cells in which the nucleic acid is introduced or administered. The coding sequence may be codon-optimized for expression in target cells.
[0035] Complementarity: As used herein, in the context of nucleic acid molecules, the terms “complement” and “complementary” refer to the ability of a nucleic acid molecule to form Watson-Crick (e.g., AT / U and CG) or Hoogsteen base pairs between nucleotides or nucleotide analogues. “Complementarity” refers to a property shared between two nucleic acid sequences such that when they are antiparallel aligned, the nucleotide bases at each position will be complementary.
[0036] Corresponding to (or corresponding to): The terms “corresponding to” or “corresponding to”, as used to refer to the amino acid sequence of the B-GEn polypeptide of any one of SEQ ID NO: 1, SEQ ID NO: 2 and SEQ ID NO: 3, or the RuvC I, RuvC II and RuvC III subdomains of the nuclease domain of any one of SEQ ID NO: 8 or SEQ ID NO: 11 (RuvC I), SEQ ID NO: 9 or SEQ ID NO: 12 (RuvC II) and SEQ ID NO: 10 or SEQ ID NO: 13 (RuvC II), are the sequence positions in the query sequence that appear at the same position in the sequence alignment of the reference and query sequences, as shown in Figure 1, for example. Sequence alignment algorithms, such as Clustal Omega (ClustalW; available at www.ebi.ac.uk / Tools / msa / clustalo / ), can be used to align reference and query sequences using the software's default parameters as of October 1, 2023.
[0037] Electroporation: As used herein, the term “electroporation” refers to a transfection technique in which electrical pulses are used to create temporary pores in the cell membrane to allow the introduction of ribonucleoproteins or nucleic acid molecules, such as DNA or RNA (e.g., mRNA), into the cell.
[0038] Encoding: The term “encoding” in relation to nucleic acids (DNA or RNA) refers to the nucleotide sequence of the nucleic acid that contains amino acids that encode polypeptides or nucleotides that encode RNA.
[0039] Engineered B-GEn peptides: As used herein, the term "engineered B-GEn peptide" refers to a variant peptide containing at least one mutation (e.g., amino acid insertion, deletion, or substitution) compared to the amino acid sequences of B-GEn.1 (SEQ ID NO: 1), B-GEn.1.2 (SEQ ID NO: 2), and B-GEn.2 (SEQ ID NO: 3). The engineered B-GEn peptides of this disclosure encompass type V endonucleases containing an amino acid other than aspartic acid at position D504 corresponding to B-GEn.1 (SEQ ID NO: 1), B-GEn.1.2 (SEQ ID NO: 2), and B-GEn.2 (SEQ ID NO: 3). In some embodiments, the amino acid at position 501 is arginine.
[0040] Expression cassette: As used herein, the term “expression cassette” refers to a DNA coding sequence that is operatively linked to a promoter.
[0041] Guide RNA: As used herein, the term “guide RNA” refers to a ribonucleic acid having a DNA-targeting sequence (also known as a “spacer region” or “DNA-targeting segment”) and a protein-binding sequence (also known as a “protein-binding segment”). The DNA-targeting sequence is sufficiently complementary to the target DNA (e.g., genomic DNA) sequence to hybridize with the target DNA sequence and guide the nucleic acid-targeting complex to sequence-specific binding to the target DNA sequence. The DNA-targeting sequence typically includes a “protospacer-like” sequence as described herein. The protein-binding sequence interacts with a site-specific modifying enzyme (e.g., the B-GEn polypeptide described in Section 6.2 below). Site-specific cleavage of the target DNA occurs at a location determined by two factors: (i) base pairing complementarity between the guide RNA and the target DNA; and (ii) a short motif in the target DNA (called a protospacer adjacent motif (PAM)). The protein-binding segment portion of the guide RNA contains two complementary nucleotide segments that hybridize to form a double-stranded RNA duplex (dsRNA duplex). In some embodiments, the guide RNA is a single-stranded guide RNA (sgRNA).
[0042] Guide RNAs and site-specific modifying enzymes (e.g., B-GEn peptides) can form ribonucleoprotein complexes (e.g., by non-covalent interactions). The guide RNA provides targeting specificity to the complex by including a nucleotide sequence complementary to the target DNA sequence. The site-specific modifying enzyme provides endonuclease activity. In other words, the site-specific modifying enzyme is guided to the target DNA sequence (e.g., target sequences in chromosomal nucleic acids; target sequences in extrachromosomal nucleic acids, such as free nucleic acids, microcircles, etc.; target sequences in mitochondrial nucleic acids; target sequences in chloroplast nucleic acids; target sequences in plasmids; etc.) by means of its binding to the protein-binding segment of the guide RNA.
[0043] Heterologous: As used herein, the term "heterologous" refers to a nucleic acid or protein that cannot be found separately in native nucleic acids or peptides. The B-GEn.1, or B-GEn.1.2, or B-GEn.2 fusion proteins described herein may, in some embodiments, comprise an RNA-binding domain of a B-GEn.1, or B-GEn.1.2, or B-GEn.2 peptide (or a variant thereof) fused to a heterologous peptide sequence (e.g., a peptide sequence from a protein other than B-GEn.1 or B-GEn.2). The heterologous peptide may exhibit activity (e.g., enzymatic activity), and the B-GEn.1, or B-GEn.1.2, or B-GEn.2 fusion proteins also exhibit this activity (e.g., methyltransferase activity, acetyltransferase activity, kinase activity, ubiquitinating activity, etc.). Heteronucleotides can be ligated to naturally occurring nucleic acids (or variants thereof) (e.g., through genetic engineering) to produce fusion nucleic acids encoding fusion peptides. In another example, in B-GEn.1 or B-GEn.2 peptide fusion variants, B-GEn.1, or B-GEn.1.2, or B-GEn.2 peptide variants can be fused to heteronucleotides (e.g., peptides other than B-GEn.1, or B-GEn.1.2, or B-GEn.2) that exhibit the activities exhibited by B-GEn.1, or B-GEn.1.2, or B-GEn.2 peptide fusion variants. Heteronucleotides can be ligated to B-GEn.1, or B-GEn.1.2, or B-GEn.2 peptide variants (e.g., through genetic engineering) to produce nucleic acids encoding B-GEn.1, or B-GEn.1.2, or B-GEn.2 peptide fusion variants. As used herein, "heteronucleotide" also refers to nucleotides or peptides that are not native to cells.
[0044] Host cell: As used herein, the terms "host cell" and "recombinant host cell" refer to a genetically engineered cell, such as by introducing a heterologous polypeptide or nucleic acid (such as the vector or system of this disclosure). It should be understood that such terms are intended to refer not only to a specific subject cell but also to the progeny of such cells. In some embodiments, the host cell carries a vector of this disclosure as an extrachromosomal heterologous expression vector. In some embodiments, the host cell contains any of the engineered B-GEn polypeptides disclosed herein, such as, introduced as an RNP complex. In other embodiments, the host cell has been genetically edited using the engineered B-GEn polypeptide of this disclosure.
[0045] iPSC: As used herein, the terms “induced pluripotent stem cell” and “iPSC” refer to a type of pluripotent stem cell artificially prepared from non-pluripotent cells (such as adult somatic cells, partially differentiated cells, or terminally differentiated cells (such as fibroblasts, hematopoietic lineage cells, myocytes, neurons, epidermal cells, etc.) by introducing or exposing cells to one or more reprogramming factors. iPSCs can be derived from a variety of different cell types, including terminally differentiated cells. iPSCs have an embryonic stem (ES) cell-like morphology, grow as flat colonies, have a high nucleus-to-cytoplasmic ratio, well-defined borders, and prominent nuclei. In addition, iPSCs express one or more key pluripotency markers known to those skilled in the art, including but not limited to alkaline phosphatase, SSEA3, SSEA4, Sox2, Oct3 / 4, Nanog, TRA160, TRA181, TDGF1, Dnmt3b, Fox03, GDF3, Cyp26al, TERT, and zfp42.
[0046] Examples of methods for generating and characterizing iPSCs can be found, for example, in US Patent Publications US20090047263, US20090068742, US20090191159, US20090227032, US20090246875, and US20090304646 and PCT Patent Publications WO2013177133 and WO2022204567, the disclosures of which are incorporated herein by reference. Typically, to generate iPSCs, somatic cells need to be reprogrammed with reprogramming factors known in the art (e.g., Oct4, SOX2, KLF4, MYC, Nanog, Lin28, etc.) to transform them into pluripotent stem cells.
[0047] Nucleic acid: As used herein, the terms “nucleic acid,” “oligonucleotide,” and “nucleic acid” refer to at least two nucleotides covalently linked together. Nucleic acids can be single-stranded or double-stranded, or may contain portions of both double-stranded and single-stranded sequences. Nucleic acids can be DNA, genomic DNA and cDNA, RNA, or hybrids, wherein the nucleic acid may contain combinations of deoxyribose- and ribose-nucleotides, and combinations of bases including uracil, adenine, thymine, cytosine, guanine, inosine, xanthine, hypoxanthine, isocytosine, and isoguanine. Nucleic acids can be obtained by chemical synthesis or by recombinant methods. The sequence of the complementary strand is also defined for the description of a single strand. Therefore, the single-stranded nucleic acid mentioned herein also encompasses the complementary strand of the single strand described.
[0048] Nuclear localization signal: As used herein, the terms “nuclear localization signal” and “NLS” refer to amino acid sequences that can facilitate the localization of peptides to the nucleus of eukaryotic cells.
[0049] Nuclease: As used herein, the terms “nuclease” and “endonuclease” are used interchangeably and refer to an enzyme having endonuclease catalytic activity for the cleavage of nucleic acids, as well as its nuclease-inactivated variants.
[0050] Nuclease Domain: As used herein, the term “nuclease domain,” “cleavage domain,” or “active domain” of a nuclease refers to an amino acid sequence or domain within a nuclease that possesses catalytic activity for cleaving DNA. A cleavage domain may be contained within a single polypeptide chain or may generate cleavage activity through the combination of two (or more) polypeptides. Within a given polypeptide, a single nuclease domain may consist of more than one amino acid segment. In some embodiments, the boundary of the nuclease domain of the B-GEn polypeptide is determined by aligning the B-GEn polypeptide to BthCas12b (Wu et al., 2017, Cell Research 27: 705-708) and identifying amino acids aligned with the BthCas12b RuvC nuclease domain, which contains the RuvC I, RuvC II, and RuvC III subdomains. The RuvC I domain of B-GEn.1 is shown in SEQ ID NO: 8, and the RuvC I domains of B-GEn.1.2 and B-GEn.2 are shown in SEQ ID NO: 11. The RuvC II domain of B-GEn.1 is shown in SEQ ID NO: 9, and the RuvC II domains of B-GEn.1.2 and B-GEn.2 are shown in SEQ ID NO: 12. The RuvC III domain of B-GEn.1 is shown in SEQ ID NO: 10, and the RuvC III domains of B-GEn.1.2 and B-GEn.2 are shown in SEQ ID NO: 13.
[0051] Nuclear transfection: As used herein, the term “nucleofection” refers to an electroporation-based transfection method that uses a combination of electrical parameters and cell type-specific reagents to directly transfer nucleic acids (such as DNA or RNA) and RNPs into the nucleus of target cells.
[0052] Operatically linked: As used herein, the term “operably linked” refers to a functional relationship between two or more peptide or polypeptide domains or nucleic acid (e.g., DNA) segments. In the context of transcriptional regulation, the term refers to a functional relationship between a transcriptional regulatory sequence and a transcriptional sequence. For example, if a promoter or enhancer sequence stimulates or regulates transcription of a coding sequence in a suitable host cell or other expression system, then the promoter or enhancer sequence is operably linked to the coding sequence.
[0053] Sequence identity percentage (%): As used herein, the terms “percent sequence identity”, “% sequence identity”, etc., relating to two amino acid sequences refer to the percentage of sequence identity determined using the BLASTP algorithm (Tatusova and Madden, 1999, FEMS Microbiol. Lett. 174: 247-250), which is available from the National Center for Biotechnology Information (NCBI) website (www.ncbi.nlm.nih.gov) with the following settings: Matrix = Blosum62; Open gap = 11; Extension gap = 1; Penalties gap x_dropoff = 50; Expect = 10; Word size = 3; Filter on. The BLAST algorithm performs a two-step operation. First, it aligns a reference sequence (e.g., the B-GEn polypeptide of any one of SEQ ID NO: 1, SEQ ID NO: 2, and SEQ ID NO: 3, or the RuvC I, RuvC II, and RuvC III subdomains of the nuclease domain of any one of SEQ ID NO: 8 or SEQ ID NO: 11 (RuvC I), SEQ ID NO: 9 or SEQ ID NO: 12 (RuvC II), and SEQ ID NO: 10 or SEQ ID NO: 13 (RuvC II)) with the query sequence. Then, it measures the sequence identity percentage within the overlap range between the two aligned sequences. In addition to the sequence identity percentage, BLASTP also measures the sequence similarity percentage based on settings. To characterize identity, the target sequence is aligned to obtain the highest order homology (match).
[0054] Polypeptides, peptides, and proteins: As used herein, the terms “polypeptide,” “peptide,” and “protein” refer to a polymer of amino acids of any length. In various embodiments, the polymer may be linear or branched, may contain modified amino acids, and may be interrupted by non-amino acid components.
[0055] Pluripotent: As used herein, the term “pluripotent” or “pluripotency” refers to the ability of a cell to self-renew and differentiate into cells from any of the three germ layers: endoderm, mesoderm, or ectoderm. “Pluripotent stem cells” or “PSCs” include, for example, embryonic stem cells derived from the inner cell mass of the blastocyst or obtained through somatic cell nuclear transfer, and iPSCs derived from non-pluripotent cells.
[0056] Promoter: As used herein, the term “promoter” refers to a nucleotide sequence that is recognized by a cellular or introduced synthetic mechanism and is required to initiate specific transcription of a nucleic acid sequence. A promoter can be a constitutively active promoter (e.g., a constitutive promoter in an active “on” state), an inducible promoter (e.g., a promoter whose active / “on” or inactive / “off” state is controlled by an external stimulus, such as the presence of a specific temperature, compound, or protein), a spatially restricted promoter (e.g., transcriptional control elements, enhancers, etc.) (e.g., tissue-specific promoters, cell-type-specific promoters, etc.), and a time-restricted promoter (e.g., the promoter is in an “on” or “off” state at a specific stage of embryonic development or a specific stage of a biological process, such as the hair follicle cycle in mice).
[0057] Protospace adjacent motif: As used herein, the term "protospace adjacent motif" or "PAM" refers to a DNA sequence downstream (e.g., immediately downstream) of the target sequence on a non-target strand recognized by the Cas protein. The PAM sequence is located at the 3' end of the target sequence on the non-target strand.
[0058] Recombinant: As used herein, the term “recombinant” in relation to nucleic acids, peptides, or cells refers directly or indirectly as a product of genetic engineering (e.g., a progeny or replica of a nucleic acid, peptide, or cell produced by genetic engineering methods). For example, a recombinant vector can be the product of various combinations of steps of cloning, restriction, polymerase chain reaction (PCR), and / or ligation, resulting in a construct with a structure-coding or non-coding sequence that is distinct from endogenous nucleic acids found in natural systems. DNA sequences encoding peptides can be assembled from cDNA fragments or a series of synthetic oligonucleotides to provide synthetic nucleic acids capable of expression in recombinant transcription units contained in cells or in cell-free transcription and translation systems. Genomic DNA containing relevant sequences can also be used to form recombinant genes or transcription units. Untranslated DNA sequences may appear at the 5' or 3' of an open reading frame, wherein these sequences do not interfere with the operation or expression of the coding region and can reliably regulate the production of the desired product through various mechanisms (see “DNA Regulatory Sequences” below). Additionally or alternatively, DNA sequences encoding untranslated RNA (e.g., guide RNA) may also be considered recombinant. Therefore, the term "recombinant" nucleic acid refers to non-naturally occurring nucleic acids, such as those artificially combined from two other isolated sequence segments through human intervention. This artificial combination is typically accomplished through chemical synthesis or by manipulating the isolated nucleic acid segments, such as through genetic engineering. This usually involves replacing codons with those encoding the same, conserved, or non-conserved amino acids. Additionally or optionally, nucleic acid segments with the desired function are linked together to produce the desired functional combination. This artificial combination is typically accomplished through chemical synthesis or by manipulating the isolated nucleic acid segments, such as through genetic engineering. When recombinant nucleic acids encode polypeptides, the sequence of the encoded polypeptide can be naturally occurring ("wild-type") or can be a variant of a naturally occurring sequence (such as a mutant). Therefore, the term "recombinant" polypeptide does not necessarily refer to a polypeptide whose sequence is not naturally occurring. Rather, a "recombinant" polypeptide is encoded by a recombinant DNA sequence, but the sequence of the polypeptide can be naturally occurring ("wild-type") or non-natural (such as variants, mutants, etc.). Therefore, a "recombinant" polypeptide is the result of human intervention, but can also be a naturally occurring amino acid sequence. The term "non-natural" includes molecules that are distinctly different from their naturally occurring counterparts, including chemically modified or mutated molecules.
[0059] Regulatory sequence: As used herein, the term "regulatory sequence" refers to a nucleic acid sequence required for the expression of an operable linker target sequence, such as a guide RNA or an engineered B-GEn polypeptide sequence. In some cases, the regulatory sequence may be a promoter sequence, and in others, it may include promoter and enhancer sequences and / or other regulatory elements required for the expression of pol. For example, a regulatory sequence may be a sequence that constitutively or tissue-specifically drives the expression of an operable linker sequence.
[0060] Ribonucleoprotein (RNP) complex, ribonucleoprotein (RNP) particle: As used herein, the terms "ribonucleoprotein complex" and "ribonucleoprotein particle" refer to a complex or particle comprising a nucleoprotein and ribonucleic acid (RNA). As used herein, "nucleoprotein" refers to a protein capable of binding nucleic acids (e.g., RNA, DNA). When a nucleoprotein binds to ribonucleic acid, it is called a "ribonucleoprotein." The interaction between the ribonucleoprotein and ribonucleic acid can be direct, such as through covalent bonds, or indirect, such as through non-covalent bonds (e.g., electrostatic interactions (e.g., ionic bonds, hydrogen bonds, halogen bonds), van der Waals interactions (e.g., dipole-dipole, dipole-induced dipole, London dispersion), ring stacking (pi effect), hydrophobic interactions, etc.). In embodiments, the ribonucleoprotein includes an RNA-binding motif that binds non-covalently to ribonucleic acid. For example, positively charged aromatic amino acid residues (e.g., lysine residues) in an RNA-binding motif can form electrostatic interactions with the negatively charged phosphate backbone of RNA, thereby forming a ribonucleoprotein complex. In some embodiments, any of the engineered B-GEn peptides disclosed herein reside in an RNP containing guide RNA.
[0061] Spacer: As used herein, the term "spacer" refers to a region of a gRNA molecule that is partially or completely complementary to the target sequence found in the + or - strand of the genomic DNA. When complexed with a Cas protein, the gRNA guides the Cas protein to the target sequence in the genomic DNA. The spacer is typically 15 to 30 nucleotides long (e.g., 20 to 25 nucleotides). The nucleotide sequence of the spacer may, but is not necessary, be completely complementary to the target sequence. For example, in some embodiments, the spacer may contain one or more mismatches with the target sequence; e.g., the spacer may contain one, two, or three mismatches with the target sequence.
[0062] Stem-loop structure: As used herein, the term "stem-loop structure" refers to a nucleic acid having a secondary structure comprising nucleotide regions known or predicted to form a double strand (stem portion), which is joined on one side by a region primarily composed of single-stranded nucleotides (loop portion). The terms "hairpin" and "fold-back" structures are also used herein to refer to stem-loop structures. Such structures are well known in the art, and these terms are consistent with their known meanings in the art. As is known in the art, stem-loop structures do not require precise base pairing. Therefore, a stem structure may include one or more base mismatches. Alternatively, base pairing may be precise, e.g., without any mismatches.
[0063] Target cell: As used herein, the term "target cell" refers to a cell in which a nuclease (e.g., the B-GEn system of this disclosure) is introduced, for example, a cell whose genome contains target DNA. It should be understood that such a term is intended to refer not only to a specific subject cell but also to the progeny of such cells. Because gene editing can occur in cells as a result of a nuclease system, such progeny need not be identical to the parent cell in which the system was originally introduced, but rather include the gene-edited counterpart of the cell. Such gene-edited progeny are still included within the scope of the term "target cell" as used herein.
[0064] Target DNA: As used herein, the term “target DNA” refers to a polydeoxyribonucleotide including a “target site” or “target sequence.” The terms “target site,” “target sequence,” “target prototypical spacer sequence DNA,” or “prototypical spacer-like sequence” are used interchangeably herein and refer to a nucleic acid sequence present in the target DNA to which the guide RNA’s DNA-targeting segment (also called a “spacer region”) will bind, provided sufficient binding conditions are provided. For example, the target site (or target sequence) 5'-GAGCATATC-3' within the target DNA is targeted (or binds to, hybridizes with), or complements the RNA sequence 5'-GAUAUGCUC-3'. Suitable DNA / RNA binding conditions include physiological conditions commonly present in cells. Other suitable DNA / RNA binding conditions (e.g., conditions in cell-free systems) are known in the art; see, for example, Sambrook, J. and Russell, W., 2001. Molecular Cloning: A Laboratory Manual, Third Edition, Cold Spring Harbor Laboratory Press. The target DNA strand that is complementary to and hybridizes with the guide RNA is called the "complementary strand," and the target DNA strand that is complementary to the "complementary strand" (and therefore not complementary to the guide RNA) is called the "non-complementary strand." In some embodiments, the target DNA is genomic DNA.
[0065] Transfection: As used herein, the term “transfection” refers to the introduction of a nucleic acid molecule, such as DNA or RNA (e.g., mRNA), into a cell, such as into the nucleus of a target cell or a production cell. In the context of this disclosure, the term “transfection” encompasses any method known to those skilled in the art for introducing a nucleic acid molecule into a cell (e.g., into a eukaryotic cell, such as into a mammalian cell). Such methods include, for example, electroporation, lipid transfection (e.g., based on cationic lipids and / or liposomes), calcium phosphate precipitation, nanoparticle-based transfection, virus-based transfection, or transfection based on cationic polymers (e.g., DEAE-glucan or polyethyleneimine). In some embodiments, the nucleic acid molecule is bound to or complexed with a polypeptide (e.g., in the form of a ribonucleoprotein).
[0066] Vector: As used herein, the term "vector" refers to a nucleic acid molecule capable of transporting another nucleic acid to which it is linked. One type of vector is a "plasmid," which is a circular double-stranded DNA loop into which an additional DNA segment can be incorporated. Another type of vector is a viral vector, in which an additional DNA segment can be linked to a viral genome. Some vectors are capable of autonomous replication in the host cells to which they are introduced (e.g., bacterial vectors with bacterial origins of replication and free mammalian vectors). Other vectors (e.g., non-attachment mammalian vectors) can integrate into the host cell's genome after introduction and thus replicate along with the host genome. Furthermore, some vectors are capable of directing the expression of a nucleotide sequence operatively linked to them. Such vectors are referred to herein as "expression vectors." In some embodiments, the vector is a viral vector, such as an adenovirus vector or an adeno-associated virus (AAV) vector.
[0067] 6.2. Engineered Type V Endonucleases and B-GEn Peptides This disclosure provides engineered type V endonucleases, such as engineered Bacillales type V endonucleases. The engineered type V endonucleases of this disclosure typically contain an amino acid other than aspartic acid at position D504 of SEQ ID NO: 1 (B-GEn.1), or position D501 of SEQ ID NO: 2 (B-GEn.1.2), or position D501 of SEQ ID NO: 3 (B-GEn.2).
[0068] In some aspects, this disclosure also provides engineered type V endonucleases comprising an OBD (e.g., OBD-II) containing a target interaction sequence motif GX1X2X3X4NX5X6X7DX8 (SEQ ID NO: 204), wherein each of X1 to X8 is any amino acid. In some embodiments, (a) X1 is selected from D, E, S, P, K, and R; (b) X2 is independently selected from V, I, and A; (c) X3 is independently selected from Y and F; (d) X4 is independently selected from L and F; (e) X5 is independently selected from I, L, F, V, and M; (f) X6 is independently selected from S, V, T, and A; (g) X7 is independently selected from V, L, and I; and (h) X8 is independently selected from V, F, L, and I. In exemplary embodiments, the target interaction sequence motif is any one of SEQ ID NO: 201, SEQ ID NO: 202, and SEQ ID NO: 3.
[0069] In some aspects, this disclosure also provides engineered type V endonucleases, for example, engineered type V endonucleases comprising an OBD containing the target-interacting sequence motif as described above and / or containing an amino acid other than aspartic acid at position D504 of the type V endonuclease corresponding to SEQ ID NO: 1 (B-GEn.1) and position D501 of the type V endonucleases corresponding to SEQ ID NO: 2 (B-GEn.1.2) and SEQ ID NO: 3 (B-GEn.2). In some embodiments, the amino acid corresponding to position D504 or D501 is arginine.
[0070] The engineered type V endonucleases disclosed herein typically comprise an amino acid sequence having at least 50% sequence identity with an amino acid sequence of a type V endonuclease of the order Bacillus, such as any amino acid sequence having at least 50%, at least 60%, at least 70%, or at least 80% sequence identity with any of the amino acid sequences of SEQ ID NO: 1, 2, 3, and 179-199, and comprise an OBD (e.g., OBD-II) containing the target interaction sequence motif GX1X2X3X4NX5X6X7DX8 (SEQ ID NO: 204), wherein each of X1 to X8 is any amino acid. In some embodiments, (a) X1 is selected from D, E, S, P, K, and R; (b) X2 is independently selected from V, I, and A; (c) X3 is independently selected from Y and F; (d) X4 is independently selected from L and F; (e) X5 is independently selected from I, L, F, V, and M; (f) X6 is independently selected from S, V, T, and A; (g) X7 is independently selected from V, L, and I; and (h) X8 is independently selected from V, F, L, and I. In an exemplary embodiment, the target interaction sequence motif is any one of SEQ ID NO: 201, SEQ ID NO: 202, and SEQ ID NO: 203.
[0071] In some aspects, the engineered type V endonuclease is an engineered B-GEn polypeptide. In some embodiments, the engineered B-GEn polypeptide comprises an amino acid sequence having at least 50% sequence identity with the full length of B-GEN 1 (SEQ ID NO: 1), B-GEN 1.2 (SEQ ID NO: 2), or B-GEn.2 (SEQ ID NO: 3) and / or differing from any one of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3 by up to 25 amino acids. The amino acid sequence preferably comprises: (a) a target interaction sequence motif of any one of SEQ ID NO: 201, SEQ ID NO: 202, SEQ ID NO: 203, and SEQ ID NO: 204; (b) a RuvC I domain comprising an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% sequence identity with the RuvC I domain of SEQ ID NO: 8 or SEQ ID NO: 11; (c) a RuvC II domain comprising an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% sequence identity with the RuvC II domain of SEQ ID NO: 9 or SEQ ID NO: 12; (d) a RuvC III domain comprising an amino acid sequence having at least 80%, at least 85%, or at least 90% sequence identity with the RuvC III domain of SEQ ID NO: 10 or SEQ ID NO: 13; or (e) Any combination of two, three, or all four from (a), (b), (c), and (d).
[0072] In some embodiments, the engineered B-GEn polypeptide of this disclosure comprises an amino acid sequence having at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% sequence identity with the full length of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3. In some embodiments, the engineered B-GEn polypeptide comprises an amino acid sequence having at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity with the full length of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3. The sequence preferably comprises: (a) a target interaction sequence motif of any one of SEQ ID NO: 201, SEQ ID NO: 202, SEQ ID NO: 203, and SEQ ID NO: 204; (b) a RuvC I domain comprising an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% sequence identity with the RuvC I domain of SEQ ID NO: 8 or SEQ ID NO: 11; (c) a RuvC II domain comprising an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% sequence identity with the RuvC II domain of SEQ ID NO: 9 or SEQ ID NO: 12; (d) a RuvCIII domain comprising an amino acid sequence having at least 80%, at least 85%, or at least 90% sequence identity with the RuvC III domain of SEQ ID NO: 10 or SEQ ID NO: 13; or (e) Any combination of two, three, or all four from (a), (b), (c), and (d).
[0073] In some aspects, the engineered B-GEn polypeptide of this disclosure comprises (a) a target-interacting sequence motif of any one of SEQ ID NO: 201, SEQ ID NO: 202, SEQ ID NO: 203, and SEQ ID NO: 204; and (b) RuvC I, RuvC II, and RuvC III amino acid sequences that differ from the corresponding sequences of B-GEn.1, B-GEn.1.2, or B-GEn.2 in a total of up to 25 amino acids in the three nuclease domains (SEQ ID NO: 8-10 for B-GEn.1, SEQ ID NO: 11-13 for B-GEn.1.2 and B-GEn.2), optionally wherein there are no more than 3 amino acid differences in the RuvC III domain. In some embodiments, the engineered B-GEn polypeptide comprises an amino acid sequence having an overall sequence identity of at least 70%, at least 80%, or at least 90% with the amino acid sequence of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3.
[0074] In another aspect, the engineered B-GEn polypeptide of this disclosure may comprise an amino acid sequence that differs from the entire sequence of SEQ ID NO: 1, SEQ ID NO: 2 or SEQ ID NO: 3 by up to 25 amino acids, including an amino acid substitution D504R relative to SEQ ID NO: 1 or an amino acid substitution D501R relative to SEQ ID NO: 2 or SEQ NO: 3. In some embodiments, the engineered B-GEn peptide comprises an amino acid sequence that differs from the full length of any one of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3 by up to 25, up to 20, up to 15, up to 14, up to 13, up to 11, up to 10, up to 9, up to 8, up to 7, up to 6, or up to 5 amino acids, in each case including an amino acid substitution D504R relative to SEQ ID NO: 1 or an amino acid substitution D501R relative to SEQ ID NO: 2 or SEQ NO: 3.
[0075] This disclosure provides polypeptides, for example, polypeptides having one or more features as described in embodiments 1 to 26, having at least 80% sequence identity with any one of SEQ ID NO: 1, SEQ ID NO: 2 or SEQ ID NO: 3, and in each case, the polypeptide includes the amino acid arginine at position D504 corresponding to SEQ ID NO: 1 or at position D501 corresponding to SEQ ID NO: 2 or SEQ ID NO: 3; or provides a nucleic acid comprising a nucleotide sequence encoding a polypeptide, said polypeptide comprising a nuclease sequence having at least 80% sequence identity with any one of SEQ ID NO: 1, SEQ ID NO: 2 or SEQ ID NO: 3, and in each case, the polypeptide includes the amino acid arginine at position D504 corresponding to SEQ ID NO: 1 or at position D501 corresponding to SEQ ID NO: 2 or SEQ ID NO: 3.
[0076] Some embodiments of this disclosure are polypeptides, for example, polypeptides having one or more features as described in embodiments 1 to 26, having at least 85% sequence identity with any one of SEQ ID NO: 1, SEQ ID NO: 2 or SEQ ID NO: 3, and in each case, the polypeptide includes the amino acid arginine at position D504 corresponding to SEQ ID NO: 1 or at position D501 corresponding to SEQ ID NO: 2 or SEQ ID NO: 3, or provides a nucleic acid comprising a nucleotide sequence encoding the polypeptide, said polypeptide comprising a nuclease sequence having at least 85% sequence identity with any one of SEQ ID NO: 1, SEQ ID NO: 2 or SEQ ID NO: 3, and in each case, the polypeptide includes the amino acid arginine at position D504 corresponding to SEQ ID NO: 1 or at position D501 corresponding to SEQ ID NO: 2 or SEQ ID NO: 3.
[0077] Other embodiments of this disclosure are polypeptides, for example, polypeptides having one or more of the features described in embodiments 1 to 26, having at least 90% sequence identity with any one of SEQ ID NO: 1, SEQ ID NO: 2 or SEQ ID NO: 3, and in each case, the polypeptide includes the amino acid arginine at position D504 corresponding to SEQ ID NO: 1 or at position D501 corresponding to SEQ ID NO: 2 or SEQ ID NO: 3, or provides a nucleic acid comprising a nucleotide sequence encoding the polypeptide, said polypeptide comprising a nuclease sequence having at least 90% sequence identity with any one of SEQ ID NO: 1, SEQ ID NO: 2 or SEQ ID NO: 3, and in each case, the polypeptide includes the amino acid arginine at position D504 corresponding to SEQ ID NO: 1 or at position D501 corresponding to SEQ ID NO: 2 or SEQ ID NO: 3.
[0078] A further embodiment of this disclosure is a polypeptide, for example, a polypeptide having one or more features as described in embodiments 1 to 26, having at least 95% sequence identity with any one of SEQ ID NO: 1, SEQ ID NO: 2 or SEQ ID NO: 3, and in each case, the polypeptide includes the amino acid arginine at position D504 corresponding to SEQ ID NO: 1 or at position D501 corresponding to SEQ ID NO: 2 or SEQ ID NO: 3, or provides a nucleic acid comprising a nucleotide sequence encoding the polypeptide, said polypeptide comprising a nuclease sequence having at least 95% sequence identity with any one of SEQ ID NO: 1, SEQ ID NO: 2 or SEQ ID NO: 3, and in each case, the polypeptide includes the amino acid arginine at position D504 corresponding to SEQ ID NO: 1 or at position D501 corresponding to SEQ ID NO: 2 or SEQ ID NO: 3.
[0079] A further embodiment of this disclosure is a polypeptide, for example, a polypeptide having one or more features as described in embodiments 1 to 26, having at least 96% sequence identity with any one of SEQ ID NO: 1, SEQ ID NO: 2 or SEQ ID NO: 3, and in each case, the polypeptide includes the amino acid arginine at position D504 corresponding to SEQ ID NO: 1 or at position D501 corresponding to SEQ ID NO: 2 or SEQ ID NO: 3, or provides a nucleic acid comprising a nucleotide sequence encoding the polypeptide, said polypeptide comprising a nuclease sequence having at least 96% sequence identity with any one of SEQ ID NO: 1, SEQ ID NO: 2 or SEQ ID NO: 3, and in each case, the polypeptide includes the amino acid arginine at position D504 corresponding to SEQ ID NO: 1 or at position D501 corresponding to SEQ ID NO: 2 or SEQ ID NO: 3.
[0080] Additional embodiments of this disclosure are polypeptides, for example, polypeptides having one or more of the features described in embodiments 1 to 26, having at least 97% sequence identity with any one of SEQ ID NO: 1, SEQ ID NO: 2 or SEQ ID NO: 3, and in each case, the polypeptide includes the amino acid arginine at position D504 corresponding to SEQ ID NO: 1 or at position D501 corresponding to SEQ ID NO: 2 or SEQ ID NO: 3, or provides a nucleic acid comprising a nucleotide sequence encoding the polypeptide, said polypeptide comprising a nuclease sequence having at least 97% sequence identity with any one of SEQ ID NO: 1, SEQ ID NO: 2 or SEQ ID NO: 3, and in each case, the polypeptide includes the amino acid arginine at position D504 corresponding to SEQ ID NO: 1 or at position D501 corresponding to SEQ ID NO: 2 or SEQ ID NO: 3.
[0081] Other additional embodiments according to this disclosure are polypeptides, for example, polypeptides having one or more of the features described in embodiments 1 to 26, having at least 98% sequence identity with any one of SEQ ID NO: 1, SEQ ID NO: 2 or SEQ ID NO: 3, and in each case, the polypeptide includes the amino acid arginine at position D504 corresponding to SEQ ID NO: 1 or at position D501 corresponding to SEQ ID NO: 2 or SEQ ID NO: 3, or provides a nucleic acid comprising a nucleotide sequence encoding the polypeptide, said polypeptide comprising a nuclease sequence having at least 98% sequence identity with any one of SEQ ID NO: 1, SEQ ID NO: 2 or SEQ ID NO: 3, and in each case, the polypeptide includes the amino acid arginine at position D504 corresponding to SEQ ID NO: 1 or at position D501 corresponding to SEQ ID NO: 2 or SEQ ID NO: 3.
[0082] Other embodiments of this disclosure are polypeptides, for example, polypeptides having one or more features as described in embodiments 1 to 26, having at least 99% sequence identity with any one of SEQ ID NO: 1, SEQ ID NO: 2 or SEQ ID NO: 3, and in each case, the polypeptide includes the amino acid arginine at position D504 corresponding to SEQ ID NO: 1 or at position D501 corresponding to SEQ ID NO: 2 or SEQ ID NO: 3, or provides a nucleic acid comprising a nucleotide sequence encoding the polypeptide, said polypeptide comprising a nuclease sequence having at least 99% sequence identity with any one of SEQ ID NO: 1, SEQ ID NO: 2 or SEQ ID NO: 3, and in each case, the polypeptide includes the amino acid arginine at position D504 corresponding to SEQ ID NO: 1 or at position D501 corresponding to SEQ ID NO: 2 or SEQ ID NO: 3.
[0083] Other embodiments of this disclosure are polypeptides, for example, polypeptides having one or more features as described in embodiments 1 to 26, having at least 99.5% sequence identity with any one of SEQ ID NO: 1, SEQ ID NO: 2 or SEQ ID NO: 3, and in each case, the polypeptide includes the amino acid arginine at position D504 corresponding to SEQ ID NO: 1 or at position D501 corresponding to SEQ ID NO: 2 or SEQ ID NO: 3, or provides a nucleic acid comprising a nucleotide sequence encoding the polypeptide, said polypeptide comprising a nuclease sequence having at least 99.5% sequence identity with any one of SEQ ID NO: 1, SEQ ID NO: 2 or SEQ ID NO: 3, and in each case, the polypeptide includes the amino acid arginine at position D504 corresponding to SEQ ID NO: 1 or at position D501 corresponding to SEQ ID NO: 2 or SEQ ID NO: 3.
[0084] Exemplary B-GEn common sequences B-GEn common sequence I and B-GEn common sequence II are provided herein as SEQ ID NO: 7 and SEQ ID NO: 205, respectively.
[0085] Exemplary engineered B-GEn sequences referred to herein as B-GEn.1 D504R (SEQ ID NO: 4), B-GEn.1.2 D501R (SEQ ID NO: 5), and B-GEn.2 D501R (SEQ ID NO: 6) are listed in Table 1 below. In some embodiments, the engineered B-GEn polypeptide contains the same amino acid sequence as the exemplary engineered B-GEn polypeptide of SEQ ID NO: 4, SEQ ID NO: 5, or SEQ ID NO: 6.
[0086] In some embodiments, the polypeptide further comprises a nuclear localization signal (e.g., as described in section 6.3) and / or a linker sequence (e.g., as described in section 6.4).
[0087] 6.3. Nuclear Localization Signals The engineered type V endonuclease and B-GEn polypeptide of this disclosure may further include one or more nuclear localization signals (NLS). In some embodiments, the engineered type V endonuclease and B-GEn polypeptide of this disclosure include one or more NLS at their N-terminus. In some embodiments, the engineered type V endonuclease and B-GEn polypeptide of this disclosure include one or more NLS at their C-terminus. In some embodiments, the engineered type V endonuclease and B-GEn polypeptide of this disclosure include one or more NLS at both their N-terminus and C-terminus. The NLS can be separated from the endonuclease sequence and from each other via adapter sequences.
[0088] Therefore, this disclosure provides engineered type V endonucleases and B-GEn peptides comprising: (i) an engineered type V endonuclease or B-GEn peptide sequence, as described in section 6.2; (ii) one or more NLS sequences as described herein (e.g., in section 6.3); and (iii) one or more adapter sequences, as described in section 6.4.
[0089] In some respects, engineered type V endonucleases or B-GEn peptides comprise a nuclease sequence and a first NLS sequence at the C-terminus of the nuclease sequence. Engineered type V endonucleases or B-GEn peptides comprising the nuclease sequence and the first NLS sequence may further comprise a first linker sequence between the nuclease sequence and the first NLS sequence.
[0090] In some respects, engineered type V endonucleases or B-GEn peptides contain more than one NLS sequence (e.g., more than one NLS sequence at the C-terminus of the nuclease sequence).
[0091] In some embodiments, the engineered type V endonuclease or B-GEn polypeptide includes a second NLS sequence at the C-terminus of the first NLS sequence. The engineered type V endonuclease or B-GEn polypeptide containing the second NLS sequence may further include a linker sequence between the first and second NLS sequences.
[0092] In a further embodiment, the engineered type V endonuclease or B-GEn polypeptide comprises a third NLS sequence at the C-terminus of the second NLS sequence. The engineered type V endonuclease or B-GEn polypeptide comprising the third NLS sequence may further comprise a linker sequence between the second and third NLS sequences.
[0093] In an additional embodiment, the engineered type V endonuclease or B-GEn polypeptide comprises a fourth NLS sequence at the C-terminus of the third NLS sequence. The engineered type V endonuclease or B-GEn polypeptide comprising the fourth NLS sequence may further comprise a linker sequence between the third and fourth NLS sequences.
[0094] In some embodiments, the engineered type V endonuclease or B-GEn polypeptide further comprises an N-terminal NLS sequence in addition to one or more NLS sequences at the C-terminus of the nuclease sequence. Therefore, in some aspects, the engineered type V endonuclease or B-GEn polypeptide comprises an N-terminal NLS sequence of the nuclease sequence. An engineered type V endonuclease or B-GEn polypeptide comprising an N-terminal NLS sequence of the nuclease sequence may further comprise a linker sequence between the NLS sequence and the nuclease sequence. In some embodiments, the engineered type V endonuclease or B-GEn polypeptide comprises more than one N-terminal NLS sequence (e.g., more than one N-terminal NLS sequence of a nuclease sequence, which may be linked via one or more linkers).
[0095] Non-limiting examples of nuclear location signals are listed in Table 2. In some embodiments, the engineered B-GEn peptide contains one or more NLS sequences only at the N-terminus of the B-GEn.2 protein sequence. In other embodiments, the fusion protein of this disclosure contains one or more NLS sequences only at the C-terminus of the B-GEn.2 protein sequence. An exemplary engineered B-GEn peptide with a C-terminal NLS is shown in SEQ ID NO: 200.
[0096] In some embodiments, the engineered B-GEn peptide contains the same multiple NLS sequences. In other embodiments, the fusion protein of this disclosure contains different multiple NLS sequences.
[0097] In some embodiments, the engineered B-GEn polypeptide comprises nucleoplasmic protein NLS (SEQ ID NO: 15) at the N-terminus of the B-GEn polypeptide sequence and SV40 large T protein NLS (SEQ ID NO: 14) at the C-terminus, as described in section 6.2.
[0098] In some embodiments, the engineered B-GEn polypeptide contains SV40 large T protein NLS (SEQ ID NO: 14) at the C-terminus of the B-GEn polypeptide sequence, as described in section 6.2.
[0099] In some embodiments, the engineered B-GEn polypeptide contains the nucleoplasmic protein NLS (SEQ ID NO: 15) at the N-terminus of the B-GEn polypeptide sequence, as described in section 6.2.
[0100] 6.4. Connector Sequence This disclosure provides engineered B-GEn polypeptides in the form of fusion proteins, the fusion proteins comprising a B-GEn polypeptide sequence, such as having an amino acid sequence as described in section 6.2, optionally fused to one or more NLS sequences via a peptide linker.
[0101] In some embodiments, the B-GEn polypeptide sequence is linked to an NLS sequence at its N- and / or C-terminus via a peptide linker. In other embodiments, the B-GEn polypeptide sequence is linked to one or more NLS sequences using a peptide linker that links a pair of individual polypeptide sequences, for example, between two NLS sequences or between an NLS sequence and a B-GEn polypeptide sequence.
[0102] In some embodiments, the B-GEn polypeptide of this disclosure comprises multiple NLS sequences linked by the same linker. In other embodiments, the B-GEn polypeptide of this disclosure comprises multiple NLS sequences linked by different linkers.
[0103] The peptide linkers applicable to B-GEn peptides in this disclosure include those disclosed in Chenet al., 2013, Adv DrugDeliv Rev. 65(10):1357-1369. Non-limiting examples of such linkers are reproduced in Table 3 below. 6.5. B-Gen V-type CRISPR-Cas System This disclosure provides engineered V-type endonuclease and B-GEn peptide systems or the nucleic acids encoding them, which are incorporated herein by reference.
[0104] In some embodiments, an engineered type V endonuclease or B-GEn polypeptide system comprises the following components: (a) an engineered type V endonuclease or B-GEn polypeptide, as described in Section 6.2, or a nucleic acid encoding such an engineered B-GEn polypeptide, as described in Section 6.8; and (b) a heterologous guide RNA (gRNA), as described in Section 6.6, or a nucleic acid that allows for the in situ generation of such gRNA (e.g., a vector as described in Section 6.9), wherein the gRNA comprises: i. an engineered DNA targeting segment composed of RNA and capable of hybridizing with a target sequence at a nucleic acid locus; ii. a tracr pairing sequence composed of RNA; and iii. a tracr RNA sequence composed of RNA, wherein the tracr pairing sequence hybridizes with the tracr sequence, and wherein (i), (ii), and (iii) are aligned in a 5' to 3' orientation. The gRNA may be a single guide RNA (sgRNA). In sgRNA, the tracr pairing sequence and the tracr sequence are typically linked by a suitable loop sequence to form a stem-loop structure.
[0105] In a specific implementation, the engineered type V endonuclease or B-GEn polypeptide system comprises the following components: (a) an engineered type V endonuclease or B-GEn polypeptide, as described in Section 6.2; and (b) a heterologous guide RNA (gRNA), as described in Section 6.6, comprising: i. an engineered DNA targeting segment composed of RNA and capable of hybridizing with a target sequence at a nucleic acid locus; ii. a tracr pairing sequence composed of RNA; and iii. a tracr RNA sequence composed of RNA, wherein the tracr pairing sequence hybridizes with the tracr sequence, and wherein (i), (ii), and (iii) are arranged in a 5' to 3' orientation. The gRNA may be a single guide RNA (sgRNA). In sgRNA, the tracr pairing sequence and the tracr sequence are typically linked by a suitable loop sequence to form a stem-loop structure.
[0106] In some embodiments, such engineered type V endonucleases or B-GEn peptides are delivered to target cells in the form of a composition, also referred to as a ribonucleoprotein (RNP) complex, as described in section 6.7.
[0107] In a specific implementation, the engineered type V endonuclease or B-GEn polypeptide system comprises the following components: (a) a nucleic acid (e.g., as described in Section 6.8), wherein the nucleic acid encodes an engineered type V endonuclease or B-GEn polypeptide (e.g., as described in Section 6.2); or (b) a nucleic acid (e.g., a vector as described in Section 6.9) that allows the production of heterologous guide RNA (gRNA) (e.g., as described in Section 6.6), wherein the gRNA comprises: i. an engineered DNA targeting segment composed of RNA and capable of hybridizing with a target sequence in a nucleic acid locus; ii. a tracr pairing sequence composed of RNA; and iii. a tracr RNA sequence composed of RNA, wherein the tracr pairing sequence hybridizes with the tracr sequence, and wherein (i), (ii), and (iii) are aligned in a 5' to 3' orientation. The gRNA may be a single guide RNA (sgRNA). In sgRNA, the tracr pairing sequence and the tracr sequence are typically linked by a suitable loop sequence to form a stem-loop structure.
[0108] Any reference to “RNA” or “guide RNA” covers RNA molecules containing both non-natural and natural nucleobases, such as one or more nucleic acid modifications described in section 6.8.3.
[0109] In some implementations, the nucleic acid and / or sgRNA encoding an engineered type V endonuclease or B-GEn polypeptide contains a suitable promoter for expression in a cellular or in vitro environment.
[0110] In some implementations, the nucleic acid encoding an engineered type V endonuclease or B-GEn polypeptide and / or sgRNA is in the form of a viral vector, such as as described in section 6.9.3.
[0111] 6.6. Guide RNA (gRNA) and Single Guide RNA (sgRNA) The systems, compositions, and methods described in some embodiments use genome-targeting nucleic acids that can guide the activity of engineered B-GEn peptides to specific target sequences within the target nucleic acid. In some embodiments, the genome-targeting nucleic acid is RNA. The genome-targeting RNA is referred to herein as "guide RNA" or "gRNA". Guide RNA has a spacer sequence (such a CRISPR repeat sequence is also called a "tracr-pairing sequence") that can hybridize with at least the target nucleic acid sequence and a CRISPR repeat sequence. In type II systems, gRNA also has a second RNA, called a tracrRNA sequence. In type II guide RNA (gRNA), the CRISPR repeat sequence and the tracrRNA sequence hybridize to form a double strand. In type V guide RNA (gRNA), the crRNA forms a double strand. In both systems, the double strand binds to a site-specific peptide, causing the guide RNA and the site-directing peptide to form a complex. The genome-targeting nucleic acid provides targeting specificity to this complex through its association with the site-specific peptide. The genome-targeting nucleic acid thus guides the activity of the site-specific peptide.
[0112] In some implementations, the genome-targeting nucleic acid is a bimolecular guide RNA. In some implementations, the genome-targeting nucleic acid is a single-molecule guide RNA or a single guide RNA (sgRNA). A bimolecular guide RNA has two RNA strands. The first strand, in the 5' to 3' direction, has an optional spacer extension sequence, a spacer sequence, and a minimal CRISPR repeat sequence. The second strand has a minimal tracrRNA sequence (complementary to the minimal CRISPR repeat sequence), a 3' tracrRNA sequence, and an optional tracrRNA extension sequence. In the type II system, the single-molecule guide RNA (sgRNA) has an optional spacer extension sequence, a spacer sequence, a minimal CRISPR repeat sequence, a single-molecule guide connector, a minimal tracrRNA sequence, a 3' tracrRNA sequence, and an optional tracrRNA extension sequence in the 5' to 3' direction. The optional tracrRNA extension may have elements that provide additional functionality (e.g., stability) for the guide RNA. The single-molecule guide connector links the minimal CRISPR repeat sequence and the minimal tracrRNA sequence together, forming a hairpin structure. The optional tracrRNA extension has one or more hairpins.
[0113] In the V-type system, the single-molecule guide RNA (sgRNA) has a minimal CRISPR repeat sequence and a spacer sequence in the 5' to 3' direction. Alternatively, the single-molecule guide RNA (sgRNA) in the V-type system has an optional tracr extension sequence, a tracr RNA sequence, a single-molecule guide adapter, a minimal CRISPR repeat sequence, a spacer sequence, and an optional spacer extension sequence in the 5' to 3' direction.
[0114] Alternatively, single-molecule guide RNAs (sgRNAs) in the V-type system have optional extension sequences, minimal CRISPR repeat sequences, spacer sequences, and optional spacer extension sequences in the 5' to 3' direction.
[0115] In some implementations, the sgRNA in the V-type system includes an optional extension sequence, an artificial nuclease-binding RNA sequence, and a spacer sequence in the 5' to 3' direction, as well as an optional spacer extension sequence.
[0116] The B-GEn.2 CRISPR Cas nuclease described in this disclosure, and the sgRNAs that are particularly useful for other potential type V nucleases, are disclosed in Table 4. Exemplary genome-targeting nucleic acids are described, for example, in WO2018002719. Typically, CRISPR repeat sequences comprise any sequence having sufficient complementarity to the tracr sequence to facilitate one or more of the following: (1) excision of DNA targeting regions flanking the CRISPR repeat sequence in cells containing the corresponding tracr sequence; and (2) formation of a CRISPR complex at the target sequence, wherein the CRISPR complex comprises a CRISPR repeat sequence hybridized to the tracr sequence. Typically, complementarity refers to the optimal alignment of the CRISPR repeat sequence and the tracr sequence along the shorter of the two sequences. Optimal alignment can be determined by any suitable alignment algorithm, and secondary structures, such as self-complementarity within the tracr sequence or the CRISPR repeat sequence, can be further considered. In some embodiments, when the tracr sequence and the CRISPR repeat sequence are optimally aligned, the complementarity along the shorter 30 nucleotide length of both is approximately or greater than 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97.5%, 99%, or higher. In some embodiments, the length of the tracr sequence is approximately or more than 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, or more nucleotides. In some embodiments, the tracr sequence and the CRISPR repeat sequence are contained in a single transcript, such that hybridization between the two produces a transcript with secondary structure, such as a hairpin. In some embodiments, the transcript or the transcribed nucleic acid sequence has at least two or more hairpins.
[0117] Suitable tracr sequences for B-GEn.2 or B-GEn.1 in the CRISPR-Cas system are listed in Table 5. Alternatively, variants of these sequences may be used. Variants may contain partial or truncated versions of these sequences and / or sequences with base modifications at one or more positions in these sequences. The corresponding RNA sequences are disclosed in SEQ ID NOs:176 and 177, respectively. The spacer region of a guide RNA comprises a nucleotide sequence complementary to a sequence in the target DNA. In other words, the spacer region of the guide RNA interacts with the target DNA in a sequence-specific manner through hybridization (such as base pairing). Therefore, the nucleotide sequence of the spacer region can vary and determines the location of the interaction between the guide RNA and the target DNA. The DNA-targeting segment of the guide RNA can be modified (e.g., through genetic engineering) to hybridize with any desired sequence within the target DNA.
[0118] In some embodiments, the spacer region has a length of 10 to 30 nucleotides. In some embodiments, the spacer region has a length of 13 to 25 nucleotides. In some embodiments, the spacer region has a length of 15 to 23 nucleotides. In some embodiments, the spacer region has a length of 18 to 22 nucleotides, such as 20 to 22 nucleotides.
[0119] In some implementations, the percentage of complementarity between the DNA target sequence of the spacer region and the prototype spacer region of the target DNA is at least 60% over 20-22 nucleotides (e.g., at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%).
[0120] In some implementations, the prototype spacer region is directly adjacent to a suitable PAM sequence at its 3' end, or such a PAM sequence is part of a DNA-targeting sequence in its 3' portion.
[0121] Suitable PAM sequences are listed in Table 6, wherein the engineered DNA targeting segment is located directly adjacent to the PAM sequence on the targeted DNA segment at its 3' end, or such a PAM sequence is part of the targeted DNA sequence in its 5' portion. Modification of the guide RNA can be used to enhance the formation or stability of the CRISPR-Cas genome editing complex, which contains the guide RNA and a Cas endonuclease (e.g., B-GEn.1 or B-GEn.2). Guide RNA modification can also, or alternatively, be used to enhance the initiation, stability, or kinetics of the interaction between the genome editing complex and the target sequence in the genome, for example, to enhance on-target activity. Guide RNA modification can also, or alternatively, be used to enhance specificity, such as the relative ratio of genome editing at the target site compared to effects at other (off-target) sites.
[0122] Modifications can also be used, or alternatively, to increase the stability of the guide RNA, for example, by increasing its resistance to degradation by ribonucleases (RNases) present in the cell, thereby prolonging its half-life in the cell. Modifications that enhance the half-life of guide RNA are particularly useful in implementations that introduce Cas endonucleases such as B-GEn.1, B-GEn.1.2, or B-GEn.2 into the edited cell via the RNA to be translated, in order to produce B-GEn.1, or B-GEn.1.2, or B-GEn.2 endonucleases, because the increased half-life of the introduced guide RNA allows the RNA encoding the endonuclease to be used to increase the time the guide RNA and the encoded Cas endonuclease coexist in the cell.
[0123] 6.6.1. Additional Sequences In some embodiments, at least one additional segment is included at the 5' or 3' end of the RNA. For example, suitable additional segments may include a 5' cap (e.g., a 7-methylguanylate cap (m7G)); a 3' polyadenylated tail (e.g., a 3' poly(A) tail); a ribo-switching sequence (e.g., allowing regulation of stability and / or accessibility via protein and protein complex regulation); a sequence that forms a dsRNA duplex (e.g., a hairpin); a sequence that targets the RNA to subcellular locations (e.g., the nucleus, mitochondria, chloroplasts, etc.); a modification or sequence that provides tracking functionality (e.g., direct conjugation to fluorescent molecules, partial conjugation to molecules that promote fluorescence detection, sequences that allow fluorescence detection, etc.); a modification or sequence that provides a protein binding site (e.g., proteins that act on DNA, including transcription activators, transcription inhibitors, DNA methyltransferases, DNA demethylases, histone acetyltransferases, histone deacetylases, etc.); a modification or sequence that provides increased, decreased, and / or controllable stability; and combinations thereof.
[0124] 6.6.1.1. Stability Control Sequences. Stability control sequences affect the stability of RNA (such as guide RNA). A non-restrictive example of a suitable stability control sequence is a transcription terminator region (such as a transcription termination sequence). The total length of a transcription terminator region in a guide RNA can range from 10 nucleotides to 100 nucleotides, such as from 10 nucleotides (nt) to 20 nt, from 20 nt to 30 nt, from 30 nt to 40 nt, from 40 nt to 50 nt, from 50 nt to 60 nt, from 60 nt to 70 nt, from 70 nt to 80 nt, from 80 nt to 90 nt, or from 90 nt to 100 nt. For example, the length of a transcription terminator region can be from 15 nucleotides (nt) to 80 nt, from 15 nt to 50 nt, from 15 nt to 40 nt, from 15 nt to 30 nt, or from 15 nt to 25 nt.
[0125] In some embodiments, the transcription termination sequence is a sequence that functions in eukaryotic cells. In some embodiments, the transcription termination sequence is a sequence that functions in prokaryotic cells.
[0126] Nucleotide sequences that can be included in stability control sequences (such as transcription termination segments, or any segment of the guide RNA, to provide enhanced stability) include, for example, Rho-independent trp termination sites.
[0127] 6.7. Ribonucleoprotein (RNP) Complex In some embodiments, the engineered type V endonuclease and B-GEn polypeptide are delivered as a complex, referred to as a ribonucleoprotein or RNP complex. The RNP complex is assembled by combining a Cas endonuclease (e.g., an engineered type V or B-GEn endonuclease) with ribonucleic acid (e.g., guide RNA (gRNA)).
[0128] In some embodiments, the ribonucleoprotein complex comprises an engineered B-GEn endonuclease (e.g., as described in Section 6.2) complexed with a suitable ribonucleic acid. In some embodiments, the ribonucleic acid is a gRNA or sgRNA, which is further described in Section 6.6. In some embodiments, the RNP complex comprises an engineered B-GEn polypeptide and an sgRNA listed in Table 4 or another suitable sgRNA.
[0129] In some embodiments, the engineered B-GEn peptide and sgRNA are present in a molar ratio of about 1:1 to about 1:4. In some embodiments, the molar ratio is about 1:1 to about 1:3. In some embodiments, the molar ratio is about 1:1 to about 1:2.5. In some embodiments, the molar ratio is about 1:1 to about 1:2. In some embodiments, the molar ratio is about 1:1 to about 1:1.5. In some embodiments, the molar ratio of peptide to sgRNA is 1:1.
[0130] One of the most common techniques used for RNP delivery is electroporation, which creates pores in the cell membrane to allow RNPs to enter the cytoplasm. Further, in a technique called “nuclear transfection,” electroporation can be combined with reagents targeted to specific cell types, which create pores in the nuclear membrane to allow DNA templates to enter.
[0131] In some implementations, the engineered type V nuclease or B-GEn peptide in the RNP complex is delivered to target cells via nuclear transfection.
[0132] 6.8. Nucleic Acids This disclosure provides nucleic acids (e.g., DNA or RNA) encoding B-GEn type V CRISPR-Cas proteins (e.g., engineered B-GEn peptides), nucleic acids encoding gRNA or sgRNA of this disclosure, nucleic acids simultaneously encoding engineered B-GEn peptides and gRNA or sgRNA, and various nucleic acids, such as those containing nucleic acids encoding engineered B-GEn peptides and gRNA or sgRNA.
[0133] Nucleic acids encoding B-GEn polypeptides can be codon-optimized, for example, by replacing at least one uncommon or rare codon with a codon common in the host or target cells. For instance, optimized nucleic acids can guide the synthesis of optimized messenger mRNAs, such as those optimized for expression in mammalian expression systems.
[0134] 6.8.1. B-GEn Encoding Sequence In some embodiments, the nucleic acid described herein includes one or more modifications as further described herein and known in the art, such modifications as those for enhancing activity, stability, or specificity, altering delivery, reducing the innate immune response in host cells, further reducing protein size, or for other enhancements. In some embodiments, such modifications will engineer the B-GEn polypeptide, the nuclease sequence component of which has at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or 100% amino acid sequence identity with the sequence of SEQ ID NOs: 4, 5, or 6.
[0135] 6.8.2. Codon Optimization In some embodiments, modified nucleic acids used in the CRISPR-B-GEn.1, B-GEn.1.2, or B-GEn.2 systems described herein, wherein the guide RNA and / or DNA or RNA containing a nucleic acid sequence encoding an engineered B-GEn polypeptide can be modified, as described below. Such modified nucleic acids can be used in the CRISPR-B-GEn.1, B-GEn.1.2, or B-GEn.2 systems to edit any one or more genomic loci. In some embodiments, such modifications in the nucleic acids of this disclosure are achieved through codon optimization, e.g., codon optimization based on the specific host cell in which the encoded polypeptide is expressed. Those skilled in the art will understand that any nucleotide sequence and / or recombinant nucleic acid of this disclosure can be codon-optimized for expression in any target species. Codon optimization is well known in the art and involves modifying nucleotide sequences with codon usage bias using species-specific codon usage tables. Codon usage tables are generated based on sequence analysis of the most expressed genes in the target species. In a non-limiting example, when the nucleotide sequence will be expressed in the cell nucleus, the codon usage table is generated based on sequence analysis of highly expressed nuclear genes in the target species. Modifications to the nucleotide sequence are determined by comparing the species-specific codon usage table with codons present in the natural nucleic acid sequence.
[0136] In some embodiments, the engineered B-GEn peptides described herein are expressed by codon-optimized nucleic acid sequences. For example, if the intended host or target cell is a human cell, a human codon-optimized nucleic acid sequence encoding the engineered B-GEn peptide, comprising the amino acid sequence of B-GEn.1, or B-GEn.1.2, or B-GEn.2 (or a variant of B-GEn.1, or B-GEn.1.2, or B-GEn.2, such as an enzyme-inactivating variant). As another non-limiting example, if the intended host or target cell is a mouse cell, a mouse codon-optimized nucleic acid sequence encoding the engineered B-GEn peptide, comprising the amino acid sequence of B-GEn.1, or B-GEn.1.2, or B-GEn.2 (or a variant of B-GEn.1, or B-GEn.1.2, or B-GEn.2, such as an enzyme-inactivating variant).
[0137] Codon optimization strategies and methods are known in the art and have been described for various systems, including but not limited to yeast (Outchkourov et al., Protein Expr Purif, 24(1):18-24 (2002)) and Escherichia coli (E. coli) (Feng et al., Biochemistry, 39(50):15399-15409 (2000)). In some embodiments, codon optimization is performed using GeneGPS® Expression Optimization Technology (ATUM) and using a manufacturer-recommended expression optimization algorithm. In some embodiments, the nucleic acids of this disclosure are codon-optimized to increase expression in human cells. In some embodiments, the nucleic acids of this disclosure are codon-optimized to increase expression in E. coli cells. In some embodiments, the nucleic acids of this disclosure are codon-optimized to increase expression in insect cells. In some embodiments, the nucleic acids of this disclosure are codon-optimized to increase expression in Sf9 insect cells. In some implementations, the expression optimization algorithm used in the codon optimization process is defined as avoiding presumed polyA signals (e.g., AATAAA and ATTAAA) and long (greater than 4) A segments that may cause polymerase slippage.
[0138] As is known in the art, codon optimization of nucleotide sequences results in a nucleotide sequence having less than 100% identity with the natural nucleotide sequence (e.g., less than 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%), but the encoded polypeptide still has the same function as the polypeptide encoded by the original natural nucleotide sequence. Therefore, in representative embodiments of this disclosure, the nucleotide sequence and / or recombinant nucleic acid of this disclosure may have optimized codons for expression in a specific target species.
[0139] In some embodiments, the codon-optimized nucleic acid sequence has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.2%, 99.5%, 99.8%, 99.9%, or 100% sequence identity with SEQ ID NO: 4, SEQ ID NO: 5, or SEQ ID NO: 6. In some embodiments, the nucleic acid of this disclosure is codon-optimized to increase the expression of the engineered B-GEn polypeptide in target cells or host cells. In some embodiments, the nucleic acid of this disclosure is codon-optimized to increase expression in human cells. Generally, the nucleic acid of this disclosure is codon-optimized to increase expression in any human cell. In some embodiments, the nucleic acid of this disclosure is codon-optimized to increase expression in *E. coli* cells. In some embodiments, the nucleic acid of this disclosure is codon-optimized to increase expression in insect cells. Generally, the polynucleotide of this disclosure is codon-optimized to increase expression in any insect cell. In some implementations, the polynucleotides of this disclosure are codon-optimized to increase expression in Sf9 insect cell expression systems.
[0140] Polyadenylation signaling can also be selected to optimize expression in the intended host.
[0141] 6.8.3. Nucleic acid modification In some embodiments, nucleic acids (e.g., guide RNA, nucleic acids containing a nucleotide sequence encoding a guide RNA; nucleic acids encoding site-specific modifying enzymes, such as engineered B-GEn peptides of this disclosure; etc.) contain modifications or sequences that provide additional desired properties (e.g., modified or regulated stability; subcellular targeting; tracking, such as fluorescent tags; binding sites for proteins or protein complexes; etc.). Non-limiting examples include: 5' caps (e.g., 7-methylguanylate cap, m7G); 3' poly(A) tails (e.g., 3' poly(A) tails); riboswitch sequences (e.g., those that allow regulation of stability and / or accessibility of regulated proteins and / or protein complexes); stability control sequences; sequences that form dsRNA duplexes (e.g., hairpin sequences); modifications or sequences that target RNA to subcellular locations (e.g., the nucleus, mitochondria, chloroplasts, etc.); modifications or sequences that provide tracking (e.g., direct conjugation to fluorescent molecules, conjugation to parts that facilitate fluorescence detection, sequences that allow fluorescence detection, etc.); modifications or sequences that provide protein binding sites (e.g., proteins that act on DNA, including transcriptional activators, transcriptional repressors, DNA methyltransferases, DNA demethylases, histone acetyltransferases, histone deacetylases, etc.); and combinations thereof.
[0142] In some implementations, the guide RNA includes an additional segment at the 5' or 3' end providing any of the features described above. For example, a suitable third segment may include a 5' cap (e.g., a 7-methylguanosine cap (m7G); a 3' polyadenylated tail (e.g., a 3' poly(A) tail); a ribo-switching sequence (e.g., allowing regulation of stability and / or accessibility of regulated proteins and / or protein complexes); a stability control sequence; a sequence forming the dsRNA duplex (e.g., a hairpin sequence); a sequence targeting the RNA to subcellular locations (e.g., the nucleus, mitochondria, chloroplasts, etc.); modifications or sequences providing tracking (e.g., direct conjugation to fluorescent molecules, conjugation to portions favorable for fluorescence detection, sequences allowing fluorescence detection, etc.); modifications or sequences providing protein binding sites (e.g., proteins acting on DNA, including transcriptional activators, transcriptional repressors, DNA methyltransferases, DNA demethylases, histone acetyltransferases, histone deacetylases, etc.); and combinations thereof.
[0143] Modifications can also be used, or alternatively, to reduce the likelihood or extent to which RNA introduced into cells triggers an innate immune response. Such responses, as described below and in the art, have been well characterized in the context of RNA interference (RNAi), including small interfering RNAs (siRNAs), which tend to be associated with a shortened RNA half-life and / or the activation of cytokines or other factors related to the immune response.
[0144] The RNA encoding the engineered B-GEn peptide introduced into the cell can also be modified by one or more types of modifications, including, but not limited to, modifications that enhance RNA stability (e.g., by reducing degradation by RNases present in the cell), modifications that enhance the translation of the resulting product (e.g., endonucleases), and / or modifications that reduce the likelihood or extent to which the RNA introduced into the cell elicits an innate immune response. Combinations of modifications such as those described above and others can also be used. In the case of engineered B-GEn peptides, for example, the guide RNA can be modified by one or more types of modifications (including those exemplified above), and / or the RNA encoding the engineered B-GEn peptide can be modified by one or more types of modifications (including those exemplified above).
[0145] By way of example, the guide RNA or other smaller RNAs used in the CRISPR-B-GEn system can be readily synthesized chemically, allowing for the easy incorporation of numerous modifications, as described below and in the art. While chemical synthesis procedures have expanded, the significant increase in nucleic acid length—exceeding approximately 100 nucleotides—often makes the purification of these RNAs using procedures such as high-performance liquid chromatography (HPLC, which avoids the use of gels such as PAGE) more challenging. One approach to generating longer, chemically modified RNAs is to produce two or more molecules linked together. Longer RNAs, such as those encoding B-GEn.1, or B-GEn.1.2, or B-GEn.2 endonucleases, are more readily produced enzymatically. While the types of modifications typically available for enzymatically produced RNA are limited, some modifications can be used to enhance stability, reduce the likelihood or extent of innate immune responses, and / or enhance other properties, as further described below and in the art; and novel modifications are constantly being developed.
[0146] By illustrating various types of modifications, particularly those frequently used for smaller chemically synthesized RNAs, modifications can include one or more nucleotides modified at the 2' position of the sugar, in some embodiments as 2'-O-alkyl, 2'-O-alkyl-O-alkyl, or 2'-fluoro modified nucleotides. In some embodiments, RNA modifications include 2'-fluorine, 2'-amino, and 2'-O-methyl modifications on the pyrimidine ribose, base residues, or a reverse base at the 3' end of the RNA. Such modifications are often incorporated into oligonucleotides, which have been shown to have a higher Tm (e.g., higher target binding affinity) than 2'-deoxy oligonucleotides targeting a given target.
[0147] Many nucleotide and nucleoside modifications have been shown to make the oligonucleotides incorporating them more resistant to nuclease digestion than the native oligonucleotides; these modified oligonucleotides also survive longer intact than unmodified oligonucleotides. Specific examples of modified oligonucleotides include those containing a modified backbone, such as phosphosulfate, phosphotriester, methyl phosphonate, short-chain alkyl or cycloalkyl sugar bonds, or short-chain heteroatom or heterocyclic sugar bonds. Some oligonucleotides are oligonucleotides with a phosphate backbone and oligonucleotides with a heteroatom backbone, particularly CH2-NH-O-CH2, CH,-N(CH3)-O-CH2 (called methylene(methylimino) or MMI backbone), CH2-ON(CH3)-CH2, CH2-N(CH3)-N(CH3)-CH2 and ON(CH3)-CH2-CH2; amide backbone [see De Mesmaeker et al., 1995, Ace. Chem. Res., 28:366-374]; morpholine backbone structure (see Summerton and Weller, U.S. Patent No. 5,034,506); peptide nucleic acid (PNA) backbone (in which the phosphodiester backbone of the oligonucleotide is replaced by a polyamide backbone, and the nucleotide is directly or indirectly linked to the aza-nitrogen atom of the polyamide backbone, see Nielse et al., 1991, Science 254:1497). Phosphorus-containing bonds include, but are not limited to: thiophosphates, chiral thiophosphates, dithiophosphates, phosphate triesters, aminoalkyl phosphate triesters, methyl and other alkyl phosphonates, including 3'-alkyl phosphonates and chiral phosphonates, phosphonates, phosphoramidates, including 3'-aminophosphates and aminoalkyl phosphates, thiophosphoramidates, thioalkyl phosphonates, thioalkyl phosphate triesters, and borophosphates having a normal 3'-5' bond, analogs of these with 2'-5' linkages, and analogs with reverse polarity, wherein adjacent nucleoside unit pairs are linked in a 3'-5' to 5'-3' or 2'-5' to 5'-2' manner; see U.S. Patent Nos. 3,687,808; 4,469,8 63; 4,476,301; 5,023,243; 5,177,196; 5,188,897; 5,264,423; 5,276,019; 5,278,302; 5,286,717; 5,321,131; 5,399,676; 5,405,939; 5,453,496; 5,455,233; 5,466,677; 5,476,925; 5,519,126; 5,536,821; 5,541,306; 5,550,111; 5,563,253; 5,571,799; 5,587,361; and 5,625,050.
[0148] Oligomeric compounds based on morpholino groups are described in Braasch and Corey, Biochemistry, 41(14):4503-4510 (2002); Genesis, Vol. 30, No. 3, (2001); Heasman, Dev. Biol., 243:209-214 (2002); Nasevicius et al., Nat. Genet., 26:216-220 (2000); Lacenraetc., Proc. Nat / . Acad. Sci., 97: 9591-9596 (2000); and U.S. Patent No. 5,034,506, Grant Announcement: July 23, 1991. Cyclohexenyl nucleic acid oligonucleotide mimics are described in Wang et al., J. Am. Chem. Soc., 122: 8595-8602 (2000). Among them, the modified oligonucleotide backbone excluding phosphorus atoms has a backbone formed by short-chain alkyl or cycloalkyl nucleoside inter-bonds, mixed heteroatoms and alkyl or cycloalkyl nucleoside inter-bonds, or one or more short-chain heteroatoms or heterocyclic nucleoside inter-bonds. These include skeletons containing morpholine bonds (partially formed from the sugar moiety of nucleosides); siloxane skeletons; sulfide, sulfoxide, and sulfone skeletons; formacetyl and thioformacetyl skeletons; methylene formacetyl and thioformacetyl skeletons; olefin-containing skeletons; aminosulfonate skeletons; methyleneimine and methylenehydrazine skeletons; sulfonate and sulfonamide skeletons; amide skeletons; and other skeletons having a mixture of N, O, S, and CH2 component moieties; see U.S. Patent Nos. 5,034,506; 5,166,315; 5,185,444; 5,214,134; 5,216,141; 5,235,033; 5,264,562; 5,264,564; 5,405,938; 5,434,257; 5,466,677; 5,470,967; 5,489,677; 5,541,307; 5,561,225; 5,596,086; 5,602,240; 5,610,289; 5,602,240; 5,608,046; 5,610,289; 5,618,704; 5,623,070; 5,663,312; 5,633,360; 5,677,437; and 5,677,439, each of which is incorporated herein by reference.
[0149] It may also include one or more substituted glycosyl groups, such as one of the following at the 2' position: OH, SH, SCH3, F, OCN, OCH3, OCH3O(CH2)nCH3, O(CH2)nNH2 or O(CH2)nCH3, where n is 1 to 10; C1-C10 low alkyl, alkoxyalkoxy, substituted low alkyl, alkylaryl or aralkyl; Cl; Br; CN; CF3; OCF3; O-, S- or N-alkyl; O-, S- or N-alkenyl; SOCH3; SO2CH3; ONO2; NO2; N3; NH2; heterocycloalkyl; heterocycloalkaryl; aminoalkylamino; polyalkylamino; substituted silyl; RNA cleaving group; reporter group; intercalator; group that improves the pharmacokinetic properties of oligonucleotides; or group used to improve the pharmacodynamic properties of oligonucleotides and other substituents with similar properties. In some embodiments, modifications include 2'-methoxyethoxy (2'-O-CH2CH2OCH3, also known as 2'-O-(2-methoxyethyl)) (Martinet a / , Helv. Chim. Acta, 1995, 78, 486). Other modifications include 2'-methoxy (2'-O-CH3), 2'-propoxy (2'-OCH2CH2CH3), and 2'-fluorine (2'-F). Similar modifications can also be made at other positions on the oligonucleotide, particularly at the 3' position of the sugar at the 3' end of the nucleotide and the 5' position of the 5' end of the nucleotide. The oligonucleotide can also have sugar mimics, such as a cyclobutyl group replacing the pentofuranosyl group. In some embodiments, the sugar and nucleoside internucleotide bonds of the nucleotide unit, such as the backbone, are replaced by new groups. The retained base units are used for hybridization with a suitable nucleic acid target compound. One such oligomer (an oligonucleotide mimic that has been shown to have excellent hybridization properties) is called peptide nucleic acid (PNA). In PNA compounds, the sugar backbone of the oligonucleotide is replaced by an amide-containing backbone, such as an aminoethylglycine backbone. Nucleotides are retained and directly or indirectly attached to the nitrogen atom of the amide moiety of the backbone. Representative U.S. patents teaching the preparation of PNA compounds include, but are not limited to, U.S. Patent Nos. 5,539,082; 5,714,331; and 5,719,262. Further teachings on PNA compounds are described in Nielsen et al., Science, 254: 1497-1500 (1991).
[0150] Guide RNAs may also contain additional or alternative nucleobase modifications or substitutions (often simply referred to in the art as "bases"). As used herein, "unmodified" or "native" nucleobases include adenine (A), guanine (G), thymine (T), cytosine (C), and uracil (U). Modified nucleobases include those found only rarely or transiently in native nucleic acids, such as hypoxanthine, 6-methyladenine, 5-Mepyrimidines, particularly 5-methylcytosine (also known as 5-methyl-2'deoxycytosine, commonly referred to in the art as 5-Me-C), 5-hydroxymethylcytosine (HMC), glycosyl HMC, and gentobiosyl HMC. HMC), and synthetic nucleobases such as 2-aminoadenine, 2-(methylamino)adenine, 2-(imidazolidinyl)adenine, 2-(aminoalkyl)adenine or other heterosubstituted alkyladenines, 2-thiouracil, 2-thiothymidine, 5-bromopyrimidine, 5-hydroxymethyluracil, 8-azaguanine, 7-deazaguanine, N6(6-aminohexyl)adenine and 2,6-diaminopurine. Kornberg, A, DNAReplication, WH Freeman & Co., San Francisco, pp75-77 (1980); Gebeyehu et al., Nucl. Acids Res. 15:4513 (1997). It may also include “universal” bases known in the art, such as inosine. 5-Me-C substitution has been shown to improve the stability of nucleic acid duplexes by 0.6–1.2 degrees Celsius (Sanghvi, YS, in Crooke, ST and Lebleu, B., eds., Antisense Research and Applications, CRCPress, Boca Raton, 1993, pp. 276–278), and is an implementation scheme for base substitution.
[0151] Modified nucleobases include other synthetic and natural nucleobases, such as 5-methylcytosine (5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine and 2-mercaptocytosine, 5-halouracil and cytosine, 5-propynyluracil and cytosine, 6-azouracil, and cytosine. And thymine, 5-uracil (pseudouracil), 4-thionuridine, 8-halogenated, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxy and other α-substituted adenine and guanine, 5-halogenated, especially 5-bromo, 5-trifluoromethyl and other 5-substituted uracil and cytosine, 7-methylquinoline and 7-methyladenine, 8-azaguanine and 8-azaadenine, 7-deazaguanine and 7-deazaadenine, and 3-deazaguanine and 3-deazaadenine.
[0152] Other useful nucleobases include those disclosed in U.S. Patent No. 3,687,808, those disclosed in "The Concise Encyclopedia of Polymer Science and Engineering," pp. 858-859, edited by Kroschwitz, Jl. John Wiley & Sons, 1990, those disclosed in Englisch et al., Angewandte Chemie, International Edition, 1991, 30, p. 613, and those disclosed in Sanghvi, YS., Chapter 15, Antisense Research and Applications, pp. 289-302, Crooke, ST. and Lebleu, B. ea., CRC Press, 1993. Some of these nucleobases are particularly useful for increasing the binding affinity of the oligomers of this disclosure. These include 5-substituted pyrimidines, 6-azapyrimidines, and N-2, N-6, and -O-6 substituted purines, including 2-aminopropyladenine, 5-propynyluracil, and 5-propynylcytosine. 5-methylcytosine substitution has been shown to improve the stability of nucleic acid double strands by 0.6–1.2 °C (Sanghvi, YS, Crooke, ST, and Lebleu, B., eds, “Antisense Research and Applications”, CRC Press, Boca Raton, 1993, pp. 276–278), and is an embodiment of base substitution, even more specifically when combined with 2'-O-methoxyethyl sugar modification. The modified nucleobases are described in U.S. Patent Nos. 3,687,808, 4,845,205; 5,130,302; 5,134,066; 5,175,273; 5,367,066; 5,432,272; 5,457,187; 5,459,255; 5,484,908; 5,502,177; 5,525,711; 5,552,540; 5,587,469; 5,596,091; 5,614,617; 5,681,941; 5,750,692; 5,763,588; and 5,830,653. 6,005,096; and U.S. Patent Application Publication 20030158403.
[0153] Not all positions in a given oligonucleotide need to be uniformly modified, and in fact, more than one of the above modifications can be incorporated into a single oligonucleotide, or even into a single nucleoside within the oligonucleotide.
[0154] In some embodiments, the guide RNA and / or mRNA encoding a nuclease (e.g., B-GEn.1, or B-GEn.1.2, or B-GEn.2 of this disclosure) is capped using any of the currently available capping methods, such as mCAP, ARCA, or enzymatic capping, to create a viable mRNA construct that maintains biological activity and avoids self / non-self intracellular responses. In some embodiments, the guide RNA and / or mRNA encoding a nuclease (e.g., B-GEn.1, or B-GEn.1.2, or B-GEn.2 of this disclosure) is capped using the CleanCap™ (TriLink) co-transcriptional capping method.
[0155] In some embodiments, the guide RNA and / or the mRNA encoding the endonuclease of this disclosure includes one or more modifications selected from pseudouracil, N1-methylpseudouracil, and 5-methoxyuracil. In some embodiments, one or more N1-methylpseudouracils are incorporated into the guide RNA and / or the mRNA encoding the endonuclease of this disclosure to provide enhanced RNA stability and / or reduced protein expression and immunogenicity in animal cells (e.g., mammalian cells such as human and mouse cells). In some embodiments, the N1-methylpseudouracil modification is incorporated in combination with one or more 5-methylcytosines.
[0156] In some implementations, the guide RNA and / or mRNA (or DNA) encoding a nuclease (e.g., B-GEn.1, or B-GEn.1.2, or B-GEn.2) are chemically linked to one or more moieties or conjugates, which enhance the activity, cellular distribution, or cellular uptake of the oligonucleotide. These moieties include, but are not limited to, lipid moieties, such as cholesterol moieties (Letsinger et al., 1989, Proc. Nat / . Acad. Sci. USA 86: 6553-6556); bile acids (Manoharan et al., 1994, Bioorg. Med. Chem. Let. 4: 1053-1060); thioethers, such as hexyl-S-triphenylmethanethiol (Manoharan et al., 1992, Ann. N. Y Acad. Sci. 660:306-309 and Manoharan et al., 1993, Bioorg. Med. Chem. Let. 3:2765-2770); thiocholesterol (Oberhauser et al., 1992, Nucl. Acids Res. 20: 533-538); aliphatic chains, such as dodecyl glycol or undecyl residues (Kabanov et al., 1990, FEBS Lett., 259: 327-330 and Svinarchuk et al., 1993, Biochimie, 75: 49-54); phospholipids, such as di-hexadecyl-rac-glycerol or triethylammonium 1,2-di-O-hexadecyl-rac-glycerol-3-H-phosphonate triethylammonium (Manoharan et al., 1995, Tetrahedron Lett. 36:3651-3654 and Shea et al., 1990, Nucl.Acids Res. 18: 3777-3783); polyamines or polyethylene glycol chains (Mancharan et al., 1995, Nucleosides & Nucleotides 14:969-973); adamantaneacetic acid (Manoharan et al., 1995, Tetrahedron Lett.). 36:3651-3654); palmitic moiety (Mishra et al., 1995, Biochim. Biophys. Acta 1264:229-237); or octadecylamine or hexylamino-carbonyl-t-oxocholesterol moiety (Crooke et al., 1996, J. Pharmacol. Exp. Ther., 277: 923-937).See also U.S. Patent Nos. 4,828,979; 4,948,882; 5,218,105; 5,525,465; 5,541,313; 5,545,730; 5,552, 538; 5,578,717; 5,580,731; 5,580,731; 5,591,584; 5,109,124; 5,118,802; 5,138,045; 5,414,077; 5,486,603; 5,512,439; 5,578,718; 5,608,046; 4,587,044; 4,605,735; 4,667,025; 4,762,779; 4,789,737; 4,824,941; 4,835,263; 4,876,335; 4,904,582; 4,958,013; 5,082,830; 5,112,963; 5,214,136; 5,082,830; 5,112,963; 5,214,136; 5,245,022; 5,254,469; 5,258,506; 5,262,536; 5,272,250; 5,292,873; 5,317,098; 5,371,241, 5,391,723; 5,416,203; 5,451,463; 5,510,475; 5,512,667; 5,514,785; 5,565,552; 5,567,810; 5,574,142; 5,585,481; 5,587,371; 5,595,726; 5,597,696; 5,599,923; 5,599,928 and 5,688,941.
[0157] Sugars and other components can be used to target proteins and nucleotide-containing complexes, such as cationic polyribosomes and liposomes, to specific sites. For example, hepatocyte-directed transfer can be mediated by the asialoglycoprotein receptor (ASGPR); see, e.g., Hu et al., 2014, Protein Pept Lett. 21(1 0):1025-30. Other systems known in the art and under continuous development can be used to target biomolecules and / or their complexes used in this application to specific target cells.
[0158] These targeting moieties or conjugates may include conjugated groups covalently linked to functional groups such as primary or secondary hydroxyl groups. Suitable conjugated groups include intercalators, reporter molecules, polyamines, polyamides, polyethylene glycols, polyethers, groups that enhance the pharmacodynamic properties of oligomers, and groups that enhance the pharmacokinetic properties of oligomers. Typical conjugated groups include cholesterol, lipids, phospholipids, biotin, phenazine, folic acid, phenanthridine, anthraquinones, acridine, fluorescein, rhodamine, coumarin, and dyes. Groups that can enhance pharmacodynamic properties include groups that improve uptake, enhance degradation resistance, and / or enhance specific hybridization with target nucleic acid sequences. Groups that can enhance pharmacokinetic properties include groups that improve the uptake, distribution, metabolism, or secretion of the compounds disclosed herein. Representative conjugated groups are disclosed in International Patent Application No. PCT / US92 / 09196, filed October 23, 1992, and U.S. Patent No. 6,287,860, which are incorporated herein by reference. The conjugated portion includes, but is not limited to, lipid groups such as cholesterol moieties, bile acids, thioethers (e.g., hexyl-5-triphenylmethanethiol), mercaptocholesterol, aliphatic chains (e.g., dodecyl glycol or undecyl residues), phospholipids (e.g., dihexadecyl-rac-glycerol or triethylammonium 1,2-bis-O-hexadecyl-rac-glycerol-3-H-phosphonate), polyamines or polyethylene glycol chains, or adamantaneacetic acid, palmitate moieties or octadecylamine or hexylamine-carbonyl-oxycholesterol moieties.See, for example, U.S. Patent Nos. 4,828,979; 4,948,882; 5,218,105; 5,525,465; 5,541,313; 5,545,730; 5,552,538; 5,578,717; 5,580,731; 5,580,731; 5,591,584; 5,109,124; 5,118,802; 5,138,045; 5,414,077; 5,486,603; 5,512,439; 5,578,718; 5,608,046; 4,587,044; 4,605,735; 4,667,025; 4,762,779; 4,789,737; 4,824,941; 4,835,263; 4,876,335; 4,904,582; 4,958,013; 5,082,830; 5,112,963; 5,214,136; 5,082,830; 5,112,963; 5,214,136; 5,245,022; 5,254,469; 5,258,506; 5,262,536; 5,272,250; 5,292,873; 5,317,098; 5,371,241, 5,391,723; 5,416,203; 5,451,463; 5,510,475; 5,512,667; 5,514,785; 5,565,552; 5,567,810; 5,574,142; 5,585,481; 5,587,371; 5,595,726; 5,597,696; 5,599,923; 5,599,928 and 5,688,941.
[0159] Longer nucleic acids that are less suited to chemical synthesis and are typically produced through enzymatic synthesis can also be modified in various ways. Such modifications may include, for example, the introduction of certain nucleotide analogs, the inclusion of specific sequences or other portions at the 5' or 3' end of the molecule, and other modifications. For example, the mRNA encoding B-GEn.1, or B-GEn.1.2, or B-GEn.2 is approximately 4 kb in length and can be synthesized via in vitro transcription. Modifications to mRNA can be used, for example, to increase its translation or stability (e.g., by increasing its resistance to degradation in cells), or to reduce the tendency of RNA to elicit an innate immune response, a tendency typically observed in cells after the introduction of exogenous RNA, particularly longer RNAs such as those encoding B-GEn.1 or B-GEn.2.
[0160] Many such modifications have been described in the art, such as poly-A tails, 5' cap analogs (e.g., anti-reverse cap analogs (ARCA) or m7G(5')ppp(5')G (mCAP)), modified 5' or 3' untranslated regions (UTRs), the use of modified bases (e.g., pseudo-UTP, 2-thio-UTP, 5-methylcytidine-5'-triphosphate (5-methyl-CTP), or N6-methyl-ATP), or treatment with phosphatases to remove the 5' terminal phosphate. These and other modification methods are known in the art, and new RNA modification methods are constantly being developed.
[0161] There are many commercial suppliers of modified RNA, including companies such as TriLink Biotech, Axolabs, Bio-Synthesis Inc., Dharmacon, and many others. As described by TriLink, 5-methyl-CTP, for example, can be used to confer desired properties, such as increased nuclease stability, enhanced translation, or reduced interaction between innate immune receptors and in vitro transcribed RNA. 5'-methylcytosine-5'-triphosphate (5-methyl-CTP), N6-methyl-ATP, and pseudo-UTP and 2-thio-UTP have also been shown to reduce innate immune stimulation in cultures and in vivo while enhancing translational capacity, as described in the publications of Konmann et al. and Warren et al., as mentioned below.
[0162] It has been shown that chemically modified mRNAs delivered in vivo can achieve better therapeutic effects; see, for example, Kormann et al., Nature Biotechnology 29, 154-157 (2011). Such modifications can be used, for example, to improve the stability of RNA molecules and / or reduce their immunogenicity. By using chemical modifications such as pseudo-U, N6-methyl-A, 2-thio-U, and 5-methyl-C, it has been found that simply replacing one-quarter of the uridine and cytidine residues with 2-thio-U and 5-methyl-C, respectively, can significantly reduce toll-like receptor (TLR)-mediated recognition of mRNA in mice. Therefore, by reducing the activation of the innate immune system, these modifications can be used to effectively improve the stability and lifespan of mRNA in vivo; see, for example, Konmann et al., ibid.
[0163] It has also been shown that repeated administration of synthetic messenger RNA with incorporated modifications designed to bypass innate antiviral responses can reprogram differentiated human cells into pluripotent cells. See, for example, Warren et al., Cell StemCell, 7(5):618-30 (2010). Such modified mRNAs, acting as primary reprogramming proteins, are an effective means of reprogramming various human cell types known as induced pluripotent stem cells (iPSCs). Furthermore, it has been found that enzymatically synthesized RNA incorporating 5-methyl-CTP, pseudo-UTP, and anti-reverse cap analogs (ARCA) can be used to effectively evade cellular antiviral responses; see, for example, Warren et al., ibid. Other modifications to nucleic acids described in this art include, for example, the use of poly-A tails, the addition of 5' cap analogs (such as m7G(5')ppp(5')G (mCAP)), modifications to the 5' or 3' untranslated region (UTR), or treatment with phosphatases to remove the 5' phosphate—and new methods are constantly being developed.
[0164] Numerous compositions and techniques have been developed for generating the modified RNA used in this paper, relating to RNA interference (RNAi) modifications, including small interfering RNA (siRNA). siRNAs face particular challenges in vivo because their effects on gene silencing via mRNA interference are often transient, potentially requiring repeated administration. Furthermore, siRNAs are double-stranded RNAs (dsRNAs), and mammalian cells have evolved immune responses to detect and neutralize dsRNAs, which are often byproducts of viral infections. Therefore, there are some mammalian enzymes, such as PKR (dsRNA response kinase) and potential retinoic acid-inducible gene I (RIG-I), which can mediate cellular responses to dsRNA, as well as toll-like receptors (such as TLR3, TLR7, and TLR8) that can trigger cytokine responses to these molecules; see, for example, Angart et al., Pharmaceuticals (Basel) 6(4): 440-468 (2013); Kanasty et al., Molecular Therapy 20(3): 513-524 (2012); Burnett et al., Biotechnol J. 6(9):1130-46 (2011); Judge and Maclachlan, Hum GeneTher 19(2):111-24 (2008); and the references cited therein.
[0165] Numerous modifications have been developed and applied to improve RNA stability, reduce innate immune responses, and / or achieve other benefits that may be useful in introducing nucleic acids into human cells as described herein; see, for example, Whitehead KA et al., Annual Review of Chemical and Biomolecular Engineering, 2:77-96 (2011); Gaglione and Messere, Mini Rev Med Chem, 10(7):578-95 (2010); Chernolovskaya et al., Curr Opin Mol Ther., 12(2):158-67 (2010); Deleavey et al., Curr Protoc Nucleic Acid Chem Chapter 16:Unit 16.3 (2009); Behlke, Oligonucleotides 18(4):305-19 (2008); Fucini et al., Nucleic Acid Ther 22(3):205-210 (2012); Bremsen et al., Front Genet 3:154 (2012).
[0166] As mentioned above, there are many commercial suppliers of modified RNA, many of which specialize in modifications aimed at improving the effectiveness of siRNA. Various approaches are offered based on findings reported in the literature. For example, Dharmacon notes that replacing non-bridging oxygen with sulfur (phosphorothioate (PS)) has been widely used to improve the nuclease resistance of siRNA, as reported in Kale, Nature Reviews Drug Discovery 11:125-140 (2012). Modification at the 2' position of the ribose has been reported to improve nuclease resistance to internucleotide phosphate bonds while increasing duplex stability (Tm), which has also been shown to provide immune-activating protection. Moderate PS backbone modifications combined with small, well-tolerated 2'-substitutions (2'-O-, 2'-fluorine, 2'-hydrogen) have been associated with the in vivo application of highly stable siRNAs, as reported by Soutschek et al., Nature 432:173-178 (2004); and 2'-O-methyl modifications have been reported to be effective in improving stability, as reported by Volkov, Oligonucleotides 19:191-202 (2009). Regarding the reduction of induction of innate immune responses, modifications of specific sequences with 2'-O-methyl, 2'-fluorine, 2'-hydrogen have been reported to reduce TLR7 / TLR8 interactions while generally maintaining silencing activity; see, for example, Judge et al., Mol. Ther. 13:494-505 (2006); and Cekaite et al., J. Mol. Biol. 365:90-108 (2007). Additional modifications, such as 2-thiouracil, pseudouracil, 5-methylcytosine, 5-methyluracil, and N6-methyladenosine, have also been shown to minimize TLR3, TLR7, and TLR8-mediated immune effects; see, for example, Kariko et al., Immunity 23:165-175 (2005).
[0167] As is known in the art and commercially available, many conjugates can be applied to nucleic acids, such as RNA, used herein to enhance cellular delivery and / or uptake of them, including, for example, cholesterol, tocopherol and folic acid, lipids, peptides, polymers, linkers and aptamers; see, for example, the review in Winkler, Ther. Deliv. 4:791-809 (2013), and the references cited therein.
[0168] 6.9. Vectors This disclosure provides vectors for the nucleic acids contained herein, such as those described in section 6.8. In some embodiments, the nucleic acid comprises a nucleic acid encoding an engineered B-GEn polypeptide as described in section 6.2. In some embodiments, at least for the portion encoding the nuclease component of the engineered B-GEn polypeptide, the engineered B-GEn polypeptide coding sequence is codon-optimized.
[0169] The vector (or nucleotide sequence) may further encode gRNA.
[0170] In some implementations, the vector containing the nucleotide sequence may be an expression vector.
[0171] In some embodiments, the expression vector is a production vector for engineered B-GEn peptides, for example, that can be used to express / produce engineered B-GEn peptides in host cells. After expressing / producing the engineered B-GEn peptide in host cells, the engineered B-GEn peptide can be incorporated into an RNP for nuclear transfection of target cells.
[0172] Alternatively, the expression vector containing the nucleotide sequence can be a delivery vector for an engineered B-GEn peptide, for example, used to introduce the coding sequence of an engineered B-GEn peptide into target cells intended for gene editing. After expression / production of the engineered B-GEn peptide in the target cells, the engineered B-GEn peptide, together with a guide RNA molecule, is capable of editing the target cells. In some embodiments, the delivery vector further includes the coding sequence of gRNA. In other embodiments, a separate nucleic acid encoding gRNA is introduced into the target cells.
[0173] The expression vectors considered include, but are not limited to, viral vectors based on the following viruses: vaccinia virus, poliovirus, adenovirus, adeno-associated virus, SV40, herpes simplex virus, human immunodeficiency virus, retroviruses (e.g., Murine Leukemia Virus, Spleen necrosis virus); and retroviruses (e.g., Rous Sarcoma Virus, Harvey Sarcoma Virus, Avian leukosis virus, lentivirus, human immunodeficiency virus, myeloproliferative sarcoma virus, and mammary tumor virus). Vectors derived from the virus and other recombinant vectors. Other vectors considered for use in eukaryotic target cells include, but are not limited to, vectors pXT1, pSG5, pSVK3, pBPV, pMSG, and pSVLSV40 (Pharmacia). Additional vectors anticipated for use in eukaryotic cells include, but are not limited to, vectors pCTx-1, pCTx-2, and pCTx-3. Other vectors may be used, provided they are compatible with the intended host or target cells.
[0174] In some implementations, the expression vector has one or more transcriptional and / or translational control elements. Depending on the expression cell / vector system used, many suitable transcriptional and translational control elements can be used in the vector, including constitutive and inducible promoters, transcriptional enhancer elements, transcription terminators, etc. The vector may also contain a ribosome binding site for translation initiation and a transcription terminator.
[0175] Non-limiting examples of suitable eukaryotic promoters (i.e. promoters that function in eukaryotic cells) include promoters derived from the immediate early stage of cytomegalovirus (CMV), promoters of thymidine kinase from herpes simplex virus (HSV), early and late SV40 promoters, promoters of long terminal repeats (LTRs) from retroviruses, the human elongation factor-1 promoter (EF1), promoters with hybrid structures of cytomegalovirus (CMV) enhancers fused to the chicken beta-actin promoter (CAG), the stem cell virus promoter (MSCV), the phosphoglycerate kinase-1 locus promoter (PGK), and the mouse metallothionein-I promoter.
[0176] In some embodiments, the promoter is an inducible promoter (e.g., heat shock promoter, tetracycline-regulated promoter, steroid-regulated promoter, metal-regulated promoter, estrogen receptor-regulated promoter, etc.). In some embodiments, the promoter is a constitutive promoter (e.g., CMV promoter, UBC promoter). In some embodiments, the promoter is a spatially restricted and / or temporally restricted promoter (e.g., tissue-specific promoter, cell type-specific promoter, etc.). In some embodiments, if the gene will be expressed under an endogenous promoter present in the genome after insertion into the genome, the vector does not have a promoter for the expression of at least one gene in the host cell.
[0177] For the expression of small RNAs, including guide RNAs, various promoters, such as RNA polymerase III promoters, including, for example, U6 and H1, can be advantageous. Therefore, such promoters can be advantageously incorporated into delivery vectors. Descriptions and parameters for enhancing the use of such promoters are known in the art, and additional information and methods are frequently described; see, for example, Ma, H. et al., Molecular Therapy - Nucleic Acids 3, e161 (2014) doi:10.1038 / mtna.2014.12.
[0178] In some implementations, the vector is a self-inactivating vector that inactivates the viral sequence or components or other elements of the CRISPR mechanism. Self-inactivating vectors are particularly suitable for delivering vectors to screen cells that retain the engineered B-GEn polypeptide coding sequence after gene editing.
[0179] In some embodiments, the expression vector is an RNA vector. In other embodiments, the expression vector is a DNA vector.
[0180] 6.9.1. RNA Vector The expression vector of this disclosure may be an RNA vector.
[0181] Particularly suitable vectors are viral replicons based on RNA viruses, such as alphaviruses and paramyxoviruses. Alphavirus and paramyxovirus replicons do not involve DNA intermediates used for replication, and therefore provide a safer alternative to several other commonly used viral vectors, including lentiviruses and retroviruses (Yoshioka et al., 2013, Cell Stem Cell. 13(2):246-54; Yoshioka and Dowdy, 2017, PLOS ONE 12:e0182018). Alphaviruses are lipid-enveloped, positive-sense RNA viruses that constitute more than 30 genera of viruses in the Togaviridae family, including Eastern, Western, and Venezuelan equine encephalitis viruses (EEEV, WEEV, and VEEV, respectively), chikungunya (CHIK), Sindbis, Ross River, and O'nyong-nyong viruses, etc. Sendai viruses (SeV) are enveloped, single-stranded antisense paramyxoviruses that replicate freely in the cytoplasm of host cells.
[0182] Therefore, in some implementations, the RNA vector is derived from an RNA virus, such as alphavirus, paramyxovirus, flavivirus, rhabdovirus, measles virus, or picornavirus.
[0183] In some embodiments, the RNA vector is a single-stranded RNA replicon. In some embodiments, the single-stranded RNA replicon is positive-stranded. In some other embodiments, the single-stranded RNA replicon is negative-stranded. In some embodiments, the RNA vector comprises one or more coding sequences of one or more engineered B-GEn peptides and autonomous replication elements.
[0184] The RNA replicones disclosed herein typically include regulatory elements and a subgenomic (SG) promoter operatively linked to an engineered B-GEn polypeptide coding sequence. The sequence containing the engineered B-GEn polypeptide coding sequence is typically flanked by 5' and 3' UTR sequences, with the 3' UTR sequence typically followed by a polyadenylation signal.
[0185] The RNA vector construct can be generated from a DNA template (DNA plasmid construct). For example, the RNA construct can be transcribed from a DNA template using an SP6 or T7 in vitro transcription kit.
[0186] RNA vectors are particularly useful as delivery vectors.
[0187] 6.9.2. DNA Vectors In some embodiments, the expression vector of this disclosure is a DNA vector. This disclosure provides two types of DNA vectors: (1) DNA vectors as production or delivery vectors, and (2) DNA vectors from which RNA vectors of the disclosure can be transcribed (as described in section 6.9.1), as described in section 6.9.2.2. DNA vectors are sometimes referred to herein as “template vectors” from which RNA replicons of the disclosure can be transcribed.
[0188] In some embodiments, the DNA vector of this disclosure is a non-integrating DNA vector. For example, the vector may be a free vector. For example, plasmids contained in many DNA viruses, such as adenovirus, simian vacuolating virus 40 (SV40), bovine papillomavirus (BPV), or budding yeast ARS (autonomous replication sequence), can be used without genome integration.
[0189] In some embodiments, the DNA vector of this disclosure includes a replication origin. Examples of replication origins that can be incorporated into the DNA vector of this disclosure include replication origins of lymphoherpesviruses, gamma herpesviruses, adenoviruses, bovine papillomaviruses, or yeast. In some embodiments, the replication origin is derived from a lymphoherpesvirus or gamma herpesvirus corresponding to oriP of EBV, as an autonomous replication element. In some embodiments, the lymphoherpesvirus is Epstein-Barr virus (EBV), Kaposi's sarcomaherpes virus (KSHV), herpesvirus saimiri (HS), or Marek's disease virus (MDV). Epstein-Barr virus (EBV) and Kaposi's sarcomaherpes virus (KSHV) are also examples of gamma herpesviruses.
[0190] In some embodiments, the vector of this disclosure contains the origin of EBV replication, OriP. OriP is a site at or near the DNA replication initiation point and consists of two cis-acting sequences approximately 1 kbp apart, referred to as a repeat family (FR) and a binary symmetry (DS). The FR consists of 21 imperfect copies of a 30 bp repeat and contains 20 high-affinity EBNA-1 binding sites. When the FR binds to EBNA-1, it acts as both a transcriptional enhancer for the cis-promoter and a transcriptional enhancer up to 10 kb away. The DS is sufficient to initiate DNA synthesis in the presence of EBNA-1, with initiation occurring at or near the DS.
[0191] One or more expression cassettes in the replication DNA vector may further contain a nucleotide sequence encoding a trans-acting factor that binds to the origin of replication to replicate an extrachromosomal template. Optionally or additionally, somatic cells may express this trans-acting factor.
[0192] In other embodiments, the DNA vector of this disclosure lacks a replication origin.
[0193] The DNA vectors disclosed herein typically contain one or more promoters, SP6 or T7, to drive the expression of engineered B-GEn peptides in the case of a DNA vector intended for production, or to drive the expression of RNA replicons in the case of a DNA vector intended for template.
[0194] 6.9.2.1. Expression Vector In some embodiments, the expression vector is a DNA vector containing an expression cassette for expressing one or more target proteins, the expression cassette being operatively linked to a regulatory element containing a promoter suitable for driving the expression of an engineered B-GEn polypeptide in a target cell type. Examples of promoters suitable for driving protein expression in mammalian cells include the cytomegalovirus (CMV) promoter, EF1a promoter, SV40 promoter, Ubc promoter, human β-actin promoter, PGK1 promoter, and CAG promoter.
[0195] DNA vectors used for the direct expression of engineered B-GEn peptides (rather than as templates for RNA expression vectors, as described in Section 6.9.2.2) do not need to include autonomously replicating sequences of RNA replicons, such as the nsP1-nsP4 proteins of VEEV or the NP, P, and L proteins of Sendai virus.
[0196] In some implementations, the DNA expression vector is a non-replicating DNA vector. In other implementations, the DNA expression vector is a replicating DNA vector.
[0197] 6.9.2.2. Template Vector The DNA vector of this disclosure can also serve as a template for transcription of RNA replicons as described herein. Therefore, the "expression cassette" contained in the template vector is intended for transcription of RNA replicons generated from the transcription of RNA replicons.
[0198] Therefore, the template vector of this disclosure contains a nucleotide sequence that encodes an RNA replicon as described herein under the control of a regulatory element (e.g., an SP6 or T7 promoter).
[0199] In some implementations, the template DNA vector is a non-replicating DNA vector. In other implementations, the template DNA vector is a replicating DNA vector.
[0200] In some embodiments, the template vector is used for in vitro transcription of an RNA replicon, which is then introduced into cells to drive the expression of an engineered B-GEn peptide.
[0201] 6.9.3. Viral Vectors Recombinant adeno-associated virus (AAV) vectors can be used for delivery. Known techniques in the art for producing rAAV particles involve providing a cell with the polynucleotide to be delivered between two AAV inverted terminal repeat (ITR) sequences, the AAV rep and cap genes, and helper viral function. Production of rAAV requires the following components in a single cell (referred to herein as a packaging cell): the target polynucleotide between the two ITRs, the AAV rep and cap genes separated from the AAV genome (i.e., not in the AAV genome), and the helper viral function. The AAV rep and cap genes can be derived from any recombinant virally derived AAV serotype, or from an AAV serotype different from the ITR on the packaging polynucleotide, including but not limited to AAV serotypes AAV-1, AAV-2, AAV-3, AAV-4, AAV-5, AAV-6, AAV-7, AAV-8, AAV-9, AAV-10, AAV-11, AAV-12, AAV-13, and AAV rh.74. For example, WO 01 / 83692 disclosed the production of the fake virus rAAV. The method for generating packaging cells involves creating cell lines that stably express all the components required for AAV particle production. For example, a plasmid (or multiple plasmids) is integrated into the cell's genome, containing the target polynucleotide between multiple AAV ITRs, AAV rep and cap genes independent of the AAV genome, and selectable markers such as neomycin resistance genes. The AAV genome has been introduced into bacterial plasmids via procedures such as GC tailing (Samulski et al., 1982, Proc. Natl. Acad. Sci. USA, 79:2077-2081), adding synthetic adapters containing restriction endonuclease cleavage sites (Laughlin et al., 1983, Gene, 23:65-73), or direct blunt-end joining (Senapathy & Carter, 1984, J. Biol. Chem., 259:4661-4666). The packaging cell lines are then infected with helper viruses, such as adenovirus. The advantage of this method is that the cells are selectable and suitable for large-scale production of rAAV. Examples of other suitable methods include using adenovirus or baculovirus, rather than plasmids, to introduce the rAAV genome and / or rep and cap genes into packaging cells.
[0202] The general principles of rAAV production are summarized in the following references, for example, Carter, 1992, Current Opinions in Biotechnology, 1533-539; and Muzyczka, 1992, Curr. Topics in Microbial. and lmmunol., 158:97-129). Several methods are described in Ratschin et al., Mol. Cell. Biol. 4:2072 (1984); Hermonat et al., Proc. Natl. Acad. Sci. USA, 81:6466 (1984); Tratschin et al., Mol. Cell. Biol. 5:3251 (1985); Mclaughlin et al., J. Virol., 62:1963 (1988); and Lebkowski et al., 1988 Mol. Cell. Biol., 7:349 (1988). Samulski et al. (1989, J. Virol., 63:3822-3828); U.S. Patent Nos. 5,173,414; WO 95 / 13365 and corresponding U.S. Patent Nos. 5,658,776; WO 95 / 13392; WO 96 / 17947; PCT / US98 / 18600;WO97 / 09441 (PCT / US96 / 14423); WO 97 / 08298 (PCT / US96 / 13872); WO 97 / 21825 (PCT / US96 / 20777); WO 97 / 06243 (PCT / FR96 / 01064); WO 99 / 11764; Perrin et al. (1995) Vaccine 13:1244-1250; Paul et al. (1993) Human Gene Therapy 4:609-615; Clark et al. (1996) Gene Therapy 3:1124-1132; U.S. Patent No. 5,786,211; U.S. Patent No. 5,871,982; and U.S. Patent No. 6,258,595. The AAV vector serotype used for transduction depends on the target cell type. For example, the following exemplary cell types are known to be transducible by indicated AAV serotypes, etc. Many suitable expression vectors are known to those skilled in the art, and many are commercially available. The following vectors are provided by way of example for eukaryotic host cells: pXT1, pSG5 (Stratagene), pSVK3, pBPV, pMSG, and pSVLSV40 (Pharmacia). However, any other vector may be used as long as it is compatible with the host cell.
[0203] 6.10. Host Cells and Recombinant Expression In some embodiments, host cells may be used to express gRNA, sgRNA, or engineered B-GEn peptides of this disclosure. Suitable host cells include naturally occurring cells; genetically modified cells (e.g., cells genetically modified in a laboratory); and cells manipulated in vitro in any way. In some embodiments, the host cells are isolated.
[0204] The host cell can be a eukaryote or a prokaryote, and includes, for example, yeast (e.g., Pichia pastoris or Saccharomyces cerevisiae), bacteria (e.g., E. coli or Bacillus subtilis), insect Sf9 cells (e.g., Sf9 cells infected with baculovirus), or mammalian cells (e.g., human embryonic kidney (HEK) cells, Chinese hamster ovary cells, HeLa cells, human 293 cells, and monkey COS-7 cells).
[0205] Host cells may be derived from established cell lines, or they may be primary cells, where “primary cells,” “primary cell lines,” and “primary cultures” are used interchangeably herein to refer to cells and cell cultures derived from the subject and permitted to grow in vitro for a limited number of passages (e.g., culture division). For example, primary cultures include cultures that have undergone 0, 1, 2, 4, 5, 10, or 15 passages, but have not experienced a sufficient number of crisis stages. Primary cell lines may be maintained in vitro for fewer than 10 passages. In some embodiments, the host cells are PSCs (e.g., iPSCs or ESCs) or PSC-derived cells (e.g., PSC-derived neurons, PSC-derived microglia, PSC-derived cardiomyocytes, PSC-derived eye cells).
[0206] If the cells are primary cells, they can be harvested from an individual using any suitable method. The harvested cells can be dispersed or suspended using appropriate solutions. The harvested cells can be used immediately, or they can be stored long-term, frozen, and reused after thawing. In this case, cells are typically frozen in 10% dimethyl sulfoxide (DMSO), 50% serum, 40% buffered medium, or some other such solutions commonly used in the art to preserve the cells at such freezing temperatures, and thawed according to the methods commonly used in the art for thawing frozen cultured cells.
[0207] 6.11. Target Cells In some embodiments, engineered type V endonucleases or B-GEn CRISPR-Cas systems are introduced into target cells or target cell populations. Methods for introducing proteins and nucleic acids into target cells are further described in section 6.12.
[0208] The target cells and target cell populations of this disclosure may be cells in which gene editing has been performed by the system of this disclosure, or cells in which components of the system of this disclosure have been introduced or expressed but have not yet undergone gene editing, or combinations thereof. In various embodiments, the cell population may include, for example, a population in which at least 1%, at least 5%, at least 10%, at least 15%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, or at least 70% of the cells have been gene-edited by the system of this disclosure.
[0209] In some embodiments, the methods of this disclosure can be used to induce transcriptional regulation of mitotic or postmitotic cells in vivo and / or in vitro and / or in vitro. In some embodiments, the methods of this disclosure can be used to induce DNA cleavage, DNA modification, and / or transcriptional regulation (e.g., to produce genetically modified cells that can be reintroduced into an individual) in vivo and / or in vitro and / or in vitro.
[0210] Because guide RNA provides specificity by hybridizing with target DNA, mitotic and / or post-mitotic cells can be any of a variety of target cells and can be modified in vivo or in vitro. Suitable target cells include, but are not limited to, bacterial cells; archaea; single-celled eukaryotic organisms; plant cells; algal cells, such as *Botryococcus braunii*, *Chlamydomonas reinhardtii*, *Nannochloropsis gaditana*, *Chlorela pyrenoidosa*, and *Sargassum patens (C. Agardh)*; fungal cells; animal cells; cells derived from invertebrates (such as insects, cnidarians, echinoderms, nematodes, etc.); eukaryotic parasites (such as *Plasmodium* parasites, such as *Plasmodium falciparum*, helminths, etc.); cells derived from vertebrates (such as fish, amphibians, reptiles, birds, mammals); and mammalian cells (such as rodent cells, human cells, and non-human primate cells). In some embodiments, the target cell can be any human cell. Suitable target cells include naturally occurring cells; genetically modified cells (such as those genetically modified in the laboratory, for example, by the “human hand”); and cells manipulated in vitro in any way. In some embodiments, the target cells are isolated.
[0211] Any type of cell can serve as a host cell or target cell for in vivo or in vitro modification. In various embodiments, the host cell or target cell is a stem cell (e.g., PSC, such as embryonic stem (ES) cells or induced pluripotent stem cells (iPSCs)); germ cells; somatic cells, such as fibroblasts, hematopoietic cells, immune cells (e.g., T lymphocytes, B lymphocytes, dendritic cells, or macrophages), neurons, muscle cells, osteocytes, hepatocytes, pancreatic cells; embryonic cells of an embryo at any stage, such as zebrafish embryos at 1-cell, 2-cell, 4-cell, 8-cell, etc.; etc.). Cells may be derived from established cell lines, or they may be primary cells, wherein “primary cells,” “primary cell lines,” and “primary cultures” are used interchangeably herein to refer to cells and cell cultures derived from a subject and permitted to grow in vitro for a limited number of passages (e.g., cell division). For example, primary cultures include those that have undergone 0, 1, 2, 4, 5, 10, or 15 passages, but have not experienced a sufficient number of crisis stages. Primary cell lines can be maintained in vitro for fewer than 10 passages. In some embodiments, the target cells are single-celled organisms or are grown in a culture. In some embodiments, the host cells are the same as the target cells. In some embodiments, the target cells are modified to a different cell type, resulting in host cells that are different from the target cells. As an example, the target cells may be PSCs (e.g., iPSCs), which then differentiate into PSC-derived cells (e.g., PSC-derived neurons), such that the host cells are neurons. Alternatively, the target cells may not be PSCs, but are subsequently dedifferentiated or reprogrammed to be PSCs.
[0212] If the cells are primary cells, they can be harvested from an individual using any suitable method. For example, leukocytes can be harvested via apheresis, leukocytapheresis, or density gradient separation, while cells from tissues (e.g., skin, muscle, bone marrow, spleen, liver, pancreas, lung, intestine, stomach, etc.) are best harvested via biopsy. Harvested cells can be dispersed or suspended using appropriate solutions. These solutions are typically balanced salt solutions, such as physiological saline, phosphate-buffered saline (PBS), Hank's balanced salt solution, etc., supplemented appropriately with fetal bovine serum or other naturally occurring factors, combined with acceptable low concentrations of buffer such as 5-25 mM. Suitable buffers include HEPES, phosphate buffer, lactate buffer, etc. Cells can be used immediately, or they can be stored long-term, frozen, and reused after thawing. In this case, cells are typically frozen in 10% dimethyl sulfoxide (DMSO), 50% serum, 40% buffered medium, or some other such solutions commonly used in the art to preserve the cells at such freezing temperatures and thaw them in accordance with the methods commonly used in the art for thawing frozen cultured cells.
[0213] In some embodiments, the target cells are located in a subject, and the methods of this disclosure include administering a B-GEn CRISPR-Cas system (e.g., a ribonucleoprotein comprising an engineered B-GEn peptide of this disclosure) to a subject (e.g., a mammalian subject such as a human or livestock subject). Exemplary in vivo cells to which the B-GEn CRISPR-Cas system can be delivered include, but are not limited to, fibroblasts, hematopoietic cells, immune cells (e.g., T lymphocytes, B lymphocytes, dendritic cells, or macrophages), neurons, glial cells, muscle cells (e.g., cardiomyocytes), osteoblasts, hepatocytes, and pancreatic cells.
[0214] 6.11.1. Pluripotent Stem Cells (PSCs) In some embodiments, the target cells are stem cells, such as, in particular, pluripotent stem cells (PSCs) such as induced pluripotent stem cells (iPSCs) or human embryonic stem cells (hESCs), which can be differentiated and used to generate large numbers of specific cell types that can be delivered to patients with many different diseases for regenerative medicine. In the context of PSCs, differentiation is a lineage-specific process and can be achieved using cell-specific protocols.
[0215] After PSCs (e.g., hESCs or iPSCs) are modified with engineered type V endonucleases or B-GEn peptides incorporating the present disclosure, PSCs can differentiate into target cell types for cell therapy. In some embodiments, the genome of the PSCs is edited with engineered B-GEn peptides of the present disclosure prior to differentiation.
[0216] The PSCs (e.g., PSCs containing engineered type V endonucleases or B-GEn peptides of this disclosure, or PSCs whose genomes have been edited by engineered type V endonucleases or B-GEn peptides of this disclosure) can differentiate into cells suitable for therapeutic use, including cells in the lineages of endoderm (e.g., lung, thyroid, or pancreatic cells or their progenitors), ectoderm (e.g., skin, neurons, or pigment cells or their progenitors), and mesoderm (e.g., cardiac cells, skeletal muscle cells, erythrocytes, smooth muscle cells or their progenitors or precursors).
[0217] In some embodiments, the PSCs of this disclosure differentiate into cardiac cells. In various embodiments, the cardiac cells are cardiac progenitor cells or mature or immature (atrial or ventricular) cardiomyocytes.
[0218] In other embodiments, the PSCs of this disclosure differentiate into oligodendrocyte progenitor cells or oligodendrocytes.
[0219] In other embodiments, the PSCs of this disclosure differentiate into neural lineage cells, such as neural crest cells, astrocytes, dopaminergic neuron progenitor cells, dopaminergic neuron cells, midbrain dopaminergic neuron progenitor cells, midbrain dopaminergic neurons, real midbrain dopaminergic (DA) neurons, dopaminergic neuron precursor cells, lamina midbrain progenitor cells, and lamina midbrain DA neurons.
[0220] In other embodiments, the PSCs of this disclosure differentiate into photoreceptor cells, photoreceptor precursor cells, retinal pigment epithelial cells, neuroretinal cells, or neuroretinal progenitor cells.
[0221] In some implementations, the PSCs of this disclosure differentiate into microglia or microglial progenitor cells.
[0222] In some implementations, the PSCs of this disclosure differentiate into macrophages.
[0223] In some implementations, the PSCs of this disclosure differentiate into intestinal progenitor cells or intestinal cells.
[0224] In some implementations, the PSCs of this disclosure differentiate into immune cells, such as T lymphocytes, B lymphocytes, dendritic cells, or macrophages.
[0225] In some implementations, PSCs can be genetically engineered (e.g., to produce functional proteins that are defective in patients, to produce therapeutic proteins, to include a shut-off switch, or to evade immune detection, thereby supporting allogeneic applications) before differentiating into the target cell type.
[0226] 6.12. Methods for Introducing Nucleic Acids into Host and Target Cells In some embodiments, the methods of this disclosure include introducing one or more nucleic acids into a host or target cell (or a host or target cell population), said nucleic acid comprising a nucleotide sequence encoding a guide RNA and / or a nucleotide sequence encoding an engineered B-GEn polypeptide (e.g., a codon-optimized nucleotide sequence). In some embodiments, the methods of this disclosure include introducing a guide RNA encoding an engineered B-GEn polypeptide and / or a nucleotide sequence (e.g., a codon-optimized nucleotide sequence) into a host or target cell (or a host or target cell population).
[0227] In some embodiments, the target cells (e.g., cells containing DNA targeted by guide RNA for editing with engineered B-GEn peptides) are in vitro cells, such as in cell cultures. In some embodiments, the target cells are in vivo, such as in a subject (e.g., a mammal, such as a human) intended to receive gene therapy.
[0228] In some embodiments, the nucleotide sequence encoding the guide RNA and / or the engineered B-GEn polypeptide is operatively linked to an inducible promoter. In some embodiments, the nucleotide sequence encoding the guide RNA and / or the engineered B-GEn polypeptide is operatively linked to a constitutive promoter.
[0229] Guide RNA or nucleic acids containing nucleotide sequences encoding guide RNA can be introduced into host or target cells using any of the various known methods. Similarly, when a method involves introducing a nucleic acid containing a nucleotide sequence encoding an engineered B-GEn polypeptide (e.g., a codon-optimized nucleotide sequence) into a host or target cell, such nucleic acid can be introduced into the host or target cell using any of the various known methods. Guide nucleic acids (RNA or DNA; e.g., guide RNA or one or more DNA molecules encoding guide RNA) and / or nucleic acids (RNA or DNA) encoding engineered B-GEn polypeptides can be delivered using viral or non-viral delivery vectors known in the art.
[0230] Methods for introducing nucleic acids into host or target cells are known in the art, and any known method can be used to introduce nucleic acids (such as expression constructs) into stem cells or progenitor cells. Suitable methods include, for example, viral or phage infection, transfection, conjugation, protoplast fusion, lipid infection, electroporation, nuclear transfection, calcium phosphate precipitation, polyethyleneimine (PEI)-mediated transfection, diethylaminoethyl dextran (DEAE-dextran)-mediated transfection, liposome-mediated transfection, particle gun technology, calcium phosphate precipitation, direct microinjection, nanoparticle-mediated nucleic acid delivery (see, for example, Panyam et al., Adv Drug Deliv Rev. Sep 13, 2012. pii: 50169-409X(12)00283-9. doi: 10.1016 / j.addr.2012.09.023), etc., including but not limited to exosome delivery.
[0231] Polynucleotide sequences can be delivered via non-viral delivery vectors, including but not limited to nanoparticles, liposomes, ribonucleoproteins, positively charged peptides, small RNA conjugates, aptamer-RNA chimeras, and RNA fusion protein complexes. Some exemplary non-viral delivery vectors are described in Peer and Lieberman, Gene Therapy, 18:1127-1133 (2011) (which focuses on non-viral delivery vectors for siRNA, which can also be used to deliver other nucleic acids).
[0232] Suitable systems and techniques for delivering nucleic acids (such as mRNA and sgRNA) of this disclosure for gene editing include lipid nanoparticles (LNPs). As used herein, the term "lipid nanoparticle" includes liposomes, regardless of their lamellarity, shape, or structure, and also includes lipid complexes used to introduce nucleic acids and / or peptides into cells. These lipid nanoparticles can be complexed with bioactive compounds (such as nucleic acids and / or peptides) and can be used as in vivo delivery carriers. Generally, any method known in the art can be used to prepare lipid nanoparticles comprising one or more nucleic acids of this disclosure, as well as to prepare complexes of bioactive compounds with said lipid nanoparticles. Examples of such methods have been widely published, such as in Biochim Biophys Acta 1979, 557:9; Biochim et Biophys Acta 1980, 601:559; Liposomes: A practical approach (Oxford University Press, 1990); Pharmaceutica Acta Helvetiae 1995, 70:95; Current Science 1995, 68:715; Pakistan Journal of Pharmaceutical Sciences 1996, 19:65; Methods in Enzymology 2009, 464:343. Particularly suitable systems and techniques for preparing LNP formulations comprising one or more nucleic acids and / or peptides comprising the present disclosure include, but are not limited to, those developed by Intellia (see, e.g., WO2017173054A1), Alnylam (see, e.g., WO2014008334A1), Modernatx (see, e.g., WO2017070622A1 and WO2017099823A1), TranslateBio, Acuitas (see, e.g., WO2018081480A1), Genevant Sciences, Arbutus Biopharma, Tekmira, Arcturus, Merck (see, e.g., WO2015130584A2), Novartis (see, e.g., WO2015095340A1), and Dicerna; all of these documents are incorporated herein by reference in their entirety.
[0233] Suitable nucleic acids containing nucleotide sequences encoding engineered B-GEn peptides and / or guide RNA include expression vectors. In some embodiments, the expression vector is a viral construct, such as a recombinant adeno-associated virus construct (see, e.g., U.S. Patent No. 7,078,387), a recombinant adenovirus construct, a recombinant lentivirus construct, a recombinant retrovirus construct, etc.Suitable expression vectors include, but are not limited to, viral vectors (e.g., vaccine-based viral vectors; poliovirus-based viral vectors; adenovirus-based viral vectors (see, for example, Li et al., Invest Opthalmol Vis Sci 35:2543 2549, 1994; Borras et al., Gene Ther 6:515 524, 1999; Li and Davidson, PNAS92:7700 7704, 1995; Sakamoto et al., H Gene Ther 5:10881097, 1999; WO 94 / 12649, WO 93 / 03769; WO 93 / 19191; WO 94 / 28938; WO 95 / 11984 and WO 95 / 00655); adeno-associated viruses (see, for example, Ali et al., Hum Gene Ther 9:81). 86, 1998, Flannery et al., PNAS 94:6916 6921, 1997; Bennett et al., Invest Opthalmol Vis Sci 38:2857 2863, 1997; Jomary et al., Gene Ther 4:683-690, 1997; Rolling et al., Hum Gene Ther 10:641 648, 1999; Ali et al., Hum Mol Genet 5:591 594, 1996; Srivastava in WO 93 / 09239, Samulski et al., J. Vir. (1989) 63:3822-3828; Mendelson et al., Viral. (1988) 166:154-165; and Flotte et al., PNAS (1993) 90:10613-10617); SV40; herpes simplex virus; human immunodeficiency virus (see, e.g., Miyoshi et al., PNAS 94:10319-23, 1997; Takahashi et al., J Virol 73:7812 7816, 1999); retroviral vectors (e.g., murine leukosis virus, spleen necrosis virus, retroviral-derived vectors such as Rous sarcoma virus, Harvey sarcoma virus, avian leukosis virus, lentivirus, human immunodeficiency virus, myeloproliferative sarcoma virus, and mammary tumor virus), etc.
[0234] In some implementations, B-GEn ribonuclease and sgRNA are delivered to target cells via nuclear transfection, wherein nuclear transfection is a method of delivering nucleic acids to cells by creating transient pores in the cell membrane using cell-specific reagents and electrical parameters.
[0235] Engineered B-Gen V CRISPR-Cas systems can be delivered to target cells via delivery vectors (e.g., viral vectors). Engineered B-Gen V CRISPR-Cas systems can also be delivered to target cells via non-viral delivery vectors, including but not limited to nanoparticles, liposomes, ribonucleoproteins, positively charged peptides, small RNA conjugates, aptamer-RNA chimeras, and RNA-fusion protein complexes. Some exemplary non-viral delivery vectors are described in Peer and Lieberman, GeneTherapy, 18: 1127-1133 (2011).
[0236] In some implementations, the engineered B-Gen V CRISPR-Cas system is delivered to target cells via a delivery vector as described in Section 6.9.
[0237] 6.13. Methods for Gene Editing This disclosure also provides methods for gene editing using engineered type V endonucleases or B-GEn peptide systems. Methods for gene editing may include: targeting, editing, modifying, or manipulating target DNA at one or more locations in the genome of a host or target cell (or multiple host or target cells), said operations may be performed in vitro, ex vivo, or in vivo, or performed on the target DNA in a cell-free environment. Typically, methods for gene editing include: introducing an engineered type V endonuclease or B-GEn peptide system into a host or target cell (or a population of host or target cells) or placing it in a cell-free environment containing the target DNA sequence under conditions suitable for one or more modifications (e.g., nicking or cutting or base editing) of the engineered type V endonuclease or B-GEn peptide on the target DNA, wherein the engineered type V endonuclease or B-GEn peptide, in its processed or unprocessed form, is guided to the target DNA by a guide RNA.
[0238] Methods of gene editing may include introducing the engineered type V endonuclease or B-GEn polypeptide system of this disclosure into a host or target cell (or host or target cell population) via an RNP complex and / or introducing one or more nucleic acids (e.g., via a virus such as AAV or via LNP) into a host or target cell (or host or target cell population), said one or more nucleic acids comprising a guide RNA or a nucleotide sequence encoding a guide RNA and a nucleotide sequence encoding an engineered type V endonuclease or B-GEn polypeptide (e.g., a codon-optimized nucleotide sequence). The engineered type V endonuclease or B-GEn polypeptide system of this disclosure (e.g., in an RNP complex; one or more nucleic acid molecules encoding a guide RNA and one or more nucleotide sequences (e.g., one or more codon-optimized nucleotide sequences) encoding an engineered type V endonuclease or B-GEn polypeptide; or a guide RNA encoding an engineered type V endonuclease or B-GEn polypeptide and one or more nucleotide sequences (e.g., one or more codon-optimized nucleotide sequences) can be introduced into a host or target cell (or host or target cell population) by any of a variety of well-known viral or non-viral delivery methods. The gene editing methods can be used in vivo, in vitro, or in vitro in a host or target cell (or host or target cell population).
[0239] In some embodiments, the gene editing method comprises introducing an engineered type V endonuclease or B-GEn polypeptide system of the present disclosure (e.g., in an RNP complex; one or more nucleic acid molecules encoding a guide RNA and one or more nucleotide sequences encoding an engineered type V endonuclease or B-GEn polypeptide (e.g., one or more codon-optimized nucleotide sequences); or a guide RNA and one or more nucleotide sequences encoding an engineered type V endonuclease or B-GEn polypeptide (e.g., one or more codon-optimized nucleotide sequences) into the target DNA of a subject (e.g., a mammalian subject, such as a human subject) for one or more modifications (e.g., nicking or cutting or base editing), wherein gene editing (e.g., for gene therapy purposes) is desired.
[0240] In some embodiments, the gene editing method includes introducing an engineered type V endonuclease or B-GEn polypeptide system of the present disclosure (e.g., in an RNP complex; one or more nucleic acid molecules encoding a guide RNA and one or more nucleotide sequences encoding an engineered type V endonuclease or B-GEn polypeptide (e.g., one or more codon-optimized nucleotide sequences); or a guide RNA and one or more nucleotide sequences encoding an engineered type V endonuclease or B-GEn polypeptide (e.g., one or more codon-optimized nucleotide sequences) to perform one or more modifications (e.g., nicking or cutting or base editing) in ex vivo target DNA.
[0241] In some embodiments, the gene editing method includes introducing an engineered type V endonuclease or B-GEn polypeptide system of the present disclosure (e.g., in an RNP complex; one or more nucleic acid molecules encoding a guide RNA and one or more nucleotide sequences encoding an engineered type V endonuclease or B-GEn polypeptide (e.g., one or more codon-optimized nucleotide sequences); or a guide RNA and one or more nucleotide sequences encoding an engineered type V endonuclease or B-GEn polypeptide (e.g., one or more codon-optimized nucleotide sequences) to perform one or more modifications (e.g., nicking or cutting or base editing) on target DNA in vitro.
[0242] In some embodiments, the gene editing method involves contacting cells with an engineered type V endonuclease or B-GEn peptide system to form an RNP complex comprising an engineered type V or B-GEn peptide and guide RNA. Illustrative details regarding the preparation, composition, and delivery of the RNP complex are described in section 6.7. In some embodiments, the RNP complex can also be delivered into cells using delivery methods that increase the cell's plasma membrane porosity (e.g., via electroporation or nuclear transfection).
[0243] In some embodiments, the gene editing method includes contacting cells with an engineered type V endonuclease or B-GEn polypeptide system in the form of an LNP, such as an LNP containing an RNP (e.g., an LNP as described in the preceding paragraph) or an LNP containing a nucleic acid encoding an engineered type V endonuclease or B-GEn polypeptide and a guide RNA, or a nucleic acid encoding a guide RNA. LNPs can be prepared using any method known in the art, such as that described in section 6.12.
[0244] In some embodiments, gene editing methods include contacting cells with an engineered type V endonuclease or B-GEn polypeptide system, in the form of one or more viruses, such as one or more AAVs, whose genome contains one or more nucleotide sequences encoding an engineered type V endonuclease or B-GEn polypeptide and a nucleic acid encoding a guide RNA. The use of viruses to introduce transgenes, such as nucleic acids encoding engineered type V endonucleases or B-GEn polypeptides and guide RNA coding sequences, is known in the art and described in section 6.9.3. Other methods and delivery vectors for introducing nucleic acids encoding engineered type V endonucleases, B-GEn polypeptides, or guide RNA are described in sections 6.9 and 6.12.
[0245] 6.14. Pharmaceutical Compositions This document also discloses pharmaceutical formulations and pharmaceutical products comprising B-GEn protein, gRNA, nucleic acid or multiple nucleic acids, systems, particles or multiple particles, and pharmaceutically acceptable excipients.
[0246] Suitable excipients include, but are not limited to, salts, diluents (e.g., Tris-HCl, acetates, phosphates), preservatives (e.g., thimerosal, benzyl alcohol, parabens), binders, fillers, solubilizers, disintegrants, adsorbents, solvents, pH adjusters, antioxidants, anti-infectives, suspending agents, wetting agents, viscosity modifiers, tension agents, stabilizers, and other components and combinations thereof. Suitable pharmaceutically acceptable excipients may be selected from materials that are generally considered safe (GRAS) and can be administered to individuals without causing undesirable biological side effects or undesirable interactions. Suitable excipients and their formulations are described in Remington's Pharmaceutical Sciences, 16th ed. 1980, Mack Publishing Co. Furthermore, such compositions may be complexed with polyethylene glycol (PEG), metal ions, or incorporated into polymeric compounds such as polyacetic acid, polyglycolic acid, hydrogels, etc., or incorporated into liposomes, microemulsions, micelles, monolayer or multilayer vesicles, erythrocyte ghosts, or spherocytes. Suitable dosage forms for application (e.g., parenteral administration) include solutions, suspensions, and emulsions.
[0247] The components of the pharmaceutical formulation can be dissolved or suspended in a suitable solvent, such as water, Ringer's solution, phosphate-buffered saline (PBS), or isotonic sodium chloride. The formulation may also be a sterile solution, suspension, or emulsion in a non-toxic, parenteral diluent or solvent such as 1,3-butanediol.
[0248] In some cases, the formulation may include one or more tonic agents to adjust the isotonic range of the formulation. Suitable tonic agents are well known in the art and include glycerol, mannitol, sorbitol, sodium chloride, and other electrolytes. In some cases, the formulation may be buffered with an effective amount of buffer solution necessary to maintain a pH suitable for parenteral administration. Suitable buffer solutions are well known to those skilled in the art, and some examples of useful buffer solutions are acetate, borate, carbonate, citrate, and phosphate buffers.
[0249] In some embodiments, the formulation may be distributed or packaged in liquid form, or alternatively as a solid, for example, obtained by lyophilizing a suitable liquid formulation, which can be reconstituted with a suitable carrier or diluent prior to administration. In some embodiments, the formulation may contain a pharmaceutically effective amount of guide RNA and type II Cas protein sufficient to edit genes in cells. The pharmaceutical composition may be formulated for medical and / or veterinary use.
[0250] In some embodiments, the B-GEn endonuclease complex can be introduced into host cells (such as iPSCs) to generate genetically modified cells that can be reintroduced into an individual. The iPSC-derived cells described herein can be provided in a pharmaceutical composition containing the cells and a pharmaceutically acceptable carrier. The pharmaceutically acceptable carrier may be a cell culture medium optionally free of any animal-derived components. For storage and transport, the cells can be cryopreserved at <-70°C (e.g., on dry ice or in liquid nitrogen). Prior to use, the cells can be thawed and diluted in a sterile cell culture medium supporting the desired cell type.
[0251] The cells can be administered to a patient systemically (e.g., by intravenous injection or infusion) or locally (e.g., by direct injection into local tissues, such as the heart, brain, and sites of damaged tissue). Various methods for administering cells to a patient's tissues or organs are known in the art, including but not limited to, intracoronary administration, intramyocardial administration, endocardial administration, or intracranial administration.
[0252] A therapeutically effective number of iPSC-derived cells is administered to the patient. As used herein, the term "therapeutically effective" refers to the number of cells or the amount of pharmaceutical composition sufficient to treat, prevent, and / or delay the onset or progression of symptoms of a disease, condition, and / or symptom when administered to a human subject who has or is susceptible to the disease, condition, and / or symptom. Those skilled in the art will understand that a therapeutically effective amount is typically administered via a dosing regimen comprising at least one unit dose. In some embodiments, at least 10 iPSC-derived cells are administered to the subject at one or more sites at a time. 3 (e.g., at least 10) 4 At least 10 5 At least 10 6 At least 10 7 At least 10 8 At least 10 9 At least 10 10 At least 10 11 One or at least 10 12 (Number) cells. In some implementations, 10 3 -10 18 (e.g., 10) 3 -10 4 10 3 -10 5 10 3 -10 6 10 3 -10 7 10 3 -10 8 10 3 -10 9 10 3-10 10 10 3 -10 11 10 3 -10 12 10 6 -10 7 10 6 -10 8 10 6 -10 9 10 6 -10 10 10 6 -10 11 10 6 -10 12 10 9- 10 10 10 9 -10 11 10 9 -10 12 Cells are administered to a subject at one or more sites at a time. In some embodiments, more than 10 cells are administered to a subject at one or more sites at a time. 12 (e.g., more than 10) 12 More than 10 13 More than 10 14 More than 10 15 More than 10 16 More than 10 17 More than 10 18 (One or more) cells.
[0253] 7. While various specific embodiments have been shown and described in the numbered implementation schemes, it should be understood that various changes may be made without departing from the spirit and scope of this disclosure. This disclosure is illustrated by the numbered implementation schemes listed below. Unless otherwise stated, any concepts, aspects, and / or features of the embodiments described in the above detailed description, with necessary modifications, apply to any embodiment of the following numbered implementation schemes.
[0254] 1. A polypeptide comprising an amino acid other than aspartic acid at position D504 of SEQ ID NO: 1 (B-GEn.1) or position D501 of SEQ ID NO: 2 (B-GEn.1.2) or position D501 of SEQ ID NO: 3 (B-GEn.2).
[0255] 2. The polypeptide described in Implementation Scheme 1 is an engineered type V endonuclease polypeptide.
[0256] 3. The polypeptide described in Implementation Scheme 2 is an engineered Bacillales type V endonuclease polypeptide.
[0257] 4. The polypeptide according to any one of embodiments 1 to 3, wherein it is a B-GEn polypeptide.
[0258] 5. The polypeptide according to any one of embodiments 1 to 4, wherein it contains arginine at position D504 corresponding to SEQ ID NO: 1 (B-GEn.1) or D501 corresponding to SEQ ID NO: 2 (B-GEn.1.2) or SEQ ID NO: 3 (B-GEn.2).
[0259] 6. The polypeptide according to any one of embodiments 1 to 5, comprising a target interaction sequence motif of any one of SEQ ID NO: 201, SEQ ID NO: 202, SEQ ID NO: 203 and SEQ ID NO: 204.
[0260] 7. The polypeptide according to any one of embodiments 1 to 6, comprising a RuvC I domain, said RuvC I domain comprising an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% sequence identity with the RuvC I domain of SEQ ID NO: 8 or SEQ ID NO: 11.
[0261] 8. The polypeptide of embodiment 7, wherein the RuvC I domain comprises an amino acid sequence having at least 40% sequence identity with the RuvC I domain of SEQ ID NO: 8 or SEQ ID NO: 11.
[0262] 9. The polypeptide of embodiment 7, wherein the RuvC I domain comprises an amino acid sequence having at least 70% sequence identity with the RuvC I domain of SEQ ID NO: 8 or SEQ ID NO: 11.
[0263] 10. The polypeptide of embodiment 7, wherein the RuvC I domain comprises an amino acid sequence having at least 80% sequence identity with the RuvC I domain of SEQ ID NO: 8 or SEQ ID NO: 11.
[0264] 11. The polypeptide of embodiment 7, wherein the RuvC I domain comprises an amino acid sequence having at least 90% sequence identity with the RuvC I domain of SEQ ID NO: 8 or SEQ ID NO: 11.
[0265] 12. The polypeptide according to any one of embodiments 1 to 11, comprising a RuvC II domain, said RuvC II domain comprising an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% sequence identity with the RuvC I domain of SEQ ID NO: 9 or SEQ ID NO: 12.
[0266] 13. The polypeptide of embodiment 12, wherein the RuvC II domain comprises an amino acid sequence having at least 60% sequence identity with the RuvC II domain of SEQ ID NO: 9 or SEQ ID NO: 12.
[0267] 14. The polypeptide of embodiment 12, wherein the RuvC II domain comprises an amino acid sequence having at least 70% sequence identity with the RuvC II domain of SEQ ID NO: 9 or SEQ ID NO: 12.
[0268] 15. The polypeptide of embodiment 12, wherein the RuvC II domain comprises an amino acid sequence having at least 80% sequence identity with the RuvC II domain of SEQ ID NO: 9 or SEQ ID NO: 12.
[0269] 16. The polypeptide of embodiment 12, wherein the RuvC II domain comprises an amino acid sequence having at least 90% sequence identity with the RuvC II domain of SEQ ID NO: 9 or SEQ ID NO: 12.
[0270] 17. The polypeptide according to any one of embodiments 1 to 16, comprising a RuvC III domain, said RuvC III domain comprising an amino acid sequence having at least 80%, at least 85%, or at least 90% sequence identity with the RuvC III domain of SEQ ID NO: 10 or SEQ ID NO: 13.
[0271] 18. The polypeptide of embodiment 17, wherein the RuvC III domain comprises an amino acid sequence having at least 80% sequence identity with the RuvC III domain of SEQ ID NO: 10 or SEQ ID NO: 13.
[0272] 19. The polypeptide of embodiment 17, wherein the RuvC III domain comprises an amino acid sequence having at least 85% sequence identity with the RuvC III domain of SEQ ID NO: 10 or SEQ ID NO: 13.
[0273] 20. The polypeptide of embodiment 17, wherein the RuvC III domain comprises an amino acid sequence having at least 90% sequence identity with the RuvC III domain of SEQ ID NO: 10 or SEQ ID NO: 13.
[0274] 21. The polypeptide according to any one of embodiments 1 to 20, comprising the amino acid sequence of SEQ ID NO: 8.
[0275] 22. The polypeptide according to any one of embodiments 1 to 21, comprising the amino acid sequence of SEQ ID NO: 9.
[0276] 23. The polypeptide according to any one of embodiments 1 to 22, comprising the amino acid sequence of SEQ ID NO: 10.
[0277] 24. The polypeptide according to any one of embodiments 1 to 20, comprising the amino acid sequence of SEQ ID NO: 11.
[0278] 25. The polypeptide according to any one of embodiments 1 to 20 and 24, comprising the amino acid sequence of SEQ ID NO: 12.
[0279] 26. The polypeptide according to any one of embodiments 1 to 20 and 24 to 25, comprising the amino acid sequence of SEQ ID NO: 13.
[0280] 27. A polypeptide, optionally a polypeptide according to any one of embodiments 1 to 26, comprising the following amino acid sequence: (a) having a substitution at position D501 corresponding to SEQ ID NO: 3, wherein the substitution: (i) is a substitution D501R compared to the amino acid sequence of SEQ ID NO: 3; or (ii) increases gene editing efficiency activity compared to the corresponding amino acid sequence without a substitution at position D501; and (b) having: (i) at least 80%, at least 85%, at least 90%, or at least 95% sequence identity with SEQ ID NO: 1 (B-GEn.1), SEQ ID NO: 2 (B-GEn.1.2), or SEQ ID NO: 3 (B-GEn.2); (ii) (A) a target interaction sequence motif of any one of SEQ ID NO: 201, SEQ ID NO: 202, SEQ ID NO: 203, and SEQ ID NO: 204; (B) a RuvC I domain comprising a substitution D501R corresponding to SEQ ID NO: 8 or SEQ ID NO: 204. (A) The RuvC I domain of ID NO: 11 has an amino acid sequence that is at least 40%, at least 45%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% sequence identity with the RuvC II domain of SEQ ID NO: 9 or SEQ ID NO: 12; (B) The RuvC III domain contains an amino acid sequence that is at least 80%, at least 85%, or at least 90% sequence identity with the RuvC III domain of SEQ ID NO: 10 or SEQ ID NO: 13; or (C) Any combination of two, three, or all four of (A), (B), (C), and (D).
[0281] (iii) Up to 25 amino acid insertions, substitutions, and / or deletions compared to the amino acid sequences of SEQ ID NO: 1 (B-GEn.1), SEQ ID NO: 2 (B-GEn.1.2), or SEQ ID NO: 3 (B-GEn.2); (iv) (b)(i) and (b)(ii); (v) (b)(i) and (b)(iii); (vi) (b)(ii) and (b)(iii); or (vii) (b)(i), (b)(ii), and (b)(iii).
[0282] 28. The polypeptide according to any one of embodiments 1 to 27, wherein the gene editing efficiency is at least 50% higher than that of the corresponding polypeptide lacking an amino acid substitution at position D504 of SEQ ID NO: 1 (B-GEn.1) or D501 of SEQ ID NO: 2 (B-GEn.1.2) or SEQ ID NO: 3 (B-GEn.2).
[0283] 29. The polypeptide of any one of embodiments 1 to 27, wherein the gene editing efficiency is at least 70% higher than that of the corresponding polypeptide lacking an amino acid substitution at position D504 of SEQ ID NO: 1 (B-GEn.1) or D501 of SEQ ID NO: 2 (B-GEn.1.2) or SEQ ID NO: 3 (B-GEn.2).
[0284] 30. The polypeptide according to any one of embodiments 1 to 27, wherein the gene editing efficiency is at least 90% higher than that of the corresponding polypeptide lacking an amino acid substitution at position D504 of SEQ ID NO: 1 (B-GEn.1) or position D501 of SEQ ID NO: 2 (B-GEn.1.2) or position D501 of SEQ ID NO: 3 (B-GEn.2).
[0285] 31. The polypeptide according to any one of embodiments 28 to 30, wherein gene editing efficiency is assessed by an intracellular gene editing assay, optionally wherein the gene editing assay is as described in Example 1.
[0286] 32. The polypeptide of any one of embodiments 27 to 31, wherein the sequence identity is relative to SEQ ID NO: 1, and the amino acid insertions, substitutions, and / or deletions are associated with SEQ ID NO: 1.
[0287] 33. The polypeptide of embodiment 32, wherein the amino acid sequence has at least 98% sequence identity with the amino acid sequence of SEQ ID NO: 1.
[0288] 34. The polypeptide of embodiment 32, wherein the amino acid sequence has at least 99% sequence identity with the amino acid sequence of SEQ ID NO: 1.
[0289] 35. The polypeptide of embodiment 32, wherein the amino acid sequence has at least 99.5% sequence identity with the amino acid sequence of SEQ ID NO: 1.
[0290] 36. The polypeptide of any one of embodiments 27 to 31, wherein the sequence identity is relative to SEQ ID NO: 2, and the amino acid insertions, substitutions, and / or deletions are associated with SEQ ID NO: 2.
[0291] 37. The polypeptide of embodiment 36, wherein the amino acid sequence has at least 98% sequence identity with the amino acid sequence of SEQ ID NO: 2.
[0292] 38. The polypeptide of embodiment 36, wherein the amino acid sequence has at least 99% sequence identity with the amino acid sequence of SEQ ID NO: 2.
[0293] 39. The polypeptide of embodiment 36, wherein the amino acid sequence has at least 99.5% sequence identity with the amino acid sequence of SEQ ID NO: 2.
[0294] 40. The polypeptide of any one of embodiments 27 to 31, wherein the sequence identity is relative to SEQ ID NO: 3, and the amino acid insertions, substitutions, and / or deletions are associated with SEQ ID NO: 3.
[0295] 41. The polypeptide of embodiment 40, wherein the amino acid sequence has at least 98% sequence identity with the amino acid sequence of SEQ ID NO: 3.
[0296] 42. The polypeptide of embodiment 40, wherein the amino acid sequence has at least 99% sequence identity with the amino acid sequence of SEQ ID NO: 3.
[0297] 43. The polypeptide of embodiment 40, wherein the amino acid sequence has at least 99.5% sequence identity with the amino acid sequence of SEQ ID NO: 3.
[0298] 44. A polypeptide comprising the amino acid sequence of SEQ ID NO: 4.
[0299] 45. A polypeptide comprising the amino acid sequence of SEQ ID NO: 5.
[0300] 46. A polypeptide comprising the amino acid sequence of SEQ ID NO: 6.
[0301] 47. The polypeptide according to any one of embodiments 1 to 46, further comprising at least one nuclear localization signal (“NLS”).
[0302] 48. The polypeptide of embodiment 47, comprising at least one NLS located at the C-terminus of an amino acid sequence, optionally wherein: (a) the polypeptide lacks any NLS at the N-terminus of the amino acid sequence; or (b) the polypeptide comprises at least one NLS located at the N-terminus of an amino acid sequence.
[0303] 49. The polypeptide of embodiment 47 or embodiment 48, wherein any NLS is missing at the N-terminus of the amino acid sequence.
[0304] 50. The polypeptide according to any one of embodiments 47 to 49, comprising the amino acid sequence and a linker sequence between each NLS.
[0305] 51. The polypeptide according to any one of embodiments 47 to 50, wherein each NLS comprises an amino acid sequence independently selected from the NLS sequences listed in Table 2.
[0306] 52. The polypeptide according to any one of embodiments 47 to 51, wherein each adapter sequence is independently selected from the adapter sequences listed in Table 3.
[0307] 53. A nucleic acid comprising a nucleotide sequence encoding a polypeptide according to any one of embodiments 1 to 52.
[0308] 54. The nucleic acid of embodiment 53, wherein the nucleotide sequence encoding the polypeptide of any one of embodiments 1 to 52 is operatively linked to a promoter.
[0309] 55. The nucleic acid of embodiment 53 or embodiment 54, wherein the nucleic acid further encodes a guide RNA.
[0310] 56. The nucleic acid described in any one of implementation schemes 53 to 55, wherein it is in the form of a vector.
[0311] 57. The nucleic acid described in Implementation Scheme 56, wherein the vector is an expression vector.
[0312] 58. The nucleic acid described in Implementation Scheme 57, wherein the expression vector is a production vector.
[0313] 59. The nucleic acid of embodiment 57, wherein the expression vector is a delivery vector.
[0314] 60. The nucleic acid according to any one of embodiments 56 to 59, wherein the vector is an RNA vector.
[0315] 61. The nucleic acid according to any one of embodiments 56 to 59, wherein the vector is a DNA vector.
[0316] 62. The nucleic acid described in Implementation Scheme 61, wherein the DNA vector is a plasmid.
[0317] 63. A cell comprising the nucleic acid described in any one of embodiments 53 to 62.
[0318] 64. A cell engineered to express a nucleotide sequence encoding a polypeptide according to any one of embodiments 1 to 52.
[0319] 65. The cells described in Implementation Scheme 63 or Implementation Scheme 64 are eukaryotic cells.
[0320] 66. The cell described in Implementation Scheme 65 is an insect cell.
[0321] 67. The cell described in Implementation Scheme 65 is a plant cell.
[0322] 68. The cell described in Implementation Scheme 65 is a mammalian cell.
[0323] 69. The cell described in Implementation Scheme 68 is a human cell.
[0324] 70. A method for producing a polypeptide according to any one of embodiments 1 to 52, comprising culturing cells according to any one of embodiments 63 to 69 under conditions for producing said polypeptide.
[0325] 71. The method of embodiment 70, further comprising isolating and / or purifying the polypeptide.
[0326] 72. A composition comprising: (a) a polypeptide according to any one of embodiments 1 to 52; and (b) a guide RNA.
[0327] 73. The composition of embodiment 72, wherein the composition is a ribonucleoprotein complex.
[0328] 74. The composition of embodiment 72 or embodiment 73, wherein the molar ratio of peptide to guide RNA is in the range of 1:1 to 1:4.
[0329] 75. The composition of embodiment 72 or embodiment 73, wherein the molar ratio of peptide to guide RNA is in the range of 1:1 to 1:3.
[0330] 76. The composition of embodiment 72 or embodiment 73, wherein the molar ratio of peptide to guide RNA is in the range of 1:1.5 to 1:2.5.
[0331] 77. The composition of embodiment 72 or embodiment 73, wherein the molar ratio of polypeptide to guide RNA is 1:2.
[0332] 78. A method for editing the genome of a cell, comprising introducing into the cell: (a) a polypeptide according to any one of embodiments 1 to 52; and (b) a guide RNA.
[0333] 79. A method for editing the genome of a cell, comprising introducing one or more nucleic acids into the cell, said nucleic acids encoding: (a) a polypeptide as described in any one of embodiments 1 to 52; and (b) a guide RNA, optionally, wherein at least one of said one or more nucleic acids is a nucleic acid according to any one of embodiments 53 to 62.
[0334] 80. A method for editing the genome of a cell, comprising introducing into the cell: (a) one or more nucleic acids encoding a polypeptide as described in any one of embodiments 1 to 52; and (b) a guide RNA.
[0335] 81. The method of embodiment 80, comprising contacting the cell with lipid nanoparticles comprising the one or more nucleic acids and the guide RNA.
[0336] 82. A method for editing the genome of a cell, comprising introducing one or more nucleic acids into the cell, said nucleic acids encoding: (a) a polypeptide of any one of claims 1 to 52; and (b) a guide RNA, optionally, wherein at least one of said one or more nucleic acids is a nucleic acid according to any one of embodiments 53 to 62.
[0337] 83. The method of embodiment 82, comprising contacting the cell with one or more recombinant AAV particles containing one or more of the nucleic acids.
[0338] 84. The method of embodiment 82, comprising contacting cells with one or more lipid nanoparticles comprising one or more nucleic acids.
[0339] 85. A method for editing the genome of a cell, comprising introducing the composition of any one of embodiments 72 to 77 into the cell.
[0340] 86. The method of any one of embodiments 78 to 80, wherein the cell is a mammalian cell.
[0341] 87. The method of embodiment 86, wherein the mammalian cell is a human cell.
[0342] 88. The method of embodiment 86 or embodiment 87, wherein the mammalian cell is an immune cell, optionally selected from T cells, T cells expressing chimeric antigen receptor (CAR) or recombinant TCR, regulatory T cells, myeloid cells, dendritic cells and immunosuppressive macrophages.
[0343] 89. The method according to any one of embodiments 78 to 87, wherein the cell is a hematopoietic stem cell, erythroid progenitor cell, lymphoid progenitor cell, peripheral blood monocyte, T lymphocyte, B lymphocyte, macrophage, monocyte, neutrophil, eosinophil, dendritic cell, or a cell reprogrammed therefrom.
[0344] 90. The method according to any one of embodiments 78 to 89, wherein the cell is a stem cell or a cell differentiated therefrom.
[0345] 91. The method of embodiment 90, wherein the stem cell is a pluripotent stem cell (PSC) or a cell differentiated from it.
[0346] 92. The method of embodiment 90 or embodiment 91, wherein the cells differentiated therefrom are human immune cells, optionally selected from T cells, T cells expressing chimeric antigen receptors (CARs) or recombinant TCRs, regulatory T cells, myeloid cells, dendritic cells, and immunosuppressive macrophages.
[0347] 93. The method of embodiment 90 or embodiment 91, wherein the cells differentiated therefrom are cells of the human nervous system, optionally selected from dopaminergic neurons, microglia, oligodendrocytes, astrocytes, cortical neurons, spinal cord or oculomotor neurons, enteric neurons, basal plate-derived cells, Schwann cells, and trigeminal nerve or sensory neurons.
[0348] 94. The method of embodiment 90 or embodiment 91, wherein the cells differentiated therefrom are cells of the human cardiovascular system, optionally selected from cardiomyocytes, endothelial cells and lymph node cells.
[0349] 95. The method of embodiment 90 or embodiment 91, wherein the cells differentiated therefrom are cells of the human metabolic system, optionally selected from hepatocytes, bile duct cells and pancreatic β cells.
[0350] 96. The method of embodiment 90 or embodiment 91, wherein the cells differentiated therefrom are cells of the human eye system, optionally selected from retinal pigment epithelial cells, cone cells, rod cells, bipolar cells or ganglion cells. 97. The method of any one of embodiments 78 to 80, wherein the cells are plant cells.
[0351] 98. A cell comprising: (a) a composition according to any one of embodiments 72 to 77; or (b) a nucleic acid according to any one of embodiments 53 to 62.
[0352] 99. The cell described in Implementation Scheme 98 is a mammalian cell.
[0353] 100. The cell of embodiment 99, wherein the mammalian cell is a human cell.
[0354] 101. The cell described in embodiment 99 or embodiment 100, wherein the mammalian cell is an immune cell, optionally selected from T cells, T cells expressing chimeric antigen receptor (CAR) or recombinant TCR, regulatory T cells, myeloid cells, dendritic cells and immunosuppressive macrophages.
[0355] 102. The cell according to any one of embodiments 98 to 100 is a hematopoietic stem cell, erythroid progenitor cell, lymphoid progenitor cell, peripheral blood monocyte, T lymphocyte, B lymphocyte, macrophage, monocyte, neutrophil, eosinophil, dendritic cell, or a cell reprogrammed therefrom.
[0356] 103. The cell described in any one of embodiments 98 to 102 is a stem cell or a cell differentiated from it.
[0357] 104. The cell described in embodiment 103, wherein the stem cell is a pluripotent stem cell (PSC) or a cell differentiated from it.
[0358] 105. The cells described in embodiment 103 or embodiment 104, wherein the cells differentiated from them are human immune cells, optionally selected from T cells, T cells expressing chimeric antigen receptors (CARs) or recombinant TCRs, regulatory T cells, myeloid cells, dendritic cells, and immunosuppressive macrophages.
[0359] 106. The cell described in embodiment 103 or embodiment 104, wherein the cell differentiated from it is a cell of the human nervous system, optionally selected from dopaminergic neurons, microglia, oligodendrocytes, astrocytes, cortical neurons, spinal cord or oculomotor neurons, enteric neurons, basal plate-derived cells, Schwann cells and trigeminal nerve or sensory neurons.
[0360] 107. The cells described in embodiment 103 or embodiment 104, wherein the cells differentiated from them are cells of the human cardiovascular system, optionally selected from cardiomyocytes, endothelial cells and lymph node cells.
[0361] 108. The cell described in embodiment 103 or embodiment 104, wherein the cell differentiated from it is a cell of the human metabolic system, optionally selected from hepatocytes, bile duct cells and pancreatic β cells.
[0362] 109. The cell described in embodiment 103 or embodiment 104, wherein the cell differentiated therefrom is a cell of the human eye system, optionally selected from retinal pigment epithelial cells, cone cells, rod cells, bipolar cells or ganglion cells.
[0363] 110. The cell described in Implementation Scheme 98 is a plant cell.
[0364] 8. Examples 8.1 Materials and Methods 8.1.1 Structural Modeling of B-GEn.2 The structure of B-GEn.2 has not been characterized to date. The structure of B-GEn.2 was predicted using AlphaFold2 software and compared with a protein database with known crystal structures. Based on the hit results of 14 lead crystal structures, a set of candidate structures for B-GEn.2 was generated. This data shows that the predicted structure of B-GEn.2 correlates very well with several crystal structures derived from Bacilli.
[0365] Based on the alignment, the OBD II and RuvC I domains of B-GEn.2 were used for NCBIBLASTp searches within the Bacillus order. Approximately 20 Cas nucleases from relevant source organisms, sharing approximately 40-70% sequence identity, were selected for multiple sequence alignment using the MUSCLE alignment algorithm. The phylogenetic relationships of these different nucleases are depicted in the tree diagram in Figure 2, while the alignment itself is shown in Figure 4.
[0366] 8.1.2 Design of B-GEn.2 Variants with Amino Acid Substitutions Two parallel approaches were used to identify amino acid substitutions that could lead to enhanced insertion / deletion activity of B-GEn.2. In the first approach, DNAStar Lasergene 17 software was used to identify amino acid residues in the AacC2c1 crystal structure within a 3 angstroms range of the target DNA substrate, guide spacer RNA, or guide tracr RNA. The same action was performed on the predicted structure of B-GEn.2, paying particular attention to those cases where the selected amino acid residues differed from the corresponding residues in AacC2c1.
[0367] The second approach involves using Swiss-PdbViewer, also known as Deepview (https: / / spdbv.unil.ch / ), to identify amino acids within the crystal structure of AacC2c1 (RCSB entry 5U31) that interact with the target DNA substrate, guide spacer RNA, or guide tracer RNA via predicted hydrogen bonds. Then, by comparing the sequence of AacC2c1 with the sequence of B-GEn.2, instances were searched where the predicted protein-nucleic acid contact residues in AacC2c1 differed in B-GEn.2. To narrow the focus from dozens of potential candidate mutants to a more manageable number for final protein production and purification, further evaluation focused on Arg or Lys residues: these residues form tight contacts with the target or non-target DNA strands in the AacC2c1 crystal but have no corresponding Arg or Lys residues in the predicted B-GEn.2 structure.
[0368] 8.1.3 Expression and Purification of B-GEn.2 Variants Plasmids were constructed to express selected point mutants of B.Gen.2. All constructs were expressed in BL21(DE3) cells by first chemically transforming the encoding plasmids into these *E. coli* cells and then plated on antibiotic-selective LB plates. Colonies were scraped and transferred to 250 mL of MagicMedia (Thermo Scientific). Cultures were grown at 37°C for 4 h, then switched to 16°C for 40 h for protein expression. The cell pellet was then harvested by rotating at 5000 × g for 15 min at 4°C and frozen for subsequent processing. To thaw, the pellet was resuspended in lysis buffer (500 mM NaCl, 50 mM Tris, pH 8.0, 5% glycerol, 5 mM EDTA, 0.5 mM TCEP), sonicated, and the resulting lysate was clarified by centrifugation at 50,000 × g for 30 min at 4°C. The clarified lysate was loaded onto a heparin agarose 6FF column in loading buffer (100 mM NaCl, 50 mM Tris, pH 8.0, 5% glycerol, 5 mM EDTA, 0.5 mM TCEP), followed by a washing step with the same buffer. Protein was eluted using the same buffer at 500 mM, 700, and 1000 mM NaCl. The protein in the major fraction was concentrated to approximately 20 mg / mL. Protein purity was estimated by band density determination on a Coomassie-stained SDS-PAGE gel, and aliquots of the purified protein were frozen at -80°C for later use.
[0369] 8.1.4 RNP Generation A purified and concentrated variant B-GEn.2 nuclease is combined with a guide RNA containing the following sequence targeting an intragenetic site of the B2M gene to form an RNP, where the spacer is the underlined portion: *mU*mA*GCUAUAGGCUAAUAAGAUAGUUGUGUCAAGUGCUUCGGAGACCUAACACGUCUCCAGUCACAACGGCUAAAAAUAGCCAGCAC AGUGUAGUACAAGAGAUAGA*mA*mA*mG (SEQ ID NO:178) RNPs were assembled by mixing B2M-targeting guide RNA with nucleases in a buffer containing 225 mM NaCl at a 2:1 molar ratio and incubating at room temperature for approximately 30 minutes. The recombination efficiency of the RNPs was determined by cation exchange UPLC on an Agilent BioSCX (NP1.7, SS) column using a linear elution gradient of NaCl. Buffer A consisted of 10 mM sodium phosphate, 100 mM NaCl, and pH 6.5, while buffer B consisted of 10 mM sodium phosphate, 1 M NaCl, and pH 6.5. The recombination efficiency percentage was obtained by dividing the peak area of the RNPs by the peak area of the total added nucleases (an equivalent sample without guide RNA) and multiplying by 100. Values ranged from 70% to 90% recombination, depending on the nuclease variant tested. The ability of the RNPs to cleave DNA in vitro was determined using an in vitro plasmid cutting assay. Incremental amounts of RNP were mixed with a fixed amount of plasmid containing the B2M target site in CutSmart buffer (New England Biolabs), the plasmid having been linearized using Xho-I enzyme. After incubation at 37°C for 30 min, the reaction was quenched at 50°C for 10 min in the presence of proteinase K. Samples were applied to D5000 tapes and analyzed on an Agilent 4200 TapeStation instrument. The percentage of lysis for each reaction was calculated by dividing the peak area of the product by the peak area of the total substrate added (equivalent to the RNP), and multiplying by 100. EC50 of each variant protein (or Cpf1 control) series was determined using Prism software (version 9.3.1) in GraphPad. 50 The value ranges from 0.07 to 0.21 nM.
[0370] 8.1.5 Intracellular Gene Editing iPSCs were cultured in substrate-coated T75 flasks with Essential 8 (E8) growth medium and maintained at 37°C and 5% CO2 between passages. Passages were performed at 75%–80% confluence. On the day of nuclear transfection, iPSCs were isolated from the flasks using Accutase™ Cell Isolation Solution (Stem Cell Technologies), incubated at 37°C for 10 min, and then quenched with an equal volume of E8 medium. The cell pellet was harvested by centrifugation at 115 x g for 3 min and then resuspended in Lonza P3 primary cell nuclear transfection buffer. Ribonuclear proteins (RNPs) were assembled with the sgRNA of each nuclease construct using a 1:2 protein:sgRNA (IDT) ratio. The complexed RNPs were then nuclear transfected into the P3-resuspended iPSCs using LONZA 4D nuclear transfection. The nuclear-transfected cells were then seeded at 150,000 cells / well in substrate-coated 24-well Falcon flat-bottomed plates (Corning) and grown in E8 growth medium (with the rock inhibitor Y-27632 (Tocris)) for 72 to 96 hours. After harvest, some cells were stained with an anti-B2M APC conjugated antibody from BioLegend (for B2M-targeting experiments only) for flow cytometry. The remaining cells were resuspended in 30 μL of lysis buffer from BioRad's Singleshot Cell Lysis Kit, and crude gDNA extraction was performed by incubation at room temperature for 10 min, followed by incubation at 37°C for 5 min, and then proteinase K inactivation at 75°C for 5 min. Amplicon sequencing was performed directly with 1 μL of the crude gDNA extract per 25 μL PCR reaction, followed by end preparation and indexing and sequencing on an Illumina MiSeq sequencer.
[0371] 8.2 Example 1: Design and Expression of Variant B-GEn.2 with Amino Acid Substitution The gene-editing activity of the type V endonuclease B-GEn.2 (SEQ ID NO: 3) from the genus *Brevibacillus* was previously evaluated in several proprietary iPSC lines. With the goal of improving insertion / deletion formation, a plausible structure for wild-type B-GEn.2 was identified on a computer, as described in Section 8.1.1. Variant B-GEn.2 sequences with a single amino acid point mutation were designed and generated, as described in Section 8.1.2. The selected B-GEn.2 variant was expressed in BL21(DE3) cells and purified as described in Section 8.1.3.
[0372] The lead rational structure of B-GEn.2 was identified, and it showed good agreement with the previously disclosed crystal structures of Cas nucleases from Alicyclobacillus acidoterrestris AacC2c1 (Figs. 3A-3C) and Cas nucleases from Geobacillus thermoleovorans BthC2C1 (not shown).
[0373] Next, the amino acid sequences of AacC2c1 and B-GEn.2 were compared to determine key amino acid differences. Sequence alignment between the two nucleases revealed approximately 37% amino acid sequence identity. Using the differences in the amino acid sequences of AacC2c1 and B-GEn.2, Arg or Lys residues in the AacC2c1 crystal structure that form close contacts with both target and non-target DNA, and their corresponding non-Arg and non-Lys residues in the predicted B-GEn.2 structure, were identified as mutant targets. As described in Section 8.1.2, an initial list of 19 mutant targets was identified using DNA Star and DeepView, from which 8 mutant targets were selected for further evaluation. These mutant targets were generated and purified as described in Section 8.1.3.
[0374] The expression profiles of all eight B-GEn.2 variants were comparable to those of wild-type B-GEn.2 (Table 7). A representative Coomassie-stained SDS-PAGE image of a single-step purification based on heparin sulfate for one of the variants (B-GEn.2D501R) is shown in Figure 5. 8.3 Example 2: Characterization of RNPs with B-GEn.2 variants. Purified B-GEn.2 variants formed RNPs with guide RNA targeting intragenic sites of the B2M gene, as described in Section 8.1.4. The efficiency of RNP formation and the in vitro cleavage efficiency of the target B2M site were evaluated.
[0375] The results showed that the B-GEn.2 variant could effectively complex with the target RNA to form RNPs, with the complexation efficiency of the B-GEn.2 variant ranging from 71% to 92%.
[0376] In vitro activity assays showed no detectable significant differences between the B-GEn.2 variant and WT B-GEn.2. (e.g., by ECMO...) 50 The estimated target plasmid cleavage efficiency of the B-GEn.2 variant was found to be in the range of 0.07–0.17 nM. Figure 6 shows the cleavage efficiency of the linearized plasmid relative to WT B-GEn.2 and Cpf1 for different nucleases: the ratio of linearized plasmids.
[0377] 8.4 Example 3: Efficient intracellular gene editing of B2M loci in iPSCs with B-GEn.2 D501R Intracellular gene editing of B-GEn.2 D501R was evaluated by examining the generation of insertions and deletions at B2M target loci in iPSCs, as described in Section 8.1.4, and compared with gene editing achieved by WT B-GEn.2 and AsCpf1Ultra (IDT).
[0378] B-GEn.2D501R achieved an insertion / deletion formation efficiency greater than 80%, which is higher than that achieved by Cpf1 (Figure 7). This insertion / deletion formation efficiency achieved by B-GEn.2D501R is 2–3 times higher than that achieved by WT B-Gen.2 (Figure 7). This result indicates that replacing Asp with Arg at D501, as shown in the AacC2c1 crystal structure, forms hydrogen bonds with the +1 phosphate backbone group of the target DNA strand, enhancing the ability of the B-GEn.2 variant to cleave its target substrate.
[0379] 9. Sequence Listing Exemplary sequences of this disclosure are provided in Table 8 below (where “SEQ” refers to SEQ ID NO). 10. All publications, patents, patent applications and other documents cited in this application are incorporated herein by reference in their entirety for all purposes, and their scope is the same as that of each individual publication, patent, patent application or other document individually indicated as incorporated herein by reference for all purposes. In the event of any inconsistency between the teachings of one or more references incorporated herein and the contents of this disclosure, the teachings of this application shall prevail.
Claims
1. An engineered type V endonuclease containing an amino acid other than aspartic acid at position D504 of SEQ ID NO: 1 (B-GEn.1) or D501 of SEQ ID NO: 2 (B-GEn.1.2) or SEQ ID NO: 3 (B-GEn.2).
2. The engineered type V endonuclease according to claim 1 is an engineered B-GEn polypeptide.
3. The engineered type V endonuclease according to claim 1 or claim 2, wherein it has arginine at position D504 corresponding to SEQ ID NO: 1 (B-GEn.1) or D501 corresponding to SEQ ID NO: 2 (B-GEn.1.2) or SEQ ID NO: 3 (B-GEn.2).
4. An engineered type V endonuclease, optionally an engineered type V endonuclease according to any one of claims 1 to 3, comprising the following amino acid sequence: (a) having a substitution at position D501 corresponding to SEQ ID NO: 3, wherein the substitution: (i) is a substitution D501R compared to the amino acid sequence of SEQ ID NO: 3; or (ii) increases gene editing efficiency activity compared to the corresponding amino acid sequence without a substitution at position D501; and (b) has at least 90% sequence identity with SEQ ID NO: 1 (B-GEn.1), SEQ ID NO: 2 (B-GEn.1.2) or SEQ ID NO: 3 (B-GEn.2).
5. An engineered type V endonuclease comprising the amino acid sequence of SEQ ID NO:
4.
6. An engineered type V endonuclease comprising the amino acid sequence of SEQ ID NO:
5.
7. An engineered type V endonuclease comprising the amino acid sequence of SEQ ID NO:
6.
8. The engineered type V endonuclease according to any one of claims 1 to 7, comprising at least one NLS optionally located at the C-terminus of the amino acid sequence, further optionally wherein: (a) The engineered type V endonuclease lacks any NLS at the N-terminus of the amino acid sequence; or (b) The engineered type V endonuclease contains at least one NLS located at the N-terminus of the amino acid sequence.
9. A nucleic acid comprising a nucleotide sequence encoding an engineered type V endonuclease according to any one of claims 1 to 8.
10. The nucleic acid according to claim 9, wherein the nucleotide sequence encoding the engineered type V endonuclease according to any one of claims 1 to 8 is operatively linked to the promoter.
11. The nucleic acid according to claim 9 or claim 10, wherein the nucleic acid further encodes a guide RNA.
12. The nucleic acid according to any one of claims 9 to 11, wherein it is in vector form.
13. The nucleic acid according to claim 12, wherein the vector is a recombinant adeno-associated virus (AAV) vector.
14. A cell comprising the nucleic acid of any one of claims 9 to 13.
15. A cell engineered to express a nucleotide sequence encoding an engineered type V endonuclease of any one of claims 1 to 8.
16. The cell according to claim 14 or claim 15, wherein the cell is a eukaryotic cell, optionally wherein the eukaryotic cell is a human cell.
17. A method for producing an engineered type V endonuclease according to any one of claims 1 to 8, comprising culturing cells according to any one of claims 14 to 16 under conditions for producing the engineered type V endonuclease, the method optionally further comprising isolating and / or purifying the engineered type V endonuclease.
18. A composition comprising: (a) an engineered type V endonuclease according to any one of claims 1 to 8; and (b) a guide RNA.
19. The composition according to claim 18, wherein it is a ribonucleoprotein complex.
20. The composition according to claim 18 or claim 19, wherein the molar ratio of the engineered type V endonuclease to guide RNA is in the range of 1:1 to 1:
4.
21. A method for editing the genome of a cell, comprising introducing into the cell: (a) an engineered type V endonuclease according to any one of claims 1 to 8; and (b) a guide RNA.
22. A method for editing the genome of a cell, comprising introducing into the cell: (a) one or more nucleic acids encoding an engineered type V endonuclease as described in any one of claims 1 to 8; and (b) a guide RNA.
23. The method of claim 22, further comprising contacting the cells with lipid nanoparticles comprising the one or more nucleic acids and the guide RNA.
24. A method for editing the genome of a cell, comprising introducing one or more nucleic acids into the cell, said nucleic acids encoding: (a) an engineered type V endonuclease according to any one of claims 1 to 8; and (b) a guide RNA, optionally, wherein at least one of said one or more nucleic acids is a nucleic acid according to any one of claims 9 to 13.
25. The method of claim 24, further comprising contacting the cell with one or more recombinant AAV particles comprising the one or more nucleic acids.
26. The method of claim 24, further comprising contacting the cell with one or more lipid nanoparticles comprising the one or more nucleic acids.
27. A method for editing the genome of a cell, comprising introducing the composition of any one of claims 18 to 20 into the cell.
28. The method of claim 24, wherein the composition is a ribonucleoprotein complex.
29. The method of claim 26 or claim 27, wherein the composition is introduced into the cells via lipid nanoparticles.
30. The method according to any one of claims 21 to 29, wherein the cell is a mammalian cell, optionally a human cell.
31. The method of claim 30, wherein the cell is a stem cell or a cell differentiated therefrom, optionally wherein the stem cell is a pluripotent stem cell (PSC) or a cell differentiated therefrom.
32. A cell comprising: (a) the composition of any one of claims 18 to 20; or (b) the nucleic acid of any one of claims 9 to 13.
33. The cell according to claim 32, wherein it is a mammalian cell, optionally a human cell.
Citation Information
Patent Citations
Connector assemblies and associated methods
US11241132B2
Nuclease resistant chimeric oligonucleotides
US20030158403A1
Nuclear reprogramming factor and induced pluripotent stem cells
US20090047263A1
Nuclear Reprogramming Factor
US20090068742A1
Multipotent / pluripotent cells and methods
US20090191159A1