Engineered V-type RNA programmable endonucleases and uses thereof

By engineering the V-type CRISPR Cas nuclease B-GEn, combined with nuclear localization signals and linker sequences, the size and activity limitations of the existing CRISPR-Cas system in mammalian cells are overcome, achieving efficient and precise DNA editing, which is suitable for gene therapy.

CN120752335APending Publication Date: 2025-10-03BAYER AG
View PDF 161 Cites 0 Cited by

Patent Information

Application Number
CN202380091977.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-13
Filing Date
2023-12-12
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

The existing CRISPR-Cas system has limitations in size, activity, PAM selectivity, immune response and expression efficiency, making it difficult to effectively apply it to gene editing in mammalian cells.

Method used

We developed an engineered V-type CRISPR Cas nuclease, B-GEn, combined with a nuclear localization signal and a linker sequence to form a polypeptide construct for efficient DNA targeting and editing in mammalian cells.

Benefits of technology

It achieves efficient and precise DNA editing in mammalian cells, avoiding off-target effects and immune responses, and is suitable for gene therapy and other applications requiring high precision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120752335A_ABST
    Figure CN120752335A_ABST
Patent Text Reader

Abstract

The present disclosure provides engineered V-type nuclease enzymes suitable for editing eukaryotic genomic DNA, as well as methods of producing the nuclease enzymes, systems comprising the nuclease enzymes, and methods of editing eukaryotic genomic DNA using the nuclease enzymes and systems.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] 1. Cross-reference to related applications

[0002] This application claims priority to U.S. Provisional Application No. 63 / 432,232, filed December 13, 2022, the contents of which are incorporated herein by reference in their entirety.

[0003] 2. Sequence Listing

[0004] This application contains a sequence listing, which is filed electronically and is incorporated herein by reference in its entirety. The copy was created on November 17, 2023, is named BRT-002WO_SL, and is 160,005 bytes in size. Background Art

[0005] Clustered regularly interspaced short palindromic repeats (CRISPR) and CRISPR-associated (Cas) genes, collectively known as CRISPR-Cas or CRISPR / Cas systems, are currently believed to provide bacteria and archaea with immunity to phage infection. The CRISPR-Cas systems of prokaryotic adaptive immunity are an extremely diverse set of protein effectors, non-coding elements, and locus architectures, some of which have been engineered and adapted to produce important biotechnological applications.

[0006] The components of the system involved in host defense include one or more effector proteins capable of modifying DNA or RNA, and RNA guide elements responsible for targeting the activity of these proteins to specific sequences on phage DNA or RNA. The guide RNA consists of CRISPR RNA (crRNA) and may require additional trans-acting RNA (tracrRNA) to enable manipulation of the target nucleic acid by the effector protein. The crRNA consists of a segment called a "direct repeat sequence" that is responsible for binding the crRNA to the effector protein and a segment called a "spacer sequence" that is complementary to the desired nucleic acid target sequence. The CRISPR system can be reprogrammed to target alternative DNA or RNA targets by modifying the spacer sequence of the crRNA.

[0007] CRISPR-Cas systems can be broadly divided into two categories: Class 1 systems consist of multiple effector proteins that together form a complex around crRNA, and Class 2 systems consist of a single effector protein that, in complex with crRNA, guides the target DNA or RNA substrate. The single-subunit effector compositions of Class 2 systems provide a simpler set of components for engineering and application, and are therefore an important source of programmable effectors to date. Therefore, the discovery, engineering, and optimization of new Class 2 systems can generate widely applicable and powerful programmable technologies for genome engineering and other fields.

[0008] The RNA-guided DNA targeting principle of CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats)-Cas (CRISPR associated proteins) for genome editing has been widely used in recent years. Five types of CRISPR-Cas systems have been documented (Type I, Type II and Type IIb, Type III, Type V and Type VI). The most commonly used CRISPR-Cas for genome editing is the Type II system. The main advantage provided by the bacterial Type II CRISPR-Cas system is the minimum requirement for programmable DNA interference: the nuclease Cas9 guided by a customizable double RNA structure. As demonstrated in the original type II system of Streptococcus pyogenes, trans-activating CRISPR RNA (tracrRNA) binds to the invariant repeat sequence of precursor CRISPR RNA (pre-crRNA) to form a double RNA, which is essential for the co-maturation of crRNA by RNaseIII in the presence of Cas9 and for Cas9 to cut invading DNA. As demonstrated in Streptococcus pyogenes, Cas9 introduces site-specific double-stranded DNA (dsDNA) breaks in invading homologous DNA under the guidance of the duplex formed between mature activated tracrRNA and targeting crRNA. Cas9 is a multi-domain enzyme that uses the HNH nuclease domain to cut the target chain (defined as complementary to the spacer sequence of crRNA) and the RuvC-like domain to cut the non-target chain.

[0009] In addition to type II CRISPR Cas9 nucleases, many different type V CRISPR Cas9 nucleases have been described, such as Cas12a, Cas12b, Cas12e, Cas12f, Cas13a, Cas13b (Koonin et al., Curr Opin Microbiol. 2017 Jun; 37: 67–78, and Makarova et al., Nat Rev Microbiol. 2020 Feb; 18(2): 67–83.). Some of these systems do not require tracr RNA (Cas 12a, Cas 13a, Cas 13b), whereas Cas 12b nucleases generally require tracr RNA (Kooni et al., Curr Opin Microbiol. 2017 Jun; 37: 67–78).

[0010] Genome editing in mammalian cells has been limited to date, in part due to the size of the various Cas9 proteins. Cas9 from Staphylococcus pyogenes (SpyCas9) is the most widely used enzyme to date, containing approximately 4.2 kb of DNA (WO2013 / 176722), and direct binding to a cognate single guide RNA (sgRNA) further increases its size. Adeno-associated virus is one of the vectors used to deliver the Cas9 enzyme in gene therapy applications. However, the cargo size of AAV is limited to around 4.5 kb. Due to the size limitation, delivering Cas9 with its sgRNA and potential DNA repair template can be an obstacle to using this approach. Smaller Cas9 molecules have been characterized, but most of them have a protospacer adjacent motif (PAM) sequence that is not as well defined as the sequence used by SpyCas9. For example, Staphylococcus aureus (Sau) Cas9 uses the "NNGRR(T)" sequence, where R = A or G, and Campylobacter jejuni (Cja) Cas9 uses the "NNNACAC" / "NNNRYAC" PAMs (where Y = T or G), respectively. The ambiguity of the PAMs increases the likelihood that the enzyme will produce unwanted activity at off-target sequences that have high or complete sequence identity to the PAM. The specificity of these systems remains a concern, as accidental ("off-target") targeting of similar sites will increase the likelihood of adverse events.

[0011] Existing CRISPR-Cas systems typically have one or more of the following disadvantages:

[0012] a) They are too large to be carried within the genomes of established viral delivery systems suitable for therapeutics, such as adeno-associated virus (AAV).

[0013] b) Most of them have no appreciable activity in non-host environments, such as in eukaryotic cells, and particularly in mammalian cells.

[0014] c) Their nucleases can catalyze DNA strand cleavage when there are mismatches between the spacer sequence and the protospacer sequence, leading to undesirable off-target effects, for example, making them unsuitable for gene therapy use or other applications requiring high precision.

[0015] d) They may elicit an immune response that may limit their use in mammals.

[0016] e) They require complex and / or long PAMs that restrict target selection of the DNA targeting segment.

[0017] f) They show low expression on plasmid or viral vectors.

[0018] g) They require additional RNA sequences for activation or additional RNA sequences as part of the guide RNA.

[0019] The present disclosure provides novel engineered V-type nucleases suitable for use in systems that address one or more of the aforementioned shortcomings of existing CRISPR-Cas systems. Summary of the Invention

[0020] The present invention relates to an engineered V-type CRISPR Cas nuclease, named B-GEn, comprising nuclease sequences related to B-GEn.1 (SEQ ID NO: 4), B-GEn1.2 (SEQ ID NO: 5), and B-GEn.2 (SEQ ID NO: 6), one or more nuclear localization sequences (NLS), and optionally one or more linker sequences connecting the nuclease sequence to one or more NLS sequences. Exemplary engineered B-GEn polypeptides are described in Section 6.2, and their component nucleases, NLSs, and linker sequences are described in Sections 6.3, 6.4, and 6.5, respectively.

[0021] The present disclosure further provides a B-GEn V-type CRISPR Cas system comprising an engineered B-GEn polypeptide and a suitable guide RNA or nucleic acid encoding the same. Exemplary B-GEn V-type CRISPR Cas systems are disclosed in Section 6.6, and exemplary guide RNAs are disclosed in Section 6.7. In some embodiments, the B-GEn V-type CRISPR Cas system is a ribonucleoprotein (RNP) complex comprising an engineered B-GEn polypeptide and a guide RNA. The ribonucleoprotein complex is described in Section 6.8.

[0022] The present disclosure further provides nucleic acids encoding engineered B-GEn polypeptides, such as expression vectors for engineered B-GEn polypeptides. Exemplary nucleic acids are disclosed in Section 6.9, and exemplary vectors are disclosed in Section 6.10.

[0023] In certain aspects, provided herein is a method for targeting, editing, modifying or manipulating a target DNA at one or more sites in a cell or in vitro. The method generally requires that the B-GEn polypeptide suitable for engineering produces one or more nicks or cuts or base editing in the target DNA, and the B-GEn V type CRISPR Cas system is introduced into a cell or in vitro environment, wherein the B-GEn polypeptide of engineering is guided to the target DNA by the guide RNA existing in a processed or unprocessed form. In some embodiments, the RNP of the present disclosure (comprising engineered B-GEn polypeptides and guide RNA) is used to edit the genome of a cell. In some embodiments, the method for using RNP to carry out genomic DNA editing includes carrying out nuclear transfection with RNP to a target cell comprising genomic DNA, and the target cell is exposed to the conditions where gene editing occurs, such as by cultivating target cells under conditions suitable for carrying out genome editing by B-GEn polypeptides.

[0024] Additional features, advantages, and applications of the engineered B-GEn polypeptides of the present disclosure are described in more detail below. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figures 1A-1B Schematic diagram of exemplary engineered B-GEn polypeptide construct configurations. Figure 1A -1 is an engineered B-GEn construct with an NLS-B-GEn-NLS configuration, which has a nuclear localization signal (NLS) at its N-terminus and another NLS at its C-terminus, each linked to the B-GEn nuclease sequence via a linker (L). Figure 1A-2 The engineered B-GEn polypeptide construct has a B-GEn-NLS configuration, which has a linker and an NLS located only at its C-terminus. Figure 1A-3 is an engineered B-GEn polypeptide construct with an NLS-B-GEn configuration, which has a linker and an NLS located only at its N-terminus. Figure 1A-4 is an engineered B-GEn polypeptide construct without an NLS at either terminus. Figure 1B -1 is an engineered B-GEn polypeptide construct having a single NLS sequence connected to the B-GEn nuclease sequence via a linker. Figure 1B-2 The B-GEn construct has two NLS sequences at its C-terminus, each NLS sequence being connected via a linker. Figure 1B -3 is an engineered B-GEn polypeptide construct having three NLS sequences at its C-terminus, each NLS sequence connected by a linker. Figure 1B-4 It is an engineered B-GEn polypeptide construct having four NLS sequences at its C-terminus, each NLS sequence being connected by a linker. In some embodiments, the NLS sequences flanking either or both ends of the B-GEn nuclease sequence are nucleoplasmin NLS, SV40 NLS, c-myc NLS, or any other NLS sequence, and can be separated by the B-GEn nuclease sequence by a linker. Each NLS domain can contain more than one NLS sequence, which can be the same or different. In addition, the NLS flanking the two ends of the B-GEn nuclease sequence can be the same or different NLS. The term B-GEn generally refers to a V-type RNA programmable endonuclease having a sequence as described in Section 6.3, including but not limited to SEQ NO: 4 (B-GEn.1), SEQ ID NO: 5 (B-GEn.1.2), SEQ ID NO: 6 (B-GEn.2) and sequence variants thereof.

[0026] Figure 2 Cartoon representation of a cell-based assay that can be used for nuclear localization signal (NLS) screening.

[0027] Figures 3A-3C Shown are the amino acid sequences of exemplary engineered B-GEn.2 polypeptide constructs with different NLS configurations. Figure 3A is the amino acid sequence of a construct having an NLS-B-GEn.2-NLS configuration (SEQ ID NO: 172), wherein the N-terminal NLS comprises a nucleoplasmin NLS connected to a B-GEn.2 nuclease sequence via a linker, and the C-terminal NLS comprises an SV40 NLS connected to a B-GEn.2 nuclease sequence via another linker. Figure 3B is the amino acid sequence of a construct having a B-GEn.2-NLS configuration (SEQ ID NO: 2), wherein the NLS at the C-terminus comprises an SV40 NLS connected to a B-GEn.2 nuclease sequence via a linker. Figure 3CThe amino acid sequence of a construct having an NLS-B-GEn.2 configuration (SEQ ID NO: 173), wherein the NLS at the N-terminus comprises a nucleoplasmin NLS, which is linked to the B-GEn.2 nuclease sequence via a linker.

[0028] Figure 4 Figure 8.3 is a cartoon representation of an exemplary nuclease purification workflow for purifying engineered B-GEn.2 polypeptide (and control) constructs with various NLS configurations, as described in Section 8.3.

[0029] Figures 5A-5E Shown are the results of an exemplary 2-step purification of engineered B-GEn.2 polypeptide (and control) constructs. Figure 5A A series of heparin fast flow (FF) chromatograms are shown. The asterisk above the peak indicates the elution position of a single construct. The peak marked with a single asterisk corresponds to a B-GEn.2 polypeptide construct without an NLS, also known as B-GEn.2 (B-GEn.2-all-del) with a complete NLS deletion; the peak marked with a double asterisk corresponds to a construct with a B-GEn.2-NLS configuration, also known as B-GEn.2 (B-GEn.2-Ndel) with an N-terminal NLS deletion; and the peak marked with three stars corresponds to a construct with an NLS-B-GEn.2 configuration, also known as B-GEn.2 (B-GEn.2-Cdel) with a C-terminal NLS deletion. Figure 5B and 5C To carry Figure 5A Image of a stained SDS-polyacrylamide gel of purified fractions showing the peaks seen in Figure 5A The protein fractions of the peak in are marked with a single asterisk as B-GEn-2-homo-del, with a double asterisk as B-GEn.2-Ndel, and with a triple asterisk as B-GEn.2-Cdel. Figure 5D For Figure 5A Size Exclusion Chromatography (SEC) chromatogram of the protein fraction below the peak marked with a single asterisk in FIG. Figure 5E To carry Figure 5D Stained SDS-polyacrylamide gel of the purified fractions corresponding to Figure 5D The protein fraction of the large monomer peak is marked with a single asterisk.

[0030] Figure 6 Cartoon illustration of the in-cell gene editing workflow, which is described in detail in Section 8.4.1.

[0031] Figure 7To illustrate the use of Cpf1 / Cas12a or engineered B-GEn.2 polypeptide ( Figure 3A and capped with NLS at both its N- and C-termini; B-GEn.2Ndel: Figure 3B The B-GEn.2 constructs depicted in FIG. 1 and B-GEn.2Cdel: Figure 3C Figure 2. Histogram of B2M gene editing results in iPSCs using 4.3sgRNA using the B-GEn.2 construct depicted in Figure 2. WT is the control condition without Cas protein. The amount of RNP used is indicated below each protein. The Y-axis represents the percentage of gene editing.

[0032] Figures 8A-8B Bar graph showing the results of albumin gene editing in iPSCs using the B-GEn.2 variant. Figure 8A The results of albumin gene editing in iPSCs using sgRNAv4.3* are shown. Figure 8B Results of albumin gene editing in iPSCs using sgRNAv4.4* are shown. WT refers to the control condition without Cas protein; B-GEn.2 is Figure 3A The construct depicted in FIG, and capped with NLS at both its N- and C-termini. B-GEn.2Ndel is Figure 3B The B-GEn.2 construct depicted in FIG. 1 , B-GEn.2Cdel is Figure 3C (B-GEn.2 construct depicted in Figure 2). The amount of RNP used is indicated below each protein. The Y-axis represents the percentage of gene editing. * indicates 2'-O-methylation of the last three nucleotides at the 3' end of the sgRNA.

[0033] Figure 9 The bar graph depicts the results of albumin gene editing in HEK293-T cells using B-GEn.2 variants with sgRNA v4.3*. WT refers to the control condition without Cas protein; B-GEn.2 #1 and #2 are Figure 3A Two preparations of the construct depicted in , and capped with NLS at both the N- and C-termini. Figure 3B The B-GEn.2 construct depicted in FIG. 1 , B-GEn.2Cdel is Figure 3C The B-GEn.2 construct is depicted in Figure 2. The amount of RNP used is indicated below each protein. The Y-axis represents the percentage of gene editing. * indicates 2'-O-methylation of the last three nucleotides at the 3' end of the sgRNA. DETAILED DESCRIPTION

[0034] 6.1. Definitions

[0035] Unless otherwise defined herein, the scientific and technical terms used in conjunction with the present disclosure should have the meanings commonly understood by those of ordinary skill in the art. Exemplary methods and materials are described below, but methods and materials similar or equivalent to those described herein can also be used in the implementation or testing of the present disclosure. In the event of a conflict, the present specification (including definitions) shall prevail. Generally, the nomenclature and techniques used in connection with cell and tissue culture, molecular biology, immunology, microbiology, genetics, analytical chemistry, synthetic organic chemistry, medical and pharmaceutical chemistry, and protein and nucleic acid chemistry and hybridization as described herein are those well known and commonly used in the art. Enzymatic reactions and purification techniques, as commonly accomplished in the art or as described herein, are performed according to the manufacturer's instructions. In addition, unless the context otherwise requires, singular terms shall include the plural, and plural terms shall include the singular. Throughout this specification and embodiments, the words "have" and "comprise" or variations such as "has," "having," "comprises," or "comprising" will be understood to imply the inclusion of the stated integer or group of integers, but not to exclude any other integer or group of integers. All publications and other references mentioned herein are incorporated by reference in their entirety.Although a number of documents are cited herein, this citation does not constitute an admission that any of these documents forms part of the common general knowledge in the art.

[0036] B-GEn polypeptide: As used herein, the term "B-GEn polypeptide" refers to a polypeptide comprising at least the amino acid sequence of the nuclease domain of B-Gen.1 (SEQ ID NO: 4, whose nuclease domain comprises RuvC I, RuvC II and RuvC III subdomains corresponding to amino acids 542-624, 825-876 and 956-970, respectively), B-GEn.1.2 (SEQ ID NO: 5, whose nuclease domain comprises RuvC I, RuvC II and RuvC III subdomains corresponding to amino acids 537-621, 822-873 and 954-968, respectively), B-GEn.2 (SEQ ID NO: 6, whose nuclease domain comprises RuvC I, RuvC II and RuvC III subdomains corresponding to amino acids 537-621, 822-873 and 954-968, respectively), or a variant thereof having at least 50% sequence identity to such nuclease domain. "B-GEn polypeptide" also encompasses variants of any of B-GEn.1 (SEQ ID NO: 4), B-GEn.1.2 (SEQ ID NO: 5), and B-GEn.2 (SEQ ID NO: 6), such as variants comprising an amino acid sequence having at least 50% sequence identity to SEQ ID NO: 4, SEQ ID NO: 5, or SEQ ID NO: 6 and / or (2) variants comprising an amino acid sequence that differs from SEQ ID NO: 4, SEQ ID NO: 5, or SEQ ID NO: 6 by up to 25 amino acids. In some embodiments, the B-GEn polypeptide has nuclease activity.The term "B-GEn polypeptide" encompasses an engineered fusion polypeptide comprising the amino acid sequence of B-GEn.1 (SEQ ID NO: 4), B-GEn.1.2 (SEQ ID NO: 5), B-GEn.2 (SEQ ID NO: 6), or any variant thereof as described in Section 6.3, such as (1) an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, or 100% sequence identity to the nuclease domain or full length of any one of SEQ ID NOs: 4 to 6 and / or (2) an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% sequence identity, at least 99% sequence identity, or 100% sequence identity to the nuclease domain or full length of any one of SEQ ID NOs: 4 to 6. Any of NOs: 4 to 6 differs by up to 25 amino acids, up to 20 amino acids, up to 15 amino acids, up to 14 amino acids, up to 13 amino acids, up to 12 amino acids, up to 11 amino acids, up to 10 amino acids, up to 9 amino acids, up to 8 amino acids, up to 7 amino acids, up to 6 amino acids, or up to 5 amino acid sequences, and additional sequences (such as one or more nuclear localization sequences described in Section 6.3 and / or one or more linker sequences described in Section 6.5). Exemplary engineered B-GEn polypeptides are described in Section 6.2.

[0037] Binding: As used herein, the term "binding" (e.g., with respect to RNA binding domains of polypeptides) refers to non-covalent interactions between macromolecules (e.g., between proteins and nucleic acids). Macromolecules are said to be "associated" or "interacting" or "bound" when in a non-covalent interaction state (e.g., when molecule X is said to interact with molecule Y, this means that molecule X binds to molecule Y in a non-covalent manner). Not all components of a binding interaction need to be sequence specific (e.g., contacts with phosphate residues in a DNA backbone), but some portions of a binding interaction may be sequence specific. A binding interaction is generally characterized by a dissociation constant (Kd) that is less than 10 -6 M, less than 10 -7 M, less than 10 -8 M, less than 10 -9 M, less than 10 -10 M, less than 10 -11 M, less than 10 -12 M, less than 10 -13 M, less than 10 -14 M or less than 10 -15 M. "Affinity" refers to the strength of binding, with increases in binding affinity correlating with decreases in Kd.

[0038] Cell therapy: As used herein, the term "cell therapy" refers to a therapy in which cellular material is administered to a patient. The cellular material can be whole, living cells. For example, T cells capable of fighting cancer cells through cell-mediated immunity can be administered during immunotherapy. Cell therapy is also known as cellular therapy or cytotherapy.

[0039] Coding sequence: As used herein, the term "coding sequence" or "encoding nucleic acid" refers to a sequence within a nucleic acid (RNA or DNA) molecule that encodes a protein or RNA molecule. The coding sequence may further include start and stop signals operably linked to regulatory elements, including a promoter and a polyadenylation signal, which direct expression of the nucleic acid in the cells of the individual or mammal into which it is introduced or administered. Codon-optimized coding sequences can be used for expression in the desired cells.

[0040] Complementarity: As used herein, in the context of nucleic acid molecules, the terms "complement" and "complementary" refer to the ability to form Watson-Crick (e.g., AT / U and CG) or Hoogsteen base pairing between nucleotides or nucleotide analogs of a nucleic acid molecule. "Complementarity" refers to a property shared between two nucleic acid sequences such that when they are aligned antiparallel to each other, the nucleotide bases at every position will be complementary.

[0041] Encodes: The term "encoding" in reference to a nucleic acid (DNA or RNA) means that the nucleic acid comprises a nucleotide sequence that encodes the amino acids of a polypeptide or nucleotides of an RNA.

[0042] Expression cassette: As used herein, the term "expression cassette" refers to a DNA coding sequence operably linked to a promoter.

[0043] Guide RNA: As used herein, the term "guide RNA" refers to a ribonucleic acid having a DNA targeting sequence (also referred to as a "spacer" or "DNA targeting segment") and a protein binding sequence (also referred to as a "protein binding segment"). The DNA targeting sequence has sufficient complementarity to the target DNA (e.g., genomic DNA) sequence to hybridize to the target DNA sequence and guide sequence-specific binding of the nucleic acid targeting complex to the target DNA sequence. The DNA targeting sequence generally includes a "protospacer-like" sequence as described herein. The protein binding sequence interacts with a site-specific modifying enzyme (e.g., a B-GEn polypeptide as described in Sections 6.2 and 6.3 below). Site-specific cleavage of the target DNA occurs at a location determined by: (i) base pairing complementarity between the guide RNA and the target DNA; and (ii) a short motif in the target DNA (referred to as a protospacer adjacent motif (PAM)). The protein binding segment portion of the guide RNA comprises two complementary nucleotide stretches that hybridize to form a double-stranded RNA duplex (dsRNA duplex). In some embodiments, the guide RNA is a single-stranded guide RNA (sgRNA).

[0044] Guide RNA and site-specific modification enzymes, such as B-GEn polypeptides, can form ribonucleoprotein complexes (e.g., bound by non-covalent interactions). Guide RNA provides the targeting specificity of the complex by comprising a nucleotide sequence complementary to the target DNA sequence. The site-specific modification enzyme of the complex provides endonuclease activity. In other words, the site-specific modification enzyme is guided to the target DNA sequence (e.g., target sequence in chromosomal nucleic acid; target sequence in extrachromosomal nucleic acid, such as free nucleic acid, minicircle, etc.; target sequence in mitochondrial nucleic acid; target sequence in chloroplast nucleic acid; target sequence in plasmid; etc.) by combining it with the protein binding segment of the guide RNA.

[0045] Heterologous: As used herein, the term "heterologous" refers to a nucleic acid or peptide that is not found in a naturally occurring nucleic acid or polypeptide, respectively. The B-GEn.1, or B-GEn.1.2, or B-GEn.2 fusion proteins described herein may, in some embodiments, comprise the DNA or RNA binding domain of a B-GEn.1, or B-GEn.1.2, or B-GEn.2 polypeptide (or variant thereof) fused to a heterologous polypeptide sequence (e.g., a polypeptide sequence from a protein other than B-GEn.1 or B-GEn.2). The heterologous polypeptide may exhibit an activity (e.g., an enzymatic activity) (e.g., methyltransferase activity, acetyltransferase activity, kinase activity, ubiquitination activity, etc.) also exhibited by the B-GEn.1, or B-GEn.1.2, or B-GEn.2 fusion protein. A heterologous nucleic acid may be linked to a naturally occurring nucleic acid (or variant thereof) (e.g., by genetic engineering) to produce a fusion nucleic acid encoding the fusion polypeptide. In another example, in a B-GEn.1, or B-GEn.1.2, or B-GEn.2 polypeptide fusion variant, the B-GEn.1, or B-GEn.1.2, or B-GEn.2 polypeptide variant can be fused to a heterologous polypeptide (e.g., a polypeptide other than B-GEn.1 or B-GEn.2) that exhibits the activity exhibited by the B-GEn.1, or B-GEn.1.2, or B-GEn.2 polypeptide fusion variant. A heterologous nucleic acid can be linked to the B-GEn.1, or B-GEn.1.2, or B-GEn.2 polypeptide variant (e.g., by genetic engineering) to produce a nucleic acid encoding the B-GEn.1, or B-GEn.1.2, or B-GEn.2 polypeptide fusion variant. As used herein, "heterologous" additionally refers to a nucleotide or polypeptide in a cell that is not a native cell.

[0046] Host cell: As used herein, the terms "host cell" and "recombinant host cell" refer to cells that have been genetically engineered, such as by introducing a heterologous polypeptide or nucleic acid (such as a vector or system of the present disclosure). It should be understood that such terms refer not only to specific subject cells, but also to the progeny of such cells. In some embodiments, the host cell carries the vector of the present disclosure as an extrachromosomal heterologous expression vector. In some embodiments, the host cell comprises any one of the engineered B-GEn polypeptides disclosed herein, such as introduced as an RNP complex. In other embodiments, the host cell has been gene-edited by the engineered B-GEn polypeptides of the present disclosure.

[0047] iPSC: As used herein, the terms "induced pluripotent stem cell" and "iPSC" refer to a type of pluripotent stem cell artificially prepared from non-pluripotent cells (e.g., adult somatic cells, partially differentiated cells, or terminally differentiated cells, such as fibroblasts, hematopoietic lineage cells, muscle cells, neurons, epidermal cells, etc.) by introducing or contacting the cells with one or more reprogramming factors. iPSC can be derived from a variety of different cell types, including terminally differentiated cells. iPSC has an embryonic stem (ES) cell-like morphology, grows as flat colonies, has a large nuclear-to-cytoplasmic ratio, has clear boundaries, and has prominent nuclei. In addition, iPSC expresses one or more key pluripotency markers known to those of ordinary skill in the art, including but not limited to alkaline phosphatase, SSEA3, SSEA4, Sox2, Oct3 / 4, Nanog, TRA160, TRA181, TDGF 1, Dnmt3b, Fox03, GDF3, Cyp26al, TERT, and zfp42.

[0048] Generate and characterize the method example of iPSC and can refer to, for example, US Patent Publication No. US20090047263, US20090068742, US20090191159, US20090227032, US20090246875 and US20090304646 and PCT Patent Publication No. WO2013177133 and WO2022204567, each of which is incorporated herein by reference. Generally, in order to generate iPSC, it is necessary to provide somatic cells with reprogramming factors known in the art (such as Oct4, SOX2, KLF4, MYC, Nanog, Lin28, etc.), somatic cells are reprogrammed to become pluripotent stem cells.

[0049] Nucleic acid: As used herein, the terms "nucleic acid," "oligonucleotide," and "nucleic acid" refer to at least two nucleotides covalently linked together. Nucleic acids can be single-stranded or double-stranded, or can contain portions of double-stranded and single-stranded sequences. Nucleic acids can be DNA (both genomic and cDNA), RNA, or hybrids, wherein nucleic acids can contain a combination of deoxyribo- and ribo-nucleotides, as well as a combination of bases, including uracil, adenine, thymine, cytosine, guanine, inosine, xanthine, hypoxanthine, isocytosine, and isoguanine. Nucleic acids can be obtained by chemical synthesis methods or by recombinant methods. The description of a single strand also defines the sequence of the complementary strand. Therefore, reference to a single-stranded nucleic acid herein also encompasses the complementary strand of the described single strand.

[0050] Nuclear localization signal: As used herein, the terms "nuclear localization signal" and "NLS" refer to an amino acid sequence that promotes the localization of a polypeptide to the nucleus of a eukaryotic cell.

[0051] Nuclease: As used herein, the terms "nuclease" and "endonuclease" are used interchangeably herein to refer to an enzyme having endonucleolytic activity for nucleic acid cleavage, and nuclease-inactive variants thereof.

[0052] Nuclease Domain: As used herein, the terms "nuclease domain" and "cleavage domain" or "active domain" of a nuclease refer to the amino acid sequence or domain within a nuclease that possesses DNA cleavage catalytic activity. A cleavage domain can be contained within a single polypeptide chain, or the cleavage activity can result from the binding of two (or more) polypeptides. Within a given polypeptide, a single nuclease domain can be composed of more than one discrete stretch of amino acids. For example, the nuclease domain of B-GEn endonuclease comprises RuvC I, RuvC II, and RuvC III subdomains corresponding to amino acids 542-624, 825-876, and 956-970, respectively, of SEQ ID NO: 4 (B-GEn. 1), amino acids 537-621, 822-873, and 954-968, respectively, of SEQ ID NO: 5 (B-GEn. 1.2), and amino acids 537-621, 822-873, and 954-968, respectively, of SEQ ID NO: 6 (B-GEn. 2). In some embodiments, the boundaries of the nuclease domain of the B-GEn polypeptide are determined by aligning the B-GEn polypeptide with AacC2C1 (Cas12b) (Uniprot # T0D7A2) and identifying the amino acid aligned with the AacC2C1 (Cas12b) RuvC nuclease domain corresponding to amino acids R519 to S628. For B.GEn.2, the RuvC domain boundary corresponds to amino acids R514 to S620 and has D566 as an active site residue. In other embodiments, the boundaries of the nuclease domain of the B-GEn polypeptide are determined by aligning the B-GEn polypeptide with BthCas12b (Wu et al., 2017, Cell Research 27: 705-708) and identifying the amino acid aligned with the BthCas12b RuvC nuclease domain comprising RuvC I, RuvC II, and RuvC III subdomains.

[0053] Nucleofection: As used herein, the term "nucleofection" refers to an electroporation-based transfection method that uses a combination of electrical parameters and cell type-specific reagents to transfer nucleic acids (such as DNA or RNA) and RNPs directly to the nucleus of target cells.

[0054] Operably linked: As used herein, the term "operably linked" refers to a functional relationship between two or more peptide or polypeptide domains or nucleic acid (e.g., DNA) segments. In the context of transcriptional regulation, the term refers to the functional relationship of a transcriptional regulatory sequence to a transcribed sequence. For example, a promoter or enhancer sequence is operably linked to a coding sequence if it stimulates or regulates transcription of the coding sequence in a suitable host cell or other expression system.

[0055] Polypeptides, peptides and proteins: As used herein, the terms "polypeptide," "peptide," and "protein" refer to amino acid polymers of any length. In various embodiments, the polymers may be linear or branched, may contain modified amino acids, and may be interrupted by non-amino acids.

[0056] Pluripotency: As used herein, the term "pluripotent" or "pluripotency" refers to the ability of a cell to self-renew and differentiate into cells of any of the three germ layers: endoderm, mesoderm, or ectoderm. "Pluripotent stem cells" or "PSCs" include, for example, embryonic stem cells derived from the inner cell mass of a blastocyst or obtained by somatic cell nuclear transfer, and iPSCs derived from non-pluripotent cells.

[0057] Promoter: As used herein, the term "promoter" refers to a nucleotide sequence that is recognized by the synthetic machinery of the cell or an introduced synthetic machinery that is required to initiate the specific transcription of a nucleic acid sequence. A promoter can be a constitutively active promoter (e.g., a constitutive promoter that is in an active "on" state), it can be an inducible promoter (e.g., a promoter whose activity / "on" or inactive / "off" state is controlled by an external stimulus, such as the presence of a specific temperature, compound, or protein), it can be a spatially restricted promoter (e.g., a transcriptional control element, an enhancer, etc.) (e.g., a tissue-specific promoter, a cell type-specific promoter, etc.), and it can be a temporally restricted promoter (e.g., a promoter that is in an "on" state or an "off" state at a specific stage of embryonic development or a specific stage of a biological process, such as the hair follicle cycle in mice).

[0058] Protospacer adjacent motif: As used herein, the term "protospacer adjacent motif" or "PAM" refers to a DNA sequence downstream (e.g., immediately downstream) of a target sequence on a non-target strand recognized by a Cas protein. The PAM sequence is located 3' to the target sequence on the non-target strand.

[0059] Recombinant: As used herein, the term "recombinant" with respect to a nucleic acid, polypeptide, or cell refers to a nucleic acid (DNA or RNA), polypeptide, or cell that is, directly or indirectly, the product of genetic engineering (e.g., a progeny or replica of a nucleic acid, polypeptide, or cell produced by genetic engineering methods). For example, a recombinant vector can be the product of various combinations of cloning, restriction, polymerase chain reaction (PCR), and / or ligation steps, resulting in a construct having structural coding or non-coding sequences that are distinct from endogenous nucleic acids found in natural systems. A DNA sequence encoding a polypeptide can be assembled from cDNA fragments or a series of synthetic oligonucleotides to provide a synthetic nucleic acid that can be expressed in a recombinant transcription unit contained in a cell or in a cell-free transcription and translation system. Genomic DNA containing the relevant sequence can also be used to form a recombinant gene or transcription unit. Non-translated DNA sequences can be present 5' or 3' to the open reading frame, where these sequences do not interfere with manipulation or expression of the coding region and can indeed regulate the production of the desired product through various mechanisms (see "DNA regulatory sequences" below). Additionally or alternatively, DNA sequences encoding untranslated RNA (e.g., guide RNA) can also be considered recombinant. Therefore, the term "recombinant" nucleic acid refers to a non-naturally occurring nucleic acid, such as one that is artificially combined from two otherwise separate sequence segments through human intervention. This artificial combination is typically accomplished by chemical synthesis or by artificial manipulation of separate nucleic acid segments, such as through genetic engineering techniques. This is typically accomplished by replacing codons with codons encoding the same amino acid, conservative amino acids, or non-conservative amino acids. Additionally or alternatively, nucleic acid segments encoding the desired functions are linked together to produce the desired functional combination. This artificial combination is typically accomplished by chemical synthesis or by artificial manipulation of separate nucleic acid segments, such as through genetic engineering techniques. When a recombinant nucleic acid encodes a polypeptide, the sequence of the encoded polypeptide may be naturally occurring ("wild-type") or a variant (e.g., a mutant) of a naturally occurring sequence. Therefore, the term "recombinant" polypeptide does not necessarily refer to a polypeptide whose sequence is not naturally occurring. In contrast, a "recombinant" polypeptide is encoded by a recombinant DNA sequence, but the polypeptide sequence may be naturally occurring ("wild-type") or non-naturally occurring (e.g., a variant, mutant, etc.). Therefore, a "recombinant" polypeptide is the result of human intervention, but may also be a naturally occurring amino acid sequence. The term "non-naturally occurring" includes molecules that are significantly different from their naturally occurring counterparts, including chemically modified or mutated molecules.

[0060] Regulatory sequence: As used herein, the term "regulatory sequence" refers to a nucleic acid sequence required for expression of an operably linked sequence of interest, such as a guide RNA or an engineered B-GEn polypeptide sequence. In some instances, a regulatory sequence may be a promoter sequence, and in other instances, a regulatory sequence may include promoter and enhancer sequences and / or other regulatory elements required for expression of a polypeptide. A regulatory sequence may be, for example, a sequence that drives expression of an operably linked sequence in a constitutive or tissue-specific manner.

[0061] Ribonucleoprotein (RNP) complex, ribonucleoprotein (RNP) particle: As used herein, the terms "ribonucleoprotein complex" and "ribonucleoprotein particle" refer to a complex or particle comprising a ribonucleoprotein and a ribonucleic acid. As provided herein, "ribonucleoprotein" refers to a protein capable of binding to nucleic acids (e.g., RNA, DNA). When a ribonucleoprotein binds to a ribonucleic acid, it is referred to as a "ribonucleoprotein." The interaction between the ribonucleoprotein and the ribonucleic acid can be direct, such as through a covalent bond, or indirect, such as through a non-covalent bond (e.g., electrostatic interaction (e.g., ionic bond, hydrogen bond, halogen bond), van der Waals interaction (e.g., dipole-dipole, dipole-induced dipole, London dispersion), ring stacking (pi effect), hydrophobic interaction, etc.). In an embodiment, the ribonucleoprotein includes an RNA binding motif that is non-covalently bound to the ribonucleic acid. For example, the positively charged aromatic amino acid residues (e.g., lysine residues) in the RNA binding motif can form electrostatic interactions with the negative nucleic acid phosphate backbone of the RNA, thereby forming a ribonucleoprotein complex. In some embodiments, any of the engineered B-GEn polypeptides disclosed herein is in an RNP with a guide RNA.

[0062] Spacer: As used herein, the term "spacer" refers to a region of a gRNA molecule that is partially or fully complementary to a target sequence found in the + or - strand of genomic DNA. When complexed with the Cas protein, the gRNA guides the Cas protein to the target sequence in the genomic DNA. The length of the spacer is typically 15 to 30 nucleotides (e.g., 20 to 25 nucleotides). The nucleotide sequence of the spacer can, but is not necessarily, fully complementary to the target sequence. For example, in some embodiments, the spacer can contain one or more mismatches with the target sequence, such as, the spacer can contain one, two, or three mismatches with the target sequence.

[0063] Stem-loop structure: As used herein, the term "stem-loop structure" refers to a nucleic acid with a secondary structure comprising a region of nucleotides known or predicted to form a double strand (stem portion) that is connected on one side by a region of predominantly single-stranded nucleotides (loop portion). The terms "hairpin" and "fold-back" structures are also used herein to refer to stem-loop structures. Such structures are well known in the art, and the use of these terms is consistent with their known meanings in this field. As is known in the art, stem-loop structures do not require precise base pairing. Therefore, the stem structure may include one or more base mispairings. Alternatively, base pairing may be precise, e.g., without including any mispairings.

[0064] Target cell: As used herein, the term "target cell" refers to a cell into which a nuclease (e.g., the B-GEn system of the present disclosure) is introduced, such as a cell comprising a target DNA in its genome. It should be understood that such terms refer not only to specific subject cells, but also to the progeny of such cells. Because gene editing can occur in cells as a result of the nuclease system, such progeny do not need to be identical to the parent cell into which the system was originally introduced, but rather include gene-edited counterparts of the cell. Such gene-edited progeny are still included within the scope of the term "target cell" as used herein.

[0065] Target DNA: As used herein, the term "target DNA" refers to a polydeoxyribonucleotide comprising a "target site" or "target sequence." The terms "target site," "target sequence," "target protospacer DNA," or "protospacer-like sequence" are used interchangeably herein to refer to a nucleic acid sequence present in the target DNA to which the DNA targeting segment (also referred to as a "spacer") of a guide RNA will bind, provided sufficient binding conditions are provided. For example, the target site (or target sequence) 5'-GAGCATATC-3' within the target DNA is targeted (or bound to, hybridized to, or complementary to) the RNA sequence 5'-GAUAUGCUC-3'. Suitable DNA / RNA binding conditions include physiological conditions that are typically present in cells. Other suitable DNA / RNA binding conditions are known in the art (e.g., conditions in cell-free systems); see, for example, Sambrook, J. and Russell, W., 2001. Molecular Cloning: A Laboratory Manual, Third Edition, Cold Spring Harbor Laboratory Press. The target DNA strand that is complementary to and hybridizes with the guide RNA is referred to as the "complementary strand," and the target DNA strand that is complementary to the "complementary strand" (and therefore not complementary to the guide RNA) is referred to as the "non-complementary strand." In some embodiments, the target DNA is genomic DNA.

[0066] Transfection: As used herein, the term "transfection" refers to the introduction of a nucleic acid molecule, such as a DNA or RNA (e.g., mRNA) molecule, into a cell, such as into the nucleus of a target cell or a producer cell. In the context of the present disclosure, the term "transfection" encompasses any method known to those skilled in the art for introducing a nucleic acid molecule into a cell, such as into a eukaryotic cell, for example, into a mammalian cell. Such methods include, for example, electroporation, lipofection (e.g., based on cationic lipids and / or liposomes), calcium phosphate precipitation, nanoparticle-based transfection, virus-based transfection, or cationic polymer-based transfection (e.g., DEAE-dextran or polyethyleneimine).

[0067] Vector: As used herein, the term "vector" refers to a nucleic acid molecule capable of transporting another nucleic acid to which it has been connected. One type of vector is a "plasmid", which refers to a circular double-stranded DNA loop to which additional DNA segments can be connected. Another type of vector is a viral vector, in which additional DNA segments can be connected to the viral genome. Certain vectors are capable of autonomous replication in the host cell into which they are introduced (e.g., bacterial vectors and episomal mammalian vectors with bacterial replication origins). Other vectors (e.g., non-episomal mammalian vectors) can be integrated into the genome of the host cell after being introduced into the host cell, and thereby replicated together with the host genome. In addition, certain vectors can direct the expression of the nucleotide sequence to which they are operably connected. Such vectors are referred to herein as "expression vectors". In some embodiments, the vector is a viral vector, such as an adenoviral vector or an adeno-associated virus (AAV) vector.

[0068] 6.2 Engineered B-GEn Endonuclease

[0069] The present disclosure relates to engineered Type V CRISPR-Cas endonucleases comprising:

[0070] (i) a nuclease sequence, such as a B-GEn polypeptide sequence, as described in Section 6.3;

[0071] (ii) one or more NLS sequences, as described in Section 6.4; and

[0072] (iii) an optional linker, as described in Section 0.

[0073] An exemplary configuration of an engineered type V CRISPR-Cas endonuclease is shown in Figure 1A and 1B middle.

[0074] In certain aspects, the engineered V-type CRISPR-Cas nuclease comprises a nuclease sequence and a first NLS sequence at the C-terminus of the nuclease sequence. The engineered V-type CRISPR-Cas nuclease comprising the nuclease sequence and the first NLS sequence may further comprise a first linker sequence between the nuclease sequence and the first NLS sequence. An exemplary configuration is shown in Figure 1B of the B-1.

[0075] In certain aspects, the engineered V-type CRISPR-Cas endonuclease comprises more than one NLS sequence (e.g., more than one NLS sequence C-terminal to the nuclease sequence).

[0076] In some embodiments, the engineered V-type CRISPR-Cas nuclease comprises a second NLS sequence located at the C-terminus of the first NLS sequence. The engineered V-type CRISPR-Cas nuclease comprising the second NLS sequence may further comprise a linker sequence between the first NLS sequence and the second NLS sequence. Figure 1B of the B-2.

[0077] In a further embodiment, the engineered V-type CRISPR-Cas nuclease comprises a third NLS sequence located at the C-terminus of the second NLS sequence. The engineered V-type CRISPR-Cas nuclease comprising the third NLS sequence may further comprise a linker sequence between the second NLS sequence and the third NLS sequence. An exemplary configuration is shown in Figure 1B of B-3.

[0078] In another embodiment, the engineered V-type CRISPR-Cas nuclease comprises a fourth NLS sequence located at the C-terminus of the third NLS sequence. The engineered V-type CRISPR-Cas nuclease comprising the fourth NLS sequence may further comprise a linker sequence between the third NLS sequence and the fourth NLS sequence. An exemplary configuration is shown in Figure 1B of B-4.

[0079] In some embodiments, the engineered V-type CRISPR-Cas endonuclease comprises an N-terminal NLS sequence in addition to one or more NLS sequences at the C-terminus of the nuclease sequence. Thus, in certain aspects, the engineered V-type CRISPR-Cas endonuclease comprises an NLS sequence at the N-terminus of the nuclease sequence. The engineered V-type CRISPR-Cas endonuclease comprising an NLS sequence at the N-terminus of the nuclease sequence may further comprise a linker sequence between the NLS sequence and the nuclease sequence. Figure 1AIn certain embodiments, the engineered V-type CRISPR-Cas endonuclease comprises more than one N-terminal NLS sequence (e.g., more than one nuclease sequence N-terminal NLS sequence, which may be connected by one or more linkers).

[0080] Exemplary amino acid sequences of engineered type V CRISPR-Cas endonucleases are listed in Table 1. In some embodiments, the engineered type V CRISPR-Cas endonuclease comprises the amino acid sequence of SEQ ID NO: 1. In some embodiments, the engineered type V CRISPR-Cas endonuclease comprises the amino acid sequence of SEQ ID NO: 2. In some other embodiments, the engineered type V CRISPR-Cas endonuclease comprises the amino acid sequence of SEQ ID NO: 3.

[0081]

[0082]

[0083]

[0084] In one embodiment, the B-GEn fusion protein (SEQ ID NO: 1) comprises a nucleoplasmin NLS (SEQ ID NO: 8) at the N-terminus and a SV40 large T protein NLS (SEQ ID NO: 7) at the C-terminus of a B-GEn.2 sequence (SEQ ID NO: 6).

[0085] In another embodiment, the B-GEn fusion protein (SEQ ID NO: 2) comprises the SV40 large T protein NLS (SEQ ID NO: 7) at the C-terminus of the B-GEn.2 sequence (SEQ ID NO: 6).

[0086] In another embodiment, the B-GEn fusion protein (SEQ ID NO: 3) comprises a nucleoplasmin NLS (SEQ ID NO: 8) at the N-terminus of the B-GEn.2 sequence (SEQ ID NO: 6).

[0087] 6.3 Nuclease Sequences

[0088] The present disclosure provides engineered type V CRISPR-Cas endonucleases, which specifically comprise a B-GEn nuclease amino acid sequence.

[0089] The engineered B-GEn polypeptides of the present disclosure generally comprise an amino acid sequence that is at least 50% identical to the nuclease domain (e.g., an amino acid sequence consisting of the RuvC I, RuvC II, and RuvC III subdomains) or the full length of SEQ ID NO: 4, or SEQ ID NO: 5, or SEQ ID NO: 6 and / or differs from SEQ ID NO: 4, SEQ ID NO: 5, or SEQ ID NO: 6 by up to 25 amino acids. In various embodiments, the engineered B-GEn polypeptides of the present disclosure comprise an amino acid sequence that is at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 85%, at least 90%, or at least 95% identical to the nuclease domain (e.g., an amino acid sequence consisting of the RuvC I, RuvC II, and RuvC III subdomains) or the full length of SEQ ID NO: 4, or SEQ ID NO: 5, or SEQ ID NO: 6. In some embodiments, the engineered B-GEn polypeptide comprises an amino acid sequence that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99% identical, or at least 99.5% identical to the nuclease domain (e.g., an amino acid sequence consisting of the RuvC I, RuvC II, and RuvC III subdomains) or the full length of SEQ ID NO: 4 or SEQ ID NO: 5 or SEQ ID NO: 6. In some embodiments, the engineered B-GEn polypeptide comprises an amino acid sequence that is identical to the nuclease domain or the full length of SEQ ID NO: 4 or SEQ ID NO: 5 or SEQ ID NO: 6.

[0090] In a further aspect, the engineered B-GEn polypeptide of the present disclosure can comprise an amino acid sequence that differs by up to 25 amino acids from the nuclease domain sequence (e.g., an amino acid sequence consisting of the RuvC I, RuvC II, and RuvC III subdomains) of SEQ ID NO: 4, SEQ ID NO: 5, or SEQ ID NO: 6. In some embodiments, the engineered B-GEn polypeptide comprises an amino acid sequence that differs by up to 25 amino acids, up to 20 amino acids, up to 15 amino acids, up to 14 amino acids, up to 13 amino acids, up to 11 amino acids, up to 10 amino acids, up to 9 amino acids, up to 8 amino acids, up to 7 amino acids, up to 6 amino acids, or up to 5 amino acids from the nuclease domain sequence (e.g., an amino acid sequence consisting of the RuvC I, RuvC II, and RuvCIII subdomains) of any one of SEQ ID NOs: 4 to 6.

[0091] In yet further aspects, the engineered B-GEn polypeptides of the present disclosure may comprise an amino acid sequence that differs from the entire sequence of SEQ ID NO: 4, SEQ ID NO: 5, or SEQ ID NO: 6 by up to 25 amino acids. In some embodiments, the engineered B-GEn polypeptides differ from the full length of any one of SEQ ID NOs: 4 to 6 by up to 25 amino acids, up to 20 amino acids, up to 15 amino acids, up to 14 amino acids, up to 13 amino acids, up to 12 amino acids, up to 11 amino acids, up to 10 amino acids, up to 9 amino acids, up to 8 amino acids, up to 7 amino acids, up to 6 amino acids, or up to 5 amino acids.

[0092] Exemplary B-GEn.1, B-GEn.1.2, and B-GEn.2 nuclease sequences are shown in Table 2.

[0093]

[0094]

[0095]

[0096] One embodiment according to the present disclosure is a polypeptide comprising a nuclease sequence having at least 80% sequence identity to any one of SEQ ID NO: 4, SEQ ID NO: 5 or SEQ ID NO: 6, or a nucleic acid comprising a nucleotide sequence encoding a polypeptide comprising a nuclease sequence having at least 80% sequence identity to any one of SEQ ID NO: 4, SEQ ID NO: 5 or SEQ ID NO: 6.

[0097] Another embodiment according to the present disclosure is a polypeptide comprising a nuclease sequence having at least 85% sequence identity to any one of SEQ ID NO: 4, SEQ ID NO: 5 or SEQ ID NO: 6, or a nucleic acid comprising a nucleotide sequence encoding a polypeptide comprising a nuclease sequence having at least 85% sequence identity to any one of SEQ ID NO: 4, SEQ ID NO: 5 or SEQ ID NO: 6.

[0098] Yet another embodiment according to the present disclosure is a polypeptide comprising a nuclease sequence having at least 90% sequence identity to any one of SEQ ID NO: 4, SEQ ID NO: 5 or SEQ ID NO: 6, or a nucleic acid comprising a nucleotide sequence encoding a polypeptide comprising a nuclease sequence having at least 90% sequence identity to any one of SEQ ID NO: 4, SEQ ID NO: 5 or SEQ ID NO: 6.

[0099] A further embodiment according to the present disclosure is a polypeptide comprising a nuclease sequence having at least 95% sequence identity to any one of SEQ ID NO: 4, SEQ ID NO: 5 or SEQ ID NO: 6, or a nucleic acid comprising a nucleotide sequence encoding a polypeptide comprising a nuclease sequence having at least 95% sequence identity to any one of SEQ ID NO: 4, SEQ ID NO: 5 or SEQ ID NO: 6.

[0100] A still further embodiment according to the present disclosure is a polypeptide comprising a nuclease sequence having at least 96% sequence identity to any one of SEQ ID NO: 4, SEQ ID NO: 5 or SEQ ID NO: 6, or a nucleic acid comprising a nucleotide sequence encoding a polypeptide comprising a nuclease sequence having at least 96% sequence identity to any one of SEQ ID NO: 4, SEQ ID NO: 5 or SEQ ID NO: 6.

[0101] An additional embodiment according to the present disclosure is a polypeptide comprising a nuclease sequence having at least 97% sequence identity to any one of SEQ ID NO: 4, SEQ ID NO: 5 or SEQ ID NO: 6, or a nucleic acid comprising a nucleotide sequence encoding a polypeptide comprising a nuclease sequence having at least 97% sequence identity to any one of SEQ ID NO: 4, SEQ ID NO: 5 or SEQ ID NO: 6.

[0102] Another additional embodiment according to the present disclosure is a polypeptide comprising a nuclease sequence having at least 98% sequence identity to any one of SEQ ID NO: 4, SEQ ID NO: 5 or SEQ ID NO: 6, or a nucleic acid comprising a nucleotide sequence encoding a polypeptide comprising a nuclease sequence having at least 98% sequence identity to any one of SEQ ID NO: 4, SEQ ID NO: 5 or SEQ ID NO: 6.

[0103] Yet another embodiment according to the present disclosure is a polypeptide comprising a nuclease sequence having at least 99% sequence identity to any one of SEQ ID NO: 4, SEQ ID NO: 5 or SEQ ID NO: 6, or a nucleic acid comprising a nucleotide sequence encoding a polypeptide comprising a nuclease sequence having at least 99% sequence identity to any one of SEQ ID NO: 4, SEQ ID NO: 5 or SEQ ID NO: 6.

[0104] Yet another embodiment according to the present disclosure is a polypeptide comprising a nuclease sequence having at least 99.5% sequence identity to any one of SEQ ID NO: 4, SEQ ID NO: 5, or SEQ ID NO: 6, or a nucleic acid comprising a nucleotide sequence encoding a polypeptide comprising a nuclease sequence having at least 99.5% sequence identity to any one of SEQ ID NO: 4, SEQ ID NO: 5, or SEQ ID NO: 6.

[0105] 6.4 Nuclear localization signals

[0106] The present disclosure provides engineered type V CRISPR-Cas endonucleases in the form of fusion proteins comprising a B-GEn nuclease sequence fused to one or more additional amino acid sequences, such as one or more nuclear localization signals (NLS).

[0107] In some embodiments, the fusion proteins of the present disclosure comprise means for targeting an engineered type V CRISPR-Cas endonuclease to the nucleus via an NLS.

[0108] In some embodiments, the fusion proteins of the present disclosure comprise one or more NLS sequences located at the N-terminus and / or C-terminus of the B-GEn.2 protein sequence.

[0109] In some embodiments, the fusion proteins of the present disclosure comprise one or more NLS sequences only at the N-terminus of the B-GEn.2 protein sequence. In other embodiments, the fusion proteins of the present disclosure comprise one or more NLS sequences only at the C-terminus of the B-GEn.2 protein sequence.

[0110] In some embodiments, the fusion proteins of the present disclosure comprise multiple identical NLS sequences. In other embodiments, the fusion proteins of the present disclosure comprise multiple different NLS sequences.

[0111] Non-limiting examples of nuclear localization signals are listed in Table 3.

[0112]

[0113]

[0114]

[0115]

[0116]

[0117] 6.5 Linker Sequence

[0118] The present disclosure provides engineered type V CRISPR-Cas endonucleases in the form of fusion proteins comprising a B-GEn nuclease sequence fused to one or more NLS sequences, optionally via a peptide linker.

[0119] In some embodiments, the B-GEn nuclease sequence is linked to an NLS sequence at its N- and / or C-terminus via a peptide linker. In other embodiments, the B-GEn nuclease sequence is linked to one or more NLS sequences, such as between two NLS sequences or between an NLS sequence and a B-GEn nuclease sequence, via a peptide linker that connects a pair of separate polypeptide sequences.

[0120] In some embodiments, the B-GEn polypeptides of the present disclosure comprise multiple NLS sequences connected by the same linker. In other embodiments, the B-GEn polypeptides of the present disclosure comprise multiple NLS sequences connected by different linkers.

[0121] Suitable peptide linkers for use with the B-GEn polypeptides of the present disclosure include those disclosed in Chen et al., 2013, Adv Drug Deliv Rev. 65(10): 1357-1369. Non-limiting examples of such linkers are described in Table 4 below.

[0122]

[0123] 6.6B-GEnV-type CRISPR-Cas system

[0124] The present disclosure provides engineered B-GEN V-type CRISPR-Cas endonuclease systems that incorporate the engineered B-GEn polypeptides of the present disclosure or nucleic acids encoding them.

[0125] In some embodiments, the engineered B-GEn Type V CRISPR-Cas endonuclease system comprises the following components:

[0126] (a) an engineered B-GEn polypeptide, such as described in Section 6.2, or a nucleic acid encoding such an engineered B-GEn polypeptide, such as described in Section 6.9,

[0127] (b) a heterologous guide RNA (gRNA), as described in Section 6.7, or a nucleic acid that allows for the in situ production of such a gRNA (e.g., a vector as described in Section 6.10), wherein the gRNA comprises:

[0128] i. an engineered DNA targeting segment consisting of RNA and capable of hybridizing to a target sequence on a nucleic acid locus,

[0129] ii. a tracr partner sequence consisting of RNA, and

[0130] iii. a tracr RNA sequence consisting of RNA,

[0131] wherein the tracr mate sequence hybridizes to the tracr sequence, and wherein (i), (ii) and (iii) are arranged in a 5' to 3' direction. The gRNA can be a single guide RNA (sgRNA). In an sgRNA, the tracr mate sequence and the tracr sequence are typically connected by a suitable loop sequence to form a stem-loop structure.

[0132] In certain embodiments, the engineered B-GEn Type V CRISPR-Cas endonuclease system comprises the following components:

[0133] (a) engineered B-GEn polypeptides as described in Section 6.2;

[0134] (b) a heterologous guide RNA (gRNA), as described in Section 6.7, comprising:

[0135] i. an engineered DNA targeting segment consisting of RNA and capable of hybridizing to a target sequence on a nucleic acid locus,

[0136] ii. a tracr partner sequence consisting of RNA, and

[0137] iii. a tracr RNA sequence consisting of RNA,

[0138] wherein the tracr mate sequence hybridizes to the tracr sequence, and wherein (i), (ii) and (iii) are arranged in a 5' to 3' direction. The gRNA can be a single guide RNA (sgRNA). In an sgRNA, the tracr mate sequence and the tracr sequence are typically connected by a suitable loop sequence to form a stem-loop structure.

[0139] In some embodiments, such engineered B-GEn Type V CRISPR-Cas endonucleases are delivered to target cells in a composition referred to as a ribonucleoprotein (RNP) complex as described in Section 6.8.

[0140] In certain embodiments, the engineered B-GEn Type V CRISPR-Cas endonuclease system comprises the following components:

[0141] (a) a nucleic acid, such as described in Section 6.9, encoding an engineered B-GEn polypeptide, such as described in Section 6.2, or

[0142] (b) a nucleic acid (e.g., a vector as described in Section 6.10) that allows for the production of a heterologous guide RNA (gRNA), e.g., as described in Section 6.7, wherein the gRNA comprises:

[0143] i. an engineered DNA targeting segment consisting of RNA and capable of hybridizing to a target sequence on a nucleic acid locus,

[0144] ii. a tracr partner sequence consisting of RNA, and

[0145] iii. a tracr RNA sequence consisting of RNA,

[0146] wherein the tracr mate sequence hybridizes to the tracr sequence, and wherein (i), (ii) and (iii) are arranged in a 5' to 3' direction. The gRNA can be a single guide RNA (sgRNA). In an sgRNA, the tracr mate sequence and the tracr sequence are typically connected by a suitable loop sequence to form a stem-loop structure.

[0147] Any reference to "RNA" or "guide RNA" encompasses RNA molecules comprising non-natural as well as natural nucleobases, eg, one or more nucleic acid modifications described in Section 6.9.3.

[0148] In some embodiments, the nucleic acid encoding the engineered B-GEn polypeptide and / or sgRNA comprises a suitable promoter for expression in a cell or in vitro environment.

[0149] In some embodiments, the nucleic acid encoding the engineered B-GEn polypeptide and / or sgRNA is in the form of a viral vector, such as described in Section 6.10.3.

[0150] 6.7 Guide RNA (gRNA) and Single Guide RNA (sgRNA)

[0151] In some embodiments, the systems, compositions, and methods described herein employ genome-targeting nucleic acids that can direct the activity of engineered B-GEn polypeptides to specific target sequences within target nucleic acids. In some embodiments, the genome-targeting nucleic acid is RNA. Genome-targeting RNA is referred to herein as "guide RNA" or "gRNA." The guide RNA has a spacer sequence that can hybridize at least with the target nucleic acid sequence of interest and the CRISPR repeat sequence (such CRISPR repeat sequence is also referred to as a "tracr partner sequence"). In type II systems, the gRNA also has a second RNA, referred to as a tracrRNA sequence. In type II guide RNA (gRNA), the CRISPR repeat sequence and the tracrRNA sequence hybridize with each other to form a duplex. In type V guide RNA (gRNA), crRNA forms a duplex. In both systems, the duplex binds to a site-specific polypeptide, such that the guide RNA and the site-directed polypeptide form a complex. The genome-targeting nucleic acid provides targeting specificity for the complex by virtue of its binding to the site-specific polypeptide. The genome-targeting nucleic acid thus directs the activity of the site-specific polypeptide.

[0152] In some embodiments, the genome-targeting nucleic acid is a bimolecular guide RNA. In some embodiments, the genome-targeting nucleic acid is a single-molecule guide RNA or a single guide RNA (sgRNA). Bimolecular guide RNA has two RNA chains. The first chain has an optional spacer extension sequence, a spacer sequence, and a minimal CRISPR repeat sequence in the 5' to 3' direction. The second chain has a minimal tracrRNA sequence (complementary to the minimal CRISPR repeat sequence), a 3' tracrRNA sequence, and an optional tracrRNA extension sequence. The single-molecule guide RNA (sgRNA) in the type II system has an optional spacer extension sequence, a spacer sequence, a minimal CRISPR repeat sequence, a single-molecule guide linker, a minimal tracrRNA sequence, a 3' tracrRNA sequence, and an optional tracrRNA extension sequence in the 5' to 3' direction. The optional tracrRNA extension may have elements that provide additional functions (e.g., stability) to the guide RNA. The single-molecule guide linker connects the minimal CRISPR repeat sequence and the minimal tracrRNA sequence to form a hairpin structure. The optional tracrRNA extension has one or more hairpins.

[0153] The single-molecule guide RNA (sgRNA) in the V-type system has a minimal CRISPR repeat sequence and a spacer sequence in the 5' to 3' direction. Alternatively, the single-molecule guide RNA (sgRNA) in the V-type system has an optional tracr extension sequence, a tracr RNA sequence, a single-molecule guide linker, a minimal CRISPR repeat sequence, a spacer sequence, and an optional spacer extension sequence in the 5' to 3' direction.

[0154] Alternatively, the single-molecule guide RNA (sgRNA) in the V-type system has an optional extension sequence, a minimal CRISPR repeat sequence, a spacer sequence, and an optional spacer extension sequence in the 5' to 3' direction.

[0155] In yet another embodiment, the sgRNA in the V-type system includes an optional extension sequence, an artificial nuclease binding RNA sequence and a spacer sequence in the 5' to 3' direction, and an optional spacer extension sequence.

[0156] Table 5 discloses sgRNAs that are particularly useful for the B-GEn.2 CRISPR Cas nuclease, and for potentially other V-type CRISPR Cas nucleases, according to the present disclosure.

[0157]

[0158] For example, exemplary genome targeting nucleic acids are described in WO2018002719. Typically, CRISPR repeat sequences include any sequence that has sufficient complementarity with a tracr sequence to facilitate one or more of the following: (1) excision of the DNA targeting segments flanking the CRISPR repeat sequence in a cell containing the corresponding tracr sequence; and (2) formation of a CRISPR complex at the target sequence, wherein the CRISPR complex comprises a CRISPR repeat sequence hybridized to the tracr sequence. Typically, the degree of complementarity refers to the optimal alignment of the CRISPR repeat sequence and the tracr sequence along the shorter of the two sequences. The optimal alignment can be determined by any suitable alignment algorithm, and secondary structure, such as self-complementarity within the tracr sequence or the CRISPR repeat sequence, can be further considered. In some embodiments, when the tracr sequence and the CRISPR repeat sequence are optimally aligned, the degree of complementarity between them along the shorter of the 30 nucleotides is about or greater than 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97.5%, 99% or more. In some embodiments, the tracr sequence is about or greater than 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50 or more nucleotides in length. In some embodiments, the tracr sequence and the CRISPR repeat sequence are contained in a single transcript such that hybridization between the two produces a transcript having a secondary structure, such as a hairpin. In some embodiments, the transcript or transcribed nucleic acid sequence has at least two or more hairpins.

[0159] Suitable tracr sequences for B-GEn.2 or B-GEn.1 in the CRISPR-Cas system are listed in Table 6. Alternatively, variants of these sequences can be used. Variants can include partial or truncated versions of these sequences and / or sequences with base modifications at one or more positions in these sequences. The corresponding RNA sequences are disclosed in SEQ ID NOs: 170 and 171, respectively.

[0160]

[0161] The spacer of the guide RNA includes a nucleotide sequence that is complementary to a sequence in the target DNA. In other words, the spacer of the guide RNA interacts with the target DNA in a sequence-specific manner through hybridization (e.g., base pairing). Therefore, the nucleotide sequence of the spacer can vary and determine the position in the target DNA where the guide RNA and target DNA interact. The DNA targeting segment of the guide RNA can be modified (e.g., by genetic engineering) to hybridize with any desired sequence in the target DNA.

[0162] In some embodiments, the length of the spacer is 10 nucleotides to 30 nucleotides. In some embodiments, the length of the spacer is 13 nucleotides to 25 nucleotides. In some embodiments, the length of the spacer is 15 nucleotides to 23 nucleotides. In some embodiments, the length of the spacer is 18 nucleotides to 22 nucleotides, such as 20 to 22 nucleotides.

[0163] In some embodiments, the percent complementarity between the DNA targeting sequence of the spacer and the protospacer of the target DNA is at least 60% (e.g., at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%) over 20-22 nucleotides.

[0164] In some embodiments, the protospacer is directly adjacent to a suitable PAM sequence at its 3' end, or such a PAM sequence is part of the DNA targeting sequence at its 3' portion.

[0165] Suitable PAM sequences are listed in Table 7, where the engineered DNA targeting segment is directly adjacent to a PAM sequence on the targeted DNA segment at its 3' end, or such a PAM sequence is part of the targeted DNA sequence at its 5' portion.

[0166]

[0167] The modification of guide RNA can be used to enhance the formation or stabilization of the CRISPR-Cas genome editing complex comprising guide RNA and Cas endonuclease (e.g., B-GEn.1 or B-GEn.2). The modification of guide RNA can also or alternatively be used to enhance the initiation, stability or kinetics of the interaction between the genome editing complex and the target sequence in the genome, for example, can be used to enhance on-target activity. The modification of guide RNA can also or alternatively be used to enhance specificity, such as, compared to the effect of other (off-target) sites, the relative ratio of the genome editing at the target site.

[0168] Modifications can also or alternatively be used to increase the stability of the guide RNA, such as by increasing its resistance to degradation by ribonucleases (RNases) present in the cell, thereby extending its half-life in the cell. Modifications that enhance the half-life of the guide RNA are particularly useful in embodiments where a Cas endonuclease, such as B-GEn.1 or B-GEn.1.2 or B-GEn.2, is introduced into the edited cell via an RNA to be translated in order to produce a B-GEn.1 or B-GEn.1.2 or B-GEn.2 endonuclease, because the increased half-life of the introduced guide RNA and the RNA encoding the endonuclease can be used to increase the time that the guide RNA and the encoded Cas endonuclease coexist in the cell.

[0169] 6.7.1 Additional Sequences

[0170] In certain embodiments, the guide RNA includes at least one additional segment at the 5' end or the 3' end. For example, suitable additional segments may include a 5' cap (e.g., a 7-methylguanylate cap (m7G)); a 3' polyadenylation tail (e.g., a 3' poly (A) tail); a riboswitch sequence (e.g., a sequence that allows for regulatory stability and / or regulatory accessibility by proteins and protein complexes); a sequence that forms a dsRNA duplex (e.g., a hairpin); a sequence that targets RNA to a subcellular location (e.g., a nucleus, mitochondria, chloroplasts, etc.); a modification or sequence that provides tracking functionality (e.g., direct conjugation to a fluorescent molecule, conjugation to a portion that promotes fluorescence detection, a sequence that allows fluorescence detection, etc.); a modification or sequence that provides a protein binding site (e.g., a protein that acts on DNA, including transcriptional activators, transcriptional repressors, DNA methyltransferases, DNA demethylases, histone acetyltransferases, histone deacetylases, etc.); a modification or sequence that provides increased, decreased, and / or controllable stability; and combinations thereof.

[0171] 6.7.1.1 Stability Control Sequence

[0172] Stability control sequences affect the stability of RNA (such as guide RNA).The limiting examples of suitable stability control sequences are transcription terminators (such as transcription termination sequences).The total length of the transcription terminator segments of guide RNA can be from 10 nucleotides to 100 nucleotides, such as from 10 nucleotides (nt) to 20nt, from 20nt to 30nt, from 30nt to 40nt, from 40nt to 50nt, from 50nt to 60nt, from 60nt to 70nt, from 70nt to 80nt, from 80nt to 90nt, or from 90nt to 100nt. For example, the length of the transcription terminator segment can be from 15 nucleotides (nt) to 80nt, from 15nt to 50nt, from 15nt to 40nt, from 15nt to 30nt or from 15nt to 25nt.

[0173] In some embodiments, the transcription termination sequence is a sequence that functions in eukaryotic cells. In some embodiments, the transcription termination sequence is a sequence that functions in prokaryotic cells.

[0174] Nucleotide sequences that can be included in a stability control sequence (such as a transcription termination segment, or any segment of a guide RNA to provide enhanced stability) include, for example, a Rho-independent trp termination site.

[0175] 6.8 Ribonucleoprotein (RNP) Complex

[0176] In some embodiments, the engineered B-GEn Type V CRISPR-Cas endonucleases are delivered in a composition called a ribonucleoprotein or RNP complex. The RNP complex is assembled by combining a Cas endonuclease (e.g., an engineered B-GEn endonuclease) with a ribonucleic acid (e.g., a guide RNA (gRNA)).

[0177] In some embodiments, the ribonucleoprotein complex comprises an engineered B-GEn endonuclease complexed with a suitable ribonucleic acid, as described in Section 6.3. In some embodiments, the ribonucleic acid is a gRNA or sgRNA, which is further described in Section 6.7. In some embodiments, the RNP complex comprises an engineered B-GEn polypeptide and an sgRNA listed in Table 5 or another suitable sgRNA.

[0178] One of the most common techniques for delivering RNPs is electroporation, which creates holes in the cell membrane to allow RNPs to enter the cytoplasm. Furthermore, electroporation can be combined with cell type-specific reagents in a technique called nucleofection, which creates holes in the nuclear membrane to allow DNA templates to enter. In some embodiments, engineered B-Gen V-type CRISPR-Cas endonucleases in RNP complexes are delivered to target cells via nucleofection.

[0179] 6.9 Nucleic Acids

[0180] The present disclosure provides nucleic acids (e.g., DNA or RNA) encoding B-GEn type V CRISPR-Cas proteins (e.g., engineered B-GEn polypeptides), nucleic acids encoding gRNAs or sgRNAs of the present disclosure, nucleic acids encoding both engineered B-GEn polypeptides and gRNAs or sgRNAs, and multiple nucleic acids, e.g., comprising nucleic acids encoding engineered B-GEn polypeptides and gRNAs or sgRNAs.

[0181] Nucleic acids encoding engineered B-GEn polypeptides can be codon-optimized, e.g., wherein at least one uncommon codon or less common codon has been replaced by a codon commonly found in the host cell or target cell. For example, codon-optimized nucleic acids can direct the synthesis of optimized messenger mRNA, e.g., optimized for expression in a mammalian expression system.

[0182] 6.9.1B-GEn Coding Sequence

[0183] In some embodiments, the nucleic acids described herein comprise one or more modifications as further described herein and known in the art, such as those useful for enhancing activity, stability, or specificity, altering delivery, reducing the innate immune response in host cells, further reducing protein size, or for other enhancements. In some embodiments, such modifications will result in engineered B-GEn polypeptides having a nuclease sequence component that has at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or 100% amino acid sequence identity to the sequence of SEQ ID NO: 4, 5, or 6.

[0184] 6.9.2 Codon Optimization

[0185] In certain embodiments, the modified nucleic acids used in the CRISPR-B-GEn.1, or B-GEn.1.2, or B-GEn.2 systems described herein, wherein the guide RNA and / or the DNA or RNA comprising the nucleic acid sequence encoding the engineered B-GEn polypeptide can be modified, as described below. Such modified nucleic acids can be used in the CRISPR-B-GEn.1, or B-GEn.1.2, or B-GEn.2 systems to edit any one or more genomic loci. In some embodiments, such modifications in the nucleic acids of the present disclosure are achieved by codon optimization, such as codon optimization based on the specific host cell in which the encoded polypeptide is expressed. It will be understood by those skilled in the art that any nucleotide sequence and / or recombinant nucleic acid of the present disclosure can be codon optimized to express in any species of interest. Codon optimization is well known in the art and relates to the use of species-specific codon usage tables to modify the codon usage bias of nucleotide sequences. The codon usage table is generated based on sequence analysis of the highest expressed genes of the species of interest. In a non-limiting example, when the nucleotide sequence is to be expressed in the nucleus, a codon usage table is generated based on sequence analysis of highly expressed nuclear genes of the species of interest. Modifications to the nucleotide sequence are determined by comparing the species-specific codon usage table with the codons present in the native nucleic acid sequence.

[0186] In some embodiments, the engineered B-GEn polypeptides as described herein are expressed from codon-optimized nucleic acid sequences. For example, if the intended host cell or target cell is a human cell, the nucleic acid sequence of the engineered B-GEn polypeptides comprising the amino acid sequence of B-GEn.1, B-GEn.1.2, or B-GEn.2 (or B-GEn.1, B-GEn.1.2, or B-GEn.2 variants, such as enzymatic inactivation variants) encoded by human codon optimization will be suitable. As another non-limiting example, if the intended host cell or target cell is a mouse cell, the nucleic acid sequence of the engineered B-GEn polypeptides comprising the amino acid sequence of B-GEn.1, B-GEn.1.2, or B-GEn.2 (or B-GEn.1, B-GEn.1.2, or B-GEn.2 variants, such as enzymatic inactivation variants) encoded by mouse codon optimization will be suitable.

[0187] Strategies and methods for codon optimization are known in the art and have been described for various systems, including but not limited to yeast (Outchkourov et al., Protein Expr Purif, 24(1):18-24 (2002)) and E. coli (Feng et al., Biochemistry, 39(50):15399-15409 (2000)). In some embodiments, codon optimization is performed using Expression optimization technology (ATUM) and the expression optimization algorithm recommended by the manufacturer are used. In some embodiments, the nucleic acids of the present disclosure are codon optimized to increase expression in human cells. In some embodiments, the nucleic acids of the present disclosure are codon optimized to increase expression in Escherichia coli cells. In some embodiments, the nucleic acids of the present disclosure are codon optimized to increase expression in insect cells. In some embodiments, the nucleic acids of the present disclosure are codon optimized to increase expression in Sf9 insect cells. In some embodiments, the expression optimization algorithm used in the codon optimization process is defined as avoiding putative poly A signals (e.g., AATAAA and ATTAAA) and long (greater than 4) A segments that may cause polymerase slippage.

[0188] As is well known in the art, codon optimization of a nucleotide sequence results in a nucleotide sequence that is less than 100% identical (e.g., less than 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99%) to a native nucleotide sequence, but the encoded polypeptide still functions identically to the polypeptide encoded by the original native nucleotide sequence. Thus, in representative embodiments of the present disclosure, the nucleotide sequences and / or recombinant nucleic acids of the present disclosure can be codon-optimized for expression in a particular species of interest.

[0189] In some embodiments, the codon-optimized nucleic acid sequence has at least 90%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.2%, 99.5%, 99.8%, 99.9% or 100% sequence identity to SEQ ID NO: 4. In some embodiments, the nucleic acids of the present disclosure are codon-optimized to increase expression of the encoded engineered B-GEn polypeptide in target cells or host cells. In some embodiments, the nucleic acids of the present disclosure are codon-optimized to increase expression in human cells. Generally, the nucleic acids of the present disclosure are codon-optimized to increase expression in any human cell. In some embodiments, the nucleic acids of the present disclosure are codon-optimized to increase expression in Escherichia coli cells. In some embodiments, the nucleic acids of the present disclosure are codon-optimized to increase expression in insect cells. Generally, the nucleic acids of the present disclosure are codon-optimized to increase expression in any insect cell. In some embodiments, the nucleic acids of the disclosure are codon-optimized for increased expression in the Sf9 insect cell expression system.

[0190] The polyadenylation signal may also be selected to optimize expression in the intended host.

[0191] 6.9.3 Nucleic Acid Modification

[0192] In some embodiments, a nucleic acid (e.g., a guide RNA, a nucleic acid comprising a nucleotide sequence encoding a guide RNA; a nucleic acid encoding a site-specific modifying enzyme, such as an engineered B-GEn polypeptide of the present disclosure; etc.) comprises modifications or sequences that provide additional desired properties (e.g., modified or regulated stability; subcellular targeting; tracking, such as a fluorescent tag; a binding site for a protein or protein complex; etc.). Non-limiting examples include: a 5' cap (e.g., a 7-methylguanylate cap (m7G)); a 3' polyadenylation tail (e.g., a 3' poly (A) tail); a riboswitch sequence (e.g., one that allows for regulated stability and / or regulated accessibility of proteins and / or protein complexes); a stability control sequence; a sequence that forms a dsRNA duplex (e.g., a hairpin); a modification or sequence that targets RNA to a subcellular location (e.g., the nucleus, mitochondria, chloroplasts, etc.); a modification or sequence that provides tracking (e.g., direct conjugation to a fluorescent molecule, conjugation to a moiety that facilitates fluorescence detection, a sequence that allows fluorescence detection, etc.); a modification or sequence that provides a binding site for a protein (e.g., a protein that acts on DNA, including transcription activators, transcription repressors, DNA methyltransferases, DNA demethylases, histone acetyltransferases, histone deacetylases, etc.); and combinations thereof.

[0193] In some embodiments, the guide RNA includes an additional segment at the 5' or 3' end that provides any of the above features. For example, a suitable third segment may include a 5' cap (such as a 7-methylguanosine cap (m7G); a 3' polyadenylation tail (such as a 3' poly (A) tail); a riboswitch sequence (such as a stability that allows regulation of proteins and protein complexes and / or regulated accessibility); a stability control sequence; a sequence that forms a dsRNA duplex (such as a hairpin); a sequence that targets RNA to a subcellular location (such as a nucleus, mitochondria, chloroplasts, etc.); a modification or sequence that provides tracking (such as direct conjugation to a fluorescent molecule, conjugation to a portion that facilitates fluorescence detection, a sequence that allows fluorescence detection, etc.); a modification or sequence that provides a binding site for a protein (such as a protein that acts on DNA, including transcription activators, transcription repressors, DNA methyltransferases, DNA demethylases, histone acetyltransferases, histone deacetylases, etc.); and combinations thereof.

[0194] Modifications can also or alternatively be used to reduce the likelihood or extent that the RNA introduced into the cell will elicit an innate immune response. Such responses, as described below and in the art, have been well characterized in the context of RNA interference (RNAi), including small interfering RNA (siRNA), which is often associated with a shortened RNA half-life and / or the stimulation of cytokines or other factors associated with an immune response.

[0195] The RNA of the coding engineered B-GEn polypeptide that is introduced into the cell can also be subjected to one or more types of modification, including, but not limited to, the modification that strengthens RNA stability (for example, by reducing the degradation of the RNase present in the cell), the modification that strengthens the translation of products therefrom (such as, endonuclease), and / or the modification that reduces the possibility that the RNA that is introduced into the cell causes innate immune response or degree. Combinations such as those described above and other modifications can also be used. In the case of engineered B-GEn polypeptide, for example, one or more types of modification (including those of the above examples) can be carried out for guide RNA, and / or one or more types of modification (including those of the above examples) can be carried out to the RNA of the coding engineered B-GEn polypeptide.

[0196] By way of illustration, the guide RNA or other smaller RNA used in the CRISPR-B-GEn system can be easily synthesized by chemical means, so that many modifications can be easily incorporated, as described below and in the art. Although chemical synthesis procedures continue to expand, due to the significant increase in nucleic acid length, exceeding 100 or so nucleotides, it often becomes more challenging to purify these RNAs using procedures such as high performance liquid chromatography (HPLC, which avoids the use of gels such as PAGE). One method for producing longer chemically modified RNAs is to produce two or more molecules linked together. Longer RNAs, such as RNA encoding B-GEn.1, or B-GEn.1.2, or B-GEn.2 nucleases, are more easily produced enzymatically. Although there are fewer types of modifications that can be used for enzymatically produced RNAs, there are still some modifications that can be used, such as enhancing stability, reducing the likelihood or extent of an innate immune response, and / or enhancing other properties, as further described below and in the art; and new modifications are also constantly being developed.

[0197] By way of illustration of various types of modifications, particularly those frequently used in smaller chemically synthesized RNAs, modifications can include one or more nucleotides modified at the 2' position of the sugar, in some embodiments 2'-O-alkyl, 2'-O-alkyl-O-alkyl, or 2'-fluoro modified nucleotides. In some embodiments, RNA modifications include 2'-fluoro, 2'-amino, and 2'-O-methyl modifications on pyrimidine ribose sugars, base residues, or anti-bases at the 3' end of the RNA. Such modifications are often incorporated into oligonucleotides, and these oligonucleotides have been shown to have higher Tm (e.g., higher target binding affinity) than 2'-deoxy oligonucleotides for a given target.

[0198] Many nucleotide and nucleoside modifications have been shown to render the oligonucleotides into which they are incorporated more resistant to nuclease digestion than native oligonucleotides; these modified oligonucleotides survive intact longer than unmodified oligonucleotides. Specific examples of modified oligonucleotides include those comprising modified backbones, e.g., phosphorothioates, phosphotriesters, methylphosphonates, short-chain alkyl or cycloalkyl sugar linkages, or short-chain heteroatom or heterocyclic sugar linkages. Some oligonucleotides are oligonucleotides with phosphorothioate backbones and oligonucleotides with heteroatom backbones, particularly CH2-NH-O-CH2, CH,-N(CH3)-O-CH2 (referred to as methylene(methylimino) or MMI backbones), CH2-ON(CH3)-CH2, CH2-N(CH3)-N(CH3)-CH2 and ON(CH3)-CH2-CH2 backbones; amide backbones (see De Mesmaeker et al., 1995, Ace. Chem. Res., 28:366-374); morpholine backbone structures (see Summerton and Weller, U.S. Pat. No. 5,034,506); peptide nucleic acid (PNA) backbones (in which the phosphodiester backbone of the oligonucleotide is replaced by a polyamide backbone and the nucleotides are linked directly or indirectly to the aza nitrogen atoms of the polyamide backbone, see Nielsen et al., 1991, Science 254:1497). Phosphorus-containing linkages include, but are not limited to, phosphorothioates, chiral phosphorothioates, phosphorodithioates, phosphotriesters, aminoalkylphosphotriesters, methyl and other alkyl phosphonates, including 3' alkylene phosphonates and chiral phosphonates, phosphinates, phosphoramidates, including 3'-aminophosphoramidates and aminoalkylphosphoramidates, thionylphosphoramidates, thionylalkylphosphonates, thionylalkylphosphotriesters, and borophosphates having normal 3'-5' linkages, 2'-5' linked analogs of these, and analogs with reversed polarity, wherein adjacent pairs of nucleoside units are linked 3'-5' to 5'-3' or 2'-5' to 5'-2'; see U.S. Pat. Nos. 3,687,808; 4,464,808. and 5,625,050.

[0199] Morpholinyl-based oligomeric compounds are described in Braasch and Corey, Biochemistry, 41(14):4503-4510 (2002); Genesis, Vol. 30, No. 3, (2001); Heasman, Dev. Biol., 243:209-214 (2002); Naseviciu et al., Nat. Genet., 26:216-220 (2000); Lacenra et al., Proc. Nat / . Acad. Sci., 97:9591-9596 (2000); and U.S. Pat. No. 5,034,506, issued July 23, 1991. Cyclohexenyl nucleic acid oligonucleotide mimetics are described in Wang et al., J. Am. Chem. Soc., 122:8595-8602 (2000).

[0200] Modified oligonucleotide backbones that do not include phosphorus atoms have backbones formed from short-chain alkyl or cycloalkyl nucleoside bonds, mixed heteroatoms and alkyl or cycloalkyl nucleoside bonds, or one or more short-chain heteroatoms or heterocyclic nucleoside bonds. These include those with morpholine bonds (partially formed from the sugar portion of the nucleoside); siloxane backbones; sulfide, sulfoxide, and sulfone backbones; formacetyl and thioformacetyl backbones; methyleneformacetyl and thioformacetyl backbones; olefin-containing backbones; aminosulfonate backbones; methyleneimino and methylenehydrazine backbones; sulfonate and sulfonamide backbones; amide backbones; and other backbones with mixed N, O, S, and CH2 component moieties; see U.S. Patent Nos. 5,034,506; 5,166,315; 5,185,444; 5, 214,134; 5,216,141; 5,235,033; 5,264,562; 5,264,564; 5,405,938; 5,434,257; 5,466,677; 5,470,967; 5,489,677; 5,541,307; 5,561,225; 5,596,08 6; 5,602,240; 5,610,289; 5,602,240; 5,608,046; 5,610,289; 5,618,704; 5,623,070; 5,663,312; 5,633,360; 5,677,437; and 5,677,439, each of which is incorporated herein by reference.

[0201] One or more substituted sugar groups may also include, for example, one of the following at the 2' position: OH, SH, SCH3, F, OCN, OCH3, OCH3O(CH2)nCH3, O(CH2)nNH2, or O(CH2)nCH3, where n is 1 to 10; C1-C10 lower alkyl, alkoxyalkoxy, substituted lower alkyl, alkaryl, or aralkyl; Cl; Br; CN; CF3; OCF3; O-, S-, or N-alkyl; O-, S-, or N-alkenyl; SOCH3; SO2CH3; ONO2; N O2; N3; NH2; heterocycloalkyl; heterocycloalkaryl; aminoalkylamino; polyalkylamino; substituted silyl; RNA cleavage group; reporter group; intercalator; a group that improves the pharmacokinetic properties of an oligonucleotide; or a group for improving the pharmacodynamic properties of an oligonucleotide and other substituents with similar properties. In some embodiments, the modification includes 2'-methoxyethoxy (2'-O-CH2CH2OCH3, also known as 2'-O-(2-methoxyethyl)) (Martinet a / , Helv. Chim. Acta, 1995, 78, 486). Other modifications include 2'-methoxy (2'-O-CH3), 2'-propoxy (2'-OCH2CH2CH3) and 2'-fluoro (2'-F). Similar modifications can also be made at other positions in the oligonucleotide, particularly at the 3' position of the sugar on the 3' terminal nucleotide and the 5' position of the 5' terminal nucleotide. Oligonucleotides can also have sugar mimics, such as cyclobutyl groups that replace the pentofuranosyl group. In some embodiments, the sugar and internucleoside bonds of the nucleotide units, such as the backbone, are replaced by new groups. The retained base units are used to hybridize with suitable nucleic acid target compounds. One such oligomeric compound (an oligonucleotide mimetic that has been shown to have excellent hybridization properties) is called peptide nucleic acid (PNA). In PNA compounds, the sugar backbone of the oligonucleotide is replaced by an amide-containing backbone, such as an aminoethylglycine backbone. The nucleobases are retained and directly or indirectly attached to the nitrogen atoms of the amide portion of the backbone. Representative U.S. patents that teach the preparation of PNA compounds include, but are not limited to, U.S. Patent Nos. 5,539,082; 5,714,331; and 5,719,262. Further teachings of PNA compounds are described in Nielsen et al., Science, 254: 1497-1500 (1991).

[0202] The guide RNA may also additionally or alternatively comprise modifications or substitutions of nucleobases (often referred to in the art as "bases"). As used herein, "unmodified" or "natural" nucleobases include adenine (A), guanine (G), thymine (T), cytosine (C), and uracil (U). Modified nucleobases include nucleobases that are only rarely or transiently present in natural nucleic acids, such as hypoxanthine, 6-methyladenine, 5-Me pyrimidines, particularly 5-methylcytosine (also known as 5-methyl-2'deoxycytosine, commonly referred to in the art as 5-Me-C), 5-hydroxymethylcytosine (5-hydroxymethylcytosine, HMC), glycosyl HMC, and gentiobiosyl HMC. HMC), and synthetic nucleobases such as 2-aminoadenine, 2-(methylamino)adenine, 2-(imidazolidinyl)adenine, 2-(aminoalkylamino)adenine or other heterosubstituted alkyladenines, 2-thiouracil, 2-thiothymine, 5-bromopyrimidine, 5-hydroxymethyluracil, 8-azaguanine, 7-deazaguanine, N6 (6-aminohexyl)adenine and 2,6-diaminopurine. Kornberg, A, DNA Replication, WH Freeman & Co., San Francisco, pp75-77 (1980); Gebeyehu et al., Nucl. Acids Res. 15:4513 (1997). "Universal" bases known in the art, such as inosine, may also be included. 5-Me-C substitutions have been shown to increase nucleic acid duplex stability by 0.6-1.2 degrees Celsius (Sanghvi, YS, in Crooke, ST and Lebleu, B., eds., Antisense Research and Applications, CRC Press, Boca Raton, 1993, pp. 276-278) and are an embodiment of base substitution.

[0203] Modified nucleobases include other synthetic and natural nucleobases such as 5-methylcytosine (5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5-halouracil and cytosine, 5-propynyluracil and cytosine, 6-azouracil, cytosine and thymine, 5- Uracil (pseudouracil), 4-thiouracil, 8-halogenated, 8-amino, 8-thiol, 8-sulfanyl, 8-hydroxy and other α-substituted adenines and guanines, 5-halogenated, especially 5-bromo, 5-trifluoromethyl and other 5-substituted uracils and cytosines, 7-methylquanine and 7-methyladenine, 8-azaguanine and 8-azaadenine, 7-deazaguanine and 7-deazaadenine and 3-deazaguanine and 3-deazaadenine.

[0204] Other useful nucleobases include those disclosed in U.S. Pat. No. 3,687,808, those disclosed in "The Concise Encyclopedia of Polymer Science And Engineering", pp. 858-859, Kroschwitz, Jl, ed. John Wiley & Sons, 1990, those disclosed in Englisch et al., Angewandte Chemie, International Edition, 1991, 30, p. 613, and those disclosed in Sanghvi, YS, Chapter 15, Antisense Research and Applications, pp. 289-302, Crooke, ST and Lebleu, B.ea., CRC Press, 1993. Certain of these nucleobases are particularly useful for increasing the binding affinity of the oligomeric compounds of the present disclosure. These include 5-substituted pyrimidines, 6-azapyrimidines, and N-2, N-6, and -O-6 substituted purines, including 2-aminopropyladenine, 5-propynyluracil, and 5-propynylcytosine. Substitution with 5-methylcytosine has been shown to increase nucleic acid duplex stability by 0.6-1.20°C (Sanghvi, YS, Crooke, ST, and Lebleu, B., eds., "Antisense Research and Applications", CRC Press, Boca Raton, 1993, pp. 276-278), and is an embodiment of base substitution, even more particularly when combined with a 2'-O-methoxyethyl sugar modification. Modified nucleobases are described in U.S. Patent Nos. 3,687,808, and 4,845,205; 5,130,302; 5,134,066; 5,175,273; 5,367,066; 5,432,272; 5,457,187; 5,459,255; 5,484,908; 5,502,177; 5,525,711; 5,552,540; 5,587,469; 5,596,091; 5,614,617; 5,681,941; 5,750,692; 5,763,588; 5,830,653; 6,005,096; and U.S. Patent Application Publication 20030158403.

[0205] All positions in a given oligonucleotide need not be uniformly modified, and, in fact, more than one of the above modifications may be incorporated in a single oligonucleotide, or even within a single nucleoside within an oligonucleotide.

[0206] In some embodiments, the guide RNA and / or mRNA encoding the endonuclease (e.g., B-GEn.1, or B-GEn.1.2, or B-GEn.2 of the present disclosure) is capped using any of the current capping methods, such as mCAP, ARCA, or enzymatic capping methods, to create a viable mRNA construct that maintains biological activity and avoids self / non-self intracellular responses. In some embodiments, the guide RNA and / or mRNA encoding the endonuclease (e.g., B-GEn.1, or B-GEn.1.2, or B-GEn.2 of the present disclosure) is capped using CleaOnCap TM Capping was performed using the (TriLink) co-transcriptional capping method.

[0207] In some embodiments, the guide RNA and / or mRNA encoding the endonuclease of the present disclosure includes one or more modifications selected from pseudouracil, N1-methylpseudouracil, and 5-methoxyuracil. In some embodiments, one or more N1-methylpseudouracils are incorporated into the guide RNA and / or mRNA encoding the endonuclease of the present disclosure to provide enhanced RNA stability and / or protein expression and reduced immunogenicity in animal cells, e.g., mammalian cells (e.g., humans and mice). In some embodiments, the N1-methylpseudouracil modification is incorporated in combination with one or more 5-methylcytosines.

[0208] In some embodiments, the guide RNA and / or mRNA (or DNA) encoding the endonuclease (e.g., B-GEn.1, or B-GEn.1.2, or B-GEn.2) is chemically linked to one or more moieties or conjugates that enhance the activity, cellular distribution, or cellular uptake of the oligonucleotide. These moieties include, but are not limited to, lipid moieties, such as a cholesterol moiety (Letsinger et al., 1989, Proc. Nat / . Acad. Sci. USA 86:6553-6556); cholic acid (Manoharan et al., 1994, Bioorg. Med. Chem. Let. 4:1053-1060); thioethers, such as hexyl-S-tritylthiol (Manoharan et al., 1992, Ann. NY Acad. Sci. 660:306-309 and Manoharan et al., 1993, Bioorg. Med. Chem. Let. 3:2765-2770); thiocholesterol (Oberhauser et al., 1992, Nucl. Acids Res. 20:533-538); aliphatic chains, such as, dodecanediol or undecyl residues (Kabanov et al., 1990, FEBS Lett., 259:327-330 and Svinarchuk et al., 1993, Biochimie, 75:49-54); phospholipids, such as dicetyl-rac-glycerol or triethylammonium 1,2-bis-O-cityl-rac-glycerol-3-H-phosphonate (Manoharan et al., 1995, Tetrahedron Lett. 36:3651-3654 and Shea et al., 1990, Nucl. Acids Res. 18:3777-3783); polyamine or polyethylene glycol chains (Mancharan et al., 1995, Nucleosides & Nucleotides 14:969-973); adamantane acetic acid (Manoharan et al., 1995, Tetrahedron Lett. 36:3651-3654); a palmityl group (Mishra et al., 1995, Biochim. Biophys. Acta 1264:229-237); or an octadecylamine or hexylamine-carbonyl-t-hydroxycholesterol group (Crooke et al., 1996, J. Pharmacol. Exp. Ther., 277:923-937).See also U.S. Patent Nos. 4,828,979; 4,948,882; 5,218,105; 5,525,465; 5,541,313; 5,545,730; 5,552,538; 5,578,717; 5,580,731; 5,580,731; 5,591,584; 5,109,124; 5,118,802; 5,138,045; 5 ,414,077;5,486,603;5,512,439;5,578,718;5,608,046;4,587,044;4,605,735;4,667,025;4,762,779;4,789,737;4,824,941;4,835,263;4,876,335;4,904,582;4,958,013;5 ,082,830;5,112,963;5,214,136;5,082,830;5,112,963;5,214,136;5,245,022;5,254,469;5,258,506;5,262,536;5,272,250;5,292,873;5,317,098;5,371,241,5,391,723; 5,416,203,5,451,463; 5,510,475; 5,512,667; 5,514,785; 5,565,552; 5,567,810; 5,574,142; 5,585,481; 5,587,371; 5,595,726; 5,597,696; 5,599,923; 5,599,928 and 5,688,941.

[0209] Sugars and other groups can be used to target proteins and complexes containing nucleotides (such as cationic polysomes and liposomes) to specific sites. For example, hepatocyte-directed translocation can be mediated by the asialoglycoprotein receptor (ASGPR); see, e.g., Hu et al., 2014, Protein Pept Lett. 21(10):1025-30. Other systems known in the art and continuously being developed can be used to target the biomolecules and / or their complexes used in this application to specific target cells of interest.

[0210] These targeting groups or conjugates can include conjugated groups covalently linked to functional groups such as primary or secondary hydroxyl groups. Suitable conjugated groups include intercalators, reporter molecules, polyamines, polyamides, polyethylene glycols, polyethers, groups that enhance oligomer pharmacodynamic properties, and groups that enhance oligomer pharmacokinetic properties. Typical conjugated groups include cholesterol, lipids, phospholipids, biotin, phenazine, folic acid, phenanthridine, anthraquinone, acridine, fluorescein, rhodamine, coumarin, and dyes. Groups that can enhance pharmacodynamic properties include groups that improve uptake, enhance degradation resistance, and / or strengthen specific hybridization with target nucleic acid sequences. Groups that can enhance pharmacokinetic properties include groups that improve the uptake, distribution, metabolism, or secretion of the compounds of the present disclosure. Representative conjugated groups are disclosed in International Patent Application No. PCT / US92 / 09196, filed October 23, 1992, and U.S. Patent No. 6,287,860, which are incorporated herein by reference. Conjugation groups include, but are not limited to, lipid groups such as cholesterol groups, cholic acid, thioethers such as hexyl-5-tritylthiol, thiocholesterol, aliphatic chains such as dodecandiol or undecyl residues, phospholipids such as dihexadecyl-rac-glycerol or triethylammonium 1,2-bis-O-hexadecyl-rac-glycerol-3-H-phosphonate, polyamines or polyethylene glycol chains, or adamantaneacetic acid, palmityl or octadecylamine or hexylamino-carbonyl-hydroxycholesterol groups.See, e.g., U.S. Patent Nos. 4,828,979; 4,948,882; 5,218,105; 5,525,465; 5,541,313; 5,545,730; 5,552,538; 5,578,717; 5,580,731; 5,580,731; 5,591,584; 5,109,124; 5,118,802; 5,138,045; 5,414,077; 5,486,603; 5,512,439; 5,578,718; 5,608,046; 4,587,044; 4,605,735; 4,667,025; 4,762,779; 4,789,737; 4,824,941; 4,835,263; 4,876,335; 4,904,582; 4,958,013; 5,082,830; 5,112,963; 5,214,136; 5,082,830; 5,112,963; 5,214,136; 5,245,022; 5,254,469; 5,258,506; 5,262,536; 5,272,250; 5,292,873; 5,317,098; 5,371,241, 5,391,723; 5,416,203,5,451,463; 5,510,475; 5,512,667; 5,514,785; 5,565,552; 5,567,810; 5,574,142; 5,585,481; 5,587,371; 5,595,726; 5,597,696; 5,599,923; 5,599,928 and 5,688,941.

[0211] Longer nucleic acids that are less suitable for chemical synthesis and are typically produced by enzymatic synthesis can also be modified by various methods. Such modifications can include, for example, the introduction of certain nucleotide analogs, the incorporation of specific sequences or other moieties at the 5' or 3' end of the molecule, and other modifications. By way of example, the mRNA encoding B-GEn.1, or B-GEn.1.2, or B-GEn.2 is approximately 4 kb in length and can be synthesized by in vitro transcription. Modification of mRNA can be used, for example, to increase its translation or stability (e.g., by increasing its resistance to degradation in the cell), or to reduce the tendency of the RNA to elicit an innate immune response, which is commonly observed in cells following the introduction of exogenous RNA, particularly longer RNA, such as RNA encoding B-GEn.1 or B-GEn.2.

[0212] Many such modifications have been described in the art, such as poly A tails, 5' cap analogs (such as AntiReverse Cap Analog (ARCA) or m7G(5')ppp(5')G (mCAP)), modified 5' or 3' untranslated regions (UTRs), use of modified bases (such as pseudo-UTP, 2-thio-UTP, 5-methylcytidine-5'-triphosphate (5-methyl-CTP) or N6-methyl-ATP), or treatment with phosphatase to remove the 5' terminal phosphate. These and other modification methods are known in the art, and new RNA modification methods are constantly being developed.

[0213] There are many commercial suppliers of modified RNA, including, for example, TriLink Biotech, Axolabs, Bio-Synthesis Inc., Dharmacon, and many others. As described by TriLink, for example, 5-methyl-CTP can be used to impart desired properties, such as increased nuclease stability, increased translation, or reduced interaction of innate immune receptors with in vitro transcribed RNA. 5'-Methylcytidine-5'-triphosphate (5-methyl-CTP), N6-methyl-ATP, as well as pseudo-UTP and 2-thio-UTP have also been shown to reduce innate immune stimulation in culture and in vivo while increasing translational capacity, as described in the publications by Konmann et al. and Warren et al. mentioned below.

[0214] It has been shown that chemical modifications can be used to deliver mRNA in vivo to achieve improved therapeutic effects; see, for example, Kormann et al., Nature Biotechnology 29, 154-157 (2011). Such modifications can be used, for example, to improve the stability of RNA molecules and / or reduce their immunogenicity. By using chemical modifications such as pseudo-U, N6-methyl-A, 2-thio-U and 5-methyl-C, it has been found that as long as 2-thio-U and 5-methyl-C replace a quarter of the uridine and cytidine residues, respectively, the recognition of mRNA mediated by toll-like receptors (TLRs) in mice can be significantly reduced. Therefore, by reducing the activation of the innate immune system, these modifications can be used to effectively improve the stability and lifespan of mRNA in vivo; see, for example, Konmann et al., supra.

[0215] It has also been shown that repeated administration of synthetic messenger RNAs incorporating modifications designed to bypass the innate antiviral response can reprogram differentiated human cells to pluripotent cells. See, e.g., Warren et al., Cell Stem Cell, 7(5):618-30 (2010). Such modified mRNAs as primary reprogramming proteins may be an effective means of reprogramming a variety of human cell types. Such cells are known as induced pluripotent stem cells (iPSCs). It has also been found that enzymatically synthesized RNAs incorporating 5-methyl-CTP, pseudo-UTP, and anti-reverse cap analogs (ARCA) can be used to effectively evade the antiviral response of cells; see, e.g., Warren et al., supra. Other modifications of nucleic acids described in the art include, for example, the use of poly A tails, the addition of 5' cap analogs (e.g., m7G(5')ppp(5')G (mCAP)), modifications of the 5' or 3' untranslated regions (UTRs), or treatment with phosphatases to remove the 5' terminal phosphate - and new methods are constantly being developed.

[0216] Many compositions and techniques suitable for generating modified RNA for use herein have been developed, which are related to modifications of RNA interference (RNAi), including small interfering RNA (siRNA). siRNAs face particular challenges in vivo because their effects on gene silencing through mRNA interference are often transient, which may require repeated administration. In addition, siRNAs are double-stranded RNA (dsRNA), and mammalian cells have immune responses that have evolved to detect and neutralize dsRNA, which is often a byproduct of viral infection. Thus, there are mammalian enzymes such as PKR (dsRNA-responsive kinase) and potentially retinoic acid-inducible gene 1 (RIG-I), which can mediate cellular responses to dsRNA, as well as Toll-like receptors (such as TLR3, TLR7, and TLR8), which can trigger the induction of cytokines in response to these molecules; see, for example, reviews by Angart et al., Pharmaceuticals (Basel) 6(4):440-468 (2013); Kanasty et al., Molecular Therapy 20(3):513-524 (2012); Burnett et al., Biotechnol J. 6(9):1130-46 (2011); Judge and Maclachlan, Hum Gene Ther 19(2):111-24 (2008); and references cited therein.

[0217] A large number of modification methods have been developed and applied to improve RNA stability, reduce innate immune responses and / or achieve other benefits that may be useful in introducing nucleic acids into human cells as described herein; see, e.g., reviews by Whitehead KA et al., Annual Review of Chemical and Biomolecular Engineering, 2:77-96 (2011); Gaglione and Messere, Mini Rev Med Chem, 10(7):578-95 (2010); Chernolovskaya et al., Curr Opin Mol Ther., 12(2):158-67 (2010); Deleavey et al., Curr Protoc Nucleic Acid Chem Chapter 16:Unit 16.3 (2009); Behlke, Oligonucleotides 18(4):305-19 (2008); Fucini et al., Nucleic Acid Chem. Ther22(3):205-210(2012); Bremsen et al., FrontGenet 3:154(2012).

[0218] As mentioned above, there are many commercial suppliers of modified RNA, many of which are specialized in the modification for improving siRNA effectiveness. According to the various findings reported in the literature, a variety of methods are provided. For example, Dharmacon has noted that replacing non-bridging oxygen with sulfur (phosphorothioate (phosphorothioate, PS)) has been widely used in improving the nuclease resistance of siRNA, as reported by Kale, Nature Reviews Drug Discovery 11:125-140 (2012). It has been reported that the modification of ribose 2'-position improves the nuclease resistance of phosphate bond between nucleotides, while increasing duplex stability (Tm), which also shows that immune activation protection is provided. Moderate PS backbone modifications combined with small, well-tolerated 2'-substitutions (2'-O-, 2'-fluoro, 2'-hydrogen) have been associated with the use of highly stable siRNAs in vivo, as reported by Soutschek et al., Nature 432: 173-178 (2004); and 2'-O-methyl modifications have been reported to be effective in improving stability, as reported by Volkov, Oligonucleotides 19: 191-202 (2009). With regard to reducing the induction of innate immune responses, modification of specific sequences with 2'-O-methyl, 2'-fluoro, 2'-hydrogen has been reported to reduce TLR7 / TLR8 interactions while generally maintaining silencing activity; see, for example, Judge et al., Mol. Ther. 13: 494-505 (2006); and Cekaite et al., J. Mol. Biol. 365: 90-108 (2007). Additional modifications, such as 2-thiouracil, pseudouracil, 5-methylcytosine, 5-methyluracil, and N6-methyladenosine have also been shown to minimize TLR3-, TLR7-, and TLR8-mediated immune effects; see, e.g., Kariko et al., Immunity 23:165-175 (2005).

[0219] As is known in the art and commercially available, a number of conjugates can be applied to nucleic acids, such as the RNA used herein, which can enhance their delivery and / or uptake by cells, including, for example, cholesterol, tocopherol and folic acid, lipids, peptides, polymers, linkers and aptamers; see, e.g., review by Winkler, Ther. Deliv. 4:791-809 (2013), and references cited therein.

[0220] 6.10 Carrier

[0221] The present disclosure provides vectors comprising nucleic acids of the present disclosure, such as those described in Sections 6.9. In some embodiments, the nucleic acids comprise nucleic acids encoding engineered B-GEn polypeptides as described in Sections 6.2. In some embodiments, the engineered B-GEn polypeptide coding sequence is codon-optimized at least for the portion encoding the nuclease component of the engineered B-GEn polypeptide.

[0222] The vector (or nucleotide sequence) may further encode a gRNA.

[0223] In some embodiments, the vector comprising the nucleotide sequence may be an expression vector.

[0224] In some embodiments, the expression vector is a production vector for an engineered B-GEn polypeptide, for example, which can be used to express / produce an engineered B-GEn polypeptide in a host cell. After expression / production of the engineered B-GEn polypeptide in the host cell, the engineered B-GEn polypeptide can be incorporated into RNP for nuclear transfection of target cells.

[0225] Alternatively, the expression vector comprising the nucleotide sequence can be a delivery vector for the engineered B-GEn polypeptide, for example, which can be used to introduce the engineered B-GEn polypeptide coding sequence into a target cell intended for gene editing. After the engineered B-GEn polypeptide is expressed / produced in the target cell, the engineered B-GEn polypeptide, together with the guide RNA molecule, can edit the target cell. In some embodiments, the delivery vector further comprises a coding sequence for a gRNA. In other embodiments, a separate nucleic acid encoding the gRNA is introduced into the target cell.

[0226] Expression vectors contemplated include, but are not limited to, viral vectors based on vaccinia virus, poliovirus, adenovirus, adeno-associated virus, SV40, herpes simplex virus, human immunodeficiency virus, retroviruses (e.g., mouse leukemia virus, spleen necrosis virus, and vectors derived from retroviruses, such as Rous sarcoma virus, Harvey sarcoma virus, avian leukosis virus, lentivirus, human immunodeficiency virus, myeloproliferative sarcoma virus, and mammary tumor virus), and other recombinant vectors. Other vectors contemplated for use in eukaryotic target cells include, but are not limited to, vectors pXT1, pSG5, pSVK3, pBPV, pMSG, and pSVLSV40 (Pharmacia). Other vectors contemplated for use in eukaryotic cells include, but are not limited to, vectors pCTx-1, pCTx-2, and pCTx-3. Other vectors may be used as long as they are compatible with the intended host or target cell.

[0227] In some embodiments, the expression vector has one or more transcription and / or translation control elements. Depending on the expression cell / vector system used, many suitable transcription and translation control elements can be used in the vector, including constitutive and inducible promoters, transcription enhancer elements, transcription terminators, etc. The vector may also contain a ribosome binding site for translation initiation and a transcription terminator.

[0228] Non-limiting examples of suitable eukaryotic promoters (i.e., promoters that function in eukaryotic cells) include those derived from cytomegalovirus (CMV) immediate early, herpes simplex virus (HSV) thymidine kinase, early and late SV40, retroviral long terminal repeats (LTRs), human elongation factor-1 promoter (EF1), a hybrid construct of the cytomegalovirus (CMV) enhancer fused to the chicken beta-actin promoter (CAG), the murine stem cell virus promoter (MSCV), the phosphoglycerate kinase-1 locus promoter (PGK), and mouse metallothionein-I.

[0229] In some embodiments, the promoter is an inducible promoter (e.g., a heat shock promoter, a tetracycline-regulated promoter, a steroid-regulated promoter, a metal-regulated promoter, an estrogen receptor-regulated promoter, etc.). In some embodiments, the promoter is a constitutive promoter (e.g., a CMV promoter, a UBC promoter). In some embodiments, the promoter is a spatially restricted and / or temporally restricted promoter (e.g., a tissue-specific promoter, a cell type-specific promoter, etc.). In some embodiments, if the gene will be expressed under an endogenous promoter present in the genome after insertion into the genome, the vector does not have a promoter for expression of at least one gene in the host cell.

[0230] For expression of small RNAs, including guide RNAs, various promoters, such as RNA polymerase III promoters, including, for example, U6 and H1, can be advantageous. Thus, such promoters can be advantageously incorporated into delivery vectors. Descriptions and parameters for enhancing the use of such promoters are known in the art, and additional information and methods are frequently described; see, e.g., Ma, H. et al., Molecular Therapy-Nucleic Acids 3, e161 (2014) doi:10.1038 / mtna.2014.12.

[0231] In some embodiments, the vector is a self-inactivating vector that inactivates viral sequences or components or other elements of the CRISPR machinery. Self-inactivating vectors are particularly suitable for delivery vectors to select for cells that retain the engineered B-GEn polypeptide coding sequence after gene editing is completed.

[0232] In some embodiments, the expression vector is an RNA vector. In other embodiments, the expression vector is a DNA vector.

[0233] 6.10.1 RNA Vectors

[0234] The expression vector of the present disclosure may be an RNA vector.

[0235] Particularly suitable vectors are viral replicons based on RNA viruses, such as alphaviruses and paramyxoviruses. Alphavirus and paramyxovirus replicons do not involve DNA intermediates for replication and therefore provide safer alternatives to several other commonly used viral vectors, including lentiviral and retroviral vectors (Yoshioka et al., 2013, Cell Stem Cell. 13(2): 246-54; Yoshioka and Dowdy, 2017, PLOS ONE 12: e0182018). Alphaviruses are positive-sense RNA viruses with lipid envelopes that constitute a genus of more than 30 species of viruses in the Togaviridae family, including Eastern, Western, and Venezuelan equine encephalitis viruses (EEEV, WEEV, and VEEV, respectively), Chikungunya (CHIK), Sindbis, Ross River, and O'nyong-nyong viruses, among others. Sendai viruses (SeV) are enveloped, single-stranded, negative-sense paramyxoviruses that replicate episomally in the cytoplasm of host cells.

[0236] Thus, in some embodiments, the RNA vector is derived from an RNA virus, such as an alphavirus, paramyxovirus, flavivirus, rhabdovirus, measles virus, or picornavirus.

[0237] In some embodiments, the RNA vector is a single-stranded RNA replicon. In some embodiments, the single-stranded RNA replicon is a plus strand. In some other embodiments, the single-stranded RNA replicon is a minus strand. In some embodiments, the RNA vector comprises one or more coding sequences and self-replicating elements of one or more engineered B-GEn polypeptides.

[0238] The RNA replicons of the present disclosure typically include regulatory elements, a subgenomic (SG) promoter operably linked to an engineered B-GEn polypeptide coding sequence. The sequence comprising the engineered B-GEn polypeptide coding sequence is typically flanked by 5' and 3' UTR sequences, and the 3' UTR sequence is typically followed by a polyadenylation signal.

[0239] RNA vector constructs can be generated from DNA templates (DNA plasmid constructs). For example, RNA constructs can be transcribed from DNA templates using SP6 or T7 in vitro transcription kits.

[0240] RNA vectors are particularly useful as delivery vehicles.

[0241] 6.10.2 DNA Vectors

[0242] In some embodiments, the expression vectors of the present disclosure are DNA vectors. The present disclosure provides two types of DNA vectors: (1) DNA vectors that are production vectors or delivery vectors, and (2) DNA vectors from which RNA vectors of the present disclosure (as described in Section 6.10.1) can be transcribed, as described in Section 6.10.2.2. DNA vectors from which RNA replicons of the present disclosure can be transcribed are sometimes referred to herein as "template vectors."

[0243] In some embodiments, the DNA vectors of the present disclosure are non-integrating DNA vectors. For example, the vectors can be episomal vectors. For example, many DNA viruses, such as adenoviruses, Simian vacuolating virus 40 (SV40), bovine papilloma virus (BPV), or plasmids containing budding yeast ARS (autonomously replicating sequence) can be used without genomic integration.

[0244] In some embodiments, the DNA vector of the present disclosure includes an origin of replication. The example of the origin of replication that can be incorporated into the DNA vector of the present disclosure includes a lymphotropic herpes virus (lymphotropic herpes virus), a gamma herpes virus (gammaherpesvirus), adenovirus, bovine papilloma virus (bovine papilloma virus) or the origin of replication of yeast. In some embodiments, the replication origin is from a lymphotropic herpes virus or a gamma herpes virus corresponding to the oriP of EBV, as a self-replicating element. In some embodiments, the lymphotropic herpes virus is Epstein Barr virus (Epstein Barr virus, EBV), Kaposi's sarcoma herpes virus (Kaposi's sarcoma herpes virus, KSHV), squirrel monkey herpes virus (Herpes virus saimiri, HS) or Marek's disease virus (Marek's disease virus, MDV). Epstein Barr virus (EBV) and Kaposi's sarcoma herpes virus (KSHV) are also examples of gamma herpes viruses.

[0245] In certain embodiments, the vectors of the present disclosure comprise the EBV origin of replication, OriP. OriP is a site at or near the site of DNA replication initiation, and it consists of two cis-acting sequences approximately 1 kilobase pair apart, called the repeat family (FR) and the binary symmetry (DS). The FR consists of 21 imperfect copies of 30bp repeats and contains 20 high-affinity EBNA-1 binding sites. When the FR binds to EBNA-1, both act as transcription enhancers for cis-acting promoters up to 10kb away. The DS is sufficient to initiate DNA synthesis in the presence of EBNA-1, and this initiation occurs at or near the DS.

[0246] One or more expression cassettes in the replicating DNA vector may further comprise a nucleotide sequence encoding a trans-acting factor that binds to the replication origin to replicate the extrachromosomal template. Alternatively or additionally, the somatic cell may express such a trans-acting factor.

[0247] In other embodiments, the DNA vectors of the present disclosure lack an origin of replication.

[0248] The DNA vectors of the present disclosure typically comprise one or more promoters, such as SP6 or T7, to drive expression of an engineered B-GEn polypeptide in the case of a DNA vector intended as a production vector, or an RNA replicon in the case of a DNA vector intended as a template vector.

[0249] 6.10.2.1 Expression Vector

[0250] In some embodiments, the expression vector is a DNA vector comprising an expression cassette for expressing one or more proteins of interest, the expression cassette being operably linked to a regulatory element comprising a promoter suitable for driving expression of the engineered B-GEn polypeptide in the cell type of interest. Examples of promoters suitable for driving protein expression in mammalian cells include the cytomegalovirus (CMV) promoter, the EF1a promoter, the SV40 promoter, the Ubc promoter, the human β-actin promoter, the PGK1 promoter, and the CAG promoter.

[0251] DNA vectors used to directly express engineered B-GEn polypeptides (rather than as templates for expression of RNA vectors, as described in Section 6.10.2.2) need not include RNA replicon self-replicating sequences, such as the nsP1-nsP4 proteins of VEEV or the NP, P, and L proteins of Sendai virus.

[0252] In some embodiments, the DNA expression vector is a non-replicating DNA vector.In some embodiments, the DNA expression vector is a replicating DNA vector.

[0253] 6.10.2.2 Template Vector

[0254] The DNA vectors of the present disclosure may also serve as templates for transcribing RNA replicons as described herein. Thus, the "expression cassette" contained in the template vector is intended for transcription of the RNA replicon produced by transcription of the RNA replicon.

[0255] Thus, the template vector of the present disclosure comprises a nucleotide sequence encoding an RNA replicon as described herein under the control of a regulatory element, an SP6 or T7 promoter.

[0256] In some embodiments, the template DNA vector is a non-replicating DNA vector. In some embodiments, the template DNA vector is a replicating DNA vector.

[0257] In some embodiments, the template vector is used for in vitro transcription of an RNA replicon, which is subsequently introduced into cells to drive expression of an engineered B-GEn polypeptide.

[0258] 6.10.3 Viral Vectors

[0259] Recombinant adeno-associated virus (AAV) vectors can be used for delivery. A known technique for producing rAAV particles in the art is to provide a cell with a polynucleotide to be delivered between two AAV inverted terminal repeats (ITRs), AAV rep and cap genes, and helper virus functions. The production of rAAV requires the following components in a single cell (referred to herein as a packaging cell): the polynucleotide of interest between the two ITRs, the AAV rep and cap genes separated from the AAV genome (i.e., not in the AAV genome), and helper virus functions. The AAV rep and cap genes can be from any AAV serotype from which a recombinant virus can be derived, or from an AAV serotype different from the ITRs on the packaging polynucleotide, including but not limited to AAV serotypes AAV-1, AAV-2, AAV-3, AAV-4, AAV-5, AAV-6, AAV-7, AAV-8, AAV-9, AAV-10, AAV-11, AAV-12, AAV-13, and AAV rh.74. For example, WO 01 / 83692 discloses the production of pseudovirus rAAV.

[0260]

[0261]

[0262] The method for producing packaging cells is to create a cell line that stably expresses all the components required for the production of AAV particles. For example, a plasmid (or multiple plasmids) is integrated into the genome of the cell, and the plasmid contains a plurality of AAV ITRs between the target polynucleotides, the AAV rep and cap genes separated from the AAV genome, and a selectable marker, such as a neomycin resistance gene. The AAV genome has been introduced into bacterial plasmids by procedures such as GC tailing (Samulski et al., 1982, Proc. Natl. Acad. Sci. USA, 79: 2077-2081), addition of synthetic linkers containing restriction endonuclease cleavage sites (Laughlin et al., 1983, Gene, 23: 65-73) or direct blunt end connection (Senapathy & Carter, 1984, J. Biol. Chem., 259: 4661-4666). The packaging cell line is then infected with a helper virus, such as adenovirus. The advantage of this method is that the cell has selectivity and is suitable for large-scale production of rAAV. Other examples of suitable methods are the use of adenovirus or baculovirus rather than plasmids to introduce the rAAV genome and / or rep and cap genes into packaging cells.

[0263] General principles of rAAV production are reviewed in, for example, Carter, 1992, Current Opinions in Biotechnology, 1533-539; and Muzyczka, 1992, Curr. Topics in Microbial. and Immunol., 158:97-129). Various methods are described in Ratschin et al., Mol. Cell. Biol. 4:2072 (1984); Hermonat et al., Proc. Natl. Acad. Sci. USA, 81:6466 (1984); Tratschin et al., Mol. Cell. Biol. 5:3251 (1985); Mclaughlin et al., J. Virol., 62:1963 (1988); and Lebkowski et al., 1988 Mol. Cell. Biol., 7:349 (1988). Samulski et al. (1989, J. Virol., 63:3822-3828); U.S. Patent No. 5,173,414; WO 95 / 13365 and corresponding U.S. Patent No. 5,658.776; WO 95 / 13392; WO 96 / 17947; PCT / US98 / 18600; WO 97 / 09441 (PCT / US96 / 14423); WO 97 / 08298 (PCT / US96 / 13872); WO 97 / 21825 (PCT / US96 / 20777); WO 97 / 06243 (PCT / FR96 / 01064); WO 99 / 11764; Perrin et al. (1995) Vaccine 13:1244-1250; Paul et al. (1993) Human Gene Therapy 4:609-615; Clark et al. (1996) Gene Therapy 3:1124-1132; U.S. Patent No. 5,786,211; U.S. Patent No. 5,871,982; and U.S. Patent No. 6,258,595.

[0264] The AAV vector serotype used for transduction depends on the target cell type. For example, the following exemplary cell types are known to be transducible by the indicated AAV serotypes, etc.

[0265] Tissue / cell type Serotype liver AAV8, AAV9 skeletal muscle AAV1, AAV7, AAV6, AAV8, AAV9 central nervous system AAV5, AAV1, AAV4 RPE AAV5, AAV4 photoreceptor cells AAV5 lung AAV9 Heart AAV8 pancreas AAV8 kidney AAV2

[0266] Many suitable expression vectors are known to those skilled in the art, and many are commercially available. The following vectors are provided by way of example; for eukaryotic host cells: pXT1, pSG5 (Stratagene), pSVK3, pBPV, pMSG, and pSVLSV40 (Pharmacia). However, any other vector may be used as long as it is compatible with the host cell.

[0267] 6.11 Host Cells and Recombinant Expression

[0268] In some embodiments, host cells can be used to express gRNA, sgRNA or engineered B-GEn polypeptides of the present disclosure. Suitable host cells include naturally occurring cells; genetically modified cells (e.g., cells genetically modified in the laboratory); and cells manipulated in vitro in any manner. In some embodiments, the host cell is isolated.

[0269] Host cells can be eukaryotic or prokaryotic, and include, for example, yeast (e.g., Pichia pastoris or Saccharomyces cerevisiae), bacteria (e.g., Escherichia coli or Bacillus subtilis), insect Sf9 cells (e.g., baculovirus-infected SF9 cells), or mammalian cells (e.g., human embryonic kidney (HEK) cells, Chinese hamster ovary cells, HeLa cells, human 293 cells, and monkey COS-7 cells).

[0270] Host cells can be derived from established cell lines, or they can be primary cells, wherein "primary cells (primary cells)", "primary cell lines (primary cell lines)" and "primary cultures (primary cultures)" are used interchangeably herein and refer to cells and cell cultures derived from a subject and allowed to grow in vitro for a limited number of passages (e.g., division of the culture). For example, primary cultures include cultures that have been passaged 0, 1, 2, 4, 5, 10 or 15 times, but not enough times to go through a crisis stage (crisis stage). Primary cell lines can be maintained in vitro for less than 10 passages. In some embodiments, the host cell is a PSC (e.g., iPSC or ESC) or a PSC-derived cell (e.g., PSC-derived neurons, PSC-derived microglia, PSC-derived cardiomyocytes, PSC-derived ocular cells).

[0271] If the cell is a primary cell, such cells can be harvested from an individual by any suitable method. Suitable solutions can be used to disperse or suspend the harvested cells. The harvested cells can be used immediately, or they can be stored for a long time, frozen, and reused after thawing. In this case, cells are typically frozen in 10% dimethyl sulfoxide (DMSO), 50% serum, 40% buffered culture medium, or other such solutions commonly used in the art to preserve cells at such freezing temperatures and thawed according to methods well known in the art for thawing frozen cultured cells.

[0272] 6.12 Target Cells

[0273] In some embodiments, the B-GEn CRISPR-Cas system is introduced into a target cell or a population of target cells. Methods for introducing proteins and nucleic acids into target cells are further described in Section 6.13.

[0274] The target cells and target cell populations of the present disclosure can be cells in which gene editing by the system of the present disclosure has occurred, or cells in which components of the system of the present disclosure have been introduced or expressed but gene editing has not yet occurred, or a combination thereof. In various embodiments, the cell population can include, for example, a population in which at least 1%, at least 5%, at least 10%, at least 15%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, or at least 70% of the cells have been gene-edited by the system of the present disclosure.

[0275] In some embodiments, the methods of the present disclosure can be used to induce transcriptional regulation in mitotic or post-mitotic cells in vivo and / or ex vivo and / or in vitro. In some embodiments, the methods of the present disclosure can be used to induce DNA cleavage, DNA modification, and / or transcriptional regulation in mitotic or post-mitotic cells in vivo and / or ex vivo and / or in vitro (e.g., to produce genetically modified cells that can be reintroduced into an individual).

[0276] Because the guide RNA provides specificity by hybridizing with the target DNA, the mitotic and / or post-mitotic cells can be any of a variety of target cells, wherein suitable target cells include, but are not limited to, bacterial cells; archaeal cells; unicellular eukaryotic organisms; plant cells; algal cells, such as Botryococcus braunii, Chlamydomonas reinhardtii, Nannochloropsis gaditana, Chlorela pyrenoidosa, Sargassum patens, C. agardh, etc.; fungal cells; animal cells; cells derived from invertebrates (such as insects, cnidarians, echinoderms, nematodes, etc.); eukaryotic parasites (such as malarial parasites, such as Plasmodium, worms, etc.); cells derived from vertebrates (such as fish, amphibians, reptiles, birds, mammals); mammalian cells, such as rodent cells, human cells, non-human primate cells, etc. In some embodiments, the target cell can be any human cell. Suitable target cells include naturally occurring cells; genetically modified cells (e.g., cells genetically modified in a laboratory, such as by "human hands"); and cells manipulated in vitro in any way. In some embodiments, the target cell is isolated.

[0277] Any type of cell can be used as a host cell or target cell (e.g., stem cells, such as embryonic stem cells (ES), induced pluripotent stem cells (iPSC), germ cells; somatic cells, such as fibroblasts, hematopoietic cells, neural cells, muscle cells, bone cells, liver cells, pancreatic cells; embryonic cells of any stage embryos in vitro or in vivo, such as embryonic cells of zebrafish embryos at stages such as 1-cell, 2-cell, 4-cell, 8-cell, etc.). Cells can be derived from established cell lines, or they can be primary cells, wherein "primary cells", "primary cell lines", "primary cultures" are used interchangeably herein and refer to cells and cell cultures derived from a subject and allowed to grow a limited number of passages in vitro (e.g., division of a culture). For example, primary cultures include cultures that have been passaged 0, 1, 2, 4, 5, 10, or 15 times, but not enough times to undergo a crisis stage (crisis stage). Primary cell lines can be maintained in vitro for less than 10 passages. In some embodiments, the target cell is a unicellular organism or is grown in culture. In some embodiments, the host cell is the same as the target cell. In some embodiments, the target cell is modified to become another cell type so that the resulting host cell is different from the target cell. As an example, the target cell can be a PSC (e.g., iPSC), which is then differentiated into a PSC-derived cell (e.g., a PSC-derived neuron) so that the host cell is a neuron.

[0278] If cell is primary cell, then such cell can be gathered in the crops from individual by any suitable method.For example, leukocyte can be suitably gathered in the crops by methods such as apheresis, leukocytapheresis, density gradient separation, and the cell from tissue (such as skin, muscle, bone marrow, spleen, liver, pancreas, lung, intestine, stomach etc.) is then most suitable for gathering by biopsy. Suitable solution can be used to disperse or suspend the cell of gathering. This solution is normally balanced salt solution, such as normal saline, phosphate buffered saline (PBS), Hank's balanced salt solution etc., suitably supplements fetal bovine serum or other naturally occurring factors, in conjunction with the buffer of acceptable low concentration such as 5-25mM. Suitable buffer includes HEPES, phosphate buffer, lactate buffer etc. Cell can be used immediately, or they can be stored for a long time, freezing, and it can be reused after thawing. In such cases, the cells are typically frozen in 10% dimethyl sulfoxide (DMSO), 50% serum, 40% buffered culture medium, or some other such solution commonly used in the art to preserve the cells at such freezing temperatures, and thawed in a manner well known in the art for thawing frozen cultured cells.

[0279] 6.12.1 Induced Pluripotent Stem Cells (iPSCs)

[0280] In some embodiments, target cell is induced pluripotent stem cell (iPSC), which is a specific cell type starting point for the regenerative medicine treatment of a large number of patients with different diseases that can be potentially generated. In the context of iPSC, differentiation is the process of carrying out pedigree specialization using a cell specific scheme starting from iPSC. The iPSC of the present disclosure can be differentiated into the purpose cell type for cell therapy, including endoderm (such as, lung, thyroid or pancreatic cell or its progenitor cell), ectoderm (such as, skin, neuron or pigment cell or its progenitor cell) and mesoderm (such as, cardiac cell, skeletal muscle cell, erythrocyte, smooth muscle cell or its progenitor cell) pedigree.

[0281] In some embodiments, the iPSCs of the present disclosure are differentiated into cardiac cells. In various embodiments, the cardiac cells are cardiac progenitor cells or mature or immature (atrial or ventricular) cardiomyocytes.

[0282] In other embodiments, iPSCs of the present disclosure are differentiated into oligodendrocyte progenitor cells or oligodendrocytes.

[0283] In other embodiments, the iPSCs of the present disclosure are differentiated into cells of the neural lineage, such as neural crest cells, astrocytes, dopaminergic neuron progenitor cells, dopaminergic neuron cells, midbrain dopaminergic neuron progenitor cells, midbrain dopaminergic neurons, uthentic midbrain dopamine (DA) neurons, dopaminergic neuron precursor cells, floor plate midbrain progenitor cells, floor plate midbrain DA neurons.

[0284] In other embodiments, the iPSCs of the present disclosure are differentiated into photoreceptor cells, photoreceptor precursor cells, retinal pigment epithelial cells, neural retinal cells, or neural retinal progenitor cells.

[0285] In further embodiments, the iPSCs of the present disclosure are differentiated into microglia or microglial progenitor cells.

[0286] In a further embodiment, the iPSCs of the present disclosure are differentiated into macrophages.

[0287] In further embodiments, the iPSCs of the present disclosure are differentiated into intestinal progenitor cells or intestinal cells.

[0288] In some embodiments, iPSCs can be genetically engineered prior to differentiation into the cell type of interest (e.g., to produce a functional protein that is deficient in the patient, to produce a therapeutic protein, to include an off switch, or to evade immune detection to support allogeneic applications).

[0289] 6.13 Methods for Introducing Nucleic Acids into Host and Target Cells

[0290] In some embodiments, the method of the present disclosure includes and relates to one or more nucleic acids being introduced into a host or target cell (or a host group or a target cell group), the nucleic acid comprising a nucleotide sequence encoding a guide RNA and / or a codon-optimized nucleotide sequence encoding an engineered B-GEn polypeptide. In some embodiments, the target cell (e.g., comprising a cell targeted by a guide RNA for editing a DNA sequence by an engineered B-GEn polypeptide) is external. In some embodiments, the target cell is in vivo. In some embodiments, the nucleotide sequence encoding a guide RNA and / or an engineered B-GEn polypeptide is operably connected to an inducible promoter. In some embodiments, the nucleotide sequence encoding a guide RNA and / or an engineered B-GEn polypeptide is operably connected to a constitutive promoter.

[0291] Guide RNA or the nucleic acid comprising the nucleotide sequence of coding guide RNA can be introduced into host or target cell by any various well-known methods. Similarly, when the method relates to the nucleic acid comprising the codon-optimized nucleotide sequence of coding engineered B-GEn polypeptide and is introduced into host or target cell, this nucleic acid can be introduced into host or target cell by any various well-known methods. Guide nucleic acid (RNA or DNA) and / or the nucleic acid (RNA or DNA) of coding engineered B-GEn polypeptide can be delivered by viral or non-viral delivery vectors known in the art.

[0292] Methods for introducing nucleic acids into host or target cells are known in the art, and any known method can be used to introduce nucleic acids (such as expression constructs) into stem or progenitor cells. Suitable methods include, for example, viral or bacteriophage infection, transfection, conjugation, protoplast fusion, lipid infection, electroporation, nucleofection, calcium phosphate precipitation, polyethyleneimine (PEI)-mediated transfection, DEAE-dextran-mediated transfection, liposome-mediated transfection, particle gun technology, calcium phosphate precipitation, direct microinjection, nanoparticle-mediated nucleic acid delivery (see, e.g., Panyam et al., Adv Drug Deliv Rev. 2012 Sep 13. pii: 50169-409X(12)00283-9. doi: 10.1016 / j.addr.2012.09.023), etc., including but not limited to exosome delivery.

[0293] Polynucleotides can be delivered by non-viral delivery vectors, including but not limited to nanoparticles, liposomes, ribonucleoproteins, positively charged peptides, small molecule RNA conjugates, aptamers-RNA chimeras and RNA fusion protein complexes. Some exemplary non-viral delivery vectors are recorded in Peer and Lieberman, Gene Therapy, 18:1127-1133 (2011) (it focuses on non-viral delivery vectors of siRNA, which can also be used to deliver other nucleic acids).

[0294] Suitable systems and technologies for delivering nucleic acids (such as mRNA and sgRNA) of the present disclosure for gene editing include lipid nanoparticles (LNP). As used herein, the term "lipid nanoparticle" includes liposomes, regardless of their lamellarity, shape or structure, and also includes lipid complexes described for introducing nucleic acids and / or polypeptides into cells. These lipid nanoparticles can be compounded with bioactive compounds (such as nucleic acids and / or polypeptides) and can be used as in vivo delivery vectors. Generally, any method known in the art can be used to prepare lipid nanoparticles comprising one or more nucleic acids of the present disclosure, as well as complexes of bioactive compounds and the lipid nanoparticles. Examples of such methods have been widely disclosed, for example, in Biochim Biophys Acta 1979, 557:9; Biochim et Biophys Acta 1980, 601:559; Liposomes: A practical approach (Oxford University Press, 1990); Pharmaceutica Acta Helvetiae 1995, 70:95; Current Science 1995, 68:715; Pakistan Journal of Pharmaceutical Sciences 1996, 19:65; Methods in Enzymology 2009, 464:343). Particularly suitable systems and technologies for preparing LNP formulations comprising one or more nucleic acids and / or polypeptides of the present disclosure include, but are not limited to, those developed by Intellia (see, e.g., WO2017173054A1), Alnylam (see, e.g., WO2014008334A1), Modernatx (see, e.g., WO2017070622A1 and WO2017099823A1), TranslateBio, Acuitas (see, e.g., WO2018081480A1), Genevant Sciences, Arbutus Biopharma, Tekmira, Arcturus, Merck (see, e.g., WO2015130584A2), Novartis (see, e.g., WO2015095340A1), and Dicerna; all of which are incorporated herein by reference in their entirety.

[0295] Suitable nucleic acids comprising nucleotide sequences encoding engineered B-GEn polypeptides and / or guide RNAs include expression vectors. In some embodiments, the expression vector is a viral construct, such as a recombinant adeno-associated virus construct (see, e.g., U.S. Patent No. 7,078,387), a recombinant adenoviral construct, a recombinant lentiviral construct, a recombinant retroviral construct, or the like. Suitable expression vectors include, but are not limited to, viral vectors (e.g., vaccinia virus-based viral vectors; poliovirus-based viral vectors; adenovirus-based viral vectors (see, e.g., Li et al., Invest Opthalmol Vis Sci 35:2543-2549, 1994; Borras et al., Gene Ther 6:515-524, 1999; Li and Davidson, PNAS 92:7700-7704, 1995; Sakamoto et al., H Gene Ther 5:1088-1097, 1999; WO 94 / 12649, WO 93 / 03769; WO 93 / 19191; WO 94 / 28938; WO 95 / 11984 and WO 95 / 00655); adeno-associated virus (see, e.g., Ali et al., Hum Gene Ther 9:81-86, 1998; Flannery et al., PNAS 94:6916-6921, 1997; Bennett et al., Invest Opthalmol Vis Sci 38:2857-2863, 1997; Jomary et al., Gene Ther 4:683-690, 1997, Rolling et al., Hum Gene Ther 10:641-648, 1999; Ali et al., Hum Mol Genet 5:591-594, 1996; Srivastava in WO 93 / 09239, Samulski et al., J. Vir. (1989) 63:3822-3828; Mendelson et al., Viral. (1988) 166:154-165; and Flotte et al., PNAS (1993) 90:10613-10617); SV40; herpes simplex virus; human immunodeficiency virus (see, e.g., Miyoshi et al., PNAS 94:10319-23, 1997; Takahashi et al., J Virol 73:7812-7816, 1999); retroviral vectors (e.g., murine leukemia virus, spleen necrosis virus, retrovirus-derived vectors such as Rous sarcoma virus, Harvey sarcoma virus, avian leukosis virus, lentivirus, human immunodeficiency virus, myeloproliferative sarcoma virus, and mammary tumor virus), etc.

[0296] In some embodiments, B-GEn ribonucleoprotein comprising a B-GEn endonuclease and sgRNA is delivered to target cells by nucleofection, a method of delivering nucleic acids to cells by creating transient pores in the cell membrane using cell-specific reagents and electrical parameters.

[0297] The engineered B-Gen V type CRISPR-Cas system can be delivered to the target cell by a delivery vector (such as a viral vector). The engineered B-GEn V type CRISPR Cas system can also be delivered to the target cell by a non-viral delivery vector, including but not limited to nanoparticles, liposomes, ribonucleoproteins, positively charged peptides, small molecule RNA conjugates, aptamers-RNA chimeras, and RNA fusion protein complexes. Some exemplary non-viral delivery vectors are described in Peer and Lieberman, Gene Therapy, 18: 1127-1133 (2011).

[0298] In some embodiments, the engineered B-Gen Type V CRISPR-Cas system is delivered to target cells using a delivery vector as described in Section 6.10.

[0299] 6.14 Pharmaceutical Compositions

[0300] Also disclosed herein are pharmaceutical formulations and medicaments comprising the B-GEn protein, gRNA, nucleic acid or nucleic acids, system, particle or particles of the present disclosure and a pharmaceutically acceptable excipient.

[0301] Suitable excipients include, but are not limited to, salts, diluents (e.g., Tris-HCl, acetates, phosphates), preservatives (e.g., thimerosal, benzyl alcohol, parabens), binders, fillers, solubilizers, disintegrants, adsorbents, solvents, pH regulators, antioxidants, anti-infectives, suspending agents, wetting agents, viscosity regulators, tonicity agents, stabilizers and other components and combinations thereof. Suitable pharmaceutically acceptable excipients can be selected from materials that are generally recognized as safe (GRAS) and can be applied to an individual without causing undesirable biological side effects or undesirable interactions. Suitable excipients and their formulations are described in Remington's Pharmaceutical Sciences, 16th ed. 1980, Mack Publishing Co. In addition, such compositions can be complexed with polyethylene glycol (PEG), metal ions, or incorporated into polymeric compounds such as polyacetic acid, polyglycolic acid, hydrogels, etc., or incorporated into liposomes, microemulsions, micelles, unilamellar or multilamellar vesicles, erythrocyte ghosts, or spheroplasts. Suitable dosage forms for administration (e.g., parenteral administration) include solutions, suspensions, and emulsions.

[0302] The components of the pharmaceutical preparation can be dissolved or suspended in a suitable solvent, such as water, Ringer's solution, phosphate buffered saline (PBS) or isotonic sodium chloride. The preparation can also be a sterile solution, suspension or emulsion in a non-toxic parenterally acceptable diluent or solvent (such as 1,3-butanediol).

[0303] In some cases, the formulation may include one or more tonicity agents to adjust the isotonic range of the formulation. Suitable tonicity agents are well known in the art and include glycerol, mannitol, sorbitol, sodium chloride and other electrolytes. In some cases, the formulation may be buffered with an effective amount of buffer necessary to maintain a pH suitable for parenteral administration. Suitable buffers are well known to those skilled in the art, and some examples of useful buffers are acetate, borate, carbonate, citrate and phosphate buffers.

[0304] In some embodiments, the formulation can be dispensed or packaged in liquid form, or alternatively as a solid, for example, obtained by lyophilizing a suitable liquid formulation, which can be reconstituted with a suitable carrier or diluent before administration. In some embodiments, the formulation can include a pharmaceutically effective amount of a guide RNA and a type II Cas protein sufficient to edit genes in a cell. The pharmaceutical composition can be formulated for medical and / or veterinary use.

[0305] In some embodiments, the B-GEn endonuclease complex can be introduced into a host cell (e.g., iPSC) to produce genetically modified cells that can be reintroduced into an individual. The iPSC-derived cells described herein can be provided in a pharmaceutical composition containing cells and a pharmaceutically acceptable carrier. A pharmaceutically acceptable carrier can be a cell culture medium that is optionally free of any animal-derived components. For storage and transportation, the cells can be cryopreserved at <-70°C (e.g., on dry ice or in liquid nitrogen). Prior to use, the cells can be thawed and diluted in a sterile cell culture medium that supports the cell type of interest.

[0306] The cells can be administered to the patient systemically (e.g., by intravenous injection or infusion) or locally (e.g., by direct injection into local tissues, such as the heart, brain, and sites of damaged tissue). Various methods are known in the art for administering cells to a patient's tissue or organ, including, but not limited to, intracoronary administration, intramyocardial administration, transendocardial administration, or intracranial administration.

[0307] Administering a therapeutically effective amount of iPSC-derived cells to a patient. As used herein, the term "therapeutically effective" refers to a number of cells or an amount of a pharmaceutical composition that, when administered to a human subject suffering from or susceptible to a disease, disorder, and / or condition, is sufficient to treat, prevent, and / or delay the onset or progression of symptoms of the disease, disorder, and / or condition. One of ordinary skill in the art will understand that a therapeutically effective amount is typically administered via a dosing regimen comprising at least one unit dose. In some embodiments, at least 10 cells are administered to a subject at one time at one or more sites. 3 (e.g., at least 10 4 At least 10 5 At least 10 6 At least 10 7 At least 10 8 At least 10 9 At least 10 10 At least 10 11 or at least 10 12 In some embodiments, 10 cells are administered to a subject at one or more sites at a time. 3 -10 18 (e.g., 10 3 -10 4 10 3 -10 5 10 3 -10 6 10 3 -10 7 10 3 -10 8 10 3 -109 10 3 -10 10 10 3 -10 11 10 3 -10 12 10 6 -10 7 10 6 -10 8 10 6 -10 9 10 6 -10 10 10 6 -10 11 10 6 -10 12 10 9 -10 10 10 9 -10 11 10 9 -10 12 In some embodiments, more than 10 cells are administered to a subject at one time at one or more sites. 12 (e.g., more than 10 12 More than 10 13 More than 10 14 More than 10 15 More than 10 16 More than 10 17 More than 10 18 or more) cells.

[0308] 7. Numbering Implementation Plan

[0309] Although various specific embodiments have been shown and described, it will be appreciated that various changes may be made without departing from the spirit and scope of the present disclosure. The present disclosure is illustrated by the numbered embodiments listed below. Unless otherwise indicated, any concept, aspect, and / or feature of the embodiments described in the detailed description above apply, mutatis mutandis, to any of the embodiments in the numbered embodiments below.

[0310] 1. A fusion polypeptide comprising:

[0311] (a) Nuclease sequence:

[0312] (i) as defined in Section 4.3;

[0313] (ii) an amino acid sequence that has at least 80% sequence identity to SEQ ID NO: 4 (B-GEn. 1), SEQ ID NO: 5 (B-GEn. 1.2), or SEQ ID NO: 6 (B-GEn. 2);

[0314] (iii) which is the same as SEQ ID NO: 4 (B-GEn. 1), SEQ ID NO: 5

[0315] (B-GEn.1.2) or an amino acid sequence having at least 80% sequence identity to the nuclease domain of SEQ ID NO: 6 (B-GEn.2);

[0316] (iv) which is the same as SEQ ID NO: 4 (B-GEn. 1), SEQ ID NO: 5

[0317] (B-GEn.1.2) or an amino acid sequence that differs by no more than 25 amino acids from the amino acid sequence of SEQ ID NO: 6 (B-GEn.2); or

[0318] (v) any combination of two, three or all four of (i) to (iv),

[0319] (b) a first nuclear localization signal ("NLS") sequence located at the C-terminus of the nuclease sequence.

[0320] 2. The fusion polypeptide according to embodiment 1, further comprising a first linker sequence between the nuclease sequence and the first NLS sequence.

[0321] 3. The fusion polypeptide of embodiment 1 or embodiment 2, further comprising a second NLS sequence located at the C-terminus of the first NLS sequence.

[0322] 4. The fusion polypeptide of embodiment 3, comprising a linker sequence between the first NLS sequence and the second NLS sequence.

[0323] 5. The fusion polypeptide of embodiment 3 or embodiment 4, further comprising a third NLS sequence located at the C-terminus of the second NLS sequence.

[0324] 6. The fusion polypeptide of embodiment 5, comprising a linker sequence between the second NLS sequence and the third NLS sequence.

[0325] 7. The fusion polypeptide of embodiment 5 or embodiment 6, further comprising a fourth NLS sequence located at the C-terminus of the third NLS sequence.

[0326] 8. The fusion polypeptide of embodiment 7, comprising a linker sequence between the third NLS sequence and the fourth NLS sequence.

[0327] 9. The fusion polypeptide of any one of embodiments 1 to 8, which lacks an NLS sequence located at the N-terminus of the nuclease sequence.

[0328] 10. The fusion polypeptide of any one of embodiments 1 to 8, comprising an NLS sequence located at the N-terminus of the nuclease sequence.

[0329] 11. The fusion polypeptide of embodiment 10, comprising a linker sequence between an NLS sequence located at the N-terminus of the nuclease sequence and the nuclease sequence.

[0330] 12. The fusion polypeptide of any one of embodiments 1 to 11, wherein each NLS sequence is independently selected from the NLS sequences listed in Table 3.

[0331] 13. The fusion polypeptide of any one of embodiments 1 to 12, wherein each linker sequence is independently selected from the nuclease sequences listed in Table 4.

[0332] 14. The fusion polypeptide of any one of embodiments 1 to 13, wherein the nuclease sequence has at least 85% sequence identity to the amino acid sequence of any one of SEQ ID NO: 4, SEQ ID NO: 5, or SEQ ID NO: 6.

[0333] 15. The fusion polypeptide of any one of embodiments 1 to 13, wherein the nuclease sequence has at least 90% sequence identity to the amino acid sequence of any one of SEQ ID NO: 4, SEQ ID NO: 5, or SEQ ID NO: 6.

[0334] 16. The fusion polypeptide of any one of embodiments 1 to 13, wherein the nuclease sequence has at least 95% sequence identity to the amino acid sequence of any one of SEQ ID NO: 4, SEQ ID NO: 5, or SEQ ID NO: 6.

[0335] 17. The fusion polypeptide of any one of embodiments 1 to 13, wherein the nuclease sequence has at least 98% sequence identity to the amino acid sequence of any one of SEQ ID NO: 4, SEQ ID NO: 5, or SEQ ID NO: 6.

[0336] 18. The fusion polypeptide of any one of embodiments 1 to 13, wherein the nuclease sequence has at least 99% sequence identity to the amino acid sequence of any one of SEQ ID NO: 4, SEQ ID NO: 5, or SEQ ID NO: 6.

[0337] 19. The fusion polypeptide of any one of embodiments 1 to 13, wherein the nuclease sequence has at least 99.5% sequence identity to the amino acid sequence of any one of SEQ ID NO: 4, SEQ ID NO: 5, or SEQ ID NO: 6.

[0338] 20. The fusion polypeptide of any one of embodiments 1 to 19, comprising or consisting of the amino acid sequence of SEQ ID NO: 1.

[0339] 21. The fusion polypeptide of any one of embodiments 1 to 19, comprising or consisting of the amino acid sequence of SEQ ID NO: 2.

[0340] 22. The fusion polypeptide of any one of embodiments 1 to 19, comprising or consisting of the amino acid sequence of SEQ ID NO: 3.

[0341] 23. A nucleic acid comprising a nucleotide sequence encoding the fusion polypeptide of any one of embodiments 1 to 22.

[0342] 24. The nucleic acid of embodiment 23, wherein the nucleotide sequence encoding the fusion polypeptide of any one of embodiments 1 to 22 is operably linked to a promoter.

[0343] 25. The nucleic acid of embodiment 23 or embodiment 24, wherein the nucleic acid further encodes a guide RNA.

[0344] 26. The nucleic acid of embodiment 25, wherein the guide RNA comprises any one of the nucleotide sequences listed in Table 6.

[0345] 27. The nucleic acid of embodiment 25 or embodiment 26, wherein the guide RNA is capable of targeting a PAM sequence listed in Table 7.

[0346] 28. The nucleic acid of any one of embodiments 25 to 27, wherein the guide RNA comprises any one of the nucleotide sequences listed in Table 5.

[0347] 29. The nucleic acid according to any one of embodiments 23 to 28, which is in the form of a vector.

[0348] 30. The nucleic acid of embodiment 29, wherein the vector is an expression vector.

[0349] 31. The nucleic acid of embodiment 30, wherein the expression vector is a production vector.

[0350] 32. The nucleic acid of embodiment 30, wherein the expression vector is a delivery vector.

[0351] 33. The nucleic acid of any one of embodiments 29 to 32, wherein the vector is an RNA vector.

[0352] 34. The nucleic acid of any one of embodiments 29 to 32, wherein the vector is a DNA vector.

[0353] 35. The nucleic acid of embodiment 34, wherein the DNA vector is a plasmid.

[0354] 36. A cell comprising the nucleic acid of any one of embodiments 23 to 35.

[0355] 37. A cell engineered to express a nucleotide sequence encoding the fusion polypeptide of any one of embodiments 1 to 22.

[0356] 38. The cell of embodiment 36 or embodiment 37, which is a eukaryotic cell.

[0357] 39. The cell of embodiment 38, which is an insect cell.

[0358] 40. The cell of embodiment 38, which is a mammalian cell.

[0359] 41. A method of producing the fusion polypeptide of any one of embodiments 1 to 22, comprising culturing the cell of any one of embodiments 36 to 40 under conditions such that the fusion polypeptide is produced.

[0360] 42. The method of embodiment 41, further comprising isolating and / or purifying the fusion polypeptide.

[0361] 43. A composition comprising:

[0362] (a) the fusion polypeptide according to any one of embodiments 1 to 22; and

[0363] (b) Guide RNA.

[0364] 44. The composition of embodiment 43, wherein the guide RNA comprises any one of the nucleotide sequences listed in Table 6.

[0365] 45. The composition of embodiment 43 or embodiment 44, wherein the guide RNA is capable of targeting a PAM sequence listed in Table 7.

[0366] 46. ​​The composition of any one of embodiments 43 to 45, wherein the guide RNA comprises any one of the nucleotide sequences listed in Table 5.

[0367] 47. The composition according to any one of embodiments 43-46, which is a ribonucleoprotein complex.

[0368] 48. A method of editing the genome of a cell, comprising introducing into the cell:

[0369] (a) the fusion polypeptide according to any one of embodiments 1 to 22; and

[0370] (b) Guide RNA.

[0371] 49. A method of editing the genome of a cell, comprising introducing into the cell one or more nucleic acids encoding:

[0372] (a) the fusion polypeptide according to any one of embodiments 1 to 22; and

[0373] (b) guide RNA,

[0374] Optionally, at least one of the one or more nucleic acids is a nucleic acid according to any one of embodiments 23 to 35.

[0375] 50. The method of embodiment 48 or embodiment 49, wherein the guide RNA comprises any one of the nucleotide sequences listed in Table 6.

[0376] 51. The method of any one of embodiments 48 to 50, wherein the guide RNA is capable of targeting a PAM sequence listed in Table 7.

[0377] 52. The method of any one of embodiments 48 to 51, wherein the guide RNA comprises any one of the nucleotide sequences listed in Table 5.

[0378] 53. A method of editing the genome of a cell, comprising introducing the composition of any one of embodiments 43 to 47 into the cell.

[0379] 54. The method of any one of embodiments 48 to 53, wherein the cell is a stem cell.

[0380] 55. The method of embodiment 54, wherein the stem cells are induced pluripotent stem cells (iPSCs).

[0381] 56. The method of any one of embodiments 48 to 55, wherein the stem cells are mammalian stem cells.

[0382] 57. A method according to embodiment 56, wherein the mammalian stem cells are human stem cells.

[0383] 58. A stem cell comprising:

[0384] (a) a composition according to any one of embodiments 43 to 44; or

[0385] (b) The nucleic acid according to any one of embodiments 23 to 35.

[0386] 59. The stem cell according to embodiment 58, which is an induced pluripotent stem cell (iPSC).

[0387] 60. The stem cell of embodiment 58 or embodiment 59, which is a mammalian stem cell.

[0388] 61. The stem cell according to embodiment 60, which is a human stem cell.

[0389] 8. Examples

[0390] 8.1 Example 1: Effect of NLS sequence and orientation on nuclear translocation of B-GEn endonuclease

[0391] Developed a kind of test of screening B-GEn nuclease-nuclear localization signal (NLS) fusion, to identify the best NLS sequence and direction of B-GEn nuclease effective nuclear translocation.NLS is a short positively charged peptide that mediates the translocation of proteins from cytoplasm to nucleus via nuclear pore complex.Therefore, endonuclease coding sequence is fused to one or more NLS sequences and can promote their delivery to nucleus.Usually, endonuclease can be labeled at N- or C-terminus, two ends and certain interdomain regions that allow NLS to bind to its cognate receptor in nuclear pore complex.

[0392] Various NLS sequences have been reported in the literature, non-limiting examples of which are listed in Table 3 of Section 6.4. Figure 2The test set shown in is used to screen the NLS sequence fused to the N-terminus or C-terminus of Cas endonuclease.In short, the plasmid containing the B-GEn sequence fused to the NLS sequence at its N-terminus or C-terminus, downstream reporter gene sequence (such as GFP sequence) and sgRNA sequence is transfected into HEK293-T cells using a standard protocol. Cells were harvested 72 hours after transfection. A portion of the harvested cells is used to confirm the transfection efficiency by evaluating the expression of the reporter gene. The remaining harvested cells are used to prepare DNA extracts to test the gene editing percentage at the targeted locus. A single NLS sequence configuration is evaluated in the N-terminus and C-terminus directions relative to the B-GEn sequence.

[0393] 8.2 Example 2: Design of B-GEn.2 variant constructs

[0394] A study was conducted to compare B-GEn.2 constructs containing two NLS sequences flanking the B-GEn.2 sequence: a nucleoplasmin NLS at the N-terminus of B-GEn.2 and an SV40 NLS at the C-terminus of B-GEn.2 ( Figure 1A -1 and 3A) with variants lacking the N-terminal NLS, the C-terminal NLS, or both NLSs (Figures 1A2-1A4). Deletion of the N-terminal NLS resulted in construct B-GEn.2Ndel, which included the SV40 NLS only at its C-terminus ( Figure 3B ). Deletion of the NLS at the C-terminus resulted in construct B-GEn.2Cdel, which includes the nucleoplasmin NLS only at its N-terminus ( Figure 3C ). Deletion of both NLS sequences resulted in construct B-GEn.2 all-del, which has no NLS sequences.

[0395] 8.3 Example 3: Expression and Purification of B-GEn.2 Variants

[0396] Figure 4 The expression and purification workflow of B-GEn.2 variants is shown in Figure 2. In short, the DNA sequences of all B-GEn.2 variants mentioned in Section 8.1 were synthesized and cloned into a custom expression vector (pBLR107) based on pET29 using NdeI and XhoI. All constructs were expressed in BL21 (DE3) E. coli cells by chemically transforming E. coli with the encoding plasmid and plating the transformed cells on antibiotic selection LB plates. Successful colonies were scraped and transferred to 250 mL MagicMedia TMEscherichia coli expression medium (Thermo Scientific) with an initial absorbance of 0.01 at 600 nm. The culture was first grown at 37 ° C for 4 hours and then at 16 ° C for 40 hours. The cells were harvested and precipitated by centrifugation at 5000 x g for 15 minutes at 4 ° C and resuspended in lysis buffer (25 mM HEPES; 500 mM NaCl, 5% glycerol, 0.5 mM TCEP and a complete protease inhibitor cocktail from Sigma). The resuspended cells were then lysed 3-4 times by a microfluidizer LM-10 (Microfluidics). The cell lysate was cleared by centrifugation at 50,000 x g for 30 minutes at 4 ° C.

[0397] All proteins were purified using a two-step purification scheme of heparin fast flow chromatography followed by size exclusion chromatography (Figure 5). In the first step, the clarified lysate was loaded onto a heparin FF column (Cytiva) attached to an Akta Avant FPLC system (Cytiva) and the protein was eluted using a salt gradient. Figure 5A The results of heparin FF chromatography are shown, where each peak corresponds to a single labeled B-GEn.2 variant construct. The peak labeled with a single asterisk corresponds to the B-GEn.2 construct without the NLS (B-GEn.2 holo-del); the peak labeled with a double asterisk corresponds to B-GEn.2 Ndel; and the peak labeled with a triple asterisk corresponds to B-GEn.2-C-del ( Figure 5A ). The SDS-PAGE images confirmed that the protein fraction corresponding to the marked peak indeed mainly contained the target B-GEn.2 variant with a relatively low contamination level ( Figure 5B and 5C ).

[0398] For the second step of the purification scheme, Figure 5A The protein fractions marked with an asterisk in -C were pooled and injected onto a Superdex 200 size exclusion column (Cytiva) for a size-based separation step. Again, the proteins below the monomer peak were pooled and concentrated to 20 mg / mL, and protein aliquots were kept frozen at -80°C until further use. Figure 5D The size exclusion chromatography purification results of B-GEn.2 holo-del are shown, with the monomer peak marked with a single asterisk. SDS-PAGE of the protein fractions obtained from the size exclusion purification step indicated further concentration and purification of the construct ( Figure 5E ).

[0399] 8.4 Example 4: Intracellular Editing of the B2M Locus in iPSCs

[0400] 8.4.1 Gene Editing Workflow

[0401] The general workflow for gene editing using B-GEn.2 variants is as follows: Figure 6 Briefly, iPSCs were cultured in substrate-coated T75 flasks with essential 8 (E8) growth medium and maintained at 37°C with 5% CO2 between passages. Passaging was performed at 75% to 80% confluency. On the day of nucleofection, cells were incubated at 37°C for 10 minutes, followed by quenching with an equal volume of E8 medium and cleaving with Accutase. TM Cell separation solution (Stem Cell Technologies) separates iPSC from flask. Cell pellets were harvested by centrifugation at 115 x g for 3 minutes and then resuspended in Lonza P3 primary cell nuclear transfection buffer. For each nuclease construct, ribonucleoprotein (RNP) was assembled with sgRNA using a 1:2 ratio of protein:sgRNA (IDT). The complex RNP nuclear transfection was completed into the resuspended iPSC of P3 using a LONZA 4D nuclear transfection device. The nuclear transfected cells were then plated in substrate-coated 24-well Falcon flat-bottom plates (Corning) at 150,000 cells / well and grown in E8 growth medium (with Rock inhibitor Y-27632 (Tocris)) for 72 to 96 hours. After harvest, some cells were stained for flow cytometry using an anti-B2M APC-conjugated antibody from BioLegend (for B2M targeting experiments only). The remaining cells were resuspended in 30 μL of lysis buffer from the singleshot cell lysis kit from BioRad, and crude gDNA extraction was completed by incubation at room temperature for 10 minutes, then at 37°C for 5 minutes, followed by inactivation with proteinase K at 75°C for 5 minutes. 1 μL of crude gDNA extract was used directly for the first step of amplicon sequencing per 25 μL PCR reaction, followed by end preparation, indexing, and sequencing on an Illumina MiSeq sequencer.

[0402] 8.4.2 Results

[0403] RNPs (1151) containing the v4.3*sgRNA and the B-GEn.2 construct flanked by NLS at the N- and C-termini showed low levels of gene editing, with an editing percentage of 5.4 when 25 pmol of RNP was used and an editing percentage of 6.83 when 50 pmol of RNP was used ( Figure 7This low level of gene editing was likely not due to suboptimal conditions or experimental inefficiency, as indicated by the relatively high level of gene editing achieved by RNPs made from Cpf1 / Cas12a and its B2M-targeting sgRNA, which showed an editing percentage of 68.70 when 25 pmol of RNP was used and 74.48 when 50 pmol of RNP was used ( Figure 7 ). The deletion of the C-terminal NLS further reduced the editing efficiency of the B-GEn.2Cdel construct to a percentage editing level that was indistinguishable from the control without Cas nuclease (percent editing was 0.93 and 1.41 for 25 and 50 pmol of RNP, respectively). This reduction is likely due to the failure of nuclear localization of the N-terminal nucleoplasmin NLS of B-GEn.2Cdel. The percentage editing of RNPs with B-GEn.2Ndel was 17.74 at 25 pmol and 25.09 at 50 pmol. Therefore, the deletion of the N-terminal NLS improved gene editing by approximately 3.5-fold relative to the editing achieved by B-GEn.2 flanked by NLS on both sides ( Figure 7 ).

[0404] 8.5 Example 5: Intracellular Editing of the Albumin Locus in iPSCs

[0405] Using the same gene editing workflow described in Section 8.4.1, evaluate B-GEn.2 constructs for editing the albumin locus in iPSCs. This time, for each nuclease construct, use a 1:2 ratio of protein:sgRNA and assemble RNPs with either v4.3* or v4.4* sgRNA.

[0406] RNPs containing v4.3*sgRNA and B-GEn.2 with NLS at both ends showed low levels of gene editing, with an editing percentage of 2.16 for 25 pmol of RNP and 2.79 for 50 pmol of RNP ( Figure 8A ). Deletion of the C-terminal NLS further reduced the editing efficiency of the B-GEn.2Cdel construct to an editing percentage level (0.17 and 0.51 for 25 and 50 pmol of RNP, respectively), which was lower than the editing percentage level obtained with the negative control (editing percentage 1.29). However, deletion of the N-terminal NLS enhanced gene editing. This time, the editing percentage of RNP with B-GEn.2Ndel was 7.30 at 25 pmol and 11.45 at 50 pmol ( Figure 8AConsistent with the observations in Section 8.4.2, deletion of the N-terminal NLS increased gene editing at the albumin locus by 3.3 to 4.1 fold relative to editing by B-GEn.2 flanked by NLSs ( Figure 8A ). Similar results were obtained when v4.4* sgRNA was used instead of v4.3*. In this case, B-GEn.2 with NLS at both ends showed an editing percentage of 1.58 for 25 pmol of RNP and 3.14 for 50 pmol of RNP, and the deletion of the C-terminal NLS further reduced the editing efficiency of the B-GEn.2Cdel construct to: 0.62 for 25 pmol and 0.68 for 50 pmol ( Figure 8B N-terminal deletion of the NLS enhanced gene editing in the range of 3- and 4-fold, resulting in an editing percentage of 6.34 at 25 pmol and 9.56 at 50 pmol ( Figure 8B ).

[0407] 8.6 Example 6: Intracellular Editing of the Albumin Locus in HEK293-T Cells

[0408] A gene editing workflow similar to the protocol described in Section 8.4.1 was used with the following differences: HEK 293-T cells (ATCC CRL-3216 TM HEK293T cells were cultured in DMEM (ATCC) growth medium supplemented with 10% FBS and Pen / Strep, and the plates / flasks were not coated with substrate (unlike iPSCs). Nucleofection of complex RNPs in HEK293T cells was performed similarly to Section 8.4.1.

[0409] Gene editing of B-GEn.2 variant constructs was evaluated using 25 to 310 pmol of RNPs containing v4.3*sgRNA. The negative control was associated with a gene editing percentage of 0.18. At each RNP concentration evaluated, B-GEn.2 (#1) flanked by NLS at both ends used in the previous examples performed similarly to B-GEn.2 (#2) with the same NLS configuration (B-GEn.2 (#1) gene editing percentage: 2.53 for 25 pmol, 4.8 for 50 pmol, 5.67 for 155 pmol, and 12.71 for 310 pmol; B-GEn.2 (#2) gene editing percentage: 2.1 for 25 pmol, 4.54 for 50 pmol, 6.26 for 155 pmol, and 15.99 for 310 pmol) ( Figure 9). Deletion of the C-terminal NLS further reduced the editing efficiency of the B-GEn.2Cdel construct relative to the editing achieved by B-GEn.2 flanked by NLSs at both ends (0.96 for 25 pmol, 2.76 for 50 pmol, 4.24 for 155 pmol, and 8.92 for 310 pmol). However, when compared to the editing achieved by B-GEn.2 flanked by NLSs at both ends, deletion of the N-terminus resulted in approximately 3.5-fold higher editing percentages for all tested concentrations (10.23 for 25 pmol, 18.16 for 50 pmol, 24.72 for 155 pmol, and 40.52 for 310 pmol) ( Figure 9 ).

[0410] 9. Incorporation by Reference

[0411] All publications, patents, patent applications, and other documents cited in this application are hereby incorporated by reference in their entirety for all purposes to the same extent as if each individual publication, patent, patent application, or other document were individually indicated to be incorporated by reference for all purposes. In the event of any inconsistency between the teachings of one or more references incorporated herein and the present disclosure, the teachings of this specification shall control.

Claims

1. A fusion polypeptide comprising: (a) a nuclease sequence which is an amino acid sequence having at least 80% sequence identity to SEQ ID NO: 6 (B-GEn. 2), SEQ ID NO: 5 (B-GEn. 1.2), or SEQ ID NO: 4 (B-GEn. 1); and (b) a first nuclear localization signal ("NLS") sequence located at the C-terminus of the nuclease sequence, The fusion polypeptide lacks an NLS sequence located at the N-terminus of the nuclease sequence.

2. The fusion polypeptide according to claim 1, further comprising a first linker sequence between the nuclease sequence and the first NLS sequence.

3. The fusion polypeptide according to claim 1 or claim 2, further comprising a second NLS sequence located at the C-terminus of the first NLS sequence, optionally comprising a linker sequence between the first NLS sequence and the second NLS sequence.

4. The fusion polypeptide according to claim 3, further comprising a third NLS sequence located at the C-terminus of the second NLS sequence, and optionally comprising a linker sequence between the second NLS sequence and the third NLS sequence.

5. The fusion polypeptide according to any one of claims 1 to 4, wherein the nuclease sequence comprises an amino acid sequence that differs from the amino acid sequence of SEQ ID NO: 6 (B-GEn. 2) by no more than 25 amino acids.

6. The fusion polypeptide according to any one of claims 1 to 4, wherein the nuclease sequence comprises an amino acid sequence that differs from the amino acid sequence of SEQ ID NO: 5 (B-GEn. 1.2) by no more than 25 amino acids.

7. The fusion polypeptide according to any one of claims 1 to 4, wherein the nuclease sequence comprises an amino acid sequence that differs from the amino acid sequence of SEQ ID NO: 4 (B-GEn. 1) by no more than 25 amino acids.

8. The fusion polypeptide according to any one of claims 1 to 4, comprising or consisting of the amino acid sequence of SEQ ID NO:

2.

9. A nucleic acid comprising a nucleotide sequence encoding the fusion polypeptide according to any one of claims 1 to 8.

10. The nucleic acid of claim 9, wherein the nucleotide sequence encoding the fusion polypeptide of any one of claims 1 to 8 is operably linked to a promoter.

11. The nucleic acid of claim 9 or claim 10, wherein the nucleic acid further encodes a guide RNA.

12. The nucleic acid according to claim 10 or claim 11 in the form of a vector.

13. The nucleic acid according to claim 12, wherein the vector is an expression vector. The nucleic acid according to claim 13 , wherein the expression vector is a production vector.

15. The nucleic acid of claim 13, wherein the expression vector is a delivery vector.

16. The nucleic acid according to any one of claims 12 to 15, wherein the vector is an RNA vector.

17. The nucleic acid according to any one of claims 12 to 15, wherein the vector is a DNA vector.

18. A cell engineered to express a nucleotide sequence encoding the fusion polypeptide of any one of claims 1 to 8 or comprising the nucleic acid of any one of claims 9 to 17. The cell according to claim 18 , which is a stem cell.

20. The cell of claim 18 or claim 19, wherein the stem cell is a human stem cell, optionally wherein the human stem cell is an induced pluripotent stem cell (iPSC).

21. A method of producing the fusion polypeptide of any one of claims 1 to 8, comprising culturing the cell of claim 18 or claim 19 under conditions whereby the fusion polypeptide is produced.

22. The method of claim 21, further comprising isolating and / or purifying the fusion polypeptide.

23. A composition comprising: (a) a fusion polypeptide according to any one of claims 1 to 8; and (b) Guide RNA. The composition according to claim 23 , which is a ribonucleoprotein complex.

25. A method for editing the genome of a cell, comprising introducing into the cell: (a)(i) the fusion polypeptide of any one of claims 1 to 8; and (ii) guide RNA; (b) one or more nucleic acids, optionally in the form of a vector according to claim 15, encoding: (i) a fusion polypeptide according to any one of claims 1 to 8; and (ii) guide RNA, (c) a composition according to claim 23 or claim 24; or (d) Any combination of two or all three of (a) to (c).

26. The method of claim 25, wherein the cells are stem cells.

27. The method of claim 25 or claim 26, wherein the cells are human stem cells, optionally wherein the human stem cells are induced pluripotent stem cells (iPSCs).

28. A stem cell comprising: (a) a composition according to claim 23 or claim 24; or (b) The nucleic acid according to any one of claims 9 to 17.

29. The stem cell according to claim 28, which is a human stem cell, optionally wherein the human stem cell is an induced pluripotent stem cell (iPSC).

Citation Information

Patent Citations

  • Connector assemblies and associated methods

    US11241132B2

  • Nuclease resistant chimeric oligonucleotides

    US20030158403A1

  • Nuclear reprogramming factor and induced pluripotent stem cells

    US20090047263A1

  • Nuclear Reprogramming Factor

    US20090068742A1

  • Multipotent / pluripotent cells and methods

    US20090191159A1