Compositions and methods for delivering transgenes
By using a small-sized CRISPR-Cas nuclease and guide RNA to target the albumin locus, the inaccuracy and low integration efficiency of genome editing in existing technologies have been solved, achieving efficient and stable GOI integration and expression, thus improving the therapeutic effect of gene therapy.
Patent Information
- Application Number
- CN202480086264.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-20
- Filing Date
- 2024-12-19
- Publication Date
- 2026-08-25
AI Technical Summary
Existing genome editing technologies suffer from problems such as inaccuracy, high off-target editing levels, reduced therapeutic effects due to vector dilution, low GOI integration efficiency, and insufficient stability. In particular, when using AAV vectors, it is difficult to effectively integrate and express transgenes.
By using a small-sized CRISPR Cas nuclease (B-Gen.1) and a polypeptide with high sequence identity, combined with guide RNA and donor nucleic acid, the GOI is specifically targeted at the albumin locus to achieve efficient and stable integration and expression.
The integration and expression of highly active GOIs were achieved in vitro and in vivo, reducing off-target events and improving the stability and safety of therapeutic effects.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This disclosure provides compositions, methods, and systems for the targeted delivery of nucleic acids to target cells (e.g., human cells). Some embodiments of the invention relate to compositions, methods, and systems for expressing transgenes in cells via genome editing.
[0002] References to sequence lists are included This application contains a sequence list, which has been submitted via Online Filing 2.0. The sequence list, named “BHC231038-FC-PCT.xml”, was created on December 16, 2024, and its full text is incorporated herein by reference. Background Technology
[0003] Recent advances in genome sequencing technologies and analytical methods have significantly accelerated the ability to classify and map genetic factors associated with various biological functions and diseases. Precise genome-targeting technologies are needed to enable systematic reverse engineering of causal genetic variations through selective perturbation of individual genetic elements, and to advance synthetic biology, biotechnology, and medical applications.
[0004] Recently, gene editing using designed site-specific nucleases has emerged as a technology for basic biomedical research and therapeutic development. In recent years, various platforms based on four main types of endonucleases have been developed for gene editing: broad-spectrum nucleases and their derivatives, zinc finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), and CRISPR-associated endonuclease 9 (CAS9). Each type of nuclease can induce DNA double-strand breaks (DSBs) at specific DNA loci, thereby triggering two DNA repair pathways. The non-homologous end joining (NHEJ) pathway generates random insertion / deletion (indel) mutations at the DSB, while the homologous directed repair (HDR) pathway repairs the DSB using genetic information carried on the donor nucleic acid. In the context of this invention, the term donor nucleic acid may be used interchangeably with the term donor template.
[0005] Therefore, these gene editing platforms can manipulate genes at specific loci in a variety of ways, such as disrupting gene function, repairing mutated genes to normal, and inserting DNA material.
[0006] However, currently available genome editing technologies may be affected by some shortcomings, such as inaccuracy and unacceptable levels of off-target editing.
[0007] Therefore, there is still a need for new genome editing platforms that can manipulate genes at specific loci in multiple ways, such as disrupting gene function, restoring mutated genes to normal, and / or inserting heterologous DNA material at specific loci (e.g., safe harbor loci within the target cell genome).
[0008] Cell and gene therapies offer unprecedented opportunities for treating incurable diseases. Particularly in pediatric patients, standard gene therapy using attached AAV vectors is hampered by organ growth. This leads to vector dilution, reducing therapeutic efficacy. To mitigate this limitation, stable gene modifications are needed for gene therapy applications.
[0009] Albumin is an attractive target for gene modification and gene therapy because it exhibits high liver-specific expression and provides an intronic safe harbor platform for the stable integration and expression of therapeutic transgenes. RNA-programmable nucleases offer a highly dynamic and versatile tool for targeted modification of genomic sequences. These nucleases are able to ensure the integration of selected templates into the framework by cleaving the target site, a prerequisite for integration.
[0010] WO2020 / 082042 discloses a method based on Streptococcus pyogenes (… S.pyogenes Methods and reagents for integrating transgenes into albumin loci using CRISPR Cas9 nucleases. WO2020 / 081843 also discloses such systems based on various Cas9 (type II) CRISPR nucleases.
[0011] WO2022258753 discloses the type V CRISPR CAS nuclease used in this invention, provided that the nuclease named B-Gen.1 (SEQ ID NO: 185) in this application is named B-Gen.2 (SEQ ID NO: 2) in WO2022258753.
[0012] To date, high-opportunity target albumin has only been associated with Bacteroides pyogenes (…). B. pyogenes The Cas9 affinity of [the CRISPR Cas] is significant. However, there is a need to apply this principle to significantly smaller CRISPR Cas effectors that can also possess additional beneficial properties compared to existing systems, such as higher nuclease activity, higher selectivity, or lower PAM complexity requirements. The recently developed V-type B-Gen-1 and 2 CRISPR Cas nucleases (WO2022258753) provide small-sized effectors with low PAM complexity requirements.
[0013] WO2023 / 220649 discloses the use of certain type V nucleases for integrating GAA into exon 1 of the albumin locus.
[0014] However, existing compositions, systems, and methods have several drawbacks.
[0015] a) Nuclease effectors are very large enzymes that are difficult or impossible to package into certain viral vectors, such as AAV.
[0016] b) The integration efficiency of the target gene (GOI) is low.
[0017] c) The stable integration efficiency of the target gene (GOI) is low.
[0018] d) In vivo, integrating GOI into the genome of target cells is not very efficient and / or stable.
[0019] e) It exhibits too many off-target events, posing a risk to patients.
[0020] f) It may be effective for one GOI but not for another. Summary of the Invention
[0021] The methods and compositions provided herein allow for the use of small-sized CRISPR-Cas nucleases (B-Gen.1 (SEQ ID NO: 185 or SEQ NO: 186) and polypeptides having at least 95% sequence identity with SEQ ID NO: 185 and 186) for the integration of a target gene (GOI) into a safe harbor locus, preferably into an albumin locus. Another aspect of the compositions and methods according to the invention is the ability to obtain unexpectedly high-activity GOIs both in vitro and in vivo.
[0022] One aspect of the present invention is a system comprising: (a) a deoxyribonuclease (DNA) or a nucleic acid encoding the DNA endonuclease, selected from SEQ ID NO: 185 or SEQ ID NO: 186, or any sequence thereof having 95%, preferably 98%, more preferably 99% identity with these sequences; and (b) a guide RNA (gRNA) or nucleic acid encoding said gRNA, said gRNA comprising a spacer region sequence selected from any one of SEQ ID NOs: 1, 2, 3, 4, 5, 6, 7, 8, 9, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 35, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 51, 52, 54, 55, 56, 58, 59, 60; and (c) Donor nucleic acid, which includes a nucleic acid sequence encoding a target gene (GOI) or a functional derivative thereof.
[0023] Another aspect of the invention is a system comprising: (a) a deoxyribonuclease (DNA) or a nucleic acid encoding the DNA endonuclease, selected from SEQ ID NO: 185 or SEQ ID NO: 186, or any sequence thereof having 95%, preferably 98%, more preferably 99% identity with these sequences; and (b) a guide RNA (gRNA) selected from any one of SEQ ID NOs: 62-122, or a nucleic acid encoding said gRNA, or any sequence thereof having 95%, preferably 98%, more preferably 99% identity with these sequences; and (c) Donor nucleic acid, which includes a nucleic acid sequence encoding a target gene (GOI) or a functional derivative thereof.
[0024] Another aspect of the present invention is a method for editing the genome in a cell, the method comprising: providing the cell with the following substance: (a) a deoxyribonuclease (DNA) or a nucleic acid encoding the DNA endonuclease, selected from SEQ ID NO: 185 or SEQ ID NO: 186, or any sequence thereof having 95%, preferably 98%, more preferably 99% identity with these sequences; and (b) A guide RNA (gRNA) selected from any one of SEQ ID NOs: 62-122, or a nucleic acid encoding said gRNA, or any sequence thereof having 95%, preferably 98%, more preferably 99% identity with these sequences; or gRNA, comprising a spacer region sequence selected from any one of SEQ ID NOs: 1, 2, 3, 4, 5, 6, 7, 8, 9, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 35, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 51, 52, 54, 55, 56, 58, 59, 60; and (c) Donor nucleic acid, which includes a nucleic acid sequence encoding a target gene (GOI) or a functional derivative thereof.
[0025] This document specifically provides compositions, methods, and systems for targeted delivery of nucleic acids (including DNA and RNA) to target cells (e.g., human cells). It also provides compositions, methods, and systems for targeted integration of transgenes into the cellular genome via genome editing, thereby expressing transgenes in cells. Certain aspects and embodiments of this disclosure relate to compositions, methods, and systems for knocking a target gene (GOI) into a specific safe harbor location in the genome, particularly into a genomic location inside or around an endogenous albumin locus. Compositions, methods, and systems are also provided for treating subjects with or suspected of having a disease or health condition using in vitro and / or in vivo genome editing.
[0026] In one aspect, this article provides a guide RNA (gRNA) sequence having a sequence complementary to a genomic sequence inside or around an endogenous albumin locus.
[0027] In some embodiments, the gRNA has a sequence selected from SEQ ID NOs: 62-122 and variants thereof having at least 85%, preferably 90%, more preferably 95%, or even more preferably 98% identity with any of these sequences.
[0028] In another aspect, this article provides a composition having any of the above-mentioned gRNAs.
[0029] In one aspect, this document provides a system comprising a guide RNA (gRNA) or a nucleic acid encoding said gRNA, said guide RNA comprising the spacer regions disclosed herein. In some embodiments, the gRNA of said system has a sequence selected from SEQ ID Nos: 62 to 122 and variants thereof having at least 85%, preferably 90%, more preferably 95%, and even more preferably 98% identity with any of these sequences.
[0030] In some preferred embodiments, the gRNA includes the preferred, more preferred, particularly preferred, more particularly preferred, or most particularly preferred guide RNA (gRNA) sequences listed in Table 1. Attached Figure Description
[0031] Figure 1 This study demonstrates the insertion / deletion activity of 61 B-GEn.1 gRNAs targeting the human albumin intron 1 locus. Experiments were performed on HEK293T cells. The figure shows the mean standard deviation of two experiments performed in three technical replicates.
[0032] Figure 2aThe RLU activity of seven synthetic gRNAs of B-GEn.1 targeting the human albumin intron 1 locus is shown. Experiments were performed on primary human hepatocyte donor HUM201801 cells. The figure shows the mean of the median range for one experiment performed in three technical replicates.
[0033] Figure 2b This study illustrates the insertion / deletion activities of seven synthetic gRNAs of B-GEn.1 targeting the human albumin intron 1 locus. Experiments were performed on primary human hepatocyte donor HUM201801 cells. The figure shows the mean of the median range for one experiment performed in three technical replicates.
[0034] Figure 3 This study demonstrates the insertion / deletion activity of 53 gRNAs of B-GEn.1 targeting the mouse albumin intron 1 locus. Experiments were performed on the mouse cell line HEPA1-6. The figure shows the mean standard deviation of one experiment performed in three technical replicates.
[0035] Figure 4 Editing and integration of the B-GEn.1 gene in mice are shown. a) Insertion / deletion activity of 10 B-GEn.1 gRNAs targeting mouse albumin intron 1 locus; b) RLU values at doses of 1 mpk and 0.75 mpk in the same experiment. The figure shows the mean of the standard deviations of one experiment performed in four technical replicates. Detailed Implementation
[0036] Exemplary Implementation In some embodiments, the system further comprises one or more of the following: a deoxyribonuclease (DNA) endonuclease or a nucleic acid encoding the DNA endonuclease; and a donor nucleic acid having a nucleic acid sequence encoding a target gene (GOI) or a functional derivative thereof.
[0037] In some embodiments, the gRNA includes a spacer sequence selected from any one of SEQ ID NOs: 1 to 61. In some preferred embodiments, the gRNA includes the preferred, more preferred, particularly preferred, more particularly preferred, or most particularly preferred spacer sequences listed in Table 2.
[0038] In some embodiments, the system includes a) a DNA endonuclease according to SEQ ID NO: 185 or 186, or any sequence thereof having at least 95% identity with such sequences; and b) a guide RNA comprising a spacer region sequence selected from any one of SEQ ID Nos: 1, 2, 3, 4, 5, 6, 7, 8, 9, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 35, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 51, 52, 54, 55, 56, 58, 59, 60; and c) a donor nucleic acid comprising a nucleic acid sequence encoding a target gene (GOI) or a functional derivative thereof.
[0039] In some embodiments, the system includes a) a DNA endonuclease according to SEQ ID NO: 185 or 186, or any sequence thereof having at least 95% identity with such sequences; and b) a guide RNA comprising a spacer sequence selected from any one of SEQ ID Nos: 3, 4, 5, 6, 9, 11, 12, 13, 14, 15, 16, 20, 21, 22, 25, 26, 32, 39, 42, 44, 46, 51, 52, 54, 55, 56, 58, 59, 60; and c) a donor nucleic acid comprising a nucleic acid sequence encoding a target gene (GOI) or a functional derivative thereof.
[0040] In some embodiments, the system includes a) a DNA endonuclease according to SEQ ID NO: 185 or 186 or any sequence thereof having at least 95% identity with such sequences; and b) a guide RNA comprising a spacer region sequence selected from any one of SEQ ID Nos: 4, 5, 6, 9, 12, 13, 14, 15, 20, 26, 32, 39, 42, 52, 54, 55, 58, 59; and c) a donor nucleic acid comprising a nucleic acid sequence encoding a target gene (GOI) or a functional derivative thereof.
[0041] In some embodiments, the system includes a) a DNA endonuclease according to SEQ ID NO: 185 or 186 or any sequence thereof having at least 95% identity with such sequences; and b) a guide RNA comprising a spacer sequence selected from any one of SEQ ID Nos: 4, 5, 6, 9, 15, 20, 42, 52, 59; and c) a donor nucleic acid comprising a nucleic acid sequence encoding a target gene (GOI) or a functional derivative thereof.
[0042] In some embodiments, the system includes a) a DNA endonuclease according to SEQ ID NO: 185 or 186 or any sequence thereof having at least 95% identity with such sequences; b) a guide RNA comprising a spacer sequence selected from any one of SEQ ID Nos: 4 or 42; and c) a donor nucleic acid comprising a nucleic acid sequence encoding a target gene (GOI) or a functional derivative thereof.
[0043] In some embodiments, the system includes a) a DNA endonuclease according to SEQ ID NO: 185 or 186; b) a guide RNA comprising a spacer sequence selected from any one of SEQ ID Nos: 4 or 42; and c) a donor nucleic acid comprising a nucleic acid sequence encoding a target gene (GOI) or a functional derivative thereof.
[0044] In some embodiments, the system includes a) a DNA endonuclease according to SEQ ID NO: 185 or 186 or any sequence thereof having at least 95% identity with such sequences; b) a guide RNA comprising the spacer sequence of SEQ ID No: 4; and c) a donor nucleic acid comprising a nucleic acid sequence encoding a target gene (GOI) or a functional derivative thereof.
[0045] (a) a deoxyribonuclease (DNA) or a nucleic acid encoding the DNA endonuclease, selected from SEQ ID NO: 185 or SEQ ID NO: 186, or any sequence thereof having 95%, preferably 98%, more preferably 99% identity with these sequences; and (b) A guide RNA selected from any one of SEQ ID NOs: 65 or 103, or a nucleic acid encoding the gRNA, or any sequence thereof having 95%, preferably 98%, more preferably 99% identity with these sequences; and (c) Donor nucleic acid, which includes a nucleic acid sequence encoding a target gene (GOI) or a functional derivative thereof.
[0046] In some embodiments, the system includes a) a DNA endonuclease according to SEQ ID NO: 185 or 186 or any sequence thereof having at least 95% identity with such sequences; b) a guide RNA comprising the spacer sequence of SEQ ID No: 42; and c) a donor nucleic acid comprising a nucleic acid sequence encoding a target gene (GOI) or a functional derivative thereof.
[0047] In some embodiments, the DNA endonuclease recognizes a prespacer adjacent motif (PAM) having the sequence “DTTN”, where “D” represents “A”, “T”, or “G”; preferably NGG or NNGG, where N is any nucleotide or its functional derivative.
[0048] In some embodiments, the deoxyribonuclease (DNA) endonuclease is a type V nuclease. In some embodiments, the DNA endonuclease is selected from SEQ ID NO: 185 or SEQ ID NO: 186, or any sequence thereof having 95%, preferably 98%, more preferably 99% identity with these sequences.
[0049] In some implementations, the nucleic acid encoding the DNA endonuclease is codon-optimized for expression in the host cell.
[0050] In some implementations, the nucleic acid sequence encoding the target gene (GOI) is codon-optimized for expression in host cells.
[0051] In some embodiments, the GOI encoding is a polypeptide selected from therapeutic and preventative polypeptides.
[0052] In some implementations, the GOI encodes α-glucosidase (GAA).
[0053] In some embodiments, the GOI encodes a protein selected from factor VIII (FVIII) protein, factor IX protein, α-1-antitrypsin, factor XIII (FXIII) protein, factor VII (FVII) protein, factor X (FX) protein, protein C, serine protease inhibitor Gl (Serpin Gl), or any functional derivative thereof. In some embodiments, the GOI encodes FVIII protein or a functional derivative thereof. In some embodiments, the GOI encodes FIX protein or a functional derivative thereof. In some embodiments, the GOI encodes serine Gl protein or a functional derivative thereof.
[0054] In some embodiments, the nucleic acid encoding the DNA endonuclease is a deoxyribonucleic acid (DNA) sequence.
[0055] In some embodiments, the nucleic acid encoding the DNA endonuclease is a ribonucleic acid (RNA) sequence.
[0056] In some embodiments, the RNA sequence encoding the DNA endonuclease is covalently linked to the gRNA.
[0057] In some embodiments, the composition further comprises liposomes or lipid nanoparticles.
[0058] In some implementations, the donor nucleic acid is encoded in a viral vector.
[0059] In some implementations, the donor nucleic acid is encoded in an adeno-associated virus (AAV) vector.
[0060] In some embodiments, the DNA endonuclease is formulated in the form of liposomes or lipid nanoparticles.
[0061] In some embodiments, the liposomes or lipid nanoparticles also include gRNA.
[0062] In some implementations, the nucleic acid encoding the effector protein is mRNA.
[0063] In some embodiments, the components of the composition (DNA endonuclease, guide RNA, and donor nucleic acid) are encoded by a single viral vector.
[0064] In some embodiments, the components of the composition (DNA endonuclease, guide RNA, and donor nucleic acid) are encoded by separate viral vectors, including embodiments in which one or two components are encoded by a single viral vector.
[0065] In some implementations, the components of the system are administered separately.
[0066] In some implementations, the components of the system are administered simultaneously.
[0067] In some embodiments, the DNA endonuclease is pre-complexed with gRNA to form a ribonucleoprotein (RNP) complex.
[0068] In another aspect, this document provides a kit having any of the above-described compositions and also having instructions for use.
[0069] In another aspect, this document provides a method for editing a genome in a cell. The method includes providing the cell with: (a) any gRNA or nucleic acid encoding said gRNA as described herein; (b) a deoxyribonuclease or nucleic acid encoding said DNA endonuclease; and (c) a donor nucleic acid having a nucleic acid sequence encoding a target gene (GOI) or a functional derivative thereof. In some embodiments, said gRNA comprises a spacer region sequence from any of SEQ ID Nos: 1-61 listed in Table 2.
[0070] In some embodiments, the gRNA has a variant selected from those sequences listed in Table 1 and that has at least 85% homology with those sequences listed in Table 1.
[0071] In some implementations, the nucleic acid sequence encoding the target gene (GOI) is codon-optimized.
[0072] In some embodiments, the GOI encoding is a polypeptide selected from therapeutic and preventative polypeptides.
[0073] In some implementations, the nucleic acid sequence encoding the target gene (GOI) is inserted into the cell's genome sequence.
[0074] In some embodiments, the insertion is located in, within, or near the albumin gene or albumin gene regulatory element in the cell's genome.
[0075] In some implementations, the insertion is located within the first intron of the albumin gene.
[0076] In some embodiments, the insertion is located at least 37 bp downstream of the end of the first exon of the human albumin gene in the genome and at least 330 bp upstream of the start of the second exon of the human albumin gene in the genome.
[0077] In some implementations, the nucleic acid sequence encoding the target gene is expressed under the control of an endogenous albumin promoter.
[0078] In some implementations, the cells are hepatocytes.
[0079] In another aspect, this article provides a genetically modified cell, wherein the genome of the cell is edited by any of the methods described above.
[0080] In some implementations, the nucleic acid sequence encoding the target gene is inserted into the cell's genome sequence.
[0081] In some embodiments, the insertion is located in, within, or near the albumin gene or albumin gene regulatory element in the cell's genome.
[0082] In some implementations, the insertion is located within the first intron of the albumin gene.
[0083] In some implementations, the nucleic acid sequence encoding the target gene is expressed under the control of an endogenous albumin promoter.
[0084] In some implementations, the nucleic acid sequence encoding the target gene is codon-optimized.
[0085] In some implementations, the cells are hepatocytes.
[0086] In another aspect, this article provides a method for treating a symptom or health condition in a subject. The method involves administering any of the aforementioned genetically modified cells to the subject.
[0087] In some implementations, the genetically modified cells are autologous.
[0088] In some embodiments, the method further includes obtaining a biological sample from a subject (where the biological sample contains hepatocytes) and editing the genome of the hepatocytes by inserting a nucleic acid sequence encoding a target gene into the genome sequence of the cells, thereby producing genetically modified cells.
[0089] In another aspect, this article provides a method for treating a condition or health status in a subject. The method includes obtaining a biological sample from the subject ( wherein the biological sample contains hepatocytes), providing the hepatocytes with: (a) any gRNA or nucleic acid encoding said gRNA as described above; (b) a deoxyribonuclease or nucleic acid encoding said DNA endonuclease; and (c) a donor nucleic acid having a nucleic acid sequence encoding a target gene (GOI) or a functional derivative thereof, thereby producing genetically modified cells, and administering said genetically modified cells to the subject.
[0090] Each aspect and implementation described herein can be used together unless explicitly or clearly excluded from the context of the implementation or aspect.
[0091] The foregoing description of the invention is merely exemplary and is not intended to be limiting in any way. Other aspects, embodiments, objects, and features of this disclosure will become fully apparent from the accompanying drawings, detailed description, and claims, in addition to the exemplary embodiments and features described herein.
[0092] definition Unless otherwise stated, all technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. Unless otherwise stated or clearly apparent from the context, the following terms have the following meanings: The terms "% identical (% identity, percent identity)" or their syntactic equivalents refer to the degree to which two sequences (nucleotides or amino acids) have identical residues at the same positions in an alignment. For example, "the amino acid sequence has X% identity with SEQ ID NO: Y" can refer to the % identity of the amino acid sequence with SEQ ID NO: Y, and is interpreted as X% of the residues in the amino acid sequence being identical to the residues in the sequence disclosed in SEQ ID NO: Y. Typically, this calculation can be performed using a computer program. Exemplary procedures for comparing and aligning sequence pairs include ALIGN (Myers and Miller, Comput Appl Biosci. 1988 Mar;4(l): 11-7), FASTA (Pearson and Lipman, Proc Natl Acad Sci US A. 1988 Apr;85(8):2444-8; Pearson, Methods Enzymol. 1990;183:63-98), and vacancy BLAST (Altschul et al., Nucleic Acids Res. 1997 Sep l;25(17):3389-40), BLASTP, BLASTN, or GCG (Devereux et al., Nucleic Acids Res. 1984 Jan 11;12(1 Pt l):387-95).
[0093] The terms “amplification” or their grammatical equivalents used in this article refer to the process of producing multiple nucleic acid molecules containing the same sequence as the original nucleic acid molecule or its distinguishable portions by enzymatic replication of a nucleic acid molecule.
[0094] As used herein, the term "base editing enzyme" refers to a protein, polypeptide, or fragment thereof capable of catalyzing the chemical modification of the nucleobases of deoxyribonucleotides or ribonucleotides. For example, such a base editing enzyme can catalyze reactions that modify nucleobases present in nucleic acid molecules such as DNA or RNA (single-stranded or double-stranded). Non-limiting examples of the types of modifications that base editing enzymes can catalyze include converting existing nucleobases to different nucleobases, such as converting cytosine to guanine or thymine, or adenine to guanine; hydrolytic deamination of adenine or adenosine; or methylation of cytosine (e.g., CpG, CpA, CpT, or CpC). Base editing enzymes may or may not bind to nucleic acid molecules containing nucleobases.
[0095] As used herein, the term "base editor" refers to a fusion protein comprising a base-editing enzyme fused to an effector protein. The base editor functions when the effector protein is coupled to a guide nucleic acid. The guide nucleic acid confers specific sequence activity to the base editor. As a non-limiting example, the effector protein may include a catalytically inactivating effector protein. Furthermore, as a non-limiting example, the base-editing enzyme may include deaminase activity. Additional base editors are described herein.
[0096] As used herein, the term "catalytically inactivating effector protein" refers to an effector protein modified relative to a naturally occurring effector protein, wherein the effector protein has reduced or eliminated catalytic activity relative to the naturally occurring effector protein, but retains its ability to interact with guide nucleic acids. The reduced or eliminated catalytic activity is typically nuclease activity. The naturally occurring effector protein can be a wild-type protein. In some embodiments, the catalytically inactivating effector protein is referred to as a catalytically inactivating variant of the effector protein, such as a Cas effector protein.
[0097] The term “cis-cleavage” as used in this article refers to the cleavage of a target nucleic acid by an effector protein complexed with the guide nucleic acid (hydrolysis of the phosphodiester bond), and specifically to the cleavage of a target nucleic acid that has hybridized with the guide nucleic acid, wherein the cleavage occurs within or directly adjacent to the region of the target nucleic acid that has hybridized with the guide nucleic acid.
[0098] The term "complementary" (or "complementarity") used in this article for nucleic acid molecules or sequences refers to the property of a polynucleotide to pair with its Watson-Crick counterpart (C to G; or A to T) in a reference nucleic acid. For example, a polynucleotide is 100% complementary to a reference nucleic acid when each nucleotide in the polynucleotide forms a base pair with the reference nucleic acid. In double-stranded DNA or RNA sequences, the upper (sense) strand is generally understood as oriented from its 5' to 3' end, and the complementary sequence is therefore understood as the lower (antisense) strand aligned in the same direction as the upper strand. Following the same logic, the reverse sequence is understood as the upper strand oriented from its 3' to its 5' end, and the "reverse complement" sequence is understood as the lower strand oriented from its 5' to its 3' end. Each nucleotide in a double-stranded DNA or RNA molecule that pairs with a Watson-Crick counterpart is called its complementary nucleotide.
[0099] The term "cleave," "cleaving," and "cleavage" used in this article to refer to the hydrolysis of the phosphodiester bonds in a nucleic acid molecule, resulting in the breakage of these bonds. This breakage can result in a cleavage (the hydrolysis of a single phosphodiester bond on one side of a double-stranded molecule), a single-strand break (the hydrolysis of a single phosphodiester bond on one side of a single-stranded molecule), or a double-strand break (the hydrolysis of two phosphodiester bonds on both sides of a double-stranded molecule), depending on whether the nucleic acid molecule is single-stranded (e.g., ssDNA or ssRNA) or double-stranded (e.g., dsDNA) and the type of nuclease activity catalyzed by the effector protein.
[0100] The term “clustered regularly spaced short palindromic repeats (CRISPR)” as used in this article refers to DNA segments found in the genomes of certain prokaryotes (including some bacteria and archaea) that consist of repeating short nucleotide sequences spaced regularly between unique nucleotide sequences derived from the DNA of pathogens (e.g., viruses) that have previously infected the organism, and whose function is to protect the organism from future infections by the same pathogen.
[0101] As used herein, the terms “CRISPR RNA” and “crRNA” refer to a class of guide nucleic acids, wherein the nucleic acid is an RNA comprising a first sequence and a second sequence, the first sequence being typically referred to herein as a spacer sequence, which hybridizes to the target sequence of the target nucleic acid, and the second sequence being capable of linking the crRNA to an effector protein by a) hybridization to a portion of tracrRNA or b) non-covalent binding to an effector protein. In some embodiments, the second sequence is referred to as a repeat sequence. In a dual-nucleic acid system, where the crRNA and tracrRNA form a complex with an effector protein, the crRNA comprises a first sequence that hybridizes to the target sequence of the target nucleic acid, and a second sequence that hybridizes to a portion of the tracrRNA.
[0102] The term "donor nucleic acid" as used in this article refers to nucleic acid integrated into the target nucleic acid or target sequence.
[0103] The term "donor nucleotide" as used in this article refers to a single nucleotide integrated into the target nucleic acid. Nucleotides are typically inserted into the cleavage site of effector proteins.
[0104] As used herein, the term "effect protein" refers to a protein, polypeptide, or peptide that non-covalently binds to a guide nucleic acid to form a complex that contacts a target nucleic acid, wherein at least a portion of the guide nucleic acid hybridizes to a target sequence of the target nucleic acid. The complex between the effector protein and the guide nucleic acid may include multiple effector proteins or a single effector protein. In some embodiments, the effector protein modifies the target nucleic acid when the complex contacts it. In some embodiments, the effector protein does not modify the target nucleic acid when the complex contacts it, but it fuses with a fusion chaperone protein that modifies the target nucleic acid. A non-limiting example of effector protein modification of a target nucleic acid is the cleavage of the phosphodiester bond of the target nucleic acid. Other examples of effector protein modification of target nucleic acids are described throughout this document.
[0105] As used herein, the term "functional acidic α-glucosidase protein" refers to an acidic α-glucosidase protein that retains at least some (if not all) of its enzymatic activity relative to the wild-type protein. Functional acidic α-glucosidase proteins may also include acidic α-glucosidase proteins with enhanced enzymatic activity relative to the wild-type protein. In some instances, the enzymatic activity is the degradation of glycogen into glucose. Detection methods are known and can be used to detect and quantify acidic α-glucosidase protein activity, for example, by causing a labeled substrate to release a signal (e.g., colorimetric and fluorescent) upon substrate cleavage. In some instances, the functional acidic α-glucosidase protein is a wild-type human acidic α-glucosidase protein. In some instances, the functional acidic α-glucosidase protein is a functional portion of the wild-type human acidic α-glucosidase protein.
[0106] As used in this article, the term "functional domain" refers to a region of a protein containing one or more amino acids that is essential for the protein to perform its activity or all of its activities (as determined in in vitro assays). Activities include, but are not limited to, nucleic acid binding, nucleic acid modification, nucleic acid cleavage, and protein binding. The loss of a functional domain (including mutations in the functional domain) will result in loss or reduction of activity.
[0107] As used in this article, the term "functional fragment" refers to a protein segment that retains certain functions relative to the whole protein. Non-limiting examples of functions include nucleic acid binding, protein binding, nuclease activity, nickase activity, deaminase activity, demethylase activity, or acetylation activity.
[0108] As used in this article, the terms "fusion effector protein," "fusion protein," and "fusion polypeptide" refer to proteins that comprise at least two heterologous polypeptides. Fusion effector proteins typically include both effector proteins and fusion chaperone proteins. Generally, fusion chaperone proteins are not effector proteins. Examples of fusion chaperone proteins are provided in this article.
[0109] As used in this article, the terms "fusion chaperone protein" and "fusion chaperone" refer to proteins, polypeptides, or peptides fused to effector proteins. Fusion chaperones typically endow fusion proteins with functions not provided by effector proteins. Fusion chaperones can modify target nucleic acids, including altering the nucleotides of the target nucleic acid and chemically modifying one or more nucleotides of the target nucleic acid.
[0110] As used herein, "genetic disease" refers to a disease, symptom, condition, or syndrome caused by one or more mutations in an organism's DNA. Mutations can be caused by several different cellular mechanisms, including but not limited to errors in DNA replication, recombination, or repair, or by environmental factors. In some implementations, genetic diseases include single-gene disorders, chromosomal disorders, or multifactorial disorders.
[0111] As used herein, the term "guide nucleic acid" refers to a nucleic acid comprising a first nucleotide sequence and a second nucleotide sequence, the first nucleotide sequence hybridizing to a target nucleic acid; and the second nucleotide sequence being capable of linking an effector protein to the nucleic acid by a) hybridizing to a portion of an additional nucleic acid that binds to an effector protein (e.g., tracrRNA) or b) non-covalently binding to the effector protein. The first sequence may be referred to herein as a spacer sequence. The second sequence may be referred to herein as a repeat sequence. In some embodiments, the first sequence is located at the 5' end of the second nucleotide sequence. In some embodiments, the first sequence is located at the 3' end of the second nucleotide sequence. In some embodiments, the first nucleotide sequence is attached to the 5' or 3' end of the second nucleotide sequence. In some embodiments, the first nucleotide sequence is attached to the second nucleotide sequence via a linker nucleic acid. In some embodiments, the linker nucleic acid comprises one, two, three, four, or five nucleotide bases. In some embodiments, the linker nucleic acid comprises a polynucleotide having two, three, four, or five nucleotide bases.
[0112] In the context of sgRNA, the term "stalk sequence" as used herein refers to a portion of sgRNA capable of non-covalently binding to an effector protein. The nucleotide sequence of the stalk sequence may comprise or be derived from tracrRNA. For example, in some aspects, the stalk sequence may comprise a portion of tracrRNA capable of non-covalently binding to an effector protein, but not all or part of tracrRNA that hybridizes to a portion of crRNA in a dinucleotide system. In some aspects, the stalk sequence may comprise a portion of tracrRNA and a portion of a repetitive sequence, which may optionally be linked by a linker. In some aspects, the stalk sequence in the context of sgRNA may also be described as a portion of the sgRNA that does not hybridize to a target sequence (e.g., a spacer region sequence) in the target nucleic acid.
[0113] As used herein, the term "heterologous" refers to a nucleotide or polypeptide sequence not found in native nucleic acids or proteins. In some embodiments, the fusion protein includes an effector protein and a fusion chaperone protein, wherein the fusion chaperone protein is heterologous to the effector protein. These fusion proteins may be referred to as "heterologous proteins." A protein heterologous to an effector protein is a protein that is not covalently linked to the effector protein in nature via an amide bond. In some embodiments, the heterologous protein is not encoded by the species encoding the effector protein. In some embodiments, the heterologous protein exhibits activity (e.g., enzymatic activity) upon fusion with the effector protein. In some embodiments, the heterologous protein exhibits enhanced or diminished activity (e.g., enzymatic activity) upon fusion with the effector protein relative to when it is not fused with the effector protein. In some embodiments, the heterologous protein exhibits activity (e.g., enzymatic activity) that it does not exhibit upon fusion with the effector protein. The guide nucleic acid may include a first sequence and a second sequence, wherein the first sequence and the second sequence are not covalently linked in nature via a phosphodiester bond. Therefore, the first sequence is considered heterologous to the second sequence, and the guide nucleic acid may be referred to as a heterologous guide nucleic acid.
[0114] As used herein, the term "in vitro" describes an event that occurs within a container containing laboratory reagents, thereby isolating them from the biological source from which the material was obtained. In vitro assays can include cell-based assays, where live or dead cells are used. In vitro assays can also include cell-free assays that do not use intact cells. The term "in vivo" describes an event that occurs within a subject. The term "ex vivo" describes an event that occurs outside the subject. Ex vivo assays are not performed on the subject themselves. Instead, they are performed on a sample separated from the subject. An example of ex vivo assays on a sample is an "in vitro" assay.
[0115] The term “linked amino acids” as used in this article refers to at least two amino acids linked by an amide bond.
[0116] As used in this article, the term "linker" refers to a bond or molecule that links a first polypeptide to a second polypeptide or a first nucleic acid to a second nucleic acid. A "peptide linker" includes at least two amino acids linked by an amide bond.
[0117] As used herein, the term "modified target nucleic acid" refers to a target nucleic acid that has been modified, for example, after contact with an effector protein. In some cases, the modification is an alteration of the target nucleic acid sequence. In others, a modified target nucleic acid, compared to an unmodified target nucleic acid, includes the insertion, deletion, substitution, or combination thereof of one or more nucleotides.
[0118] The term "disease-associated mutation" as used in this article refers to the co-occurrence of a mutation and a disease phenotype. Mutations can occur in genes in which the transcriptional or translational products of the gene appear at significantly abnormal levels or in abnormal forms in cells or subjects with the mutation, compared to non-disease control subjects without the mutation.
[0119] As used herein, the terms “non-naturally occurring” and “engineered” are used interchangeably and indicate human intervention. When referring to nucleic acids, nucleotides, proteins, polypeptides, peptides, or amino acids, these terms mean that the nucleic acid, nucleotide, protein, polypeptide, peptide, or amino acid does not at least substantially possess at least one other characteristic naturally associated with and existing in nature, and / or contains modifications (e.g., chemical modifications, nucleotide sequences, or amino acid sequences) not present in naturally occurring nucleic acids, nucleotides, proteins, polypeptides, peptides, or amino acids. When referring to compositions or systems described herein, the term means a composition or system having at least one component not naturally associated with other components of the composition or system. As a non-limiting example, a composition may include non-naturally coexisting effector proteins and guide nucleic acids. Conversely, as a further illustrative non-limiting example, “natural,” “naturally occurring,” or “existing in nature” effector proteins or guide nucleic acids include effector proteins and guide nucleic acids derived from cells or organisms that have not undergone human intervention and genetic modification.
[0120] The term "nucleic acid expression vector" used in this article refers to a plasmid that can be used to express the target nucleic acid.
[0121] The term “nuclear localization signal” as used in this article refers to an entity (e.g., a peptide) that, when present in a cell containing a nuclear compartment, helps to localize nucleic acids, proteins, or small molecules to the cell nucleus.
[0122] The term "nuclease activity" used in this article refers to the enzyme activity that allows an enzyme to cleave the phosphodiester bonds between nucleotide subunits of nucleic acids; the term "endonuclease activity" refers to the enzyme activity that allows an enzyme to cleave the phosphodiester bonds within a polynucleotide chain. Enzymes possessing nuclease activity are called "nucleases".
[0123] The term "pharmaceutically acceptable excipient, carrier, or diluent" refers to any substance formulated in a pharmaceutical composition with the active ingredient that allows the active ingredient to retain its biological activity while not reacting with the immune system of the subject. Such substances may be added to achieve long-term stability, increase the volume of solid dosage forms containing small amounts of potent active ingredients, or enhance the therapeutic properties of the active ingredient in the final formulation, such as improving absorption, reducing viscosity, or increasing solubility. The selection of suitable substances may depend on factors such as route of administration, formulation form, active ingredient, and other considerations. Compositions containing these substances can be prepared using conventional methods known in the art (see, for example, Remington's PharmaceuticalSciences, 18th edition, A. Gennaro, ed., Mack Publishing Co., Easton, Pa., 1990; and Remington, The Science and Practice of Pharmacy, 21st Ed. Mack Publishing, 2005).
[0124] The term "preseptal adjacent motif (PAM)" refers to a nucleotide sequence located in a target nucleic acid that guides an effector protein to modify the target nucleic acid at a specific site. The PAM sequence is often necessary for the hybridization and modification of the target nucleic acid by a complex formed by the effector protein and the guide nucleic acid. However, some effector proteins can achieve modification without the presence of a PAM sequence in the target nucleic acid.
[0125] The term "recombinant," when used for proteins, peptides, and nucleic acids, refers to the production of constructs with coding or non-coding sequences that differ from those of endogenous nucleic acids found in natural systems, through various combinations of cloning, restriction, and / or ligation processes. Typically, the DNA sequence encoding the coding structure can be assembled from cDNA fragments and short oligonucleotide linkers or a series of synthetic oligonucleotides, and the resulting synthetic nucleic acid can be expressed from intracellular recombinant transcription units or cell-free transcription and translation systems. This sequence can be presented as an open reading frame, not interrupted by internal untranslated sequences or introns typically present in eukaryotic genes. Recombinant genes or transcription units can also be created using genomic DNA containing the relevant sequences. The untranslated DNA sequences can be located at the 5' or 3' position of the open reading frame, provided that these sequences do not interfere with the operation or expression of the coding region and that the production of the desired product can be regulated through various mechanisms. Therefore, the terms "recombinant polynucleotide" or "recombinant nucleic acid" refer to artificially created sequences, meaning they are composed of two originally separate sequence fragments combined through artificial intervention. This artificial combination is typically achieved through chemical synthesis or manipulation of separated nucleic acid fragments using genetic engineering techniques. This typically involves replacing codons with redundant codons encoding the same or conserved amino acids, while routinely introducing or removing recognition sites. Alternatively, it may involve linking nucleic acid fragments with the desired function to create the desired functional combination. Similarly, the term "recombinant polypeptide" or "recombinant protein" refers to a protein that does not exist naturally and is formed by artificially combining two separate amino acid sequences. For example, polypeptides containing heterologous amino acid sequences would be classified as recombinant polypeptides.
[0126] The term "trans-activating RNA (tracrRNA)" refers to a nucleic acid that includes a first sequence capable of forming a non-covalent bond with an effector protein. tracrRNA may also contain a second sequence, often referred to as a repeat hybridization sequence, that hybridizes with a crRNA fragment. In some implementations, tracrRNA is covalently linked to crRNA.
[0127] The terms "treatment" and "treating" refer to a medication or other intervention aimed at achieving a beneficial or desired outcome for the recipient. Such beneficial outcomes can include therapeutic and / or preventative benefits. Therapeutic benefits typically include the elimination or relief of symptoms, or the resolution of an underlying condition. Furthermore, therapeutic benefits can manifest as a reduction or improvement in one or more physical symptoms associated with an underlying condition, even if the subject still has that condition. Preventative effects can include delaying, preventing, or eliminating the onset of a disease or symptom, as well as delaying or eliminating the appearance of symptoms, slowing, halting, or reversing the progression of a disease, or any combination of these outcomes. For preventative benefits, individuals at risk of developing a particular disease or those experiencing physical symptoms of one or more diseases may receive treatment, even if they have not yet been formally diagnosed.
[0128] The term "viral vector" refers to a nucleic acid used to deliver a virus or viral particle generated through recombination into a host cell. This nucleic acid can be single-stranded or double-stranded, linear or circular, segmented or non-segmented, and can consist of DNA, RNA, or a combination of both. Examples of viruses or viral particles that can be used as vectors include retroviruses (such as lentiviruses and gamma retroviruses), adenoviruses, arenaviruses, alphaviruses, adeno-associated viruses (AAVs), baculoviruses, vaccinia virus, herpes simplex virus, and poxviruses. Viral vectors delivered by these viruses or viral particles can be identified by the type of virus used for delivery (e.g., an AAV viral vector indicates delivery facilitated by adeno-associated virus). Viral vectors named after the type of virus used for delivery may include viral elements (e.g., nucleotide sequences) necessary for packaging the viral vector into a virus or viral particle, viral replication, or other desired viral functions. Viruses containing viral vectors may be replicating, replication-deficient, or replication-deprived.
[0129] The present invention discloses compositions, systems and methods comprising at least one of the following: (a) a polypeptide (e.g., a programmable nuclease) or a nucleic acid encoding a polypeptide; and (b) a guide nucleic acid or a nucleic acid encoding the guide nucleic acid.
[0130] In some embodiments, the compositions, systems, and methods include polypeptides or nucleic acids encoding polypeptides, wherein the polypeptide is an effector protein, also known as a programmable nuclease or programmable cleavage enzyme. Effector proteins and programmable nucleases / cleavage enzymes are described herein and throughout. Typically, a programmable nuclease (e.g., an effector protein) is a protein that binds to nucleic acids in a sequence-specific manner. In some embodiments, a programmable nuclease (e.g., an effector protein) is a protein that binds to and cleaves nucleic acids in a sequence-specific manner. A programmable nuclease (e.g., an effector protein) can bind to a target region of a nucleic acid and cleave the nucleic acid within or near the target region. In some embodiments, the programmable nuclease (e.g., an effector protein) is activated when it binds to a target region of a nucleic acid to cleave a nucleic acid region that is adjacent to but not adjacent to the target region. A programmable nuclease (e.g., an effector protein), such as a CRISPR-associated (CAS) protein, can be coupled to a guide nucleic acid that confers activity or sequence selectivity to the programmable nuclease (e.g., an effector protein). The programmable nucleases (e.g., effector proteins) described herein include effector proteins. Typically, the guide nucleic acid comprises a CRISPR RNA (crRNA) or a single guide RNA (sgRNA) that is at least partially complementary to the target nucleic acid. Therefore, in some embodiments, the effector protein and the guide nucleic acid can form a complex that recognizes the target sequence and cleaves the nucleic acid within, adjacent to, or near the target sequence. In some cases, the composition including the effector protein and the guide nucleic acid further includes a trans-activating crRNA (tracrRNA), at least a portion of which interacts with a programmable nuclease (e.g., the effector protein). In some embodiments, the composition, system, and method including the effector protein and the guide nucleic acid further includes an intermediate RNA, at least a portion of which interacts with a programmable nuclease (e.g., the effector protein). In some cases, the tracrRNA or intermediate RNA is provided separately from the guide nucleic acid. The tracrRNA can hybridize with a portion of the guide nucleic acid that does not hybridize with the target nucleic acid.
[0131] Programmable nucleases (e.g., effector proteins) can cleave nucleic acids, including single-stranded RNA (ssRNA), double-stranded DNA (dsDNA), and single-stranded DNA (ssDNA), or combinations thereof. Programmable nucleases (e.g., effector proteins) can provide binding activity, cis-cleavage activity, trans-cleavage activity, nicking enzyme activity, nuclease activity, or combinations thereof. Cis-cleavage activity is the cleavage of a target nucleic acid that hybridizes to a guide RNA (crRNA or sgRNA), wherein the cleavage occurs within or directly adjacent to the target nucleic acid region that hybridizes to the guide RNA. Trans-cleavage activity (also known as trans paraphyletic cleavage) refers to the cleavage of ssDNA or ssRNA that is close to but does not hybridize with the guide RNA. Trans-cleavage activity is initiated by the hybridization of the guide RNA to the target nucleic acid. Nicking enzyme activity is the selective cleavage of one strand of a dsDNA molecule.
[0132] Programmable CRISPR-associated (CAS) nucleases allow for precise and efficient editing of target DNA sequences through their ability to cut DNA at precise target locations in the genomes of various cells and organisms. In some implementations, CAS includes single-stranded DNA-binding (SSB) and double-stranded DNA-binding (DSB) proteins. SSB and DSB are effective methods for interfering with target genes, generating DNA or RNA modifications, and treating genetic diseases through gene correction.
[0133] The compositions, systems, and methods described herein are not naturally occurring. In some embodiments, the compositions, methods, and systems include at least one engineered effector protein and engineered guide nucleic acid, which may be simply referred to herein as effector protein and guide nucleic acid, respectively. In some embodiments, the compositions, systems, and methods include an effector protein or its use. In some embodiments, the compositions, systems, and methods include isolated effector proteins or their use. Generally, effector proteins and guide nucleic acids refer to effector proteins and guide nucleic acids that do not exist in nature, respectively. In some embodiments, the systems, methods, and compositions described herein include at least one non-naturally occurring component. For example, the disclosed compositions, methods, and systems may include guide nucleic acids, wherein the sequence of the guide nucleic acid is different from or modified from the sequence of naturally occurring guide nucleic acids.
[0134] In some embodiments, the compositions, systems, and methods include at least two non-naturally coexisting components. For example, the disclosed compositions, systems, and methods may include guide nucleic acids containing repeat regions and spacer regions that are non-naturally coexisting and / or heterologous to each other. Furthermore, as a non-limiting example, the disclosed compositions, systems, and methods may include non-naturally coexisting guide nucleic acids and effector proteins. Similarly, as a non-limiting example, the disclosed compositions, systems, and methods may include ribonucleotide-protein (RNP) complexes that include non-naturally coexisting effector proteins and guide nucleic acids. Conversely, for clarity, "natural," "naturally occurring," or "found in nature" effector proteins or guide nucleic acids include effector proteins and guide nucleic acids derived from cells or organisms that have not undergone genetic modification by humans or machines. Alternatively, in some embodiments, the compositions and systems include at least two components that do not coexist in nature, wherein the at least two components include at least one of an effector protein, a fusion chaperone, and a guide nucleic acid.
[0135] In some embodiments, the guide nucleic acid includes a non-natural nucleotide sequence. In some embodiments, the non-natural nucleotide sequence is a nucleotide sequence that does not exist in nature. The non-natural nucleotide sequence may include a portion of a naturally occurring sequence, wherein that portion of the naturally occurring sequence does not exist in nature in the absence of the remaining portion of the naturally occurring sequence. In some embodiments, the non-natural sequence is generated by chemical synthesis or by artificially manipulating isolated nucleic acid fragments (e.g., through genetic engineering). In some embodiments, the guide nucleic acid includes two naturally occurring sequences arranged in a sequence or close proximity not observed in nature. In some embodiments, the composition and system include a ribonucleotide complex comprising an effector protein and a guide nucleic acid that do not coexist in nature. Engineered guide nucleic acids may include a first and a second sequence that do not coexist naturally. For example, a guide nucleic acid may include a naturally occurring repeat region and a spacer region sequence complementary to a naturally occurring eukaryotic sequence. A guide nucleic acid may include a repeat region sequence naturally occurring in an organism and a spacer region sequence not naturally occurring in that organism. A guide nucleic acid may include a first sequence present in a first organism and a second sequence present in a second organism, wherein the first and second organisms are different. The guide nucleic acid may include a third sequence located at the 3' or 5' end of the guide nucleic acid or between the first and second sequences of the guide nucleic acid. For example, the guide nucleic acid may include naturally occurring crRNA and tracrRNA sequences linked by a linker sequence (e.g., sgRNA). In some embodiments, the guide nucleic acid includes two heterologous sequences arranged in a sequence or close proximity not observed in nature. Therefore, the compositions and systems described herein are non-natural.
[0136] In some embodiments, the compositions, systems, and methods described herein include effector proteins similar to naturally occurring effector proteins. The effector protein may lack a portion of a naturally occurring effector protein. The effector protein relative to a naturally occurring effector protein may include a mutation, wherein the mutation is not present in nature. The effector protein relative to a naturally occurring effector protein may also include at least one additional amino acid. For example, the effector protein relative to a naturally occurring effector protein may include the addition of a nuclear localization signal. In some embodiments, the nucleotide sequence encoding the effector protein is codon-optimized relative to a naturally occurring sequence (e.g., for expression in eukaryotic cells).
[0137] I. Peptide System This document provides compositions, systems, and methods comprising peptides or peptide systems, wherein the peptides or peptide systems described herein include one or more effector proteins or variants thereof, one or more effector chaperones or variants thereof, one or more linkers for the peptide, or combinations thereof. In some embodiments, a variant is a protein in a form or version different from a naturally occurring or wild-type protein. For example, a variant may have one or more amino acid substitutions, insertions, or deletions relative to a wild-type protein. In some embodiments, the variant has at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% identity with the corresponding wild-type protein. The variant may have a function or activity different from that of a naturally occurring or wild-type protein.
[0138] Peptides may include encoded and non-coding amino acids, chemically or biochemically modified or derived amino acids, and peptides having modified peptide backbones. Therefore, peptides described herein may include one or more mutations, one or more engineered modifications, or both, relative to naturally occurring or wild-type proteins. It should be understood that when describing the coding sequence of a peptide described herein, the coding sequence does not necessarily require a codon encoding an N-terminal methionine (M) or valine (V) as described for the effector protein herein. Those skilled in the art will understand that a start codon may be substituted or replaced by a start codon encoding an amino acid residue sufficient to initiate translation in a host cell. In some instances, when a heteropeptide (such as a fusion chaperone protein, protein tag, or nuclear localization signal (NLS)) is located at the N-terminus of an effector protein, the start codon of the heteropeptide also serves as the start codon of the effector protein. Therefore, a native start codon encoding an amino acid residue sufficient to initiate the translation of the effector protein (e.g., methionine (M) or valine (V)) may be removed or deleted.
[0139] effector proteins In some embodiments, this document provides compositions, systems, and methods comprising one or more effector proteins or nucleotide sequences encoding one or more effector proteins. In some embodiments, the effector protein is a protein, polypeptide, or peptide that is non-covalently bound to a guide nucleic acid to form a complex that interacts with a target nucleic acid. When the guide nucleic acid includes a nucleotide sequence complementary to a target sequence in the target nucleic acid, the effector protein can be made accessible to the target nucleic acid in its presence. The ability of the effector protein to modify the target nucleic acid can depend on the binding of the effector protein to the guide nucleic acid and the hybridization of the guide nucleic acid to the target nucleic acid. The effector protein can also recognize a pre-spacer adjacent motif (PAM) sequence present in the target nucleic acid, which can guide the modification activity of the effector protein.
[0140] The modification activity of the effector protein or engineered protein described herein can be cleavage activity, binding activity, insertion activity, substitution activity, etc. The modification activity of the effector protein can result in: cleavage of at least one strand of the target nucleic acid, deletion of one or more nucleotides of the target nucleic acid, insertion of one or more nucleotides into the target nucleic acid, replacement of one or more nucleotides of the target nucleic acid with alternative nucleotides, one or more of the foregoing, or any combination thereof. In some embodiments, modification of the target nucleic acid includes the introduction or removal of epigenetic modifications(s). In some embodiments, the ability of the effector protein to edit the target nucleic acid can depend on the complexation of the effector protein with the guide nucleic acid, hybridization of the guide nucleic acid with the target sequence of the target nucleic acid, the distance between the target sequence and the PAM sequence, or a combination thereof. The target nucleic acid includes a target strand and a non-target strand. Therefore, in some embodiments, the effector protein can edit the target strand and / or the non-target strand of the target nucleic acid. The effector protein can modify the nucleic acid by cis-cleavage or trans-cleavage. As a non-limiting example, modification of the target nucleic acid produced by the effector protein can lead to the expression of a protein encoded by the donor nucleic acid.
[0141] As a non-limiting example, modification of the target nucleic acid produced by an effector protein can lead to regulation of the expression of the target nucleic acid (e.g., increasing or decreasing nucleic acid expression) or regulation of the activity of the translation product of the target nucleic acid (e.g., inactivation of proteins bound to RNA molecules or hybridized proteins). Therefore, in some embodiments, methods are provided herein for editing target nucleic acids using effector proteins of the present disclosure or compositions or systems thereof. Methods are also provided herein for regulating the expression of target nucleic acids using effector proteins of the present disclosure or compositions or systems thereof. Furthermore, methods are provided herein for regulating the activity of the translation product of target nucleic acids using effector proteins of the present disclosure or compositions or systems thereof.
[0142] In some embodiments, a given effector protein may not require the presence of a PAM sequence in the target nucleic acid for effector protein modification. Therefore, in some embodiments, the effector protein may also recognize a sequence at least one, two, three, four, five, six, seven, eight, nine, or ten nucleotides from the 5' or 3' end of a PAM sequence present in the target nucleic acid, which can guide the effector protein's modification activity.
[0143] In some embodiments, the effector proteins disclosed herein may provide catalytic activity similar to that of naturally occurring effector proteins (e.g., cleavage activity, nickase activity, nuclease activity, other activities, or combinations thereof), such as naturally occurring effector proteins having reduced cleavage activity (including cis-cleavage activity, trans-cleavage activity, or combinations thereof). In some embodiments, the effector proteins disclosed herein may be fused with effector chaperones or fusion proteins, wherein the effector chaperones or fusion proteins are capable of performing some functions or activities not provided by the effector proteins.
[0144] Effector proteins can be CRISPR-associated (“Cas”) proteins. Effector proteins can function as single proteins, including single proteins capable of binding to guide nucleic acids and modifying target nucleic acids. Alternatively, effector proteins can function as part of a multi-protein complex, including, for example, a complex containing two or more effector proteins (including two or more identical effector proteins (e.g., dimers or multimers)). When an effector protein functions in a multi-protein complex, it may have only one functional activity (e.g., binding to the guide nucleic acid), while other effector proteins present in the multi-protein complex may have other functional activities (e.g., modifying the target nucleic acid). In some embodiments, the effector protein or its multi-protein complex binds to the guide nucleic acid through non-covalent interactions. Non-limiting examples of non-covalent interactions are ionic bonds, hydrogen bonds, van der Waals forces, and hydrophobic interactions. In some embodiments, when an effector protein functions in a multi-protein complex, it may have functional activities different from and / or complementary to those of other effector proteins in the multi-protein complex. In some embodiments, the complementary functional activity of effector proteins including modified or artificial base pairs can be based on other types of hydrogen bonding and / or hydrophobicity of bases and / or shape complementarity between bases. The multimeric complexes and their functions will be described in more detail below. In some embodiments, the effector protein can be a modified effector protein that has different modification activity and / or substrate binding activity (e.g., substrate selectivity, specificity, and / or affinity) compared to an unmodified effector protein. For example, in some embodiments, the effector protein can be a modified effector protein with reduced modification activity (e.g., catalytically deficient effector protein) or no modification activity (e.g., catalytically inactive effector protein). Therefore, the effector proteins used herein include modified or programmable nucleases that do not possess nuclease activity.
[0145] In some embodiments, the effector proteins described herein may include one or more functional domains. These functional domains may include prespacer adjacent motif (PAM) interaction domains, oligonucleotide interaction domains, various recognition domains, non-target strand interaction domains, and RuvC domains. PAM interaction domains may be classified as target strand PAM interaction domains (TPIDs) or non-target strand PAM interaction domains (NTPIDs). In some instances, a PAM interaction domain (e.g., a TPID or NTPID) refers to a polypeptide fragment that binds to a target nucleic acid.
[0146] Effector proteins may also contain a RuvC domain, which is typically located near the C-terminus of the protein. This RuvC domain can cleave target nucleic acids and, in some cases, process precursor crRNA. A single RuvC domain may include several subdomains, such as RuvC-I, RuvC-II, and RuvC-III. In some embodiments, the RuvC domain may also include RuvC-like domains. Various RuvC-like domains are well-documented and can be identified using online resources such as InterPro (…). https: / / www.ebi.ac.uk / interpro / For example, RuvC-like domains may share homology with regions found in IS605 and related transposon families of TnpB proteins, as detailed in review articles such as Shmakov et al. (Nature Reviews Microbiology volume 15, pages 169-182 (2017)) and Koonin EV and Makarova KS (2019, Phil. Trans. R. Soc., B 374:20180087). In some cases, RuvC domains can possess substrate-binding activity, catalytic activity, or both. They can be defined by a single continuous sequence or by a collection of discontinuous RuvC subdomains within a primary amino acid sequence. Effector proteins can possess multiple RuvC subdomains that work together to form a RuvC domain with substrate-binding or catalytic activity. For example, effector proteins may include three RuvC subdomains (RuvC-I, RuvC-II, and RuvC-III), which are not contiguous in the primary sequence but assemble into a functional RuvC domain during protein production and folding. Typically, effector proteins include a recognition domain (REC domain) that binds to the guide nucleic acid or a guide nucleic acid-target nucleic acid heteroduplex. In some embodiments, the REC domain may consist of an α-helical recognition region or a lobe. CRISPR / Cas proteins may contain two REC domains (RECI and REC2), which typically contribute to regulating and stabilizing the hybridization of the guide and target nucleic acids. Effector proteins may contain at least one REC domain (e.g., RECI or REC2) and may also include a zinc finger domain. In some instances, effector proteins do not possess an HNH domain.
[0147] Effector proteins can be relatively small, which is advantageous for nucleic acid detection or editing applications (e.g., their smaller size may reduce the likelihood of adsorption onto surfaces or other biological entities). The compact nature of these effector proteins can facilitate more efficient packaging and delivery in the context of genome editing, as well as their incorporation as reagents into detection. In some embodiments, the effector protein is at least 400 linked amino acid residues in length. In other cases, the length can be less than 500 linked amino acid residues. In addition, the length can be in the range of about 400 to about 500 linked amino acid residues, or about 450 to about 550, or in the following specific ranges, such as about 400 to about 420, about 420 to about 440, about 440 to about 460 linked amino acid residues, about 460 to about 480, about 480 to about 500, about 500 to about 520, about 520 to about 540, about 540 to about 560, about 560 to about 580, about 580 to about 600, about 600 to about 620, about 620 to about 640, about 640 to about 660, about 660 to about 680, about 680 to about 700 linked amino acid residues.
[0148] In some cases, effector proteins can recognize the various PAMs described herein. Furthermore, effector proteins can produce blunt ends or short staggered ends. Blunt cleavage may be advantageous compared to staggered cleavage provided by other effector proteins because it reduces the likelihood of spontaneous (or perfect) repair, potentially increasing the success rate of target nucleic acid editing and / or donor nucleic acid insertion.
[0149] In some implementations, effector proteins function as endonucleases that catalyze cleavage within the target nucleic acid. They can also catalyze non-sequence-specific cleavage of single-stranded nucleic acids. In some instances, effector proteins (e.g., those having any of the amino acid sequences listed in Table 1) are activated to perform trans-cleavage activity after the guide nucleic acid binds to the target nucleic acid. This trans-cleavage activity, also referred to as “para-ligand” or “trans-para-ligand” cleavage, may involve non-specific cleavage of nearby single-stranded nucleic acids by the activated effector protein, such as trans-cleavage of a detection nucleic acid containing a detection moiety.
[0150] The effector proteins described herein can act as endonucleases to facilitate cleavage at specific sites (e.g., specific nucleotides within a nucleic acid sequence) in a designated target nucleic acid. The target nucleic acid can be single-stranded RNA (ssRNA), double-stranded DNA (dsDNA), or single-stranded DNA (ssDNA). In some embodiments, the target nucleic acid is single-stranded DNA, while in other embodiments, it is single-stranded RNA. The effector protein can exhibit cis-cleavage activity, trans-cleavage activity, cleavage enzyme activity, or a combination of these functions.
[0151] Cis-cleavage activity refers to the cleavage of a target nucleic acid that hybridizes with a guide RNA (such as in a dual guide RNA system or a single guide RNA, sgRNA), where cleavage occurs within or directly adjacent to the target nucleic acid region that hybridizes with the guide RNA. Trans-cleavage activity, also known as trans para-cleavage, involves the cleavage of ssDNA or ssRNA that is adjacent to but does not hybridize with the guide RNA. This trans-cleavage activity is initiated by the hybridization of the guide nucleic acid with the target nucleic acid. In some instances, trans-cleavage refers to the hydrolysis of one or more nucleic acids by an effector protein that is complexed with both the guide nucleic acid and the target nucleic acid. These nucleic acids can include both target and non-target nucleic acids. Trans-cleavage can occur near, but not within, or directly adjacent to, the target nucleic acid region that hybridizes with the guide nucleic acid.
[0152] Cleavage enzyme activity involves the selective cleavage of one strand of double-stranded DNA (dsDNA).
[0153] Table 1 provides exemplary amino acid sequences of effector proteins applicable to the compositions, systems, and methods discussed herein. In some embodiments, the effector protein exhibits at least 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% identity with any of the amino acid sequences listed in Table 1. Furthermore, the nucleic acid encoding the effector protein can be operatively linked to a promoter that functions in eukaryotic or prokaryotic cells. The promoter can be one or more of the following: constitutive promoters, inducible promoters, cell type-specific promoters, or tissue-specific promoters. In some cases, the promoter is effective in a variety of cell types, including plant cells, fungal cells, animal cells, invertebrate cells, fly cells, vertebrate cells, mammalian cells, primate cells, non-human primate cells, and human cells. Additionally, in some embodiments, the nucleic acid can be part of the nucleic acid expression vector described herein.
[0154] In some embodiments, the compositions, systems, and methods described herein include an effector protein or a nucleic acid encoding an effector protein, wherein the amino acid sequence of the effector protein comprises at least about 200 consecutive amino acids, at least about 220 consecutive amino acids, at least about 240 consecutive amino acids, at least about 260 consecutive amino acids, at least about 280 consecutive amino acids, at least about 300 consecutive amino acids, at least about 320 consecutive amino acids, at least about 340, at least about 360, at least about 380 consecutive amino acids, at least about 400, or at least about 4... The amino acid sequence may contain 20 consecutive amino acids, at least about 440 consecutive amino acids, at least about 460 consecutive amino acids, at least about 480 consecutive amino acids, at least about 500 consecutive amino acids, at least about 520 consecutive amino acids, at least about 540 consecutive amino acids, at least about 560 consecutive amino acids, at least about 580 consecutive amino acids, at least about 600 consecutive amino acids, at least about 620 consecutive amino acids, at least about 640 consecutive amino acids, at least about 660 consecutive amino acids, at least about 680 consecutive amino acids, or at least about 700 consecutive amino acids or more. In some embodiments, the compositions, systems, and methods described herein comprise an effector protein or a nucleic acid encoding an effector protein, wherein the amino acid sequence of the effector protein comprises at least 200 consecutive amino acids or more of any of the amino acid sequences listed in Table 1. In some embodiments, the compositions, systems, and methods described herein comprise an effector protein or a nucleic acid encoding an effector protein, wherein the amino acid sequence of the effector protein comprises at least 300 or more consecutive amino acids of any of the amino acid sequences listed in Table 1. In some embodiments, the compositions, systems, and methods described herein comprise an effector protein or a nucleic acid encoding an effector protein, wherein the amino acid sequence of the effector protein comprises at least 400 or more consecutive amino acids of any of the amino acid sequences listed in Table 1.
[0155] In some embodiments, the compositions, systems, and methods described herein include an effector protein or a nucleic acid encoding an effector protein, wherein the effector protein includes a portion of any of the amino acid sequences listed in Table 1. In some embodiments, the effector protein includes a portion of any of the amino acid sequences listed in Table 1, wherein the portion does not include at least the first 10 amino acids, at least the first 20 amino acids, at least the first 40 amino acids, at least the first 60 amino acids, at least the first 80 amino acids, at least the first 100 amino acids, at least the first 120 amino acids, at least the first 140 amino acids, at least the first 160 amino acids, at least the first 180 amino acids, or at least the first 200 amino acids of any of the amino acid sequences listed in Table 1. In some embodiments, the effector protein includes a portion of any of the amino acid sequences listed in Table 1, wherein the portion does not include the last 10, last 20, last 40, last 60, last 80, last 100, last 120, last 140, last 160, last 180, or last 200 amino acids of any of the amino acid sequences listed in Table 1.
[0156] In some embodiments, the compositions, systems, and methods described herein include an effector protein or a nucleic acid encoding an effector protein, wherein the effector protein comprises an amino acid sequence having at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100% identity with any of the amino acid sequences listed in Table 1. In some embodiments, the effector protein provided herein comprises an amino acid sequence having at least 65% identity with any of the amino acid sequences listed in Table 1. In some embodiments, the effector protein provided herein comprises an amino acid sequence having at least 70% identity with any of the amino acid sequences listed in Table 1. In some embodiments, the effector protein provided herein comprises an amino acid sequence having at least 75% identity with any of the amino acid sequences listed in Table 1. In some embodiments, the effector protein provided herein comprises an amino acid sequence having at least 80% identity with any of the amino acid sequences listed in Table 1. In some embodiments, the effector protein provided herein comprises an amino acid sequence having at least 85% identity with any of the amino acid sequences listed in Table 1. In some embodiments, the effector protein provided herein comprises an amino acid sequence having at least 90% identity with any of the amino acid sequences listed in Table 1. In some embodiments, the effector protein provided herein comprises an amino acid sequence having at least 95% identity with any of the amino acid sequences listed in Table 1. In some embodiments, the effector protein provided herein comprises an amino acid sequence having at least 97% identity with any of the amino acid sequences listed in Table 1. In some embodiments, the effector protein provided herein comprises an amino acid sequence having at least 98% identity with any of the amino acid sequences listed in Table 1. In some embodiments, the effector protein provided herein comprises an amino acid sequence identical to any of the amino acid sequences listed in Table 1.
[0157] In some embodiments, the compositions, systems, and methods described herein may include a variant of a reference effector protein (wild-type effector protein). These embodiments may have an effector protein or a nucleic acid encoding an effector protein, wherein the amino acid sequence of the effector protein exhibits at least 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or even 100% similarity to any of the amino acid sequences listed in Table 1. The similarity of the amino acid sequence to the reference sequence is determined by calculating a similarity score divided by the alignment length.
[0158] The similarity between two amino acid sequences can be assessed using the BLOSUM62 similarity matrix (Henikoff and Henikoff, Proc. Natl. Acad. Sci. USA., 89: 10915-10919 (1992)), which has been modified so that any score greater than 1 is replaced with +1, and any score less than 0 is replaced with 0. For example, the substitution from His (H) to Leu (L) scores +2.0 using the BLOSUM62 matrix, which is converted to +1 in the modified matrix. This conversion facilitates the calculation of the percentage of similarity, rather than just the similarity score. Alternatively, when comparing two full-length protein sequences, pairwise alignment can be performed using MUSCLE alignment, which allows the determination of the percentage of similarity for each residue, divided by the total length of the alignment.
[0159] To assess the percentage similarity of specific protein domains or motifs, multilevel shared sequences (or PROSITE motif sequences) can be used to evaluate the conservation of each domain or motif. When calculating domain or motif similarity, the second and third levels of the multilevel sequence are considered equivalent to the highest level. Furthermore, a +1 point is awarded if any amino acid substitution at that position in the multilevel shared sequence is considered conserved. For example, given the multilevel shared sequences RLG and YCK, the test sequence QIQ will receive 3 points based on the following scores from the transformed BLOSUM62 matrix: QR: +1; QY: +0; IL: +1; IC: +0; QG: +0; QK: +1. When calculating overall similarity, the highest score for each position is used. The percentage similarity can also be calculated using commercially available software such as Geneious Prime, with parameters set to use the BLOSUM62 matrix and a threshold greater than 1. Therefore, the percentage similarity is defined as the similarity score divided by the alignment length.
[0160] In some embodiments, the compositions, systems, and methods described herein include an effector protein or a nucleic acid encoding an effector protein, wherein the amino acid sequence of the effector protein has at least 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% similarity to the amino acid sequences listed in Table 1. Specifically, the effector protein may contain an amino acid sequence having at least 65% similarity to any sequence listed in Table 1. Similarly, additional embodiments may specify that the effector protein has an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or even 100% similarity to the sequences listed in Table 1.
[0161] Such effector proteins may contain one or more modifications, which can be classified as conserved or non-conserved alterations (e.g., conserved or non-conserved substitutions). In some embodiments, these modifications may consist independently of one or more conserved substitutions, one or more non-conserved substitutions, or a combination of both. A conserved alteration (e.g., a conserved substitution) refers to the substitution of one amino acid with another amino acid from a family of amino acids with similar side-chain characteristics. Conversely, a non-conserved alteration (e.g., a non-conserved substitution) involves the substitution of one amino acid residue with another amino acid residue that does not belong to the same side-chain family.
[0162] Based on their associated side chains, the amino acids encoded by genes can be divided into four families: (1) acidic (negatively charged): Asp (D), Glu (E); (2) basic (positively charged): Lys (K), Arg (R), His (H); (3) nonpolar (hydrophobic): Cys (C), Ala (A), Val (V), Leu (L), Ile (I), Pro (P), Phe (F), Met (M), Trp (W), Gly (G), Tyr (Y); among which nonpolar is further divided into: (i) strongly hydrophobic: Ala (A), Val (V), Leu (L), Ile (I), Met (M), Phe (F); and (ii) moderately hydrophobic: Gly (G), Pro (P), Cys (C), Tyr (Y), Trp (W); and (4) polar uncharged: Asn (N), Gln (Q), Ser (S), Thr Amino acids can also be linked together by aliphatic side chains: Gly (G), Ala (A), Val (V), Leu (L), Ile (I), Ser (S), Thr (T), with Ser (S) and Thr (T) sometimes classified separately as aliphatic hydroxyl groups. Aromatic side chains include: Phe (F), Tyr (Y), Trp (W). Amide side chains consist of: Asn (N), Gln (Q), while sulfur-containing side chains include: Cys (C) and Met (M).
[0163] In some instances, effector proteins containing one or more amino acid modifications are considered variants of the effector proteins described herein. It is important to note that any reference to an effector protein also includes its variants. Modifications may include the deletion or insertion of an amino acid. Furthermore, amino acid changes may include non-conserved substitutions. In some cases, these modifications may consist of a combination of one or more conserved amino acid substitutions and one or more non-conserved amino acid substitutions. The effector protein or the nucleic acid encoding it may exhibit 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more amino acid changes compared to any amino acid specified in the TR.
[0164] In some embodiments, the compositions, systems, and methods described herein may include an effector protein or a nucleic acid encoding an effector protein, wherein the effector protein comprises one or more substitutions relative to any amino acid sequence listed in Table 1. These substitutions may be at least 1 to 20 or more and may be categorized into various ranges, such as having 1 to 16, 1 to 12, 1 to 8, 1 to 4, 4 to 20, 4 to 16, 4 to 12, 4 to 8, 8 to 20, 8 to 16, 8 to 12, 12 to 20, 12 to 16, or 16 to 20 substitutions relative to any sequence in Table 1.
[0165] Substitutions can consist of one or more conserved substitutions, one or more non-conserved substitutions, or a combination of both. In some cases, the effector protein or the nucleic acid encoding it may include one or more conserved substitutions compared to any of the amino acid sequences listed in Table 1. These conserved substitutions may also fall within the same range of 1 to 20 or more, as described in the various specific ranges above.
[0166] Similarly, effector proteins may also include one or more non-conserved substitutions relative to the sequences in Table 1, wherein the number of substitutions has the same possible range.
[0167] In some instances, amino acid alterations can lead to changes in the activity of an effector protein compared to its naturally occurring counterpart. For example, these alterations can enhance or reduce the catalytic activity of the effector protein or affect its binding activity compared to the naturally occurring version. In some embodiments, these alterations can result in catalytically inactive variants of the effector protein.
[0168] Furthermore, the effector proteins described herein can undergo enzymatic reactions similar to those of wild-type (WT) effector proteins. Variants of WT effector proteins may include modifications that impart beneficial characteristics, such as enhanced activity (e.g., improved insertion / deletion activity, catalytic activity, specificity, selectivity, or affinity for substrates such as target or guide nucleic acids). In some cases, the activity of effector proteins may be equal to or greater than that of WT effector proteins, meaning they exhibit one or more identical or higher activities at the same amino acid positions compared to effector proteins without any variants. For example, variants may exhibit at least a 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, or even 200% increase in activity compared to WT effector proteins. The activity of an effector protein or its variants can be assessed by a cleavage assay compared to a WT effector protein.
[0169] The activity of effector proteins or their variants compared to WT effector proteins can be assessed through cleavage experiments.
[0170] In some embodiments, the effector protein may include one or more amino acid substitutions compared to any sequence in Table 1, with the remaining amino acid sequence having at least 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% identity with the reference amino acid sequence in Table 1. Furthermore, the remaining amino acid sequence of the variant may also exhibit at least 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% similarity to the corresponding reference sequence. In some cases, the substitution may involve one or more positively charged amino acid residues, such as Lys(K), Arg(R), His(H), or combinations thereof. Modification of the effector protein includes modification of one or more nucleic acid residues of the nucleotide sequence encoding the effector protein or one or more amino acid residues of the amino acid sequence of the effector protein. In some embodiments, the modification includes chemical modification of one or more nucleobases; or chemical modification of the phosphate backbone, nucleotide, nucleobase, or nucleoside. Such modification may be made to the amino acid sequence of the effector protein or the nucleotide sequence encoding the effector protein. Methods for modifying nucleic acid or amino acid sequences are known. Those skilled in the art will understand that modifications can be located at any position in a nucleotide or amino acid sequence. In some embodiments, modifications can significantly alter the function of an effector protein, composition, or system. In some embodiments, modifications can significantly reduce the function of an effector protein, composition, or system. In some embodiments, modifications can significantly increase the function of an effector protein, composition, or system. In some embodiments, modifications can significantly alter the structural stability of an effector protein. Modified effector proteins can be prepared using any available technology, including but not limited to chemical synthesis, enzymatic synthesis (often referred to as w / ro-transcription), cloning, enzymatic catalysis, or chemical cleavage. In some embodiments, the compositions, systems, and methods described herein may include an effector protein or a nucleic acid encoding an effector protein, wherein the effector protein includes one or more substitutions relative to any amino acid sequence listed in Table 1. These substitutions can be at least 1 to 20 or more, and can be categorized into various ranges, such as containing 1 to 16, 1 to 12, 1 to 8, 1 to 4, 4 to 20, 4 to 16, 4 to 12, 4 to 8, 8 to 20, 8 to 16, 8 to 12, 12 to 20, 12 to 16, or 16 to 20 substitutions relative to the sequences in Table 1.
[0171] Substitutions may include one or more conserved substitutions, one or more non-conserved substitutions, or a combination of both. In some instances, the effector protein or the nucleic acid encoding it may include one or more conserved substitutions relative to any amino acid sequence in Table 1, wherein the number of substitutions also falls within the range described above.
[0172] Similarly, relative to any sequence in Table 1, effector proteins may also include one or more non-conserved substitutions, with the possible number of substitutions ranging from 1 to 20 or more.
[0173] In some cases, amino acid alterations can lead to changes in the activity of an effector protein compared to its naturally occurring counterpart. For example, these modifications can enhance or reduce the catalytic or binding activity of the effector protein. In some embodiments, these changes can result in catalytically inactive variants of the effector protein.
[0174] Furthermore, the effector proteins described herein can undergo enzymatic reactions similar to those of wild-type (WT) effector proteins. Variants of WT effector proteins may contain modifications that confer beneficial characteristics, such as enhanced activity (e.g., improved insertion / deletion activity, catalytic activity, specificity, selectivity, or affinity for substrates such as target or guide nucleic acids). In some cases, the activity of effector proteins may be equal to or greater than that of WT effector proteins, meaning that they exhibit one or more identical or higher activities at the same amino acid positions compared to effector proteins without any variants. For example, variants may exhibit an increase in activity of at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, or even 200% compared to WT effector proteins.
[0175] The activity of effector proteins or their variants compared to WT effector proteins can be assessed through cleavage experiments.
[0176] In some embodiments, the effector protein may have one or more amino acid substitutions compared to any sequence in Table 1, wherein the remaining amino acid sequence has at least 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% identity with the reference amino acid sequence in Table 1. Furthermore, the remaining amino acid sequence of the variant may also exhibit at least 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% similarity to the corresponding reference sequence. In some cases, the substitution may involve one or more positively charged amino acid residues, such as Lys(K), Arg(R), His(H), or combinations thereof.
[0177] In some implementations, the term "in vitro" refers to a process that occurs outside a living organism, typically in a controlled environment (such as a test tube or petri dish).
[0178] The effector proteins described herein may include encoded and non-coding amino acids, as well as chemically or biochemically modified or derived amino acids, and proteins with altered peptide backbones. These effector proteins may contain one or more mutations, engineered modifications, or both. It should be understood that the coding sequence of these effector proteins does not necessarily need to include an N-terminal codon containing either methionine (M) or valine (V). Those skilled in the art will understand that the start codon may be replaced by a codon encoding an amino acid residue sufficient to initiate translation in the host cell.
[0179] Furthermore, the compositions, systems, and methods may include heterologous peptides or polypeptides, which may be fusion proteins comprising an effector protein and one or more fusion chaperone proteins. The fusion chaperone protein may be located at the N-terminus of the effector protein. In this case, the start codon of the heterologous peptide or polypeptide may also serve as the start codon of the effector protein, allowing for the removal or deletion of the native start codon that normally encodes the amino acid residues required for translation initiation.
[0180] Heteropeptides can consist of at least two distinct polypeptide sequences that do not coexist in nature. Heterologous systems may include at least one component that does not naturally coexist with other components.
[0181] In some implementations, the heteropeptide or polypeptide may contain subcellular localization signals, such as nuclear localization signals (NLS) that facilitate the targeting of nucleic acids, proteins, or small molecules to the cell nucleus within the cell. Other types of localization signals may include nuclear export signals (NES), signals for retaining effector proteins in the cytoplasm, mitochondrial localization signals, chloroplast localization signals, endoplasmic reticulum retention signals, etc. In some cases, effector proteins may not include subcellular localization signals to prevent targeting of the cell nucleus, which may be beneficial when the target nucleic acid is RNA present in the cytoplasm.
[0182] In some embodiments, the heterologous polypeptide may be an endosome escape peptide (EEP), which is designed to rapidly disrupt endosomes to minimize the residence time of delivery molecules (e.g., effector proteins) in the endosome environment, thereby avoiding encapsulation in endosome vesicles and subsequent degradation in the lysosomal compartment.
[0183] In addition, heterologous peptides can be cell-penetrating peptides (CPPs), also known as protein transduction domains (PTDs), which facilitate the penetration of lipid bilayers, cell membranes, organelle membranes, or vesicle membranes.
[0184] In some instances, heteropeptides or polypeptides may include protein tags that can be used as purification tags or fluorescent proteins. Such tags can be detectable for the identification or purification of effector proteins. Depending on the intended application, a variety of protein tags can be used, including but not limited to fluorescent proteins, histidine tags (e.g., 6XHis), hemagglutinin (HA) tags, FLAG tags, Myc tags, and maltose-binding protein (MBP). Examples of fluorescent proteins include green fluorescent protein (GFP), yellow fluorescent protein (YFP), red fluorescent protein (RFP), cyan fluorescent protein (CFP), mCherry, and tdTomato.
[0185] Heteropeptides can be located at or near the amino terminus (N-terminus) or carboxyl terminus (C-terminus) of the effector protein. In some cases, heteropeptides can be located at suitable insertion sites within the effector protein.
[0186] II. Guided Nucleic Acid The compositions, systems, and methods disclosed herein may include guide nucleic acids or their use. These may also include compositions, systems, and methods having one or more guide nucleic acids and DNA molecules encoding those guide nucleic acids. Those skilled in the art will understand that a DNA molecule “encoding” a nucleic acid (such as a guide nucleic acid) means a DNA molecule having a nucleotide sequence that, upon transcription, produces an RNA molecule (e.g., a guide nucleic acid). It should be understood that references to guide nucleic acids also include the DNA molecule encoding that guide nucleic acid.
[0187] Guide nucleic acids and their components (e.g., spacer sequences, repeat sequences, linker nucleotide sequences, stalk sequences, intermediate RNA, etc.) may consist of one or more deoxyribonucleotides (DNA), ribonucleotides (RNA), or combinations of both (e.g., RNA containing thymine bases) and biochemically or chemically modified nucleotides. Such nucleotide sequences may be described as DNA or RNA; however, regardless of form, it should be understood that these sequences may be adapted to RNA or DNA as needed to describe sequences comprising or encoding guide nucleic acids, such as nucleotide sequences for vectors. The disclosure of the nucleotide sequence also includes its complementary sequence, reverse sequence, and reverse complementary sequence, any of which may serve as the nucleotide sequence for the guide nucleic acid.
[0188] In some implementations, the guide nucleic acid may include CRISPR RNA (crRNA) or a single guide RNA (sgRNA). sgRNA is formed by binding a spacer sequence (which hybridizes to a target sequence in the target nucleic acid) to a stalk sequence, wherein the two sequences are covalently linked. The spacer and stalk sequences can be linked by a phosphodiester bond or by one or more linked nucleotides. In some cases, the guide nucleic acid may include a spacer sequence, a repeat sequence, a stalk sequence, or a combination thereof, wherein the stalk sequence may comprise some or all of the repeat sequence.
[0189] In some embodiments, the composition may include tracrRNA. crRNA and tracrRNA may function as separate, unlinked molecules, or they may be covalently linked. Linkage may occur via a phosphodiester bond or via one or more linked nucleotides. In some embodiments, the composition may not include tracrRNA.
[0190] Guide RNAs can be naturally occurring or non-natural; non-natural guide RNAs may include chemical or biochemical modifications. Guide RNAs can be produced through chemical synthesis or recombination. The sequence of a guide RNA, or a portion thereof, may differ from the sequence of naturally occurring nucleic acids.
[0191] In some embodiments, the compositions, systems, and methods of this disclosure may include one or more additional guide nucleic acids or their uses. For example, these may include two or more additional guide nucleic acids (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, or more guide nucleic acids) that can target effector proteins to different locations within the target nucleic acid by binding to different portions of the target nucleic acid. The guide nucleic acid may bind to fragments of the target nucleic acid upstream or downstream of the target gene, thereby allowing modification at two different locations. This dual-targeting approach, referred to as "dual cleavage," may involve two effector proteins, each corresponding to one guide RNA, or a single effector protein with two different guide RNAs to achieve the dual cleavage effect. In some cases, multiple effector proteins (e.g., 2, 3, 4, 5, 6, 7, 9, 10, or more) may be used in the dual-guider system described herein.
[0192] In some embodiments, the guide nucleic acid may include 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 linked nucleotides. Typically, the guide nucleic acid will consist of at least a certain number of linked nucleotides, with some embodiments specifically specifying at least 25 linked nucleotides. The length of the guide nucleic acid can range from 10 to 50 linked nucleotides. In some cases, the guide nucleic acid may consist primarily of about 12 to about 80, about 12 to about 50, about 12 to about 45, about 12 to about 40, about 12 to about 35, about 12 to about 30, about 12 to about 25, about 12 to about 20, about 12 to about 19, about 19 to about 20, about 19 to about 25, about 19 to about 30, about 19 to about 35, about 19 to about 40, about 19 to about 45, about 19 to about 50, about 19 to about 60, about 20 to about 25, about 20 to about 30, about 20 to about 35, about 20 to about 40, about 20 to about 45, about 20 to about 50, or about 20 to about 60 linked nucleotides. In some embodiments, the guide nucleic acid may have about 10 to about 60, about 20 to about 50, or about 30 to about 40 linked nucleotides.
[0193] In some implementations, the engineered guide nucleic acid comprises at least 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 consecutive nucleotides complementary to a eukaryotic sequence. The eukaryotic sequence refers to a nucleotide sequence present in a host eukaryotic cell, distinguishing it from sequences present in prokaryotic cells or viruses. This sequence can be located within genes, exons, introns, non-coding regions (e.g., promoters or enhancers), optional markers, tags, signals, or similar elements.
[0194] In some embodiments, the guide nucleic acid or the nucleic acid encoding the guide nucleic acid includes the nucleotide sequences described herein (e.g., Tables 4, 5, 6, or 8). These nucleotide sequences may be characterized as DNA or RNA; however, it should be understood that such sequences may be adapted as needed to describe sequences within the guide nucleic acid or sequences encoding it, such as nucleotide sequences for vectors. Furthermore, the disclosure of these nucleotide sequences also includes their complementary sequences, reverse sequences, and reverse complementary sequences, any of which may be used as nucleotide sequences for use in the guide nucleic acids described herein.
[0195] In some embodiments, the compositions, systems, and methods provided herein include the sequences listed in Tables 1 and 2, or any combination thereof; and also include one or more sequence modifications or mutations. Sequence mutations or modifications may include modifications to one or more nucleic acid residues of a nucleotide sequence, such as chemical modifications to one or more nucleobases; or chemical modifications to the phosphate backbone, nucleotide, nucleobase, or nucleoside. Such modifications may be made to the guide nucleic acid nucleotide sequences disclosed herein, or any sequence (e.g., nucleic acids encoding effector proteins or nucleic acids that produce guide nucleic acids during transcription). Methods for modifying nucleic acids are known. Those skilled in the art will understand that sequence modifications can be located at any position on the nucleic acid without significantly reducing the function of the nucleic acid, composition, or system. The nucleic acids provided herein can be prepared according to any available technology, including but not limited to chemical synthesis, enzymatic synthesis (often referred to as vv / ro-transcription), cloning, enzymatic or chemical cleavage, etc. In some instances, the nucleic acids provided herein are not uniformly modified along the full length of the molecule. Different nucleotide modifications and / or backbone structures may be present at various positions within the nucleic acid.
[0196] Those skilled in the art will understand that when discussing nucleic acid molecules containing multiple residues, the terms nucleotide and / or nucleoside are used interchangeably to refer to the sugar and base components of these residues. Similarly, those skilled in the art will understand that, in the context of nucleic acids with multiple linked residues, the terms linked nucleotides and / or linked nucleosides are also interchangeable terms describing linked sugars and bases present in the nucleic acid molecule. When discussing nucleobases within or linked to nucleic acid molecules, this can be interpreted as referring to the base components of the residues in the nucleic acid, such as nucleotides, nucleosides, or linked nucleotides or nucleosides.
[0197] Those skilled in the art will also recognize the distinction between RNA and DNA, particularly the substitution of thymidine for uridine or vice versa. The presence of nucleoside analogs (such as modified uridine) does not affect the identity or complementarity between polynucleotides, provided that the relevant nucleotides (such as thymidine, uridine, or modified uridine) have corresponding complements (e.g., adenosine complements of all forms of thymidine, uridine, or modified uridine; similarly, cytosine and 5-methylcytosine have guanosine or modified guanosine as their complements). Therefore, the sequence 5'-AXG, where X represents any modified uridine (such as pseudouridine, N1-methylpseudouridine, or 5-methoxyuridine), is considered to have 100% identity with AUG because the two sequences are perfectly complementary to the same sequence (5'-CAU).
[0198] In some implementations, the guide nucleic acid or the nucleic acid encoding the guide nucleic acid may include a sequence having at least 70%, 80%, 85%, 90%, 95%, or even 100% identity with any nucleotide sequence listed in Table 2.
[0199] Exemplary PAM sequences useful for the effectors and their variants listed in Table 1 are disclosed in WO2022258753.
[0200] It should be understood that guide nucleic acids and any components thereof (e.g., crRNA, spacer regions, repeat regions, etc.) may include deoxynucleotides, ribonucleotides, biochemically or chemically modified nucleotides (e.g., one or more sequence modifications described herein) or any combination thereof.
[0201] In some implementations, the guide nucleic acid may include additional elements that impart additional functionality to the guide nucleic acid (e.g., stability, heat resistance, etc.), which may be one or more nucleotide alterations, nucleotide sequences, intermolecular secondary structures, or intramolecular secondary structures (e.g., one or more hairpin regions, one or more convex loops, etc.).
[0202] Interval Typically, a guide nucleic acid includes a spacer region that hybridizes to a target sequence of a target nucleic acid. The spacer region may include complementarity with the target sequence of the target nucleic acid (e.g., hybridization). The spacer region sequence can function to guide the guide nucleic acid to the target nucleic acid for detection and / or modification. The spacer region sequence can be engineered (e.g., through genetic engineering) to hybridize with any desired target sequence within the target nucleic acid (e.g., while also taking PAM into account).
[0203] In some implementations, the spacer region can function to guide the RNP complex, including the guide nucleic acid, to the target nucleic acid for detection and / or modification. The spacer region can function to guide the RNP complex to the target nucleic acid for detection and / or modification. The spacer region can be complementary to a target sequence adjacent to a PAM that can be recognized by the effector protein described herein.
[0204] In some embodiments, the spacer region consists of 15 to 28 linked nucleosides. In other instances, the spacer region may contain 15 to 26, 15 to 24, 15 to 22, 15 to 20, 15 to 18, 16 to 28, 16 to 26, 16 to 24, 16 to 22, 16 to 20, 16 to 18, 17 to 26, 17 to 24, 17 to 22, 17 to 20, 17 to 18, 18 to 26, 18 to 24, or 18 to 22 linked nucleosides. Furthermore, the spacer region may be specified as having a length of 18 to 24 linked nucleosides. In some cases, the length of the spacer region is at least 15 linked nucleosides. Its length may also be at least 16, 18, 20, or 22 linked nucleosides. The spacer region may include at least 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In some embodiments, the spacer region has a minimum length of 17 linked nucleotides. Its length may also be at least 18 or 20 linked nucleotides.
[0205] In some cases, the spacer region exhibits at least 80%, 85%, 90%, 95%, or even 100% complementarity with the target sequence of the target nucleic acid. The term "% complementarity" refers to the percentage of nucleotides in two equal-length sequences that can form stable base pairs in an antiparallel manner at two or more corresponding positions. Therefore, a nucleic acid sequence lacking complete complementarity throughout its entire length contains one or more mismatches, which can occur anywhere between non-complementary opposite nucleotides. The complementarity percentage is determined by dividing the total number of complementary residues by the total number of nucleotides in one of the equal-length sequences and multiplying the result by 100. A nucleotide sequence exhibiting complete or overall complementarity means that 100% of its residues are complementary to residues in the reference nucleotide sequence. Conversely, a partially complementary nucleotide sequence has at least 20% but less than 100% of its residues complementary to residues in the reference sequence. In some instances, at least 50% but less than 100% of the nucleotide sequence residues are complementary to residues in the reference nucleotide sequence. Furthermore, it is possible for at least 70%, 80%, 90%, or 95% but less than 100% of the nucleotide sequence residues to be complementary to the reference sequence. A non-complementary nucleotide sequence is defined as one in which less than 20% of the residues are complementary to residues in a reference sequence. In some embodiments, the spacer region is completely complementary to the target sequence of the target nucleic acid. Furthermore, in some instances, the spacer region comprises at least 15 consecutive nucleotides complementary to the target nucleic acid, or at least 17 consecutive nucleotides complementary to the target nucleic acid.
[0206] It should be understood that the sequence of the spacer region does not need to be 100% complementary to the target sequence of the target nucleic acid in order to hybridize with or achieve specific hybridization with the target sequence. The guide nucleic acid may include at least one uracil within 5 to 20 nucleic acid residues in the spacer region, which is not complementary to the corresponding nucleotide in the target sequence. Furthermore, the guide nucleic acid may contain at least one uracil between 5 to 9, 10 to 14, or 15 to 20 nucleic acid residues in the spacer region, which does not match the corresponding nucleotide in the target sequence. In some cases, the target nucleic acid portion complementary to the spacer region may have epigenetic or post-transcriptional modifications. Such modifications may include, but are not limited to, acetylation, methylation, or thiol modifications.
[0207] Typically, the spacer sequence consists of a spacer region that hybridizes with the target sequence of the target nucleic acid. Therefore, in some embodiments, the spacer sequence, or a portion thereof, is capable of hybridizing with the target sequence of the target nucleic acid. In some instances, the spacer sequence may be an RNA sequence having at least 65%, 70%, 80%, 90%, 92%, 95%, 97%, 99%, or even 100% identity with any of SEQ ID NO: 1-61 or 216-261.
[0208] Typically, guide nucleic acids contain repeat regions that interact with effector proteins. These repeat regions, also referred to as "protein-binding fragments," are usually located near the spacer region. For example, guide RNAs that bind to effector proteins include a repeat region located at the 5' end of the spacer region. In some instances, the repeat region is positioned before the spacer region in a 5' to 3' orientation. The repeat region can be 15 to 50 nucleotides long, with some embodiments specifying a length of 19 to 37 nucleotides. Furthermore, guide nucleic acids can contain more than one repeat region. In some embodiments, the guide nucleic acid consists of a first repeat sequence in a 5' to 3' orientation, followed by a spacer region sequence and a second repeat sequence. The first and second repeat sequences can be identical or different from each other. The spacer region sequence and the repeat sequence can be directly linked, or a short linker of 1, 2, or 3 nucleotides can exist between them. In some cases, the spacer region sequence and the repeat sequence of the guide nucleic acid can exist within a single molecule linked by base-pairing interactions.
[0209] In some implementations, the repetitive sequence is adjacent to the intermediate RNA, which may be located at the 3' end of the repetitive sequence. The intermediate RNA may then be ligated to the repetitive sequence in the 5' to 3' direction, and the repetitive sequence may then be ligated to the spacer sequence. The repetitive sequence may be ligated directly to the spacer sequence and / or the intermediate RNA, or via a suitable adapter.
[0210] In some implementations, the protein-binding fragment may consist of two complementary repeat sequences that hybridize to form a double-stranded RNA duplex (dsRNA duplex). This dsRNA duplex region may contain 5-25 base pairs (bp), and not all nucleotides in the duplex region need to be paired, allowing for the presence of convex loops. The repeating region that may include dsRNA may contain one or more convex loops. Furthermore, the repeating region may form a hairpin structure, particularly in the 3' portion, which includes a double-stranded stem loop and a single-stranded loop. In this case, one strand of the stem may contain a sequence that is at least partially complementary to the other strand.
[0211] Connectors for nucleic acids In some embodiments, the guide nucleic acid used in the compositions, systems, and methods described herein may include one or more adapters, or nucleic acids encoding one or more adapters. The guide nucleic acid may contain at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 adapters. The guide nucleic acid may also include more than one adapter, at least two of which may be identical or different.
[0212] The adapter may consist of 1 to 10, 1 to 7, 1 to 5, 1 to 3, 2 to 10, 2 to 8, 2 to 6, 2 to 4, 3 to 10, 3 to 7, 3 to 5, 4 to 10, 4 to 8, 4 to 6, 5 to 10, 5 to 7, 6 to 10, 6 to 8, 7 to 10, or 8 to 10 linked nucleotides. In some embodiments, the adapter may have a 5'-GAAA-3' nucleotide sequence.
[0213] The guide nucleic acid may include one or more adapters for linking one or more repetitive sequences. The guide nucleic acid may also include adapters for linking one or more repetitive sequences to one or more spacer sequences, and in some cases, at least two repetitive sequences may be linked by adapters.
[0214] intermediate RNA The guide nucleic acid described herein may include one or more intermediate RNAs. Intermediate RNA is generally a nucleotide sequence present in the stalk sequence that can non-covalently bind to an effector protein and form a complex (e.g., a ribonucleoprotein (RNP) complex). Typically, intermediate RNA is neither transactivated nor does it exert a transactivating effect. Intermediate RNA may also be referred to as an intermediate sequence, which may include deoxyribonucleotides in addition to ribonucleotides and / or modified bases. Intermediate RNA binds to effector proteins primarily in a non-covalent manner, and in some embodiments, such as in a cellular environment, it can form secondary structures that allow effector proteins to bind to.
[0215] The length of the intermediate RNA is variable, with some embodiments specifying a minimum length of at least 30, 50, 70, 90, 110, 130, 150, 170, 190, or 210 linked nucleotides. In other instances, the length of the intermediate RNA is limited to a maximum of 30, 50, 70, 90, 110, 130, 150, 170, 190, or 210 linked nucleotides, specifically in ranges such as approximately 30 to 210, 60 to 210, 90 to 210, 120 to 210, 150 to 210, 180 to 210, 30 to 180, 60 to 180, 90 to 180, 120 to 180, or 150 to 180 linked nucleotides.
[0216] Intermediate RNAs can also employ secondary structures (such as one or more hairpin loops), which facilitate the binding of effector proteins to guide nucleic acids and / or enhance the modification activity of effector proteins on target nucleic acids. Intermediate RNAs typically consist of a 5' region, a hairpin region, and a 3' region, where the 5' region may hybridize with the 3' region; in some cases, the 5' region may not hybridize with the 3' region.
[0217] The hairpin region may include a first sequence and a second sequence that is anticomplementary to the first sequence, connected by a stem-loop structure. The stem region may include 4 to 8 linked nucleotides, typically 5 to 6 or 4 to 5 linked nucleotides in length. The intermediate RNA may also contain a pseudoknot, a secondary structure involving hybridization with another stem or half-stem structure. Effector proteins can interact with intermediate RNAs containing a single stem region or multiple stem regions, where the nucleotide sequences of these regions may be identical or different.
[0218] In some embodiments, the terms "intermediate RNA" and "intermediate sequence" refer to a nucleotide sequence located within the stalk sequence that can non-covalently bind to an effector protein to form a complex (e.g., a ribonucleoprotein (RNP) complex). In the systems, methods, and compositions described herein, the intermediate sequence is not considered a trans-activating nucleic acid.
[0219] Single nucleic acid system In some embodiments, the compositions, systems, and methods described herein can incorporate a single nucleic acid system comprising a guide nucleic acid or a nucleotide sequence encoding the guide nucleic acid, and one or more effector proteins or nucleotide sequences encoding those effector proteins. Within this single nucleic acid system, a first region (FR) of the guide nucleic acid non-covalently interacts with the effector protein. A second region (SR) of the guide nucleic acid hybridizes with the target sequence of the target nucleic acid. Notably, in this arrangement, the effector protein is not trans-activated by the guide nucleic acid, indicating that the activity of the effector protein is independent of binding to a second non-target nucleic acid molecule. Examples of guide nucleic acids suitable for this single nucleic acid system include crRNA or sgRNA.
[0220] Guide nucleic acids and their components can originate from CRISPR arrays present in the host organism's genome. crRNA can be produced by cleaving longer precursor CRISPR RNA (precursor crRNA) within each direct repeat sequence, resulting in shorter, mature crRNA. Various mechanisms exist for crRNA production, including the action of specific endonucleases (e.g., Cas6 or Cas5d in type I and type III systems), coupling of host endonucleases (e.g., RNase III) with tracrRNA (type II system), or the intrinsic ribonuclease activity of the effector protein itself (e.g., Cpfl in type V system). Furthermore, crRNA can be produced independently of precursor crRNA processing and can be directly associated with effector proteins in vivo or in vitro.
[0221] In some embodiments, crRNA is used as a guide nucleic acid in the single nucleic acid systems of the compositions, methods, and systems described herein. In this case, the guide nucleic acid comprises crRNA, wherein a repetitive sequence is capable of linking the crRNA to an effector protein. In some cases, the guide nucleic acid may include crRNA linked to another nucleotide sequence that can non-covalently bind to an effector protein. In these examples, the repetitive sequence of the crRNA may be linked to an intermediate RNA. Thus, a single nucleic acid system may include a guide nucleic acid comprising crRNA and an intermediate RNA.
[0222] Typically, crRNA may include a spacer region that hybridizes to the target sequence of the target nucleic acid. In some embodiments, the crRNA of the guide nucleic acid includes a repeat region and a spacer region, wherein the repeat region binds to an effector protein and the spacer region hybridizes to the target sequence of the target nucleic acid.
[0223] In some embodiments, the composition includes a guide RNA and an effector protein (e.g., a single nucleic acid system) without tracrRNA, wherein the guide RNA is sgRNA. sgRNA may include deoxyribonucleosides, ribonucleosides, chemically modified nucleosides, deoxyribonucleotides, ribonucleotides, chemically modified nucleotides, or any combination thereof. sgRNA may also include a nucleotide sequence forming a secondary structure (e.g., one or more hairpin loops) that promotes the binding of the effector protein to the sgRNA and / or the modification activity of the effector protein on a target nucleic acid (e.g., a hairpin region). Such a sequence may be included within the stalk sequence described herein. In some embodiments, sgRNA includes one or more of the following: a stalk sequence, intermediate RNA, crRNA, repeat sequences, spacer sequences, adapters, or combinations thereof. For example, sgRNA includes a stalk sequence and a spacer sequence; intermediate RNA and crRNA; intermediate RNA, repeat sequences, and spacer sequences, etc.
[0224] In some embodiments, the sgRNA comprises an intermediate RNA and a crRNA. In some embodiments, the intermediate RNA is located at the 5' end of the crRNA in the sgRNA. In some embodiments, the sgRNA comprises a linked intermediate RNA and crRNA. In some embodiments, the intermediate RNA and crRNA are directly linked within the sgRNA (e.g., covalently linked, such as via a phosphodiester bond). In some embodiments, the intermediate RNA and crRNA are linked within the sgRNA via any suitable adapter, examples of which are provided herein.
[0225] In some implementations, the sgRNA consists of a stalk sequence and a spacer region sequence. In some cases, the stalk sequence is located at the 5' end of the spacer region sequence within the sgRNA. The sgRNA may contain a ligated stalk sequence and a spacer region sequence, wherein the two components are directly linked (e.g., covalently linked, such as via a phosphodiester bond). Alternatively, the stalk sequence and the spacer region sequence may be linked by any suitable adapter, examples of which are provided herein.
[0226] In some instances, sgRNA may also include intermediate RNA, repetitive sequences, and spacer sequences. The intermediate RNA may be located at the 5' end of the repetitive sequence in the sgRNA. Furthermore, the sgRNA may have linked intermediate RNA and repetitive sequences, wherein these elements are directly linked (e.g., covalently linked by phosphodiester bonds) or linked by suitable adapters, as described above.
[0227] Furthermore, in some embodiments, the repeat sequence is located at the 5' end of the spacer sequence within the sgRNA. The sgRNA may also include linked repeat sequences and spacer sequences, which may be directly linked (e.g., covalently linked via phosphodiester bonds) or linked via suitable adapters, as described in the provided examples.
[0228] Dual nucleic acid system In some embodiments, the compositions, systems, and methods described herein may relate to a dual nucleic acid system comprising crRNA or a nucleotide sequence encoding crRNA, tracrRNA or a nucleotide sequence encoding tracrRNA, and one or more effector proteins or nucleotide sequences encoding those effector proteins. In this dual nucleic acid system, crRNA and tracrRNA exist as separate, unconnected molecules. The repeat hybridization region of tracrRNA is capable of hybridizing with an isolength portion of crRNA to form a tracrRNA-crRNA duplex. Notably, the isolength portion of crRNA does not contain a spacer region sequence; instead, it is capable of hybridizing with a target sequence within the target nucleic acid.
[0229] In a dual nucleic acid system, the complex consists of a guide nucleic acid, tracrRNA, and an effector protein. The effector protein is transactivated by tracrRNA, meaning that the activity of the effector protein depends on its binding to the tracrRNA molecule.
[0230] In some implementations, the repeat hybridization sequence is located at the 3' end of the tracrRNA. The length of this repeat hybridization sequence is variable, and may be approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 14, 16, 18, or 20 linked nucleotides. In some cases, the length of the repeat hybridization sequence can be from 1 to 20 linked nucleotides.
[0231] The tracrRNA and / or tracrRNA-crRNA duplex can employ a secondary structure that enhances the binding of the effector protein to the tracrRNA or tracrRNA-crRNA complex. In some embodiments, this secondary structure can improve the activity of the effector protein on the target nucleic acid. The secondary structure can consist of a stem-loop configuration, comprising a stem region and a loop region. In some instances, the stem region can be 4 to 8 linked nucleotides in length, with specific embodiments specifying that the stem region can be 5 to 6 linked nucleotides or 4 to 5 linked nucleotides in length.
[0232] Furthermore, secondary structures can have pseudoknots, defined as a configuration in which at least part of the stem hybridizes with a second stem or half-stem structure. Effector proteins can interact with secondary structures comprising multiple stem regions. In some cases, the nucleotide sequences of these multiple stem regions may be identical, while in other instances, at least one stem region may have a different sequence.
[0233] The secondary structure may include at least 2, 3, 4, or 5 stem regions. Furthermore, the secondary structure may include one or more loops, wherein the number of loops may be at least 1, 2, 3, 4, or 5.
[0234] tracrRNA tracrRNA and / or tracrRNA-crRNA duplexes can form secondary structures in which auxiliary effector proteins bind to tracrRNA or tracrRNA-crRNA complexes. In some instances, this secondary structure can improve the activity of the effector protein on the target nucleic acid. The secondary structure can consist of a stem-loop conformation comprising a stem region and a loop region. The length of the stem region is variable, typically 4 to 8 linked nucleotides, with some embodiments specifying a length of 5 to 6 or 4 to 5 linked nucleotides.
[0235] Furthermore, tracrRNA may have pseudoknots, which are secondary structures involving at least a portion of the stem hybridizing with another stem or half-stem structure. Effector proteins can recognize tracrRNAs comprising multiple stem regions, which may have the same or different nucleotide sequences. In some embodiments, tracrRNA may comprise at least 2, 3, 4, or 5 stem regions.
[0236] The length of tracrRNA is variable, with some implementations specifying that it should not exceed 50, 56, 68, 71, 73, 95, or 105 linked nucleotides. Other implementations may specify a tracrRNA length of approximately 30 to 120 linked nucleotides, or more specifically 50 to 105, 50 to 95, 50 to 73, 50 to 71, 50 to 68, or 50 to 56 linked nucleotides. In some cases, the length of tracrRNA may fall within 56 to 105 linked nucleotides, or even within 40 to 60 nucleotides.
[0237] An exemplary tracrRNA may include a 5' region, a hairpin region, a repeat hybridization region, and a 3' region in a 5' to 3' orientation. In some instances, the 5' region may hybridize with the 3' region, while in others it may not. The 3' region may be covalently linked to the crRNA (e.g., via a phosphodiester bond). Furthermore, the tracrRNA may contain an unhybridized region at its 3' end, which may be approximately 1 to 20 linked nucleotides in length.
[0238] In some embodiments, the composition including the effector protein and guide RNA may not contain tracrRNA. In other instances, the effector protein may not require tracrRNA to locate or cleave the target nucleic acid. The crRNA in the guide nucleic acid may consist of a repeat region and a spacer region, wherein the repeat region binds to the effector protein, and the spacer region hybridizes to the target sequence within the target nucleic acid. The repeat sequence of the crRNA can interact with the effector protein, promoting the formation of the ribonucleoprotein (RNP) complex.
[0239] Furthermore, repeat regions can also be referred to as "protein-binding fragments," and they are typically adjacent to spacer regions. For example, guide RNAs that interact with effector proteins may include repeat regions located at the 5' end of the spacer region.
[0240] V. Modify The effector proteins and nucleic acids (e.g., engineered guide nucleic acids) described herein can be further modified as described throughout and further herein. Examples include targeted modifications that do not alter the primary sequence, including chemical derivatization of effector proteins, such as acylation, acetylation, carboxylation, amidation, etc. Glycosylation modifications are also included, such as modifications resulting from altering the glycosylation pattern of the effector protein during its synthesis and processing or in further processing steps; for example, by exposing the effector protein to enzymes that influence glycosylation, such as mammalian glycosylation or deglycosylation enzymes. Sequences containing phosphorylated amino acid residues, such as phosphotyrosine, phosphotyserine, or phosphotythreonine, are also included.
[0241] The modifications described herein may also include alterations to effector proteins and / or engineered guide nucleic acids by any suitable method, including molecular biology techniques and synthetic chemistry. These modifications aim to enhance resistance to proteolytic degradation, modulate target sequence specificity, optimize solubility, alter protein activity (such as transcriptional regulation or enzymatic activity), or better suit them for their intended application. Analogs of these effector proteins may include those containing residues other than naturally occurring L-amino acids, such as D-amino acids, or synthetic non-naturally occurring amino acids. D-amino acids may replace some or all of the amino acid residues. Furthermore, modifications may involve the introduction of non-naturally occurring amino acids. The specific sequence and method of preparation will depend on factors such as convenience, cost, desired purity, and other considerations.
[0242] Further modifications may involve incorporating various functional groups into the effector proteins and / or engineered guide nucleic acids described herein. For example, these groups can be added during the synthesis or expression of the effector protein to enable linkage with other molecules or surfaces. For instance, cysteine residues can be used to form thioethers, histidine residues can be used to bind to metal ion complexes, carboxyl groups can promote the formation of amides or esters, and amino groups can also be used to manufacture amides.
[0243] Furthermore, modifications may include alterations to nucleic acids (e.g., engineered guide nucleic acids) as described herein to impart new or enhanced characteristics, such as improved stability. Such modifications may involve base modifications, backbone modifications, sugar modifications, or combinations thereof affecting one or more nucleotides, nucleosides, or nucleobases within the nucleic acid.
[0244] In some embodiments, the nucleic acids described herein (e.g., engineered guide nucleic acids) include one or more modifications comprising: 2'O-methyl modified nucleotides, 2' fluorine modified nucleotides; locked nucleic acid (LNA) modified nucleotides; peptide nucleic acid (PNA) modified nucleotides; nucleotides having a thiophosphate bond; 5' caps (e.g., 7-methylguanosine cap (m7G)), thiophosphates, chiral thiophosphates, dithiophosphates, triphosphates, aminoalkyl phosphates, methylphosphonates, and other alkylphosphonates, including 3'-alkylene phosphonates, 5'-alkylene phosphonates, and chiral phosphonates, phosphonates, phosphoramides, including 3'-aminophosphonamides and aminoalkylphosphonamides, diphosphonamides, thiophosphonamides, thioalkylphosphonamides, etc. Thioalkyl phosphate triesters, selenophosphates, and borophosphates, having normal 3'-5' bonds, their 2'-5' linked analogues, and those with reverse polarity (where one or more nucleotide internal bonds are 3' to 3', 5' to 5', or 2' to 2' bonds); thiophosphate and / or heteroatomic nucleoside internal bonds, such as -CH2-NH-O-CH2-, -CH2-N(CH3)-O-CH2- (referred to as the methylene (methylimino) or MMI backbone), -CH2-ON(CH)-CH2-, -CH2-N(CH3)-N(CH3)-CH2-, and -ON(CH3)-CH2-CH2- (where the native phosphodiester nucleotide internal bond is represented as -OP(=O)(OH)-O). -CH2-); morpholine bonds (partially formed by the sugar moiety of nucleosides); morpholine skeletons; diphosphamide or other non-phosphodiester nucleoside intra-bonds; siloxane skeletons; sulfide, sulfoxide, and sulfone skeletons; formylacetyl and thioformylacetyl skeletons; methyleneformylacetyl and thioformylacetyl skeletons; ribose acetyl skeletons; olefin-containing skeletons; aminosulfonate skeletons; methyleneimine and methylenehydrazine skeletons; sulfonate and sulfonamide skeletons; amide skeletons; other skeleton modifications having mixed N, O, S, and CH2 component moieties; and combinations thereof.
[0245] Vectors and multiple expression vectors The compositions, systems, and methods described herein may include vectors or their applications. Vectors may contain a target nucleic acid. In some embodiments, the target nucleic acid comprises one or more components of the compositions or systems outlined herein. The nucleic acid may consist of nucleotide sequences encoding one or more components of the compositions or systems described herein.
[0246] Components may include effector proteins, guide nucleic acids, target nucleic acids, and donor nucleic acids. In some instances, a component may be a nucleic acid encoding an effector protein, a donor nucleic acid, a guide nucleic acid, or a nucleic acid encoding a guide nucleic acid. A vector may be part of a vector system comprising a vector library, each vector being designed to encode one or more components of the compositions or systems described herein.
[0247] In some embodiments, the components (including effector proteins, guide nucleic acids, and / or target nucleic acids) may be entirely encoded by a single vector. Alternatively, these components may be encoded by different vectors within the system. In some cases, the vector may include a nucleotide sequence encoding one or more effector proteins described herein. In specific embodiments, one or more effector proteins may consist of at least two effector proteins. These effector proteins may be identical in some instances, or different from each other in others.
[0248] The nucleotide sequence within the vector is typically operatively linked to a promoter that functions in target cells (e.g., eukaryotic cells). In some embodiments, the vector may encode 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more effector proteins.
[0249] In some embodiments, the fusion effector protein described herein may be inserted into a vector. The vector may also contain one or more promoters, enhancers, ribosome binding sites, RNA splicing sites, polyadenylation sites, origin of replication, and / or transcription termination sequences. In some embodiments, the vector may encode one or more of any system components, including but not limited to the effector protein, guide nucleic acid, donor nucleic acid, and target nucleic acid described herein. In some embodiments, the system component encoding the sequence is operatively linked to a promoter operable in target cells (e.g., eukaryotic cells). In some embodiments, the vector may encode one, two, three, four, or more of any system components. For example, the vector may encode two or more guide nucleic acids, each guide nucleic acid comprising a different sequence. The vector may encode both effector proteins and guide nucleic acids. The vector may encode effector proteins, guide nucleic acids, and donor nucleic acids.
[0250] In some embodiments, the vector comprises one or more guide nucleic acids, or nucleotide sequences encoding one or more guide nucleic acids described herein. In some embodiments, the one or more guide nucleic acids comprise at least two guide nucleic acids. In some embodiments, at least two guide nucleic acids are identical. In some embodiments, at least two guide nucleic acids are different from each other. In some embodiments, the guide nucleic acid or the nucleotide sequence encoding the guide nucleic acid is operatively linked to a promoter operable in target cells (e.g., eukaryotic cells). In some embodiments, the vector comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more guide nucleic acids. In some implementations, the vector includes nucleotide sequences encoding 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more guide nucleic acids.
[0251] In some embodiments, the vector comprises one or more donor nucleic acids as described herein. In some embodiments, the one or more donor nucleic acids comprise at least two donor nucleic acids. In some embodiments, at least two donor nucleic acids are identical. In some embodiments, at least two donor nucleic acids are different from each other. In some embodiments, the vector comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more donor nucleic acids.
[0252] In some embodiments, the vector may include or encode one or more regulatory elements. These regulatory elements are sequences that control transcription and translation, such as promoters, enhancers, polyadenylated signals, terminators, and protein degradation signals. These elements promote and / or regulate the transcription of non-coding sequences (e.g., guide nucleic acids) or coding sequences (e.g., effector proteins, fusion proteins, etc.), and they may also regulate the translation of encoded effector proteins.
[0253] In addition, the vector may include or encode one or more complementary elements, including an origin of replication, an antibiotic resistance gene (or the nucleic acid encoding it), a tag (or the nucleic acid encoding it), a selectable marker, and similar components. In some instances, the vector may also include elements such as ribosome binding sites and RNA splicing sites.
[0254] The vectors described herein can encode promoters, which are regulatory regions on nucleic acids (e.g., DNA sequences) capable of initiating transcription downstream (3' direction) of coding or non-coding sequences. The promoter can be linked at its 3' end to the nucleic acid to be expressed or transcribed, extending upstream (5' direction) to contain bases or elements required to initiate transcription or induce expression, which can be measured at a detectable level. Promoters comprise nucleotide sequences, referred to as "promoter sequences," which may include a transcription start site and one or more protein-binding domains, such as RNA polymerases, essential for binding to transcriptional mechanisms. Eukaryotic promoters may contain elements such as "TATA" boxes and "CAT" boxes. A variety of promoters, including inducible promoters, can be used to drive the expression or transcriptional activation of a target nucleic acid. Thus, in some embodiments, the target nucleic acid can be operatively linked to a promoter. The promoter can be any suitable type of promoter designed for the compositions, systems, and methods described herein. Examples include constitutively active promoters (e.g., CMV promoters), inducible promoters (e.g., heat shock promoters, tetracycline-regulated promoters, steroid-regulated promoters, metal-regulated promoters, estrogen receptor-regulated promoters, etc.), and spatially restricted and / or time-restricted promoters (e.g., tissue-specific promoters, cell-type-specific promoters, etc.). Suitable promoters include, but are not limited to: SV40 early promoter, mouse mammary tumor virus long terminal repeat (LTR) promoter; adenovirus major late promoter (Ad MLP); herpes simplex virus (HSV) promoters, cytomegalovirus (CMV) promoters, such as the CMV immediate early promoter region (CMVIE), Rous sarcoma virus (RSV) promoter, human U6 small nucleus promoter (U6), enhanced U6 promoter, and human H1 promoter (H1). Through transcriptional activation, the aim is to increase transcription in target cells by 2, 5, 10, 50, 100, 500, or 1000 times or more compared to baseline levels. In addition, a vector for delivering nucleic acids to cells, wherein the nucleic acids are transcribed to produce guide nucleic acids and / or nucleic acids encoding effector proteins, the vector may include a nucleic acid sequence encoding a selectively tagged nucleic acid in a target cell to identify cells that have already absorbed guide nucleic acids and / or effector proteins.
[0255] Typically, the plasmids and vectors described herein contain at least one promoter. In some embodiments, these promoters are constitutive promoters, while in other instances they are inducible promoters. Furthermore, some embodiments may have prokaryotic promoters that drive gene expression in prokaryotic cells, while other embodiments may include eukaryotic promoters that promote gene expression in eukaryotic cells.
[0256] Exemplary promoters include, but are not limited to, ApoE, TBG, CMV, EF1α, SV40, PGK1, Ubc, human β-actin, CAG, TRE, UAS, Ac5, polyhedrome, CaMKIIα, GALI-10, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, CaMV35S, SV40, CMV, and HSV TK promoters. In some cases, the promoter can be CMV, EF1α, ApoE, TBG, or ubiquitin.
[0257] In addition, the vector may be bicistronic or polycistronic, meaning that it contains two or more loci responsible for producing the protein, and may include internal ribosome entry sites (IRES) for cap-independent translation initiation.
[0258] Typically, the vectors provided herein comprise at least one promoter or combination of promoters that drive the expression or transcription of one or more genome editing tools described herein. In some embodiments, the vector comprises a nucleotide sequence for a promoter, and in other instances, the vector may contain two or three promoters. The length of the promoter is variable, and in some cases less than about 500, 400, 300, or 200 linked nucleotides. Conversely, in other embodiments, the length of the promoter may be at least 100, 200, 300, 400, or 500 linked nucleotides. Non-limiting examples of promoters include CMV, 7SK, EF1α, RPBSA, hPGK, EFS, SV40, PGK1, Ubc, human β-actin, CAG, TRE, UAS, Ac5, polyhedron, CaMKIIα, GALI-10, H1, TEF1, GDS, ADH1, CaMV35S, HSV TK, Ubi, U6, MNDU3, MSCV, MND, and CAG. In some embodiments, the promoter is a constitutive promoter. In some embodiments, the promoter is an inducible promoter. In some embodiments, an inducible promoter drives the expression of its corresponding coding sequence (e.g., effector protein or guide nucleic acid) only in the presence of a signal (e.g., hormone, small molecule, peptide). Non-limiting examples of inducible promoters include the T7 RNA polymerase promoter, the T3 RNA polymerase promoter, the isopropyl-β-D-thiogalactoside (IPTG) regulatory promoter, the lactose-inducible promoter, the heat shock promoter, the tetracycline-regulated promoter (tetracycline-inducible or tetracycline-inhibitory), the steroid-regulated promoter, the metal-regulated promoter, and the estrogen receptor-regulated promoter.
[0259] In some embodiments, the promoter is an activation-inducible promoter, such as the CD69 promoter; in other embodiments, the promoter used to express the effector protein is a pan-expression promoter. In some embodiments, the pan-expression promoter includes an MND or CAG promoter sequence.
[0260] In some embodiments, the vector described herein is a nucleic acid expression vector. In some embodiments, the vector described herein is a recombinant expression vector. In some embodiments, the vector described herein is a messenger RNA.
[0261] In some embodiments, the vector described herein functions as a delivery vector. This delivery vector can be a eukaryotic vector, a prokaryotic vector (e.g., a bacterial vector), a viral vector, or any combination of these vectors; in some cases, the delivery vector can be a non-viral vector. Furthermore, it can take the form of a plasmid, which can include DNA or RNA. Examples of plasmids include circular double-stranded DNA or linear plasmids.
[0262] Plasmids may contain one or more target genes as well as various regulatory elements. Typically, plasmids include a bacterial backbone with an origin of replication and an antibiotic resistance gene or another selective marker to promote plasmid amplification in bacteria. In some cases, plasmids may be microcircular plasmids. Furthermore, plasmids may contain genes that provide selective markers to promote plasmid retention in target cells. The formulation used for delivery may include formulations for injection using a syringe or for electroporation. Plasmids may also be engineered using synthetic methods or other suitable techniques known in the art. For example, genetic elements can be assembled by restrictive digestion of a donor plasmid or a desired gene sequence in an organism, resulting in ends that can be readily attached to another genetic sequence.
[0263] Typically, guide nucleic acids include repeat regions that interact with effector proteins. These repeat regions, sometimes referred to as "protein-binding fragments," are usually located near the spacer region. For example, guide RNAs that interact with effector proteins have repeat regions located at the 5' end of the spacer region. In some cases, the repeat regions are located before the spacer region in a 5' to 3' orientation. The length of the repeat regions is variable, typically 15 to 50 nucleotides, with some implementations specifying a length of 19 to 37 nucleotides. Furthermore, guide nucleic acids may contain multiple repeat regions.
[0264] In specific embodiments, the guide nucleic acid may consist of a first repeat sequence in the 5' to 3' orientation, followed by a spacer sequence and a second repeat sequence. The first and second repeat sequences may be identical or different from each other. The spacer sequence and the repeat sequence may be directly linked, or there may be a short linker of 1, 2, or 3 nucleotides between them. In some instances, the spacer sequence and the repeat sequence may be present in a single molecule linked by base pairing interactions.
[0265] Furthermore, the repetitive sequence is adjacent to the intermediate RNA, which may be located at the 3' end of the repetitive sequence. The intermediate RNA may then be ligated to the repetitive sequence in the 5' to 3' direction, and the repetitive sequence may then be ligated to the spacer sequence. The repetitive sequence may be ligated directly to the spacer sequence and / or the intermediate RNA, or via a suitable adapter.
[0266] In some instances, protein-binding fragments can consist of two complementary repeat sequences that hybridize to form a double-stranded RNA duplex (dsRNA duplex). This dsRNA duplex region can contain 5–25 base pairs (bp), and not all nucleotides in the duplex region need to be paired, allowing for the presence of convex loops. Repeat regions that can include dsRNA may contain one or more convex loops. Furthermore, repeat regions can form hairpin structures, particularly at the 3' portion, which include a double-stranded stem and a single-stranded loop. In this case, one strand of the stem may contain a sequence that is at least partially complementary to the other strand.
[0267] Connectors for nucleic acids In some embodiments, the guide nucleic acid used in the compositions, systems, and methods described herein may introduce one or more adapters, or a nucleic acid encoding one or more adapters. The guide nucleic acid may contain at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 adapters. The guide nucleic acid may also include multiple adapters, at least two of which may be identical or different.
[0268] The adapter can consist of 1 to 10, 1 to 7, 1 to 5, 1 to 3, 2 to 10, 2 to 8, 2 to 6, 2 to 4, 3 to 10, 3 to 7, 3 to 5, 4 to 10, 4 to 8, 4 to 6, 5 to 10, 5 to 7, 6 to 10, 6 to 8, 7 to 10, or 8 to 10 linked nucleotides. In some instances, the adapter can have a 5'-GAAA-3' nucleotide sequence.
[0269] The guide nucleic acid may have one or more adapters for linking various repetitive sequences. The guide nucleic acid may also include adapters for linking one or more repetitive sequences to one or more spacer sequences, and in some cases, at least two repetitive sequences may be linked by adapters.
[0270] intermediate RNA The guide nucleic acids described herein may include one or more intermediate RNAs. Intermediate RNAs are generally nucleotide sequences located in the stalk sequence that can non-covalently bind to effector proteins and form complexes (e.g., ribonucleoprotein (RNP) complexes). Typically, intermediate RNAs are neither transactivated nor do they exert a transactivating effect. Intermediate RNAs can also be referred to as intermediate sequences, which may include deoxyribonucleotides in addition to ribonucleotides and / or modified bases. Intermediate RNAs primarily bind to effector proteins non-covalently, and under certain conditions (such as in the cellular environment), they can form secondary structures that allow effector proteins to bind to them.
[0271] The length of the intermediate RNA is variable, with some embodiments specifying a length of at least 30, 50, 70, 90, 110, 130, 150, 170, 190, or 210 linked nucleotides. In other cases, the length of the intermediate RNA is limited to no more than 30, 50, 70, 90, 110, 130, 150, 170, 190, or 210 linked nucleotides, with specific length ranges such as about 30 to about 210, about 60 to about 210, about 90 to about 210, about 120 to about 210, about 150 to about 210, about 180 to about 210, about 30 to about 180, about 60 to about 180, about 90 to about 180, about 120 to about 180, or about 150 to about 180 linked nucleotides.
[0272] Intermediate RNA can also form secondary structures (such as one or more hairpin loops) that can facilitate the binding of effector proteins to guide nucleic acids and / or enhance the modification activity of effector proteins on target nucleic acids. Intermediate RNA may consist of a 5' region, a hairpin region, and a 3' region, where the 5' region may hybridize with the 3' region. In some embodiments, the 5' region does not hybridize with the 3' region.
[0273] The hairpin region may include a first sequence and a second sequence that is anticomplementary to the first sequence, connected by a stem-loop structure. The stem region may consist of 4 to 8 linked nucleotides, typically 5 to 6 or 4 to 5 linked nucleotides in length. The intermediate RNA may also include a pseudoknot, a secondary structure involving hybridization with another stem or half-stem structure. Effector proteins may interact with intermediate RNAs having a single stem region or multiple stem regions, wherein the nucleotide sequences of these stem regions may be identical or different.
[0274] In some embodiments, "intermediate RNA" and "intermediate sequence" refer to nucleotide sequences located in the stalk sequence that can non-covalently bind to effector proteins to form complexes (e.g., RNP complexes). In the systems, methods, and compositions described herein, the intermediate sequence is not a trans-activating nucleic acid.
[0275] Single nucleic acid system In some embodiments, the compositions, systems, and methods described herein may incorporate a single nucleic acid system comprising a guide nucleic acid or a nucleotide sequence encoding the guide nucleic acid, and one or more effector proteins or nucleotide sequences encoding those effector proteins. Within this single nucleic acid system, a first region (FR) of the guide nucleic acid non-covalently interacts with the effector protein. A second region (SR) of the guide nucleic acid hybridizes with the target sequence of the target nucleic acid. In this arrangement, the effector protein is not trans-activated by the guide nucleic acid, indicating that the activity of the effector protein is independent of binding to a second non-target nucleic acid molecule. Examples of guide nucleic acids suitable for this single nucleic acid system include crRNA or sgRNA.
[0276] Guide nucleic acids and their components can originate from CRISPR arrays present in the host organism's genome. crRNA can be produced by cleaving longer precursor CRISPR RNA (precursor crRNA) within each direct repeat sequence, resulting in shorter, mature crRNA. Various mechanisms exist for crRNA production, including those involving specialized endonucleases (e.g., Cas6 or Cas5d in type I and III systems), coupling of host endonucleases (e.g., RNase III) with tracrRNA (in type II systems), or the intrinsic ribonuclease activity of the effector protein itself (e.g., Cpfl in type V systems). Furthermore, crRNA can be produced independently of precursor crRNA processing and can be directly associated with effector proteins, either in vivo or in vitro.
[0277] In some instances, crRNA functions as a guide nucleic acid in the single-nucleic acid systems of the compositions, methods, and systems described herein. In these cases, the guide nucleic acid comprises crRNA, wherein a repetitive sequence facilitates the linking of the crRNA to an effector protein. In some cases, the guide nucleic acid may consist of crRNA linked to another nucleotide sequence, which can non-covalently bind to an effector protein. In these cases, the repetitive sequence of the crRNA may be linked to an intermediate RNA. Therefore, a single-nucleic acid system may include a guide nucleic acid composed of crRNA and an intermediate RNA.
[0278] The methods, systems, and compositions described herein enable the editing or modification of target nucleic acids, wherein such editing or modification can be quantified by insertion / deletion activity. Insertion / deletion activity refers to the degree of change (e.g., nucleotide deletion and / or insertion) of a target nucleic acid compared to a target nucleic acid not exposed to peptides in the compositions, systems, and methods described herein.
[0279] For example, insertion / deletion activity can be assessed using next-generation sequencing of one or more target loci within the target nucleic acid, where the insertion / deletion percentage is calculated as the ratio of sequencing reads containing insertions or deletions relative to an unedited reference sequence. In some embodiments, methods, systems, and compositions including effector proteins and guide nucleic acids can exhibit insertion / deletion activity of about 0.0001% to about 65% or higher compared to target nucleic acids not treated with said composition, system, or method upon contact with the target nucleic acid.
[0280] For example, this method, system, and composition can exhibit insertion / deletion activity levels of about 0.0001%, 0.001%, 0.01%, 0.1%, 1%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, or even higher.
[0281] Composition The compositions disclosed herein comprise one or more effector proteins or nucleic acids encoding such effector proteins, and one or more guide nucleic acids or nucleic acids encoding guide nucleic acids, or combinations of these components. In some embodiments, the effector protein may be any protein exhibiting nuclease activity that can recognize any PAM sequence described herein. Exemplary effector proteins are detailed throughout the specification.
[0282] In some cases, one or more repetitive sequences within the guide nucleic acid can interact with effector proteins. Furthermore, the spacer region sequence of the guide nucleic acid can hybridize with the target sequence of the target nucleic acid. The composition may also include one or more donor nucleic acids as described herein. These compositions are capable of editing the target nucleic acid within cells or subjects. They can facilitate the editing of the target nucleic acid or its expression in cells, tissues, organs, or in vitro, in vivo, or ex vivo. Furthermore, the composition can edit the target nucleic acid present in samples containing the target nucleic acid.
[0283] In some embodiments, the composition may consist of plasmids, viral vectors, non-viral vectors, or combinations thereof. Some embodiments specifically include viral vectors, and more specifically, may include adeno-associated virus (AAV). Furthermore, the composition may include liposomes (e.g., cationic or neutral lipids), dendritic polymers, lipid nanoparticles (LNPs), or cell-penetrating peptides. In certain cases, the composition may specifically include LNPs.
[0284] In some embodiments, the composition is used to edit the human albumin gene. Editing can result in at least partially functional expression of acid α-glucosidase. In some cases, this editing can lead to partial or complete cure of acid α-glucosidase deficiency.
[0285] Drug composition and administration method In some embodiments, the compositions described herein are pharmaceutical compositions. In some embodiments, the pharmaceutical composition comprises the compositions described herein and a pharmaceutically acceptable carrier or diluent. Non-limiting examples of pharmaceutically acceptable carriers and diluents suitable for the pharmaceutical compositions disclosed herein include buffers (e.g., neutral buffered saline, phosphate buffered saline); carbohydrates (e.g., glucose, mannose, sucrose, dextran, mannitol); peptides or amino acids (e.g., glycine); antioxidants; chelating agents (e.g., EDTA, glutathione); adjuvants (e.g., aluminum hydroxide); surfactants (polysorbate 80, polysorbate 20, or Pluronic F68); glycerol; sorbitol; mannitol; polyethylene glycol; and preservatives.
[0286] This document discloses pharmaceutical compositions for modifying target nucleic acids in cells or subjects, including any one and any combination thereof of the effector proteins, engineered effector proteins, fusion effector proteins, or guide nucleic acids described herein. This document also discloses pharmaceutical compositions comprising encoding any one and any combination thereof of the effector proteins, engineered effector proteins, fusion effector proteins, guide nucleic acids, or donor nucleic acids described herein. In some embodiments, the pharmaceutical composition includes multiple guide nucleic acids. The pharmaceutical compositions can be used to modify target nucleic acids or their expression in cells, in vitro, in vivo, or ex vivo.
[0287] In some embodiments, the pharmaceutical composition may include one or more nucleic acids encoding an effector protein, a fusion effector protein, a fusion chaperone, a guide nucleic acid, or a combination of these components, and a pharmaceutically acceptable carrier or diluent. The effector protein, fusion effector protein, fusion chaperone protein, or combination thereof may be any of those described herein. The nucleic acid may be in the form of a plasmid, a nucleic acid expression vector, or a viral vector.
[0288] In some instances, the composition (particularly a pharmaceutical composition) may consist of a viral vector encoding a fusion effector protein and a guide nucleic acid, wherein at least a portion of the guide nucleic acid binds to the effector protein within the fusion effector protein. The vector may be formulated for delivery by injection using a syringe. Other formulations may include delivery via electroporation or by chemical methods. The pharmaceutical composition may include a viral vector or a non-viral vector. In some cases, the pharmaceutical composition may contain a virus, the viral vector comprising a viral vector encoding a fusion effector protein, an effector protein, a fusion chaperone, a guide nucleic acid, or a combination of these elements, and a pharmaceutically acceptable carrier or diluent.
[0289] The pharmaceutical compositions described herein may also include salts. In some embodiments, the salt may be a sodium salt, while in other embodiments it may be a potassium or magnesium salt. Specific examples of salts include NaCl, KNO3, and Mg2+SO4.
[0290] In some embodiments, the pharmaceutical composition is in solution (e.g., liquid) form. In some embodiments, the solution may be formulated for injection, e.g., intravenous or subcutaneous injection. In some embodiments, the solution has a pH of about 7, about 7.1, about 7.2, about 7.3, about 7.4, about 7.5, about 7.6, about 7.7, about 7.8, about 7.9, about 8, about 8.1, about 8.2, about 8.3, about 8.4, about 8.5, about 8.6, about 8.7, about 8.8, about 8.9, or about 9. In some embodiments, the pH is 7 to 7.5, 7.5 to 8, 8 to 8.5, 8.5 to 9, or 7 to 8.5. In some cases, the solution has a pH less than 7. In some cases, the pH is greater than 7.
[0291] system This article discloses, in some respects, systems designed for modifying or editing target nucleic acids, including effector proteins or nucleic acids encoding these effector proteins, or polymeric complexes thereof. These systems can be used to alter or edit target nucleic acids, as well as to insert donor nucleic acids into target nucleic acids.
[0292] In some embodiments, the system comprises an effector protein or nucleic acid encoding an effector protein described herein, a guide nucleic acid or nucleic acid encoding a guide nucleic acid, reagents described herein, donor nucleic acid, support medium, or any combination of these elements. In some instances, the effector protein may be an effector protein described herein or a fusion protein.
[0293] In some embodiments, the system includes the effector protein described herein, the guide nucleic acid described herein, reagents, support media, or combinations thereof. In some embodiments, the effector protein includes the effector protein described herein or its fusion protein. In some embodiments, the effector protein includes an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% identity with any of the amino acid sequences listed in Table 1. In some embodiments, the amino acid sequence of the effector protein has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% identity with any of the amino acid sequences listed in Table 1.
[0294] In some embodiments, the systems described herein may consist of individual compositions, solutions, containers, kits, carriers, or similar articles, each comprising an effector protein, a nucleic acid encoding the effector protein, a guide nucleic acid, a nucleic acid encoding the guide nucleic acid, a donor nucleic acid, or a combination of these elements. These systems facilitate the individual delivery of the effector protein, the nucleic acid encoding the effector protein, the guide nucleic acid, the nucleic acid encoding the guide nucleic acid, or the donor nucleic acid described herein.
[0295] Furthermore, in some embodiments, the system may include a composition, solution, container, kit, carrier, or similar entity comprising two or more of the following: effector protein, nucleic acid encoding the effector protein, guide nucleic acid, nucleic acid encoding the guide nucleic acid, and donor nucleic acid. Such a system is capable of delivering multiple components, including effector protein, nucleic acid encoding the effector protein, guide nucleic acid, nucleic acid encoding the guide nucleic acid, and donor nucleic acid.
[0296] Additional system components In some embodiments, the system includes packaging, a carrier, or a container divided to contain one or more containers, such as vials, tubes, etc., each container including one of the individual elements used in the methods described herein. Suitable containers include, for example, test ports, bottles, vials, and test tubes. In one embodiment, the container is formed of a variety of materials, such as glass, plastic, or polymer. One or more systems described herein include packaging materials. Examples of packaging materials include, but are not limited to, pouches, blister packs, bottles, tubes, bags, containers, and any packaging material suitable for the intended application mode.
[0297] The system may include labels detailing the contents and / or providing instructions for use, as well as packaging inserts containing usage guidelines. Typically, a set of instruction manuals will also be included. In one embodiment, the label is affixed to or associated with the container. When letters, numbers, or other characters are directly affixed, molded, or etched onto the container itself, the label is considered to be on the container; when it is included inside a container or carrier that also holds the container (e.g., a drug instruction manual), the label is considered to be associated with the container.
[0298] In one implementation, the label may indicate that the contents are intended for a specific therapeutic application. As described herein, the label may also provide guidance on the use of the contents. After the product is packaged and wrapped or boxed to ensure a sterile barrier, it may be terminally sterilized by methods such as heat sterilization, gas sterilization, gamma radiation, or electron beam sterilization. Alternatively, the product may be prepared and packaged using aseptic processing.
[0299] Amplification reagents / components In some embodiments, the system described herein includes reagents or components for amplifying nucleic acids. Non-limiting examples of reagents for amplifying nucleic acids include polymerases, primers, and nucleotides. In some embodiments, the system includes reagents for amplifying a target nucleic acid in a sample. Nucleic acid amplification of the target nucleic acid can improve at least one of the sensitivity, specificity, or accuracy of detection of the target nucleic acid. In some embodiments, the nucleic acid amplification is isothermal nucleic acid amplification, allowing the system to be used in remote areas or low-resource environments without requiring specialized amplification equipment. In some embodiments, the amplification of the target nucleic acid increases the concentration of the target nucleic acid in the sample relative to the concentration of nucleic acids that do not correspond to the target nucleic acid.
[0300] Reagents used for nucleic acid amplification may include recombinases, oligonucleotide primers, single-stranded DNA-binding (SSB) proteins, polymerases, or combinations thereof suitable for amplification reactions. Non-restrictive examples of amplification reactions include transcription-mediated amplification (TMA), helicase-dependent amplification (HDA) or ring helicase-dependent amplification (CHDA), strand substitution amplification (SDA), recombinase polymerase amplification (RPA), loop-mediated amplification (LAMP), exponential amplification reaction (EXPAR), rolling circle amplification (RCA), ligase chain reaction (LCR), simplified RNA target amplification method (SMART), single primer isothermal amplification (SPIA), multiple substitution amplification (MDA), sequence-based amplification (NASBA), hinge-initiated primer-dependent nucleic acid amplification (HIP), nicking enzyme amplification reaction (NEAR), and modified multiple substitution amplification (IMDA).
[0301] Certain system conditions Certain conditions that may enhance the activity of effector proteins include the presence or concentration of certain salts in the solution in which the activity occurs. For example, the cis-cleavage activity of an effector protein can be inhibited or stopped by a high salt concentration. The salt can be a sodium, potassium, or magnesium salt. In some embodiments, the salt is NaCl. In some embodiments, the salt is KNO3. In some embodiments, the salt concentration is less than 150 mM, less than 125 mM, less than 100 mM, less than 75 mM, less than 50 mM, or less than 25 mM.
[0302] Certain conditions that can enhance the activity of effector proteins include the pH of the solution in which the activity occurs. For example, increasing the pH can enhance trans-cleavage activity. For example, the rate of trans-cleavage activity can increase with increasing pH up to pH 9. In some embodiments, the pH is about 7, about 7.1, about 7.2, about 7.3, about 7.4, about 7.5, about 7.6, about 7.7, about 7.8, about 7.9, about 8, about 8.1, about 8.2, about 8.3, about 8.4, about 8.5, about 8.6, about 8.7, about 8.8, about 8.9, or about 9. In some embodiments, the pH is 7 to 7.5, 7.5 to 8, 8 to 8.5, 8.5 to 9, or 7 to 8.5. In some embodiments, the pH is less than 7. In some embodiments, the pH is greater than 7.
[0303] Certain conditions that can enhance the activity of effector proteins include the temperature at which the activity occurs. In some embodiments, this temperature ranges from about 25°C to about 50°C. Other embodiments may specify a temperature of about 20°C to about 40°C, about 30°C to about 50°C, or about 40°C to about 60°C. Furthermore, the temperature may be set to about 25°C, 30°C, 35°C, 40°C, 45°C, or 50°C.
[0304] Methods and formulations for introducing systems and compositions into target cells Various established methods can be used to introduce the guide nucleic acids (or nucleic acids containing the nucleotide sequences encoding them) and / or effector proteins described herein into host cells. For example, guide nucleic acids and / or effector proteins can be bound to lipids. Alternatively, these components can be formulated as granules or bound to granules.
[0305] This document describes methods for introducing various components into a host. A host can refer to any suitable entity, such as a host cell. When mentioned, a host cell can be an in vivo or in vitro eukaryotic cell, a prokaryotic cell (e.g., bacteria or archaea), or a cell derived from a multicellular organism cultured as a single-celled entity (e.g., a cell line). These eukaryotic or prokaryotic cells can be recipients of the introduction methods described herein and can include progeny cells that have been transformed by these methods. It should be understood that progeny cells of a single cell may not be completely identical to the original parent in morphology or genomic content due to natural, accidental, or intentional mutations. If a heterologous nucleic acid (e.g., an expression vector) has been introduced into the cell, the host cell can be classified as a recombinant host cell or a genetically modified host cell. Methods for introducing nucleic acids and / or proteins into host cells are known in the art, and any convenient method can be used to introduce a subject nucleic acid (e.g., an expression construct / vector) into target cells (e.g., human cells, etc.). Suitable methods include, for example, viral infection, transfection, conjugation, protoplast fusion, lipid transfection, electroporation, calcium phosphate precipitation, polyethyleneimine (PEI)-mediated transfection, DEAE-dextran-mediated transfection, liposome-mediated transfection, particle gun technology, calcium phosphate precipitation, direct microinjection, and nanoparticle-mediated nucleic acid delivery (see, for example, Panyam et al. Adv Drug Deliv Rev. 2012 Sep 13. pii: S0169-409X(12)00283-9. doi: 10.1016 / j.addr.2012.09.023). In some embodiments, nucleic acids and / or proteins are included in a pharmaceutical composition comprising guide nucleic acids and / or effector proteins and pharmaceutically acceptable excipients for introduction into disease cells.
[0306] In some implementations, a target molecule (e.g., nucleic acid) is introduced into the host. This may also include the introduction of an effector protein into the host. Furthermore, vectors (e.g., lipid particles and / or viral vectors) may also be introduced into the host. The introduction may be intended to contact or assimilate into the host, such as by introducing it into host cells.
[0307] In some embodiments, methods for introducing one or more nucleic acids into host cells are described, said nucleic acids including nucleic acids encoding effector proteins, nucleic acids that generate engineered guide nucleic acids during transcription, and / or donor nucleic acids, or combinations thereof. A variety of suitable methods can be used to introduce nucleic acids into cells. Examples of these methods include viral infection, transfection, lipid transfection, electroporation, calcium phosphate precipitation, polyethyleneimine (PEI)-mediated transfection, DEAE-dextran-mediated transfection, liposome-mediated transfection, particle gun technology, direct microinjection, nanoparticle-mediated nucleic acid delivery, etc. Additional methods are detailed throughout this document. The introduction of one or more nucleic acids into host cells can occur in any culture medium and under any culture conditions that promote cell survival. The introduction of one or more nucleic acids into host cells can be performed in vivo or in vitro. The introduction of one or more nucleic acids into host cells can be performed in vitro.
[0308] In some embodiments, the effector protein can be delivered in the form of RNA. This RNA can be produced by direct chemical synthesis or by in vitro transcription from a DNA sequence encoding the effector protein. After synthesis, the RNA can be introduced into cells using various techniques suitable for nucleic acid delivery, such as microinjection, electroporation, transfection, etc. In some instances, the introduction of one or more nucleic acids may involve the use of a vector and / or a vector system. Therefore, in some embodiments, the compositions and systems described herein may include vectors and / or vector systems.
[0309] The vector can be directly introduced into the host. In some cases, host cells can be treated with one or more of the described vectors, and in some instances, these vectors can be absorbed by the cells. Methods for delivering the vector into cells include, but are not limited to, electroporation, calcium chloride transfection, microinjection, lipid transfection, and direct contact with cells or particles containing the target molecule.
[0310] The components described in this article can also be directly introduced into the host. For example, engineered guide nucleic acids can be delivered to the host, particularly host cells.
[0311] Methods for introducing nucleic acids (such as RNA) into cells include, but are not limited to, direct injection, transfection, or any other method suitable for nucleic acid delivery.
[0312] The effector proteins described herein can also be introduced directly into the host. In some embodiments, these effector proteins can be modified to facilitate their introduction into the host. For example, modifications can be made to enhance the solubility of the effector protein. Such modifications may involve fusing the effector protein with a peptide domain that increases solubility. This domain can be linked to the effector protein via a defined protease cleavage site (e.g., a TEV sequence cleaved by a TEV protease). The linker may also contain one or more flexible sequences, such as 1 to 10 glycine residues. In some cases, peptide cleavage is performed in a buffer that maintains product solubility, for example, in the presence of 0.5 to 2 M urea, or in the presence of a solubility-enhancing peptide and / or polynucleotide.
[0313] Target domains may include endosomolytic domains (such as the influenza HA domain) and other peptides that facilitate production (such as IF2 domains, GST domains, GRPE domains, etc.). Furthermore, effector proteins can be modified to improve their stability. For example, PEGylation of effector proteins can prolong their shelf life in the bloodstream due to the presence of polyethylene glycol groups.
[0314] Effector proteins can also be modified to facilitate uptake by the host (e.g., host cells). For example, effector proteins can be fused to cell-penetrating peptides to promote cellular uptake. Various suitable permeation domains, including peptides, peptide mimics, and non-peptide carriers, can be used in the non-integrating peptides described herein. Examples include permeation proteins derived from the third α-helix of the Drosophila melanogaster transcription factor antennalopodia; the amino acid sequence of the basic region of HIV-1 tat (e.g., amino acids 49-57 of the naturally occurring tat protein); and polyarginine motifs, such as amino acid regions 34-56, nonaarginine, octaarginine, and similar sequences of the HIV-1 rev protein. Fusion sites can be selected to optimize the biological activity, secretion, or binding properties of the effector protein, with the optimal site determined by appropriate methods.
[0315] The formulations described herein are designed for introducing systems and compositions into a host. In some embodiments, these formulations, systems, and compositions may include effector proteins and carriers (e.g., excipients, diluents, carriers, or fillers).
[0316] In some aspects of the invention, the effector protein is incorporated into a pharmaceutical composition comprising the effector protein and any pharmaceutically acceptable excipient, carrier, or diluent. A pharmaceutically acceptable excipient, carrier, or diluent is any substance formulated in the pharmaceutical composition with the active ingredient that allows the active ingredient to retain its biological activity while not reacting with the subject's immune system.
[0317] These substances can serve a variety of purposes, including long-term stability, increasing the volume of solid dosage forms containing small amounts of potent active ingredients, or enhancing the therapeutic properties of the active ingredient in the final formulation. This enhancement may involve improving absorption, reducing viscosity, or increasing solubility. The selection of suitable substances can depend on factors such as route of administration, formulation, active ingredient, and other considerations. Compositions containing these substances can be formulated using established methods.
[0318] Target gene (GOI) / Donor nucleic acid / Donor template The systems or compositions according to the present invention can be used to treat a variety of diseases by introducing a specific GOI into the albumin site.
[0319] The following provides an exemplary and non-exhaustive list of GOIs and the corresponding diseases that may be treated using the system according to the invention.
[0320] a) Factor IX (F9) - used to treat hemophilia B.
[0321] b) Factor VIII (F8) - used to treat hemophilia A.
[0322] c) α-1 antitrypsin (SERPINA1) - used to treat α-1 antitrypsin deficiency.
[0323] d) Glucocerebroside lipase (GBA) - used to treat Gaucher disease.
[0324] e) Phenylalanine hydroxylase (PAH) - used to treat phenylketonuria (PKU).
[0325] f) Urokinase plasminogen activator (PLAU) - may be used in thrombolytic therapy.
[0326] g) Lipoprotein lipase (LPL) - used to treat hyperlipoproteinemia.
[0327] h) C1 esterase inhibitor (SERPING1) - used to treat hereditary angioedema.
[0328] i) Sphingomyelinase (SMPD1) - used to treat Niemann-Pick disease.
[0329] j) N-acetylglucosamine-1-phosphotransferase (GNPTAB) - used to treat type II mucolipidosis (type I cell disease).
[0330] k) Cystic fibrosis transmembrane conductance modulator (CFTR) - used to treat cystic fibrosis.
[0331] l) Dipeptidyl peptidase IV (DPP4) - may be used to treat diabetes.
[0332] m) Adenosine deaminase (ADA) - used to treat severe combined immunodeficiency (SCID).
[0333] m) Ornithine carbamoyltransferase (OTC) - used to treat ornithine carbamoyltransferase deficiency.
[0334] n) β-glucuronidase (GUSB) - used to treat mucopolysaccharidosis type VII (MPS VII).
[0335] o) Cystine β synthase (CBS) - used to treat homocystinuria.
[0336] p) Sodium-dependent glucose transporter 1 (SGLT1) - used to treat glucose-galactose malabsorption disorders.
[0337] q) Fumarase (FH) - Used to treat fumarate deficiency.
[0338] r) Arginase (ARG1) - used to treat arginase deficiency.
[0339] s) Glutamate decarboxylase (GAD1) - used to treat autoimmune encephalitis.
[0340] t) Acid α-glucosidase (GAA) - used to treat Pompe disease.
[0341] u) Hexoaminoamino acid enzyme A (HEXA) - used to treat Ty Sachs disease.
[0342] v) UDP-glucuronyltransferase 1A1 (UGT1A1) - used to treat Krieger-Najjar syndrome.
[0343] w) Sulfatase (ARSB) - used to treat Maroto-Rami syndrome.
[0344] x) Epidermal growth factor receptor (EGFR) - used for targeted therapy of various cancers.
[0345] BRCA1 (BRCA1) is used to treat hereditary breast cancer and ovarian cancer syndromes.
[0346] z) TP53 (TP53) - for targeted therapy of cancers associated with p53 mutations.
[0347] aa) Cystic fibrosis transmembrane conductance modulator (CFTR) - used to treat cystic fibrosis.
[0348] (bb) Nerve growth factor (NGF) - may be used for neurodegenerative diseases.
[0349] cc) Insulin (INS) - used to treat diabetes.
[0350] dd) Cytosin γ-lyase (CTH) - used to treat cytosin β-synthetase deficiency.
[0351] In one aspect, the compositions and methods described herein include a donor nucleic acid encoding acid α-glucosidase, and the methods described herein include inserting the donor nucleic acid encoding acid α-glucosidase into intron 1 of the human albumin gene. In some embodiments, the donor nucleic acid is cDNA encoding acid α-glucosidase.
[0352] The term "GAA" used in this article refers to the gene encoding acid α-glucosidase. The human GAA gene is located on chromosome 17q25.2-q25.3.
[0353] An exemplary amino acid sequence of human acidic α-glucosidase encoded by the human GAA gene, UniProtKB protein P10253 (LYAG_HUMAN), accessed on 2024-12-17.
[0354] In some embodiments, the method includes inserting a donor nucleic acid encoding acid α-glucosidase into intron 1 of the human albumin gene in a human cell. In some instances, the human cell includes the GAA gene. In some embodiments, the GAA gene includes a mutation. In some embodiments, the mutation includes point mutations or single nucleotide polymorphisms (SNPs), chromosomal mutations, copy number mutations, or any combination thereof. Point mutations optionally include substitutions, insertions, or deletions. In some embodiments, the mutation includes a chromosomal mutation. Chromosomal mutations can include inversions, deletions, duplications, or translocations. In some embodiments, the mutation includes copy number variations. Copy number variations can include gene amplification or amplified trinucleotide repeats. The mutation can be located in a non-coding or coding region of the gene.
[0355] In some embodiments, the GAA gene includes mutations, where the mutation is a SNP. A single nucleotide mutation or SNP may be associated with a phenotype of the sample or an organism from which the sample was taken. In some embodiments, the SNP is associated with an altered phenotype of the wild-type phenotype. An SNP may be a synonymous substitution or a non-synonymous substitution. A non-synonymous substitution may be a missense substitution or a meaningless point mutation. A synonymous substitution may be a silent substitution. A mutation may be the deletion of one or more nucleotides. Typically, single nucleotide mutations, SNPs, or deletions are associated with diseases, such as genetic disorders. Mutations (e.g., single nucleotide mutations), SNPs, or deletions may encode a target nucleic acid sequence from the germline of an organism or may encode a target nucleic acid from diseased cells. In some embodiments, diseased cells are cells that include pathway conditions or pathway systems that are detrimental to cell survival, tissue survival, systemic survival, or organismal survival.
[0356] In some implementations, the GAA gene includes mutations, wherein the mutation is a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more nucleotides. The mutation can be a deletion of about 5, about 10, about 15, about 20, about 25, about 30, about 35, about 40, about 45, about 50, about 55, about 60, about 65, about 70, about 75, about 80, about 85, about 90, about 95, about 100, about 200, about 300, about 400, about 500, about 600, about 700, about 800, about 900 or about 1000 nucleotides. Mutations can be 1 to 5, 5 to 10, 10 to 15, 15 to 20, 20 to 25, 25 to 30, 30 to 35, 35 to 40, 40 to 45, 45 to 50, 50 to 55, 55 to 60, 60 to 65, 65 to 70, 70 to 75, 75 to 80, 80 to 85, 85 to 90, 90 to 95, 95 to 100, 100 to 2 The deletion of 00, 200 to 300, 300 to 400, 400 to 500, 500 to 600, 600 to 700, 700 to 800, 800 to 900, 900 to 1000, 1 to 50, 1 to 100, 25 to 50, 25 to 100, 50 to 100, 100 to 500, 100 to 1000, or 500 to 1000 nucleotides.
[0357] In some implementations, the GAA gene includes disease-associated mutations. In some embodiments, a disease-associated mutation refers to a mutation whose presence in a subject indicates that the subject is susceptible to or has a disease, condition, or pathological state. In some embodiments, a disease, condition, or pathological state-associated mutation refers to a mutation that causes, contributes to, or indicates the development of a disease, condition, or pathological state. A disease-associated mutation may also refer to any mutation that produces an abnormal level or form of transcriptional or translational product in disease-affected cells relative to disease-free controls.
[0358] Mutations can cause diseases. Diseases may include, at least in part, hereditary conditions. In some embodiments, a disease, symptom, or pathological condition includes a hereditary condition. A disease may include, at least in part, a hereditary disease. A disease may include, at least in part, glycogen storage diseases. In some embodiments, glycogen storage disease is Pompe disease.
[0359] Pompe disease is a genetic metabolic disorder associated with a deficiency of acid alpha-glucosidase. Pompe disease is an autosomal recessive genetic disorder. Acid alpha-glucosidase helps digest glycogen in lysosomes. However, individuals with Pompe disease are unable to digest glycogen, leading to excessive glycogen accumulation. This condition impairs the normal functioning of Pompe disease patients by causing irreversible damage to muscles. Therefore, in some embodiments, one or more mutations or abnormal expression of the acid alpha-glucosidase protein are associated with Pompe disease.
[0360] Mutations associated with Pompe disease can be point mutations, single nucleotide polymorphisms (SNPs), chromosomal mutations, copy number mutations, or any combination thereof. Mutations in the GAA gene can be associated with acid α-glucosidase expression, acid α-glucosidase activity, and acid α-glucosidase structural stability. As a result, mutations can lead to decreased acid α-glucosidase expression, decreased or absent acid α-glucosidase activity, decreased acid α-glucosidase half-life, increased lysosomal glycogen concentration, or combinations thereof. Alternatively, mutations in the region responsible for GAA gene expression can lead to abnormal or low expression of acid α-glucosidase.
[0361] This document discloses compositions, systems, methods for expressing acid α-glucosidase, and methods for treating acid α-glucosidase deficiency. Therefore, in some embodiments, the compositions, systems, and methods disclosed herein involve inserting a donor nucleic acid into a cleaved target nucleic acid, wherein the donor nucleic acid includes a nucleotide sequence encoding acid α-glucosidase.
[0362] The donor nucleic acid can be inserted into a specific (e.g., effector protein-targeted) site within the target nucleic acid. In some embodiments, the donor nucleic acid can be inserted into the target sequence, directly adjacent to the target sequence, directly adjacent, adjacent to about 1 to 20 nucleotides, adjacent to about 1 to 10 nucleotides, adjacent to about 1 to 5 nucleotides, adjacent to about 5 to 20 nucleotides, adjacent to about 5 to 10 nucleotides, or adjacent to about 10 to 20 nucleotides. In some embodiments, the donor nucleic acid encodes the amino acid sequence of a functional acidic alpha-glucosidase. In some embodiments, the functional acidic alpha-glucosidase has at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 100%, at least 110%, at least 120%, at least 130%, at least 140%, at least 150%, at least 180%, at least 200%, at least 300%, or at least 400% of the enzyme activity compared to wild-type acidic alpha-glucosidase. In some embodiments, the functional acidic alpha-glucosidase includes wild-type acidic alpha-glucosidase. In some embodiments, the wild-type acid α-glucosidase comprises the amino acid sequence of human wild-type acid α-glucosidase. In some embodiments, the donor nucleic acid encodes an amino acid sequence having at least 70%, at least 80%, at least 90%, at least 92%, at least 95%, at least 97%, at least 99%, or 100% identity with the wild-type sequence. In some embodiments, the method includes contacting the target nucleic acid with an effector protein comprising an amino acid sequence having at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100% identity with any of the amino acid sequences listed in Table 1, thereby introducing a single-strand break in the target nucleic acid; and contacting the target nucleic acid with a second effector protein, the second effector protein optionally comprising an amino acid sequence having at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100% identity with any of the amino acid sequences listed in Table 1. An amino acid sequence with at least 5%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity is used to generate a second cleavage site in the target nucleic acid, connecting regions flanking the first and second cleavage sites, optionally by NHEJ or single-strand annealing, thereby causing the excision of a portion of the target nucleic acid between the first and second cleavage sites from the target nucleic acid; and optionally by contacting the target nucleic acid with a donor nucleic acid by HDR or NHEJ for homologous recombination, thereby introducing a new sequence into the target nucleic acid (e.g., at or between the cleavage sites).
[0363] XIII. Nucleic Acid Editing Methods This document provides methods for editing target nucleic acids. Generally, editing refers to modifying the nucleotide sequence of a target nucleic acid; however, the compositions and systems disclosed herein can also perform epigenetic modifications on target nucleic acids. The effector proteins, their polymeric complexes, and systems described herein can be used to edit or modify target nucleic acids. Editing a target nucleic acid can include one or more of the following: cleaving the target nucleic acid, deleting one or more nucleotides of the target nucleic acid, inserting one or more nucleotides into the target nucleic acid, mutating one or more nucleotides of the target nucleic acid, or modifying (e.g., methylating, demethylating, deamination, or oxidation) one or more nucleotides of the target nucleic acid.
[0364] Editing methods may include contacting a target nucleic acid with an effector protein and a guide nucleic acid as described herein, wherein the effector protein comprises an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100% identity with any of the amino acid sequences listed in Table 1.
[0365] Editing can introduce mutations (e.g., point mutations, insertions, deletions) into target nucleic acids relative to the corresponding wild-type nucleotide sequence. Editing can remove or insert nucleic acid sequences to produce the corresponding wild-type protein. Editing can remove / insert tissue-specific nucleic acid sequences into target nucleic acids. Editing can be used to generate gene knock-in, gene editing, or combinations thereof. The methods of this disclosure can target any locus in the cellular genome.
[0366] Editing can include single-strand cleavage, double-strand cleavage, donor nucleic acid insertion, epigenetic modifications (e.g., methylation, demethylation, acetylation, or deacetylation), or combinations thereof. In some embodiments, the cleavage (single-strand or double-strand) is site-specific, meaning that the cleavage occurs at a specific site in the target nucleic acid, typically within a region of the target nucleic acid that hybridizes with the guide nucleic acid spacer region. In some cases, the cleavage occurs directly adjacent to or within approximately 1 to 10 nucleotides of the region of the target nucleic acid that hybridizes with the guide nucleic acid spacer region. In some cases, the effector protein introduces a single-strand break in the target nucleic acid to produce the cleaved nucleic acid. In some cases, the effector protein is capable of introducing a break in single-stranded RNA (ssRNA). The effector protein can couple to a guide nucleic acid targeting a specific region in the ssRNA. In some embodiments, the target nucleic acid and the resulting cleaved nucleic acid are contacted to perform homologous recombination (e.g., homology-directed repair (HDR)) or non-homologous end joining (NHEJ). In some cases, double-strand breaks in the target nucleic acid can be repaired without the insertion of donor nucleic acid (e.g., via NHEJ or HDR), resulting in insertion or deletion at or near the double-strand break site in the target nucleic acid.
[0367] In some implementations, insertions and deletions are sometimes referred to as insertion-deletion or insertion-deletion mutations, a type of genetic mutation caused by the insertion and / or deletion of nucleotides in a target nucleic acid. Insertions and deletions are variable in length (e.g., from 1 to 1000 nucleotides) and can be detected using methods known in the art, including sequencing. If the number of nucleotides in the insertion / deletion is not divisible by three, and it occurs in a protein-coding region, it is also a frameshift mutation.
[0368] In some implementations, the insertion / deletion percentage is based on the percentage of sequencing reads showing that at least one nucleotide has been edited into a nucleotide insertion and / or deletion, regardless of the size of the insertion or deletion, or the number of nucleotides edited. For example, if at least one nucleotide deletion is detected in a given target nucleic acid, it is counted in the percentage insertion / deletion value. As another example, if one copy of the target nucleic acid has 1 nucleotide deleted and another copy of the target nucleic acid has 10 nucleotides deleted, they are counted the same. This number reflects the percentage of the target nucleic acid edited by a given effector protein.
[0369] In some embodiments, where the compositions, systems, and methods of this disclosure include additional guide nucleic acids or their uses, the dual-guide compositions, systems, and methods described herein can modify the target nucleic acid at two sites. In some cases, dual-guide editing can include cleaving the target nucleic acid at two sites targeted by the guide RNA. In some embodiments, new nucleotide sequences can be inserted when the sequence between the guide nucleic acids is removed.
[0370] Therefore, in some embodiments, the compositions, systems, and methods described herein can edit 1 to 1,000 nucleotides, or any integer between the two, in the target nucleic acid. In some embodiments, the compositions, systems, and methods described herein can edit 1 to 1,000, 2 to 900, 3 to 800, 4 to 700, 5 to 600, 6 to 500, 7 to 400, 8 to 300, 9 to 200, or 10 to 100 nucleotides, or any integer between the two. In some embodiments, the compositions, systems, and methods described herein can edit 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotides. In some embodiments, the compositions, systems, and methods described herein can edit 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, or more nucleotides, or any integer between the two. In some embodiments, the compositions, systems, and methods described herein can edit 100, 200, 300, 400, 500, 600, 700, 800, 900, or more nucleotides, or any integer between the two.
[0371] The method may include the use of two or more effector proteins. An exemplary method for introducing a break in a target nucleic acid includes contacting the target nucleic acid with: (a) a first engineered guide nucleic acid including a region that binds to a first effector protein, wherein the effector protein includes a sequence having at least 75% identity with any of the amino acid sequences listed in Table 1; and (b) a second engineered guide nucleic acid including a region that binds to a second effector protein, wherein the effector protein includes a sequence having at least 75% identity with any of the amino acid sequences listed in Table 1, wherein the first engineered guide nucleic acid includes an additional region that binds to the target nucleic acid and wherein the second engineered guide nucleic acid includes an additional region that binds to the target nucleic acid.
[0372] In some embodiments, editing the target nucleic acid includes genome editing. Genome editing can include modifying the genome, chromosomes, plasmids, or other genetic material of a cell or organism. In some embodiments, the genome, chromosomes, plasmids, or other genetic material of a cell or organism is modified in vivo. In some embodiments, the genome, chromosomes, plasmids, or other genetic material of a cell or organism is modified in a cell. In some embodiments, the genome, chromosomes, plasmids, or other genetic material of a cell or organism is modified in vitro. For example, plasmids can be modified in vitro using the compositions described herein and introduced into cells or organisms. In some embodiments, modifying the target nucleic acid can include deleting a sequence from the target nucleic acid. In some embodiments, modifying the target nucleic acid can include replacing a sequence in the target nucleic acid with a second sequence. In some embodiments, modifying the target nucleic acid can include introducing a sequence into the target nucleic acid.
[0373] In some embodiments, the method includes editing a target nucleic acid having two or more effector proteins. Editing the target nucleic acid may include introducing two or more single-strand breaks in the target nucleic acid. In some embodiments, breaks may be introduced by contacting the target nucleic acid with an effector protein and a guide nucleic acid. The guide nucleic acid may bind to the effector protein and hybridize with a region of the target nucleic acid, thereby recruiting the effector protein to the region of the target nucleic acid. Binding of the effector protein to the regions of the guide nucleic acid and the target nucleic acid may activate the effector protein, and the effector protein may introduce breaks (e.g., single-strand breaks) into the region of the target nucleic acid. In some embodiments, modifying the target nucleic acid may include introducing a first break into a first region of the target nucleic acid and a second break into a second region of the target nucleic acid. For example, modifying the target nucleic acid may include contacting the target nucleic acid with a first guide nucleic acid and a second guide nucleic acid, the first guide nucleic acid binding to a first effector protein and hybridizing with a first region of the target nucleic acid, and the second guide nucleic acid binding to a second programmable nickase and hybridizing with a second region of the target nucleic acid. The first effector protein may introduce a first break into a first strand in the first region of the target nucleic acid, and the second effector protein may introduce a second break into a second strand in the second region of the target nucleic acid. In some embodiments, a segment of the target nucleic acid between the first and second breaks may be removed, thereby modifying the target nucleic acid. In some embodiments, the target nucleic acid fragment between the first and second breaks can be modified (e.g., by replacing it with a donor nucleic acid). In some embodiments, the effector protein comprises an amino acid sequence having at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100% identity with any of the amino acid sequences in Table 1.
[0374] In some implementations, the method includes inserting a donor nucleic acid into a cleaved target nucleic acid. The donor nucleic acid may be inserted into a specific (e.g., effector protein-targeted) site within the target nucleic acid. In some embodiments, the method includes contacting a target nucleic acid with an effector protein comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100% identity with any of the amino acid sequences listed in Table 1, thereby introducing a single-strand break in the target nucleic acid; contacting the target nucleic acid with a second effector protein, the second effector protein optionally comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% identity with any of the amino acid sequences listed in Table 1, to create a second cleavage site in the target nucleic acid, connecting regions flanking the first and second cleavage sites, optionally by NHEJ or single-strand annealing, thereby causing the excision of a portion of the target nucleic acid between the first and second cleavage sites from the target nucleic acid; and optionally contacting the target nucleic acid with a donor nucleic acid by HDR or NHEJ for homologous recombination, thereby introducing a new sequence into the target nucleic acid (e.g., at or between the cleavage sites).
[0375] In some embodiments, the method includes editing a target nucleic acid having two or more effector proteins. Editing the target nucleic acid may include introducing two or more single-strand breaks in the target nucleic acid. In some embodiments, the breaks may be introduced by contacting the target nucleic acid with an effector protein and a guide nucleic acid. The guide nucleic acid may bind to the effector protein and hybridize with a region of the target nucleic acid, thereby recruiting the effector protein to the region of the target nucleic acid. Binding of the effector protein to the regions of the guide nucleic acid and the target nucleic acid may activate the effector protein, and the effector protein may introduce a break (e.g., a single-strand break) into the region of the target nucleic acid. In some embodiments, modifying the target nucleic acid may include introducing a first break in a first region of the target nucleic acid and a second break in a second region of the target nucleic acid. For example, modifying the target nucleic acid may include contacting the target nucleic acid with a first guide nucleic acid and a second guide nucleic acid, the first guide nucleic acid binding to a first effector protein and hybridizing with a first region of the target nucleic acid, and the second guide nucleic acid binding to a second programmable nickase and hybridizing with a second region of the target nucleic acid. The first effector protein may introduce a first break in a first strand of the first region of the target nucleic acid, and the second effector protein may introduce a second break in a second strand of the second region of the target nucleic acid. In some embodiments, the target nucleic acid fragment between the first and second breaks may be removed, thereby modifying the target nucleic acid. In some embodiments, the target nucleic acid fragment between the first and second breaks can be modified by replacing it (e.g., with a donor nucleic acid). In some embodiments, the effector protein comprises an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100% identity with any of the amino acid sequences listed in Table 1.
[0376] In some embodiments, editing is achieved by fusing an effector protein to a heterologous sequence. The heterologous sequence can be a suitable fusion partner, such as a protein that provides recombinase activity by acting on a target nucleic acid. In some embodiments, the fusion protein comprises an effector protein fused to the heterologous sequence via a linker. The heterologous sequence or fusion partner can be a base-editing domain. The base-editing domain can be ADAR1 / 2 or any functional variant thereof. The heterologous sequence or fusion partner can be fused to the C-terminus, N-terminus, or internal portion of the effector protein (e.g., a portion other than the N-terminus or C-terminus). The heterologous sequence or fusion partner can be fused to the effector protein via a linker. The linker can be a peptide linker or a non-peptide linker. In some embodiments, the linker is an XTEN linker. In some embodiments, the linker comprises one or more repeating tripeptides GGS. In some embodiments, the linker length is from 1 to 100 amino acids. In some embodiments, the linker length is more than 100 amino acids. In some embodiments, the linker length is from 10 to 27 amino acids. Non-peptide connectors can be polyethylene glycol (PEG), polypropylene glycol (PPG), copolymer (ethylene / propylene) glycol, polyoxyethylene (POE), polyurethane, polyphosphazene, polysaccharides, dextran, polyvinyl alcohol, polyvinylpyrrolidone, polyvinyl ether, polyacrylamide, polyacrylate, polycyanoacrylate, lipid polymers, chitin, hyaluronic acid, heparin, or alkyl connectors.
[0377] In some embodiments, the editing or modification of the target nucleic acid can be site-specific, wherein the compositions, systems, and methods described herein can edit or modify the target nucleic acid at one or more specific loci to achieve one or more specific mutations, including sequence deletion, sequence knock-in, or any combination thereof. For example, editing or modification of a specific locus can achieve sequence knock-in. In some embodiments, sequence knock-in is a modification in which one or more sequences are inserted into the target nucleic acid relative to a target nucleic acid without sequence knock-in. In some embodiments, editing or modification of a specific locus can achieve both sequence knock-in and sequence deletion. In some embodiments, the editing or modification of the target nucleic acid can be locus-specific, modification-specific, or both. In some embodiments, the editing or modification of the target nucleic acid can be site-specific, modification-specific, or both, wherein the compositions, systems, and methods described herein include effector proteins and guide nucleic acids described herein. In some embodiments, the editing or modification of the target nucleic acid is specific to intron 1 of the mammalian albumin gene.
[0378] Methods for editing target nucleic acids or regulating their expression can be performed in vivo. Methods for editing target nucleic acids or regulating their expression can also be performed in vitro. For example, plasmids can be modified in vitro and introduced into cells or organisms using the compositions, systems, and methods described herein. Methods for editing target nucleic acids or regulating their expression can also be performed in vitro. For example, a method may include obtaining cells from a subject, modifying the target nucleic acids in the cells with the methods described herein, and returning the cells to the subject.
[0379] Transfection of donor nucleic acid Regarding viral vectors, the term "donor nucleic acid" refers to the nucleotide sequence that is about to be introduced into the cell or has already been introduced after transfection with a viral vector. Donor nucleic acid can be introduced into the cell through any mechanism of viral vector transfection, including but not limited to integration into the cell's genome or introduction of a free plasmid or viral genome. As another example, when used for effector protein activity, the term "donor nucleic acid" refers to the nucleotide sequence that is about to be inserted into or has already been inserted into the cleavage site of the effector protein (cleavage (hydrolysis of phosphodiester bonds) of nucleic acid leads to a nick or double-strand break – nuclease activity).
[0380] As another example, when used for homologous recombination, the term donor nucleic acid refers to a DNA sequence that serves as a template during homologous recombination, carrying modifications that are about to be introduced or have already been introduced into the target nucleic acid. By using the donor nucleic acid as a template, the genetic information, including the modifications, is copied into the target nucleic acid via homologous recombination.
[0381] Donor nucleic acids of any suitable size can be integrated into the target nucleic acid or the genome. In some embodiments, the length of the donor polynucleotide integrated into the genome is less than 3, approximately 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, 10, 10.5, 11, 11.5, 12, 12.5, 13, 13.5, 14, 14.5, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 100, 150, 200, 250, 300, 350, 400, 450, or 500 kilobases. In some embodiments, the length of the donor nucleic acid is greater than 500 kilobases (kb).
[0382] Donor nucleic acids may include sequences derived from animals. Animals may be human. Animals may be non-human animals, such as, by way of non-limiting examples, mice, rats, hamsters, rabbits, pigs, cattle, deer, sheep, goats, chickens, cats, dogs, ferrets, birds, and non-human primates (e.g., marmosets, rhesus monkeys). Non-human animals may be domesticated mammals or agricultural mammals.
[0383] Genetically modified cells and organisms Genetically modified cells can be generated using the editing methods described herein. Cells can be eukaryotic cells (e.g., mammalian cells) or prokaryotic cells (e.g., archaea cells). Cells can be derived from multicellular organisms and cultured as single-celled entities. Cells can include heritable genetic modifications such that progeny cells derived therefrom include heritable genetic mutations. Cells can be progeny of genetically modified cells that include genetically modified parental cells. Genetically modified cells can include deletions, insertions, mutations, or non-natural sequences relative to a wild-type version of the cell or the organism from which the cell is derived.
[0384] In some embodiments, when the target nucleic acid is modified by the compositions, systems and methods described herein, the target nucleic acid may include intron deletion, intron knock-in, or a combination thereof.
[0385] The method may include contacting cells with a nucleic acid (e.g., a plasmid or mRNA) that includes a nucleotide sequence encoding an effector protein, wherein the effector protein includes an amino acid sequence that has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% identity with any of the amino acid sequences listed in Table 1.
[0386] Methods may include contacting cells with nucleic acids (e.g., plasmids or mRNA) including nucleotide sequences encoding guide nucleic acids, tracrRNA, intermediate RNA, crRNA, or any combination thereof. Contact may include electroporation, acoustic perforation, photoperforation, viral vector-based delivery, iTOP, nanoparticle delivery (e.g., lipid or gold nanoparticle delivery), cell-penetrating peptide (CPP) delivery, DNA nanostructure delivery, or any combination thereof.
[0387] Methods may include contacting cells with nucleic acids (e.g., plasmids or mRNA) that include nucleotide sequences encoding guide nucleic acids, tracrRNA, intermediate RNA, sgRNA, or any combination thereof.
[0388] The method may include contacting cells with an effector protein or its polymeric complex, wherein the effector protein comprises an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% identity with any of the amino acid sequences listed in Table 1.
[0389] In some embodiments, the compositions, systems, and methods described herein may include an effector protein or a nucleic acid encoding an effector protein, wherein the effector protein has one or more substitutions relative to any amino acid sequence listed in Table 1. These substitutions may range from at least 1 to 20 or more and may be categorized into various ranges, such as 1 to 16, 1 to 12, 1 to 8, 1 to 4, 4 to 20, 4 to 16, 4 to 12, 4 to 8, 8 to 20, 8 to 16, 8 to 12, 12 to 20, 12 to 16, or 16 to 20 substitutions relative to the sequences present in Table 1.
[0390] Substitutions may include one or more conserved substitutions, one or more non-conserved substitutions, or a combination of both. In some instances, the effector protein or the nucleic acid encoding the effector protein may contain one or more conserved substitutions relative to any amino acid sequence in Table 1, which also falls within the scope mentioned above.
[0391] Similarly, effector proteins may also contain one or more non-conserved substitutions relative to any sequence in Table 1, wherein the substitutions may be from 1 to 20 or more. In some cases, amino acid modifications can lead to changes in the activity of the effector protein relative to its naturally occurring counterpart. For example, these modifications can enhance or reduce the catalytic activity or binding affinity of the effector protein. In some embodiments, the changes can result in catalytically inactive variants of the effector protein.
[0392] The effector proteins described herein can perform enzymatic reactions similar to those of wild-type (WT) effector proteins. Variants of WT effector proteins may include modifications that provide beneficial characteristics, such as enhanced activity (e.g., increased insertion / deletion activity, catalytic activity, specificity, selectivity, or affinity for substrates such as target or guide nucleic acids). In some cases, the activity of effector proteins may be equal to or greater than that of WT effector proteins, indicating that they exhibit one or more identical or higher activities at the same amino acid positions compared to effector proteins without any variants. For example, variants may exhibit an increase in activity of at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, or even 200% compared to WT effector proteins.
[0393] The activity of effector proteins or their variants compared to WT effector proteins can be assessed through cleavage experiments.
[0394] In some embodiments, the effector protein may have one or more amino acid substitutions compared to any of the sequences listed in Table 1, with the remaining amino acid sequence exhibiting at least 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% identity with the reference amino acid sequence in Table 1. Furthermore, the remaining amino acid sequence of the variant may also exhibit at least 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% similarity to the corresponding reference sequence. In some cases, the substitution may involve one or more positively charged amino acid residues, such as Lys(K), Arg(R), His(H), or combinations thereof.
[0395] In some implementations, the term "in vitro" refers to a process that occurs outside a living organism, typically in a controlled environment such as a test tube or petri dish.
[0396] The effector proteins described herein may include encoded and non-coding amino acids, as well as chemically or biochemically modified or derived amino acids, and proteins with altered peptide backbones. These effector proteins may contain one or more mutations, engineered modifications, or both. It should be understood that the coding sequence of these effector proteins does not necessarily need to include an N-terminal codon containing either methionine (M) or valine (V). Those skilled in the art will understand that the start codon may be replaced by a codon encoding an amino acid residue sufficient to initiate translation in the host cell.
[0397] Furthermore, the compositions, systems, and methods may include heterologous peptides or polypeptides, which may be fusion proteins comprising an effector protein and one or more fusion chaperone proteins. The fusion chaperone protein may be located at the N-terminus of the effector protein. In this case, the start codon of the heterologous peptide or polypeptide may also serve as the start codon of the effector protein, allowing for the removal or deletion of the native start codon that normally encodes the amino acid residues required for translation initiation.
[0398] Heteropeptides can consist of at least two distinct polypeptide sequences that do not coexist in nature. Heterologous systems may include at least one component that does not naturally coexist with other components.
[0399] In some embodiments, the heteropeptide or polypeptide may contain subcellular localization signals, such as nuclear localization signals (NLS) that facilitate the targeting of nucleic acids, proteins, or small molecules to the cell nucleus within the cell. Other types of localization signals may include nuclear export signals (NES), signals for retaining effector proteins in the cytoplasm, mitochondrial localization signals, chloroplast localization signals, endoplasmic reticulum retention signals, etc. In certain cases, effector proteins may not include subcellular localization signals to prevent targeting of the cell nucleus, which may be advantageous when the target nucleic acid is RNA present in the cytoplasm.
[0400] In some embodiments, the heterologous polypeptide may be an endosome escape peptide (EEP), which is designed to rapidly disrupt endosomes to minimize the residence time of delivery molecules (e.g., effector proteins) in the endosome environment, thereby avoiding encapsulation in endosome vesicles and subsequent degradation in the lysosomal compartment.
[0401] In addition, heterologous peptides can be cell-penetrating peptides (CPPs), also known as protein transduction domains (PTDs), which facilitate the penetration of lipid bilayers, cell membranes, organelle membranes, or vesicle membranes.
[0402] In some cases, heteropeptides or polypeptides may include protein tags that can be used as purification tags or fluorescent proteins. Such tags can be detectable for the identification or purification of effector proteins. Depending on the intended application, a variety of protein tags can be used, including but not limited to fluorescent proteins, histidine tags (e.g., 6XHis), hemagglutinin (HA) tags, FLAG tags, Myc tags, and maltose-binding protein (MBP). Examples of fluorescent proteins include green fluorescent protein (GFP), yellow fluorescent protein (YFP), red fluorescent protein (RFP), cyan fluorescent protein (CFP), mCherry, and tdTomato.
[0403] Heteropeptides can be located at or near the amino terminus (N-terminus) or carboxyl terminus (C-terminus) of the effector protein. In some cases, heteropeptides can be located at suitable insertion sites within the effector protein.
[0404] Guided Nucleic Acid The compositions, systems, and methods disclosed herein may include guide nucleic acids or their use. These may also include compositions, systems, and methods having one or more guide nucleic acids and DNA molecules encoding those guide nucleic acids. Those skilled in the art will understand that a DNA molecule “encoding” a nucleic acid (such as a guide nucleic acid) means a DNA molecule having a nucleotide sequence that, upon transcription, produces an RNA molecule (e.g., a guide nucleic acid). It should be understood that references to guide nucleic acids also include the DNA molecule encoding that guide nucleic acid.
[0405] Guide nucleic acids and their components (e.g., spacer sequences, repeat sequences, linker nucleotide sequences, stalk sequences, intermediate RNA, etc.) may consist of one or more deoxyribonucleotides (DNA), ribonucleotides (RNA), or combinations of both (e.g., RNA containing thymine bases) and biochemically or chemically modified nucleotides. Such nucleotide sequences may be described as DNA or RNA; however, regardless of form, it should be understood that these sequences may be adapted to RNA or DNA as needed to describe sequences within or encoding the guide nucleic acid, such as nucleotide sequences for vectors. The disclosure of the nucleotide sequence also includes its complementary sequence, reverse sequence, and reverse complementary sequence, any of which may serve as the nucleotide sequence for the guide nucleic acid.
[0406] In some implementations, the guide nucleic acid may include CRISPR RNA (crRNA) or a single guide RNA (sgRNA). sgRNA is formed by binding a spacer sequence (which hybridizes to a target sequence in the target nucleic acid) to a stalk sequence, wherein the two sequences are covalently linked. The spacer and stalk sequences can be linked by a phosphodiester bond or by one or more linked nucleotides. In some cases, the guide nucleic acid may include a spacer sequence, a repeat sequence, a stalk sequence, or a combination thereof, wherein the stalk sequence may comprise some or all of the repeat sequence.
[0407] In some embodiments, the composition may include tracrRNA. crRNA and tracrRNA may function as separate, unlinked molecules, or they may be covalently linked. Linkage may occur via a phosphodiester bond or via one or more linked nucleotides. In some embodiments, the composition may not include tracrRNA.
[0408] Guide RNAs can be naturally occurring or non-natural; non-natural guide RNAs may include chemical or biochemical modifications. Guide RNAs can be produced through chemical synthesis or recombination. The sequence of a guide RNA, or a portion thereof, may differ from the sequence of naturally occurring nucleic acids.
[0409] In some embodiments, the compositions, systems, and methods of this disclosure may include one or more additional guide nucleic acids or their uses. For example, these may include two or more additional guide nucleic acids (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, or more guide nucleic acids) that can target effector proteins to different locations within the target nucleic acid by binding to different portions of the target nucleic acid. The guide nucleic acid may bind to fragments of the target nucleic acid upstream or downstream of the target gene, thereby allowing modification at two different locations. This dual-targeting approach, referred to as "dual cleavage," may involve two effector proteins, each corresponding to a guide RNA, or a single effector protein with two different guide RNAs to achieve the dual cleavage effect. In some cases, multiple effector proteins (e.g., 2, 3, 4, 5, 6, 7, 9, 10, or more) may be used in the dual-guide system described herein.
[0410] In some embodiments, the guide nucleic acid may include 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 linked nucleotides. Typically, the guide nucleic acid will consist of at least a certain number of linked nucleotides, with some embodiments specifying at least 25 linked nucleotides. The length of the guide nucleic acid can range from 10 to 50 linked nucleotides. In some cases, the guide nucleic acid may consist primarily of about 12 to about 80, about 12 to about 50, about 12 to about 45, about 12 to about 40, about 12 to about 35, about 12 to about 30, about 12 to about 25, about 12 to about 20, about 12 to about 19, about 19 to about 20, about 19 to about 25, about 19 to about 30, about 19 to about 35, about 19 to about 40, about 19 to about 45, about 19 to about 50, about 19 to about 60, about 20 to about 25, about 20 to about 30, about 20 to about 35, about 20 to about 40, about 20 to about 45, about 20 to about 50, or about 20 to about 60 linked nucleotides. In some embodiments, the guide nucleic acid may have about 10 to about 60, about 20 to about 50, or about 30 to about 40 linked nucleotides.
[0411] In some implementations, the engineered guide nucleic acid comprises at least 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 consecutive nucleotides complementary to a eukaryotic sequence. The eukaryotic sequence refers to a nucleotide sequence present in a host eukaryotic cell, distinguishing it from sequences present in prokaryotic cells or viruses. This sequence can be located within genes, exons, introns, non-coding regions (e.g., promoters or enhancers), optional markers, tags, signals, or similar elements.
[0412] In some embodiments, the guide nucleic acid or the nucleic acid encoding the guide nucleic acid includes the nucleotide sequences described herein (e.g., Tables 4, 5, 6, or 8). These nucleotide sequences may be characterized as DNA or RNA; however, it should be understood that such sequences may be adapted as needed to describe sequences within the guide nucleic acid or sequences encoding it, such as nucleotide sequences for vectors. Furthermore, the disclosure of these nucleotide sequences also includes their complementary sequences, reverse sequences, and reverse complementary sequences, any of which may be used as nucleotide sequences for use in the guide nucleic acid.
[0413] Table 1. Exemplary CRISPR-Cas nucleases / effectants used in this paper: Table 2. Sequence List “ " represents the thiophosphate bond, and "m" represents the 2O' methyl group. The inserts in the pBLR plasmid are also called spacers.
[0414] Example Example 1 Cloning and plasmid preparation Cloning of the B-Gen.1 expression vector pBLR709, transfection in HEK293T cells, AMP-Seq, DNA library preparation and NGS analysis are described in detail in WO2022258753 (wherein the nuclease of SEQ ID NO: 2 is also disclosed in this application as SEQ ID NO: 185).
[0415] Cloning of the human albumin guide RNA expression vector was performed as follows. The receptor vector pBLR1936 was digested with BbsI-HF (New England Biolabs) according to the manufacturer's instructions. Oligonucleotides containing the respective spacer sequences of SEQ ID NOs: 1-61 were annealed to produce small double-stranded DNA fragments with suitable overhangs. The annealed oligonucleotides were incubated with pre-cut pBLR1936 in a Golden Gate reaction according to the manufacturer's instructions to produce pBLR2521-2582. A single Golden Gate reaction was performed for transformation and propagation in Top10 electroactive *E. coli*. The plasmids were Sanger sequenced, and positive clones were propagated in a mid-scale (midi) culture system and isolated using the NuceloSnap Plasmid Midi Kit (Macherey Nagel) for plasmid DNA.
[0416] HEK293T cell culture HEK293T (ATCC) ® CRL-3216™ cells were cultured in T75 flasks containing DMEM medium with 10% FCS and 1% Pen / Strep. Cells were passaged every 3 days.
[0417] HEK293T cell transfection Following the manufacturer's instructions, divide the cells as described above one day before transfection and count them using a Luna-FL™ dual-fluorescence cell counter (Biocat). Seed the cells in poly-D-lysine-coated 96-well plates at a density of 18,000 cells per well in 100 μl of culture medium.
[0418] Change the medium on the day of transfection. Adjust the LipoD293™ (SignaGen) reagent to room temperature shortly before transfection (approximately 10 min). On the day of transfection, mix 140 ng of pBLR709 DNA plasmid with 60 ng of guide RNA plasmid or guide RNA-free plasmid (nuclease control only) DNA and antibiotic-free medium, and add FCS to a total of 10 μl. In another tube, mix 10 μl of additive-free DMEM with 0.3 μl of LipoD293. Add the Lipo mixture to the DNA mixture. Seal the plate, vortex at 300 g for 10 seconds, and incubate at room temperature for 15 minutes. Then add the transfection mixture to the cells.
[0419] Cells were harvested three days after transfection. The culture medium was aspirated, and 50 μl of PBS was added to each well. Then, 25 μl of TrypLE was added to the top. Cells were trypsinized at 37°C for 15 minutes. The cells were resuspended in 60 μl of PBS. 60 μl of the cell suspension was added to a 96-well PCR tube. Cells were rotated at 300 g for 3 minutes. The supernatant was discarded, and cells were lysed in a thermal cycler containing 60 μl of lysis buffer (10 mM Tris, pH 7.0, 0.05% SDS) + 1:800 proteinase K.
[0420] AAV preparation AAV is generated carrying reporter 2, which is a bidirectional promoterless nanoluc-eGFP-polyA loading fragment (SEQ ID NO: 210).
[0421] Guide RNA (gRNA) Synthetic guides (SEQ ID NOs: 195-199, 205-209) are ordered from IDT or Axolabs for research use only (RUO). gRNA is reconstituted with IDTE at pH 7.5 and stored in 5 μl samples at -70°C for single use.
[0422] mRNA source SpCas9 mRNA is derived from Trilink. The ORF sequence can be found at the following URL: trilinkbiotech.com / media / maravai / productattachments / product_insert / cas9_catno_l-8106_l-7206_l- 7606_.txt, accessed on 2024-12-17 .
[0423] The B-GEn.1 mRNA described below includes B-GEn.1 ORF SEQ ID NO: 210. Using a codon table preferably used for humans, B-GEn.1 mRNA was artificially depleted by changing T-rich codons to T-free codons. Codons were selected to have the lowest possible uridine content while maximizing the expression of the corresponding tRNA in the liver. Reducing uridine content aims to minimize the innate immune response to the mRNA and provide other benefits. The SpCas9 and SEQ ID NO: 210 mRNAs used in this study were derived from TriLink as a product for research use only (RUO) and prepared by adding a Poly A sequence to a plasmid template via PCR amplification. All mRNAs were characterized by N1-methyl pseudo-U modification, a 120 nt polyA tail, CleanCapAG capping reagent, purification with a silica membrane, and suspension in 1 mM sodium citrate at pH 6.4.
[0424] LNP preparation scheme For the experiments described in other embodiments, the LNP was prepared using a cross-flow technique, which utilizes an impinging jet to mix lipids in ethanol with three volumes of RNA solution and one volume of water. The lipids in ethanol were then mixed with three volumes of RNA solution via cross-flow mixing. A fourth water stream was mixed with the cross-flow outlet stream through an inline tee connector. The LNP was maintained at room temperature for 1 hour and further diluted with water (approximately 1:1 volume / volume (v / v)). The diluted LNP was concentrated using tangential flow filtration on a plate filter (Sartorius, 100 kD MWCO), and then the buffer was exchanged for 50 mM Tris, 45 mM NaCl, 5% (w / v) sucrose, pH 7.5 (TSS) via dialysis. Alternatively, the final buffer exchange for TSS was performed using a PD-10 desalting column. If necessary, the formulation was concentrated by centrifugation using an Amicon 100 kDa centrifuge filter (Millipore). The resulting mixture was then filtered using a 0.2 μm sterile filter. The final LNP was stored at -80°C until further use. LNP was formulated with a molar ratio of ionizable lipids:cholesterol:DSPC:PEG2k DMG of 50:38:9:3, wherein the molar ratio of lipid amines to RNA phosphate (N:P) was approximately 6.0, and the weight ratio of gRNA to mRNA was 1:1.
[0425] Cell Culture PHH Maintenance culture medium - Williams medium E (Cat. W1878-500ML, Sigma Aldrich) Contains 5% FCS (Cat. S0615, Sigma Aldrich), 1% P / S (Cat. P4333, Sigma Aldrich), 15 mM HEPES (Cat. 15630-056, Gibco), 1x ITS (Cat. I3146-5ML, Sigma Aldrich), final 6.25 µg / mL insulin, 6.25 µg / mL transferrin, 1.25 ng / mL selenite, 1x GlutaMAX (Cat. 35050-038, Gibco), 50 µg / mL gentamicin (Cat. G1272, Sigma Aldrich), and 100 nM dexamethasone (A13449, Gibco).
[0426] Primary human hepatocytes (Lonza, multiple batches) Inoculation and treatment Thaw and resuspend hepatocytes (PHH) in 35 mL of maintenance medium. Add 90% Percoll (diluted in 10x PBS). Carefully mix the resuspended cells and centrifuge at 150 x g for 5 min. Discard the supernatant, resuspend the cell pellet in 50 mL of maintenance medium, and centrifuge at 150 x g for 5 min. Carefully resuspend the resulting cell pellet. Count the cells and seed them at a density of 50,000 cells / well in 96-well plates previously coated with 0.0006% collagen R (Cat. 47254.02, SERVA) for 30 min using tissue culture. Allow the seeded cells to settle and adhere for 1 h at room temperature, and then incubate for 3 h in a tissue culture incubator at 37°C and 5% CO2. After cell incubation, check for monolayer formation and wash twice with maintenance medium.
[0427] PHH cell transfection Prior to treatment, lipid nanoparticles (LNPs) containing Cas9 or B-Gen.1 mRNA were diluted in phosphate-buffered saline (PBS) using serial 1:2 dilutions from 40 µg RNA / mL to 0.625 µg / mL. The resulting dilutions were further diluted 1:4 in maintenance medium (e.g., 40 µL LNP dilution + 120 µL maintenance medium). AAV containing nanoluciferase DNA was diluted in maintenance medium to an MOI of 2e5 (6.25E10 vg / mL). Before transfection, medium was aspirated from the cells, and 40 µL of each LNP was added to the cells, followed by 120 µL of diluted AAV. Cells were incubated in a tissue culture incubator at 37°C and 5% CO2.
[0428] Twenty-four hours after treatment, aspirate the supernatant and wash the cells twice with serum-free medium. Add 20% collagen I solution (diluted in serum-free medium) to each well. Incubate the plate at 37°C. After two hours, check the plate to see if the gel has solidified and add 200 μL of maintenance medium.
[0429] Cells were imaged, and supernatant was collected on days 3, 6, and 7 post-transfection, including changes in culture medium. On day 7, cells were frozen to -20°C for further analysis.
[0430] RLU detection For experiments involving the detection of NanoLuc in cell culture medium, use 1 volume of Nano-Glo ® Luciferase detection substrate and 50 times the volume of Nano-Glo ® The luciferase assay buffer was combined. The assay was run on an Agilent Biotek Synergy Neo2 hybrid multimode microplate analyzer with an integration time of 0.2 seconds and a gain of 200. Undiluted sample and assay reagent were combined in a 1+1 ratio (e.g., 10 µl + 10 µl) in black 384-well plates, with the bottom closed. After combining the sample and reagent, the plate was incubated on a plate shaker at medium speed for 1 minute before measurement.
[0431] Crude lysates derived from cell lines and primary human hepatocytes Frozen cells were lysed in 30 μl (pH 7.0) or 60 μl (HEK293T and HEPA1-6) of crude lysis buffer (10 mM Tris, pH 7.0, 0.005% SDS) supplemented with proteinase K (1:800 dilution). The solution was transferred to 96-well plates and sealed. The cycling conditions were as follows: Step 1: 60 min at 37°C, 20 min at 65°C.
[0432] Mouse genomic DNA extraction The liver-derived tissue was cut into approximately 10 mg pieces using a scalpel. Tissue homogenization was performed using Qiagen TissueLyser II, following the protocol instructions. DNA extraction was performed using KingFisher Apex, following the manufacturer's instructions.
[0433] Amplicon generation To generate amplicones, endpoint PCR was performed using barcode primers. The PCR reaction preparation is as follows: 12.5 µL 2x Q5 Mastermix, 8 µL H2O, 3 µL barcoding primers, 1.5 µL crude lysed genomic DNA Use the following cycling conditions: Step 1: 98°C, 30 sec; Step 2: 98°C, 10 sec; Step 3: 66°C, 20 sec; Step 4: 72°C, 20 sec; Step 5: 32 cycles of Steps 2-4; Step 6: 72°C, 2 min; Step 7: Hold at 12°C. For gene editing evaluation, amplicon derived from human albumin intron 1 (SEQ ID NO: 188-193) and mouse albumin intron 1 (SEQ ID NOs: 194, 376-379) were established.
[0434] DNA library preparation and NGS analysis Combine all PCR reactions. For this, combine 5 μl of individual reactions. Mix the total volume of the combined PCR with an equal volume of AMPure XP Bead cleanup. After incubating on a rotor for 5 minutes, place the mixture on a magnet and extract DNA according to the manufacturer's instructions. For repairing ligation of Illumina aptamers to the ends of spliced fragments, use NEBNext. ® Ultra™ II end-repair / dA-tailing module. A total of 2 μg amplicon (up to 80 μL) was incubated with 6 μL NEBNext Ultra II end-preparation enzyme mixture and 14 μL NEBNext Ultra II end-preparation reaction buffer. Then, H2O was added to bring the volume to 100 μL. The reaction in the thermal cycler followed the protocol: 20 °C, 30 min, 65 °C, 30 min, maintained at 4 °C. Samples were purified with AMPure XP Bead Cleanup and eluted in 32 μL EB buffer. For ligation, we used Blunc / TA Ligase Master Mix from NEB and Illumina aptamers from the TruSeq DNA PCR-FreeLT Sample Prep Kit. 30 μL of end-prepared DNA was mixed with 2 μL Illumina and 32 μL Blunc / TA Ligase Master Mix. It was incubated in a thermal cycler at 21 °C for 1 h. Then, 34 μl of H2O was added to a total of 100 μl, and AMPure XP BeadCleanup containing a final elution step in 50 μl of EB buffer was performed. The double-stranded DNA library was measured using a Qubit 4 HS system according to the manufacturer's instructions. The DNA library was diluted to 1 nm with RSB buffer, then diluted to 450 pM. The DNA library was loaded into a NextSeq 1000 Sequencer (Illumina) using the manufacturer's protocol.
[0435] The original fastq file uses cutadapt version 1.18 ( https: / / cutadapt.readthedocs.io / en / v1.18 / index.html The reads (accessed on 2024-12-17) underwent quality trimming with a minimum quality score of Q30. Filtered reads were joined using fastq-join version 1.3.1 and partitioned using custom demultiplexing as described in WO2022258753 (storage.googleapis.com). The partitioned files were then analyzed using CRISPressov.1.0.13 (doi: 10.1038 / nbt.3583). The final values were plotted using GraphPad Prism 9.50 (GraphPad Software).
[0436] Example 2 A screening method for detecting nuclease activity at the human albumin intron 1 locus in mammalian cells (HEK293T).
[0437] In this embodiment, HEK293T cells were cultured and transfected with the B-GEn.1 nuclease plasmid (pBLR709) and gRNA expression plasmid (see Table 3) to evaluate nuclease-mediated DNA double-strand insertion / deletion formation as described in Example 1.
[0438] Table 3. Insertion and deletion of human albumin intron 1 in HEK293T cells Example 3 Screening methods for detecting the integration activity of B-GEn.1 nuclease and albumin intron 1 locus in primary human hepatocytes (PHH). In this embodiment, PHH cells were cultured as described, infected with AAV8, and transfected with LNP-form B-GEn.1 mRNA and synthetic gRNA (see Table 4) to assess AAV helper insertion and nuclease-mediated DNA double-strand insertion / deletion formation as described in Example 1. (See Table 4 and...) Figure 2a and Figure 2b As shown, SEQ ID 199 displays significant editing and integration signals.
[0439] Table 4. Editing and integration of B-GEn.1 in PHH Example 4 A screening method for detecting nuclease activity (HEPA1-6) at the mouse albumin intron 1 locus in mammalian cells. In this example, HEPA1-6 cells were cultured and transfected with the B-GEn.1 nuclease plasmid (pBLR709) and a gRNA expression plasmid targeting mouse albumin intron 1 (see Table 5) to evaluate nuclease-mediated DNA double-strand insertion / deletion (indel) formation as described in Example 1. (See Table 5 and...) Figure 3 As shown, some gRNAs exhibit significant editing.
[0440] Table 5. Insertion-deletion effects of targeting mouse albumin intron 1 in HEPA1-6 cells Example 5 Screening methods for detecting the integration activity of nucleases and mouse albumin intron 1 locus in C57 / BL6 mice. In this embodiment, C57 / BL6 mice were injected with AAV8 (SEQ ID NO: 213) and then transfected with LNP. The LNP loading consisted of B-GEn.1 mRNA (SEQ ID NO: 214) and synthetic gRNA (see SEQ ID NO: in Table 6). The LNP transfection dose ranged from 0.75 to 2 mg (mpk) per kg, while the AAV dose remained constant at 1E+11. AAV8 control 2 was injected at 9 am on day 1, followed by LNP injection at 3 pm on day 1. Serum was collected after 1 week and the expression of nanoluciferase (nLuc) was measured. Liver was harvested and genomic DNA preparation, PCR, and insertion / deletion assessment were performed as described in Example 1. Insertion / deletion and RLU values are shown in Table 6 and Figure 4 The description is as follows. SEQ ID NO: 208 shows the highest levels of insertions / deletions and RLU.
[0441] Table 6. Editing and integration of B-GEn.1 in C57 / BL6 mice
Claims
1. A system comprising: (a) a deoxyribonuclease (DNA) or a nucleic acid encoding the DNA endonuclease, selected from SEQ ID NO: 185 or SEQ ID NO: 186, or any sequence thereof having 95%, preferably 98%, more preferably 99% identity with these sequences; and (b) a guide RNA (gRNA) or nucleic acid encoding said gRNA, said gRNA comprising a spacer region sequence selected from any one of SEQ ID NOs: 1, 2, 3, 4, 5, 6, 7, 8, 9, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 35, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 51, 52, 54, 55, 56, 58, 59, 60; and (c) Donor nucleic acid, which includes a nucleic acid sequence encoding a target gene (GOI) or a functional derivative thereof.
2. A system comprising: (a) a deoxyribonuclease (DNA) or a nucleic acid encoding the DNA endonuclease, selected from SEQ ID NO: 185 or SEQ ID NO: 186, or any sequence thereof having 95%, preferably 98%, more preferably 99% identity with these sequences; and (b) a guide RNA (gRNA) selected from any one of SEQ ID NOs: 62-122, or a nucleic acid encoding said gRNA, or any sequence thereof having 95%, preferably 98%, more preferably 99% identity with these sequences; and (c) Donor nucleic acid, which includes a nucleic acid sequence encoding a target gene (GOI) or a functional derivative thereof.
3. The system according to claim 1, wherein the guide RNA (gRNA) comprises a spacer sequence selected from any one of SEQ ID NOs: 3, 4, 5, 6, 9, 11, 12, 13, 14, 15, 16, 20, 21, 22, 25, 26, 32, 39, 42, 44, 46, 51, 52, 54, 55, 56, 58, 59, 60, or a nucleic acid encoding the gRNA.
4. The system according to claim 1, wherein the guide RNA (gRNA) comprises a spacer sequence selected from any one of SEQ ID NOs: 4, 5, 6, 9, 12, 13, 14, 15, 20, 26, 32, 39, 42, 52, 54, 55, 58, 59, or a nucleic acid encoding the gRNA.
5. The system according to claim 1, wherein the guide RNA (gRNA) comprises a spacer sequence selected from any one of SEQ ID NOs: 4, 5, 6, 9, 15, 20, 42, 52, 59, or a nucleic acid encoding the gRNA.
6. The system of claim 1, wherein the guide RNA (gRNA) comprises a spacer sequence selected from any one of SEQ ID NOs: 4 or 42, or a nucleic acid encoding the gRNA.
7. The system according to any one of claims 1-6, wherein the GOI encodes acidic α-glucosidase (GAA).
8. A method for editing a genome in a cell, the method comprising: Provide the cells with the following substances: (a) a deoxyribonuclease (DNA) or a nucleic acid encoding the DNA endonuclease, selected from SEQ ID NO: 185 or SEQ ID NO: 186, or any sequence thereof having 95%, preferably 98%, more preferably 99% identity with these sequences; and (b) a guide RNA (gRNA) or nucleic acid encoding said gRNA, said gRNA comprising a spacer region sequence selected from any one of SEQ ID NOs: 1, 2, 3, 4, 5, 6, 7, 8, 9, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 35, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 51, 52, 54, 55, 56, 58, 59, 60; and (c) Donor nucleic acid, which includes a nucleic acid sequence encoding a target gene (GOI) or a functional derivative thereof.
9. The method according to any one of claims 1-8, wherein the cell is a hepatocyte.
10. The system of any one of claims 1-7 is used to treat a disease or health condition in a subject, wherein (a), (b) and (c) are provided to cells in the subject.
11. The system of any one of claims 1-7 or the method of claim 8, wherein the components (a) to (c) are encoded by one or a single viral vector, optionally, such viral vector is AAV.
12. The system of any one of claims 1-7 or the method of claim 8, wherein the components (a) to (c) are formulated in one or more liposomes or lipid nanoparticles.
13. A genetically modified cell, wherein the genome of the cell is edited by the method of any one of claims 8 or 9.
14. The system of any one of claims 1-7 is used to treat a disease or health condition, wherein (a), (b) and (c) are provided to cells in a subject.
15. The system of claim 7 for treating Pompe disease, wherein (a), (b) and (c) are provided to cells in a subject.
16. A kit comprising elements of the system according to any one of claims 1-7 and 11, 12, and further comprising instructions for use.
Citation Information
Patent Citations
Compositions and methods for delivering transgenes
WO2020081843A1
Compositions and methods for transgene expression from an albumin locus
WO2020082042A2
Type v RNA programmable endonuclease systems
WO2022258753A1
Effector protein compositions and methods of use thereof
WO2023220649A2