Editing systems and components and uses thereof
By designing oligonucleotide compositions of target recognition regions and protein recognition regions, the problems of low efficiency and insufficient specificity of the CRISPR-Cas system in gene editing and epigenetic regulation have been solved, achieving more efficient and precise genome editing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- IONIS PHARMACEUTICALS INC
- Filing Date
- 2024-09-13
- Publication Date
- 2026-06-09
AI Technical Summary
Existing CRISPR-Cas systems suffer from inefficiency and lack of specificity in gene editing and epigenetic regulation, especially in the identification and cleavage of targeted DNA sequences.
Guided nucleic acid binding agents such as Cas proteins, inactivated Cas proteins, and Cas fusion proteins are provided, along with editing systems that improve the specific recognition and cleavage of target DNA by designing oligonucleotide compositions of target recognition and protein recognition regions.
It significantly improves the efficiency and specificity of gene editing, enhances the ability to recognize and cut target DNA, and enables more precise genome editing.
Smart Images

Figure CN122180767A_ABST
Abstract
Description
[0001] sequence list This application is submitted together with an electronic sequence listing. The sequence listing is provided as a file named EDIT0001SEQ.xml, created on September 13, 2024, and is 2,646 KB in size. Information about the electronic sequence listing is incorporated herein by reference in its entirety. Technical Field
[0002] This embodiment provides a guided nucleic acid binder, such as a Cas protein; a guide; an editing system comprising a guided nucleic acid binder and a guide; and a method of using the same. Background Technology
[0003] Cas proteins, along with their associated clusters of regularly spaced short palindromic repeats (CRISPR) guide RNAs (gRNAs), appear to be a ubiquitous component of the prokaryotic immune system (present in approximately 45% of bacteria and 84% of archaea). These CRISPR-gRNA cleavage systems protect these microorganisms from non-self nucleic acids such as infectious viruses and plasmids. CRISPR-associated (Cas) proteins are highly diverse, containing a variety of nucleic acid-binding domains. Cas nucleic acid-binding domains can be applied to a variety of applications, including but not limited to gene editing and epigenetic regulation of gene expression. Summary of the Invention
[0004] Some embodiments provided herein are directed to guided nucleic acid binders (Cas proteins, inactivated Cas proteins, and Cas fusion proteins); guides (single-guided and dual-guided); and editing systems comprising such guided nucleic acid binders and guides. Some embodiments are directed to methods of using such editing systems.
[0005] In some embodiments, this document provides a polypeptide. In some embodiments, the polypeptide is a guided nucleic acid binder. In some embodiments, the polypeptide provided herein is a Cas protein. In some such embodiments, the Cas protein is a Cas enzyme and induces the cleavage (double-strand break) of both strands of a double-stranded DNA target. In some embodiments, the Cas protein is a Cas enzyme and induces the cleavage (cut) of one strand of a double-stranded DNA target. In some embodiments, the Cas protein is an inactivated Cas protein that does not induce the cleavage of double-stranded DNA. In some embodiments, the polypeptide comprises a Cas protein. Some Cas proteins may be fused with one or more other proteins or protein domains (also called "heterologous domains") to produce Cas fusion proteins.
[0006] In some embodiments, this document provides nucleic acids that encode the polypeptides described above. In some embodiments, such nucleic acids are DNA (e.g., template DNA) or RNA (e.g., mRNA).
[0007] In some embodiments, a guide is provided herein. In some embodiments, the guide consists of a single oligonucleotide (single guide). In some embodiments, the guide consists of two oligonucleotides (a first oligonucleotide and a second oligonucleotide) that are partially complementary to each other, thus enabling them to hybridize together (double guide). The guide includes a target recognition region and a protein recognition region. For a single guide, the target recognition region and the protein recognition region are regions of a single oligonucleotide. For a double guide consisting of a first oligonucleotide and a second oligonucleotide, the first oligonucleotide contains a first portion of the target recognition region and the protein recognition region, and the second oligonucleotide contains a second portion of the protein recognition region. In some such embodiments, the first oligonucleotide may be referred to as “crRNA”, and the second oligonucleotide may be referred to as “tracrRNA”. The target recognition region of the guide is a region of an oligonucleotide (crRNA in the case of a single guide or a double guide) having a nucleobase sequence complementary to an equal-length portion of the target sequence of the target nucleic acid. In some embodiments, the target nucleic acid is target DNA.
[0008] In some embodiments, the editing system comprises a guided nucleic acid binder and a guide. In some such embodiments, the guided nucleic acid binder is a Cas protein. In some embodiments, the guide is a single guide. In some embodiments, the editing system comprises a Cas protein and a single guide.
[0009] In some embodiments, this document provides a composition comprising template DNA encoding both a guide-bound nucleic acid binder and a guide. In some embodiments, this document provides a composition comprising template DNA encoding both a Cas protein and a guide. In some embodiments, this document provides a composition comprising a guide and mRNA encoding a guide-bound nucleic acid binder. In some embodiments, this document provides a composition comprising a guide and mRNA encoding a Cas protein. In some embodiments, this document provides a composition comprising a single guide and mRNA encoding a guide-bound nucleic acid binder. In some embodiments, this document provides a composition comprising a dual guide and mRNA encoding a guide-bound nucleic acid binder. In some embodiments, this document provides a composition comprising a single guide and mRNA encoding a Cas protein. In some embodiments, this document provides a composition comprising a dual guide and mRNA encoding a Cas protein. In such embodiments, the Cas protein is selected from Cas enzymes, inactivated Cas proteins, and Cas fusion proteins. In some embodiments, such compositions are administered to a subject to induce gene editing. Brief description of the attached diagram Figure 1 A gel demonstrating ION-CAS protein cleavage activity in HEK293T nuclear extract using a universal target recognition region with a guide is shown.
[0011] Figure 2 A gel demonstrating ION-CAS protein cleavage activity in HEK293T nuclear extract using Cas-specific target recognition region and Cas-specific protein recognition region in guide RNA is shown. Figure 3 The common PAM sequences of IONCAS008, IONCAS009 and IONCAS0016 are shown.
[0012] Figures 4a to 4c show gels exhibiting ION-CAS protein cleavage activity in HEK293T nuclear extract using a universal target recognition region guided by a wizard.
[0013] Figure 5 A schematic diagram of the CAS editing system is provided.
[0014] Figure 6 A schematic diagram of the secondary structure of the SpCas9 wizard is provided. Detailed Implementation
[0015] It should be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only, and do not limit the claimed embodiments. In this document, the singular is used to include the plural unless otherwise specified. As used herein, the word “or” means “and / or” unless otherwise specified. Furthermore, the term “including” and other forms of use such as “includes” and “included” are not limiting.
[0016] The section headings used herein are for organizational purposes only and should not be construed as limiting the subject matter described. All documents or portions thereof cited in this application, including but not limited to patents, patent applications, articles, books, papers, and GenBank and NCBI reference sequence records, are hereby expressly incorporated by reference into the portions of the documents discussed herein, and in their entirety.
[0017] It should be understood that, throughout this specification, unless otherwise stated, the first letter in a peptide sequence is the first amino acid at the N-terminus of the peptide, and the last letter in a peptide sequence is the next amino acid at the C-terminus of the peptide. Similarly, unless otherwise stated, the first nucleoside in a nucleotide sequence represents the 5' end of the nucleotide, and the last letter in the nucleotide sequence represents the 3' end.
[0018] Unless specifically defined, the nomenclature, procedures, and techniques used in relation to analytical chemistry, synthetic organic chemistry, and pharmaceutical and medicinal chemistry as described herein are those well-known and commonly used in the art. Where permitted, all patents, applications, published applications, and other publications, as well as other data, mentioned throughout this disclosure are incorporated herein by reference in their entirety.
[0019] Unless otherwise stated, the following terms have the following meanings: As used herein, “2’-deoxynucleoside” refers to a nucleoside containing a 2’-H(H)deoxyfuranose moiety. In some embodiments, the 2’-deoxynucleoside is a 2’-β-D-deoxynucleoside and contains a 2’-β-D-deoxyribosyl sugar moiety having the β-D ribosyl conformation found in naturally occurring deoxyribonucleic acid (DNA). In some embodiments, the 2’-deoxynucleoside may contain modified nucleotides or may contain RNA nucleotides (uracil).
[0020] As used herein, “2’-F” refers to a 2’-F group that replaces the 2’-OH group of the furanose moiety. “2’-fluorosugar moiety” or “2’-F sugar moiety” refers to a sugar moiety having a 2’-F group that replaces the 2’-OH group of the furanose moiety. Unless otherwise stated, the 2’-F sugar moiety is β-D-ribosyl configuration.
[0021] As used in this article, "2'-F nucleoside" refers to a nucleoside containing the 2'-F sugar moiety.
[0022] As used herein, “2’-MOE” refers to a 2’-OCH2CH2OCH3 group that replaces the 2’-OH group of the furanose moiety. “2’-MOE sugar moiety” refers to a sugar moiety having a 2’-OCH2CH2OCH3 group that replaces the 2’-OH group of the furanose moiety. Unless otherwise stated, the 2’-MOE sugar moiety is β-D-ribosyl. “MOE” means O-methoxyethyl.
[0023] As used in this article, "2'-MOE nucleoside" refers to a nucleoside containing the 2'-MOE sugar moiety.
[0024] As used herein, “2’-NMA” refers to a 2’–O-CH2-C(=O)-NH-CH3 group that replaces the 2’-OH group of the ribosyl sugar moiety. A “2’-NMA sugar moiety” is a sugar moiety having a 2’–O-CH2-C(=O)-NH-CH3 group that replaces the 2’-OH group of the ribosyl sugar moiety. Unless otherwise stated, the 2’-NMA sugar moiety is β-D configured. “NMA” refers to O-(N-methyl)acetamide.
[0025] As used in this article, "2'-NMA nucleoside" refers to a nucleoside containing the 2'-NMA sugar moiety.
[0026] As used herein, “2’-OMe” refers to a 2’-OCH3 group that replaces the 2’-OH group of the furanose moiety. “2’-O-methyl sugar moiety” or “2’-OMe sugar moiety” refers to a sugar moiety having a 2’-OCH3 group that replaces the 2’-OH group of the furanose moiety. Unless otherwise stated, the 2’-OMe sugar moiety is β-D-ribosyl configuration.
[0027] As used in this article, "2'-OMe nucleoside" refers to a nucleoside containing the 2'-OMe sugar moiety.
[0028] As used herein, “2’-substituted nucleoside” means a nucleoside that contains a 2’-substituted sugar moiety. As used herein, “2’-substituted” with respect to the sugar moiety means a sugar moiety containing at least one 2’-substituent other than H or OH.
[0029] As used herein, "5-methylcytosine" refers to cytosine modified with a methyl group attached to position 5. 5-methylcytosine is a modified nucleobase.
[0030] As used herein, "active site" refers to a region of a protein where a substrate molecule binds to carry out enzymatic catalysis of a chemical reaction. An active site contains one or more catalytic residues. As used herein, "catalytic residue" refers to an amino acid that participates in the enzymatic catalysis of the chemical reaction at its active site. Catalytic residues may directly participate in the catalytic mechanism (e.g., as a nucleophile), may exert a catalytically conducive effect on another residue directly involved in the catalytic mechanism or on a water molecule (e.g., through electrostatic or acid-base interactions), may stabilize transition intermediates, and / or may exert a catalytically conducive effect on a substrate or cofactor (e.g., by polarizing a bond about to break), including steric and electrostatic effects.
[0031] As used herein, "Cas enzyme" refers to a guided nucleic acid binder capable of cleaving one or both strands of double-stranded DNA. In some embodiments, the target nucleic acid is double-stranded "target DNA". A "Cas enzyme" can be an enzyme capable of cleaving both strands of double-stranded DNA, or it can be an enzyme capable of cleaving only one strand of double-stranded DNA ("Cas cleavage enzyme") (e.g., containing a mutation in one, but not two, catalytic residues).
[0032] As used in this article, "Cas protein" refers to "Cas enzyme" or "inactivated Cas protein".
[0033] As used in this article, "inactivated Cas protein" refers to a guided nucleic acid binder that cannot cleave either strand of double-stranded DNA.
[0034] As used herein, “cEt” or “restricted ethyl” refers to the bicyclic sugar moiety, wherein the first ring of the bicyclic sugar moiety is a ribosyl sugar moiety, and the second ring of the bicyclic sugar is formed via a bridge connecting the 4'-carbon and the 2'-carbon, the bridge having the formula 4'-CH(CH3)-O-2', and the bridge is... S Configuration. The cEt bicyclic sugar moiety is in the β-D configuration.
[0035] As used herein, “cleavable portion” means a bond or group of atoms that is broken under physiological conditions (e.g., inside a cell or in a subject).
[0036] As used herein, "complementarity" in oligonucleotides means that when the nucleotide sequence of an oligonucleotide is aligned in reverse to the nucleotide sequence of another nucleic acid, at least 70% of the nucleotide bases of that oligonucleotide are hydrogen-bonded to the nucleotide bases of the other nucleic acid or one or more regions thereof. "Complementary region" in oligonucleotide terms means that when the nucleotide sequence of an oligonucleotide is aligned in reverse to the nucleotide sequence of another nucleic acid, at least 70% of the nucleotide bases of that region are hydrogen-bonded to the nucleotide bases of the other nucleic acid or one or more regions thereof. Complementary nucleobases are nucleobases capable of forming hydrogen bonds with each other. Complementary nucleobase pairs include adenine (A) and thymine (T), adenine (A) and uracil (U), cytosine (C) and guanine (G), and 5-methylcytosine (mC) and guanine (G). Complementary oligonucleotides and / or nucleic acids do not need to have nucleobase complementarity at every nucleoside. Instead, some mismatches are tolerable. As used in this article, “perfectly complementary” or “100% complementary” for oligonucleotides means that the oligonucleotide is complementary to another oligonucleotide or nucleic acid at each nucleoside of the oligonucleotide.
[0037] As used herein, the “complementary strand” of the target DNA is a strand of double-stranded DNA complementary to the target recognition region of the guide. As used herein, the “non-complementary strand” of the target DNA is a strand of double-stranded DNA complementary to the “complementary strand.” The PAM sequence is found within the “non-complementary strand.”
[0038] As used in this article, "domain" refers to a subset of amino acids linked together in a polypeptide. A domain can be composed of multiple smaller domains called subdomains, for example... Streptococcus pyogenes The RuvC domain of SpyCas9 consists of three discontinuous subdomains: RuvC-I, RuvC-II, and RuvC-III.
[0039] As used in this article, "double-strand break" or "DSB" refers to the breakage of the two strands of double-stranded DNA.
[0040] As used herein, “exogenous mRNA” means any mRNA introduced into an organism or cell that is not synthesized by the recipient organism or cell itself. Exogenous mRNA can be isolated or purified from an organism or cell, transcribed in vitro, or produced synthetically. Exogenous mRNA contains a coding region (e.g., an open reading frame (ORF)) that encodes a polypeptide sequence.
[0041] As used herein, “gene editing,” “genome editing,” or “editing of the genome” means a process that alters the nucleobase sequence of a genome (e.g., insertion, deletion, mutation, or alteration of nucleobases, including epigenetic states (e.g., methylation of A or G)) directly or through innate cellular processes (e.g., after the introduction of double-strand breaks or nicks). In some embodiments, gene editing is mediated by a complex comprising a guided nucleic acid binder and a guide. In some embodiments, gene editing is mediated by a complex comprising a Cas protein, an inactivated Cas protein, or a Cas fusion protein and a guide.
[0042] As used herein, "guide" refers to an oligonucleotide ("single guide") or a complex ("double guide"), which consists of two oligonucleotides that partially hybridize with each other, and in both cases includes a target recognition region and a protein recognition region. In embodiments where the guide is single-guided, the target recognition region and the protein recognition region are regions of one oligonucleotide constituting the single guide. In embodiments where the guide is double-guided, the target recognition region is a region of the first oligonucleotide, and the protein recognition region is part of a complex comprising regions of each of the two oligonucleotides. In some embodiments, the guide directs a guided nucleic acid binder to a target sequence of the target DNA.
[0043] As used herein, “guided nucleic acid binder” means a polypeptide comprising: (1) a region that interacts with a guide protein recognition region; and (2) a region that interacts with a PAM of the target DNA. In some embodiments, the guide target recognition region enables the guided nucleic acid binder to specifically interact with the target DNA. The guided nucleic acid binder may comprise one or more active or inactive nuclease domains. The guided nucleic acid binder may comprise heterologous domains. Guided nucleic acid binders include Cas enzymes, inactivated Cas proteins, and Cas fusion proteins.
[0044] As used herein, "HNH domain" or "HNH" refers to the cation-dependent endonuclease domain of a protein having an active site containing a histidine (H) catalytic residue. In some embodiments, the active site of the HNH domain contains both histidine (H) and asparagine (N) catalytic residues. In some embodiments, the HNH domain of the Cas protein cleaves the complementary strand of the target DNA. The HNH domain may have one or more substitutions that inactivate its catalytic activity.
[0045] As used herein, “hybridization” means the annealing of oligonucleotides and / or nucleic acids. While not limited to a specific mechanism, the most common hybridization mechanisms involve hydrogen bonds between complementary nucleobases, which can be Watson-Crick, Hoogsteen, or reverse Hoogsteen hydrogen bonds. In some embodiments, complementary nucleic acid molecules include, but are not limited to, antisense compounds and nucleic acid targets.
[0046] As used herein, “identity” or “percentage identity” in the context of amino acid sequences refers to the percentage of identical amino acids between two sequences when the amino acid sequence alignment achieves maximum similarity. Percentage identity can be measured using methods such as BLASTP (Altschul et al.). J. Mol.Biol., 1990; Altschul, et al. Nucleic Acids Research, (1997) or Clustal Omega (Sievers, et al.) Molecular Sys.Biol. 2011; Goujon et al., Nucleic Acids Research, 2010; McWilliam, et al. Nucleic Acids Research (2013).
[0047] As used herein, the term “interaction” in relation to two macromolecules (e.g., polypeptides and oligonucleotides; two polypeptides with each other; or polypeptides and nucleic acids) means the formation of multiple non-covalent stable interactions between the two macromolecules, including but not limited to dipole interactions (hydrogen bonds), cation-π or π-π stacking interactions, electrostatic interactions (salt bridges), and hydrophobic interactions.
[0048] As used herein, the term "nucleoside bond" is a covalent bond between adjacent nucleosides in an oligonucleotide. As used herein, "modified nucleoside bond" refers to any nucleoside bond other than a phosphodiester nucleoside bond. As used herein, "mismatch" or "non-complementary" means that when the first and second oligonucleotides are aligned, the nucleobases of the first oligonucleotide are not complementary to the corresponding nucleobases of the second oligonucleotide or the target nucleic acid.
[0049] As used in this article, "motif" refers to the pattern of unmodified and / or modified sugar moieties, nucleobases, and / or nucleoside bonds in an oligonucleotide.
[0050] "Modified guide" refers to a guide that contains modified nucleotide interbonds, modified sugar moieties, and / or modified nucleobases.
[0051] As used in this article, “natural amino acid” means Gly or each of the following.L -Isomers: Ala, Arg, Asn, Asp, Cys, Gln, Glu, His, Ile, Lys, Leu, Met, Phe, Pro, Ser, Thr, Trp, Tyr, Val.
[0052] As used in this article, a “cut” refers to a break in one strand of a double-stranded DNA.
[0053] As used herein, “non-natural amino acid” means any amino acid other than the standard twenty amino acids encoded by the human genetic code, including each of the following. D -Isomers: Ala, Arg, Asn, Asp, Cys, Gln, Glu, His, Ile, Lys, Leu, Met, Phe, Pro, Ser, Thr, Trp, Tyr, Val. Non-natural amino acids may have modified or functionalized side chains (e.g., through attachment via linkers). Examples of non-natural amino acids include, but are not limited to, alloleucine, 2-amino-3-ethyl-valerate, aminoisobutyric acid, aminobutyric acid, aziridine, 7-azatryptophan, 6-azidolysine, β-cyclobutylalanine, β-methylisoleucine, 4,4-biphenylalanine, cis-hydroxyproline, cyclobutylglycine, cyclohexylglycine, cyclopentylalanine, cyclopentylglycine, 2,6-dimethyltyrosine, 3,3-diphenylalanine, 4-trans-hydroxy-L-proline, 1-naphthylalanine, 2-naphthylalanine, N-methylalanine, 1-methylhistidine, 3-methylhistidine, N-methyl-tryptophan, piperidinic acid, 4-pyridylalanine, sarcosine, tert-butylalanine, or 3-tert-butyltyrosine.
[0054] As used herein, "nucleobase" means an unmodified or modified nucleobase. A nucleobase is a heterocyclic moiety. As used herein, "unmodified nucleobase" is adenine (A), thymine (T), cytosine (C), uracil (U), or guanine (G). As used herein, "modified nucleobase" is a group of atoms other than unmodified A, T, C, U, or G that can pair with at least one other nucleobase. "5-methylcytosine" is a modified nucleobase. A universal base is a modified nucleobase that can pair with any one of the five unmodified nucleobases.
[0055] As used in this article, "nucleobase sequence" refers to the sequence of consecutive nucleobases in a nucleic acid or oligonucleotide, independent of any sugar or nucleoside inter-bond modifications.
[0056] As used herein, the term “nucleobase sequence” or “sequence” in reference to a nucleobase SEQ ID NO refers only to the nucleobase sequence provided in such SEQ ID NO, whereby, unless otherwise stated, compounds, including each sugar moiety and each nucleoside bond, may be independently modified or unmodified, regardless of whether the modifications indicated in the referenced SEQ ID NO are present or not.
[0057] As used herein, "nucleoside" refers to a compound or fragment of a compound that contains a nucleobase and a sugar moiety. The nucleobase and sugar moiety are either independently unmodified or modified.
[0058] As used herein, “oligonucleotide” means a chain of linked nucleosides connected by nucleotide bonds, wherein each nucleoside and nucleotide bond may be independently modified or unmodified. Unless otherwise stated, an oligonucleotide consists of 8–150 linked nucleosides. As used herein, “modified oligonucleotide” means an oligonucleotide in which at least one nucleoside or nucleotide bond is modified. As used herein, “unmodified oligonucleotide” means an oligonucleotide that does not contain any nucleoside or nucleotide modifications.
[0059] As used herein, "PAM interaction domain" or "PI domain" refers to a domain of a guided nucleic acid binder that can bind to or associate with one or more PAM sequences in the non-complementary strand of the target DNA. The PAM interaction domain can confer specificity for one or more PAM sequences.
[0060] As used herein, “PAM sequence,” “PAM,” or “protospacer adjacent motif” refers to 2 to 9 linked nucleosides that interact with a guided nucleic acid binding agent within the non-complementary strand of the target DNA. The PAM sequence may be adjacent to the 5' end of the protospacer sequence. The PAM sequence may also be adjacent to the 3' end of the protospacer sequence.
[0061] As used herein, the term "peptide sequence" or "peptide sequence" or "sequence" in reference to a peptide / polypeptide SEQ ID NO refers to the linear amide bond-linked amino acid sequence provided in such SEQ ID NO, even if the given peptide or polypeptide contains one or more modified side chains linked to another portion.
[0062] As used herein, “pharmaceutically acceptable carrier or diluent” means any substance suitable for administration to a subject. Certain such carriers enable pharmaceutical compositions to be formulated as, for example, tablets, pills, sugar-coated pills, capsules, liquids, gels, syrups, slurries, suspensions, and lozenges for oral ingestion by a subject. In some embodiments, pharmaceutically acceptable carriers or diluents are sterile water, distilled water for injection, sterile saline, sterile buffer solutions, or sterile artificial cerebrospinal fluid.
[0063] As used herein, "pharmaceutically acceptable salt" means a compound that is physiologically and pharmaceutically acceptable. Pharmacologically acceptable salts retain the desired biological activity of the parent compound without conferring undesirable toxicological effects.
[0064] As used herein, "pharmaceutical composition" means a mixture of substances suitable for administration to a subject. For example, a pharmaceutical composition may comprise an active agent and a sterile aqueous solution. In some embodiments, the pharmaceutical composition has shown activity in free uptake assays in certain cell lines.
[0065] As used in this article, "protein recognition region" refers to a portion of the guide that interacts with the guided nucleic acid binder.
[0066] As used herein, the “protospacer sequence” of the target DNA refers to the reverse complementary sequence of the “target sequence” and is found in the non-complementary strand of the target DNA. In some embodiments, the sequence of the guide’s target recognition region has at least 90%, 95%, 98%, or 99% identity with the protospacer sequence of the target DNA.
[0067] As used herein, "RuvC domain" or "RuvC" refers to the cation-dependent endonuclease domain of a protein having an active site containing an aspartic (D) catalytic residue. In some embodiments, the active site of the RuvC domain contains both aspartic (D) and glutamate (E) catalytic residues. In some embodiments, the RuvC domain cleaves the non-complementary strand of the target DNA. In some embodiments, the RuvC domain cleaves both strands of the target DNA. The RuvC domain may be formed from discontinuous amino acids; for example, the RuvC domain of SpCas9 is divided into RuvCI, RuvCII, and RuvCIII moieties. The RuvC domain may have one or more substitutions that inactivate its catalytic activity.
[0068] As used in this article, "standard-length nucleoside interbond" refers to a nucleoside interbond with a structure represented by formula Z1, Z2, or Z3: Wherein, for each nucleoside linker of formula Z1, Z2, or Z3, the following is independent: Each X 1 Independently selected from O and S; X 2 Selected from O, NR 1 CH2 and S; X 3 Selected from O, NR 1 CH2 and S; L does not exist, NR 1 、N(R 1SO2, -N=, O, C1-C6 alkylene or C1-C6 heteroalkylene; Each R 1 Independently selected from H, C1-C6 alkyl and substituted C1-C6 alkyl, or two R on the same atom 1 Together they form = O; and R 2 Selected from -OH, -SH, Cl-C 22 Alkyl, substituted C1-C 22 Alkyl, C2-C 22 Alkenyl, substituted C2-C 22 Alkenyl, cycloalkyl, substituted cycloalkyl, heterocyclic, substituted heterocyclic, heteroaryl, substituted heteroaryl, aryl and substituted aryl; When the group is substituted, it contains one or more elements selected from halogens, -OH, -N(R) 1 )2、-O-C1-C6 alkyl、C1-C 22 Alkyl, C2-C 22 Substituents of alkenyl, cycloalkyl, heterocyclic, heteroaryl, and aryl groups.
[0069] As used in this article, “subject” refers to human or non-human animals, including but not limited to mice, rats, rabbits, dogs, cats, pigs and non-human primates, including but not limited to monkeys and chimpanzees.
[0070] As used herein, “glycan” means an unmodified or modified sugar moiety. As used herein, “unmodified sugar moiety” means, for example, the 2'-OH(H) ribosyl sugar moiety found in RNA (“unmodified RNA sugar moiety”), or the 2'-H(H) deoxyribosyl sugar moiety found in DNA (“unmodified DNA sugar moiety”). An unmodified sugar moiety has one hydrogen atom at each of the 1', 3', and 4' positions, one oxygen atom at the 3' position, and two hydrogen atoms at the 5' position. As used herein, “modified sugar moiety” or “modified sugar” means a modified furanosyl sugar moiety or a sugar substitute.
[0071] As used herein, "sugar substitute" refers to a modified sugar moiety, other than the furanyl moiety, that allows a nucleobase to be linked to an internucleotide bond. Modified nucleosides containing sugar substitutes can be incorporated into one or more positions within an oligonucleotide, and such oligonucleotides can hybridize with complementary oligonucleotides or target nucleic acids.
[0072] As used herein, a “target recognition region” refers to a 12 to 30-linked nucleoside region of a guide that is complementary to a “target sequence” within the complementary strand of the target DNA. The target recognition region may be perfectly complementary to the target sequence within the complementary strand of the target DNA.
[0073] As used in this article, a "target sequence" refers to 12 to 30 linked nucleoside segments of a nucleic acid target that are complementary to the guide's "target recognition region." The target sequence is located within the "complementary strand" of the target DNA.
[0074] As used herein, "therapeutic effective amount" means the amount of a dose or pharmaceutical composition that provides a therapeutic benefit to a subject. For example, a therapeutic effective amount improves the symptoms of a disease.
[0075] As used herein, “treatment” means improving a subject’s disease or condition by administering the agents or compounds described herein. In some embodiments, treating a subject improves symptoms relative to the same symptoms without treatment. In some embodiments, treatment reduces the severity or frequency of symptoms, or delays the onset of symptoms, slows the progression of symptoms, or reduces the severity or frequency of symptoms.
[0076] Some embodiments This disclosure provides the following non-limiting numbered embodiments: Example 1. A polypeptide comprising a region having at least 900, at least 1000, at least 1100, at least 1200, or at least 1300 linked amino acids, wherein the amino acid sequence of the region has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with an isolength portion of any of SEQ ID NO: 12, 13, or 20; wherein the polypeptide comprises a RuvC domain and an HNH domain; optionally, wherein the polypeptide is isolated.
[0077] Example 2. A polypeptide comprising a region wherein the amino acid sequence of the region has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any of SEQ ID NO: 12, 13, or 20; optionally, wherein the polypeptide is isolated.
[0078] Example 3. A polypeptide, wherein the amino acid sequence of the polypeptide has at least 85%, at least 90%, at least 95%, or 100% identity with any of SEQ ID NO: 12, 13, or 20; optionally, wherein the polypeptide is isolated.
[0079] Example 4. The polypeptide of Example 1 or 2, which is composed of regions.
[0080] Example 5. The polypeptide of Example 1 or 2, which additionally contains a heterologous domain.
[0081] Example 6. The polypeptide of Example 5, wherein the heterologous domain is selected from transcription activators, transcription repressors, methyltransferases, demethylases, deaminases, acetyltransferases, or deacetylases.
[0082] Example 7. A polypeptide of any one of Examples 1 to 6, wherein the RuvC domain is composed of a subdomain having an amino acid sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 85, 92, and 94; or SEQ ID NO: 97, 104, and 106; or SEQ ID NO: 109, 116, and 118.
[0083] Example 8. A polypeptide of any one of Examples 1 to 7, wherein the HNH domain has an amino acid sequence that is at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to any one of SEQ ID NO: 93, 105, or 117.
[0084] Example 9. A polypeptide of any one of Examples 1 to 8, wherein the polypeptide comprises REC leaves.
[0085] Example 10. The polypeptide of Example 9, wherein the REC leaf has an amino acid sequence that is at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 91, 103, or 115.
[0086] Example 11. A polypeptide of any one of Examples 1 to 10, wherein the polypeptide comprises NUC leaves.
[0087] Example 12. The polypeptide of Example 11, wherein the NUC leaf has an amino acid sequence that is at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 96, 108, or 120.
[0088] Example 13. A polypeptide of any one of Examples 1 to 12, wherein the polypeptide comprises REC leaves and NUC leaves.
[0089] Example 14. The polypeptide of Example 13, wherein the REC leaf has an amino acid sequence that is at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 91, and the NUC leaf has an amino acid sequence that is at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 96.
[0090] Example 15. The polypeptide of Example 14, wherein the REC leaf has an amino acid sequence that is at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to that of SEQ ID NO: 103, and the NUC leaf has an amino acid sequence that is at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to that of SEQ ID NO: 108.
[0091] Example 16. The polypeptide of Example 14, wherein the REC leaf has an amino acid sequence that is at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to that of SEQ ID NO: 115, and the NUC leaf has an amino acid sequence that is at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to that of SEQ ID NO: 120.
[0092] Example 17. A polypeptide of any one of Examples 1 to 16, wherein the polypeptide has exactly two catalytically active nuclease active sites.
[0093] Example 18. A polypeptide of any one of Examples 1 to 16, wherein the polypeptide has exactly one catalytically active nuclease active site.
[0094] Example 19. A polypeptide of any one of Examples 1 to 17, wherein the polypeptide has zero catalytically active nuclease active sites.
[0095] Example 20. A polypeptide comprising a region having at least 800, at least 900, at least 1000, or at least 1050 linked amino acids, wherein the amino acid sequence has at least 85%, at least 90%, at least 95%, or 100% identity with an isometric portion of any of SEQ ID NO: 241, 242, 245, 247, 314, or 323; optionally, wherein the polypeptide is isolated.
[0096] Example 21. A polypeptide comprising a region having at least 900, at least 1000, at least 1100, or at least 1200 linked amino acids, wherein the amino acid sequence has at least 85%, at least 90%, at least 95%, or 100% identity with the isolength portion of SEQ ID NO: 260; optionally, wherein the polypeptide is isolated.
[0097] Example 22. A polypeptide comprising a region having at least 900, at least 1000, at least 1100, at least 1200, or at least 1300 linked amino acids, wherein the amino acid sequence has at least 85%, at least 90%, at least 95%, or 100% identity with an isometric portion of any one of SEQ ID NO: 266 to 268, 271, 274 to 275, 279, 282 to 284, 288 to 289, 292, 297, or 299 to 302; optionally, wherein the polypeptide is isolated.
[0098] Example 23. A polypeptide comprising a region having an amino acid sequence having at least 85%, at least 90%, at least 95%, or 100% identity with an isometric portion of any one of SEQ ID NO: 241, 242, 245, 247, 260, 266 to 268, 271, 274 to 275, 279, 282 to 284, 288 to 289, 292, 297, 299 to 302, 314, or 323; optionally, wherein the polypeptide is isolated.
[0099] Example 24. A polypeptide wherein the amino acid sequence of the polypeptide has at least 85%, at least 90%, at least 95%, or 100% identity with any one of SEQ ID NO: 241, 242, 245, 247, 260, 266 to 268, 271, 274 to 275, 279, 282 to 284, 288 to 289, 292, 297, 299 to 302, 314, or 323; optionally, wherein the polypeptide is isolated.
[0100] Example 25. A guide comprising a protein recognition element, wherein the protein recognition element binds to a polypeptide of any one of Examples 1 to 24.
[0101] Example 26. The guide of Example 25, wherein the guide is a bidirectional guide composed of two oligonucleotides connected by a double-stranded region.
[0102] Example 27. The guide of Example 25, wherein the guide is a single guide containing an oligonucleotide, and the protein recognition element is the protein recognition region of the oligonucleotide.
[0103] Example 28. The wizard of Example 27, wherein the nucleobase sequence of the protein recognition region of the wizard has at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 129, 130, 131, or 139.
[0104] Example 29. The wizard of Example 27, wherein the nucleobase sequence of the protein recognition region of the wizard has at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 520 to 521, 524, 526, 539, 545 to 547, 550, 553 to 554, 558, 561 to 563, 567, 568, 571, 576, 578 to 581, 593, or 602.
[0105] Example 30. A wizard of any one of Examples 26 to 29, wherein the wizard comprises a modified oligonucleotide.
[0106] Example 31. A wizard of any one of Examples 26 to 30, wherein the wizard comprises at least one modified sugar portion.
[0107] Example 32. The guide to Example 31, wherein the modified sugar moiety is selected from 2'-OMe and 2'-F.
[0108] Example 33. The guide of Example 32, wherein the guide contains a 2'-OMe or 2'-F sugar moiety in the first five nucleotides at the 5' end of the guide or in the 3' end of the last five nucleotides of the guide.
[0109] Example 34. A wizard of any one of Examples 26 to 33, wherein the wizard contains a modified nucleoside inter-bond.
[0110] Example 35. The guide to Example 34, wherein the modified nucleoside inter-bond is a thiophosphate nucleoside inter-bond.
[0111] Example 36. The guide of Example 34 or 35, wherein the guide contains at least one modified internucleotide bond within the first five nucleotides at the 5' end of the guide or within the last five nucleotides at the 3' end of the guide.
[0112] Example 37. A guide of any one of Examples 27 to 36, wherein the guide is composed of oligonucleotides.
[0113] Example 38. A wizard of any one of Examples 25 to 37, wherein the wizard comprises a DNA recognition element that is at least 90%, at least 95%, or 100% complementary to the target sequence.
[0114] Example 39. An editing system comprising a polypeptide of any one of Examples 1 to 24 and a wizard of any one of Examples 25 to 38.
[0115] Example 40. An editing system comprising: a. A nucleic acid encoding a polypeptide of any one of Examples 1 to 24; and b. The wizard for any of Examples 25 to 38.
[0116] Example 41. An editing system comprising: a. A nucleic acid encoding a polypeptide of any one of Examples 1 to 24; and b. A nucleic acid that encodes a wizard for any one of Examples 25 to 29.
[0117] Example 42. The editing system of Example 40 or 41, wherein the nucleic acid encoding the polypeptide is exogenous mRNA.
[0118] Example 43. The editing system of Example 40 or 41, wherein the nucleic acid encoding the polypeptide is DNA.
[0119] Example 44. An editing system of any one of Examples 41 to 43, wherein the nucleic acid encoding the guide is exogenous mRNA.
[0120] Example 45. An editing system of any one of Examples 41 to 43, wherein the nucleic acid of the coding wizard is DNA.
[0121] Example 46. A composition comprising the editing system and LNP of any one of Examples 39 to 45.
[0122] Example 47. A viral vector comprising the editing system of Example 41.
[0123] Example 48. A method for editing target nucleic acid, the method comprising administering to a subject the composition of Example 46 or the viral vector of Example 47.
[0124] Example 49. A method for generating double-strand breaks or nicks in a target nucleic acid, the method comprising administering to a subject the composition of Example 46 or the viral vector of Example 47.
[0125] Example 50. A method for editing target nucleic acids, the method comprising contacting the composition of Example 46 or the viral vector of Example 47 with cells.
[0126] Example 51. A method for generating double-strand breaks or nicks in a target nucleic acid, the method comprising contacting the composition of Example 46 or the viral vector of Example 47 with a cell.
[0127] Example 52. The composition of Example 45 or the viral vector of Example 46 is used for therapy.
[0128] Example 53. A polypeptide comprising a region having at least 800, at least 900, at least 1000, or at least 1050 linked amino acids, wherein the amino acid sequence of the region has at least 85%, at least 90%, at least 95%, or 100% identity with an isometric portion of any of SEQ ID NO: 4 to 23 or 231 to 323, optionally wherein the polypeptide is isolated.
[0129] Example 54. A polypeptide comprising a region having at least 900, at least 1000, or at least 1100 linked amino acids, wherein the amino acid sequence of the region has at least 85%, at least 90%, at least 95%, or 100% identity with the isometric portions of SEQ ID NO: 6 to 23, 250 to 308; optionally, wherein the polypeptide is isolated.
[0130] Example 55. A polypeptide comprising a region having at least 900, at least 1000, at least 1100, at least 1200, or at least 1300 linked amino acids, wherein the amino acid sequence of the region has at least 85%, at least 90%, at least 95%, or 100% identity with an isometric portion of any of SEQ ID NO: 10 to 23 or 266 to 308; optionally, wherein the polypeptide is isolated.
[0131] Example 56. A polypeptide comprising a region wherein the amino acid sequence of the region has at least 85%, at least 90%, at least 95%, or 100% identity with any one of SEQ ID NO: 4 to 23 or 231 to 323; optionally, wherein the polypeptide is isolated.
[0132] Example 57. A polypeptide, wherein the amino acid sequence of the polypeptide has at least 85%, at least 90%, at least 95%, or 100% identity with any one of SEQ ID NO: 4 to 23 or 231 to 323; optionally, wherein the polypeptide is isolated.
[0133] Example 58. The polypeptide of Example 57, wherein the amino acid sequence of the polypeptide has at least 85%, at least 90%, at least 95%, or 100% identity with any one of SEQ ID NO: 12, 13, 20, 241, 242, 245, 247, 260, 266, 267, 268, 271, 274, 275, 279, 282, 283, 284, 288, 289, 292, 297, 299, 300, 301, 302, 314, or 323.
[0134] Example 59. A polypeptide of any one of Examples 53 to 56, which is composed of regions.
[0135] Example 60. A polypeptide of any one of Examples 53 to 56, comprising a heterologous domain.
[0136] Example 61. The polypeptide of Example 60, wherein the heterologous domain is selected from transcription activators, transcription repressors, methyltransferases, demethylases, deaminases, acetyltransferases, or deacetylases.
[0137] Example 62. A polypeptide of any one of Examples 53 to 56 or 59 to 61, wherein the amino acid sequence of the region has at least 85%, at least 90%, at least 95%, or 100% identity with the isometric portion of any one of SEQ ID NO: 13, 241, 267, 275, 284, 289, or 314.
[0138] Example 63. A polypeptide of any one of Examples 53 to 62, wherein the polypeptide comprises a non-PI region and a PI domain, wherein the non-PI region of the polypeptide has an amino acid sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3186, 3188, 3190, 3192, 3194, 3196, or 3198.
[0139] Example 64. A polypeptide of any one of Examples 53 to 62, wherein the polypeptide comprises a non-PI region and a PI domain, wherein the non-PI region of the polypeptide has an amino acid sequence that is at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 3186, and the PI domain has an amino acid sequence that is at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 3187.
[0140] Example 65. A polypeptide of any one of Examples 53 to 62, wherein the polypeptide comprises a non-PI region and a PI domain, wherein the non-PI region of the polypeptide has an amino acid sequence that is at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 3188, and the PI domain has an amino acid sequence that is at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 3189.
[0141] Example 66. A polypeptide of any one of Examples 53 to 62, wherein the polypeptide comprises a non-PI region and a PI domain, wherein the non-PI region of the polypeptide has an amino acid sequence that is at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 3190, and the PI domain has an amino acid sequence that is at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 3191.
[0142] Example 67. A polypeptide of any one of Examples 53 to 62, wherein the polypeptide comprises a non-PI region and a PI domain, wherein the non-PI region of the polypeptide has an amino acid sequence that is at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 3192, and the PI domain has an amino acid sequence that is at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 3193.
[0143] Example 68. A polypeptide of any one of Examples 53 to 62, wherein the polypeptide comprises a non-PI region and a PI domain, wherein the non-PI region of the polypeptide has an amino acid sequence that is at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 3194, and the PI domain has an amino acid sequence that is at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 3195.
[0144] Example 69. A polypeptide of any one of Examples 53 to 62, wherein the polypeptide comprises a non-PI region and a PI domain, wherein the non-PI region of the polypeptide has an amino acid sequence that is at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 3196, and the PI domain has an amino acid sequence that is at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 3197.
[0145] Example 70. A polypeptide of any one of Examples 53 to 62, wherein the polypeptide comprises a non-PI region and a PI domain, wherein the non-PI region of the polypeptide has an amino acid sequence that is at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 3198, and the PI domain has an amino acid sequence that is at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 3199.
[0146] Example 71. A polypeptide of any one of Examples 53 to 70, wherein the polypeptide has exactly two catalytically active nuclease active sites.
[0147] Example 72. A polypeptide of any one of Examples 53 to 70, wherein the polypeptide has exactly one catalytically active nuclease active site.
[0148] Example 73. A polypeptide of any one of Examples 53 to 70, wherein the polypeptide has zero catalytically active nuclease active sites.
[0149] Example 74. A polypeptide of any one of Examples 53 to 73, comprising a nuclear localization signal.
[0150] Example 75. The polypeptide of Example 74, wherein the nuclear localization signal is located at the N-terminus of the polypeptide.
[0151] Example 76. The polypeptide of Example 74, wherein the nuclear localization signal is located at the C-terminus of the polypeptide.
[0152] Example 77. The polypeptide of Example 74, wherein the polypeptide contains two nuclear localization signals, one at the N-terminus and one at the C-terminus.
[0153] Example 78. A nucleic acid encoding a polypeptide according to any one of Examples 53 to 77.
[0154] Example 79. The nucleic acid of Example 78, wherein the nucleic acid is DNA.
[0155] Example 80. The nucleic acid of Example 78, wherein the nucleic acid is mRNA, the mRNA containing a coding region for a polypeptide of any one of Examples 53 to 77.
[0156] Example 81. The nucleic acid of Example 78, wherein the mRNA is exogenous mRNA containing a coding region for a polypeptide of any one of Examples 53 to 73.
[0157] Example 82. The mRNA of Example 80 or 81, which contains a 5' cap structure.
[0158] Example 83. The mRNA of any one of Examples 80 to 82, which contains a 3'-poly(A) tail.
[0159] Example 84. The mRNA of any one of Examples 80 to 83, which contains a 5'-UTR.
[0160] Example 85. The mRNA of any one of Examples 80 to 84, which contains a 3'-UTR.
[0161] Example 86. The mRNA of Example 80 or 81, which contains a 5'-cap structure and a 3'-poly(A) tail.
[0162] Example 87. mRNA of any one of Examples 80 to 86, wherein the coding region is codon-optimized for expression in eukaryotic cells.
[0163] Example 88. mRNA of Example 87, wherein the coding region has been codon-optimized for expression in mammalian cells.
[0164] Example 89. mRNA of Example 88, wherein the mammalian cell is a human cell.
[0165] Example 90. The mRNA of any one of Examples 80 to 89, wherein each nucleobase of the mRNA is selected from adenine, guanine, cytosine, uracil, thymine, and N1-methylpseuuridine.
[0166] Example 91. The mRNA of any one of Examples 80 to 89, wherein each nucleobase of the mRNA is selected from adenine, guanine, cytosine, and N1-methylpseuuridine.
[0167] Example 92. The mRNA of any one of Examples 80 to 89, wherein each nucleobase of the mRNA is selected from adenine, guanine, cytosine, and uracil.
[0168] Example 93. A guide, wherein the guide includes a protein recognition region, wherein the protein recognition region interacts with a polypeptide of any one of Examples 53 to 57.
[0169] Example 94. The wizard of Example 93, wherein the wizard is a dual wizard.
[0170] Example 95. The wizard of Example 93, wherein the wizard is a single wizard.
[0171] Example 96. A guide composed of oligonucleotides according to the following formula: T1-DXS 1a -B1-S 1b -H1-S 1b '-B2-S 1a '-L1-S2-H2-S2'-L2-S3-H3-S3'-(L3-S4-H4-S4') n -(L4-S5-H5-S5') m -T2; in: D consists of 17 to 25 linked nucleosides, wherein the nucleobase sequence of D is complementary to the nucleobase sequence of the target DNA; X is absent, or consists of one or two linked nucleosides that are not complementary to the sequence of the target DNA; Each of T1 and T2 either lacks or is independently composed of 1 to 30 linked nucleosides; S 1a and S 1a Each is composed of 6 to 10 linked nucleosides, of which S 1a The nucleobase sequence and S 1a The nucleobase sequences of ' are 100% complementary; S 1b and S 1b Each is composed of 2 to 14 linked nucleosides, of which S 1b The nucleobase sequence and S 1b The nucleobase sequences of ' are 100% complementary; S2 and S2' are each composed of 2 to 6 linked nucleosides, wherein the nucleobase sequence of S2 is 100% complementary to the nucleobase sequence of S2'; S3 and S3' are each composed of 2 to 12 linked nucleosides, wherein the nucleobase sequence of S3 is 100% complementary to the nucleobase sequence of S3'; S4 and S4' are each composed of 4 to 14 linked nucleosides, wherein the nucleobase sequence of S4 is 100% complementary to the nucleobase sequence of S4'. S5 and S5' are each composed of 4 to 11 linked nucleosides, wherein the nucleobase sequence of S5 is 100% complementary to the nucleobase sequence of S5'. n is 0 or 1; m is 0 or 1; B1 is absent, or consists of 1 to 3 linked nucleosides; B2 is absent or consists of 1 to 4 linked nucleosides; B1 is not equal to B2 unless neither of them exists; Each of H1, H2, H3, H4, and H5 is independently composed of 3 to 6 linked nucleosides; L1 consists of 2 to 3 linked nucleosides; Each of L2, L3, and L4 is independently composed of 0 to 10 linked nucleosides.
[0172] Example 97. The wizard of Example 96, wherein the wizard includes a protein recognition region that binds to a polypeptide of any one of Examples 53 to 77.
[0173] Example 98. A wizard for any of Examples 96 to 97, wherein T1 does not exist.
[0174] Example 99. A wizard for any of Examples 96 to 98, wherein T2 does not exist.
[0175] Example 100. A guide to any one of Examples 96 to 99, wherein T2 consists of 2 to 6 linked nucleosides containing uracil nucleobases.
[0176] Example 101. A wizard for any of Examples 96 to 100, where X does not exist.
[0177] Example 102. A wizard for any of Examples 96 to 101, where n is 1 and m is 0.
[0178] Example 103. A wizard for any of Examples 96 to 102, where both n and m are 1.
[0179] Example 104. A guide to any of Examples 96 to 103, wherein the number of nucleosides in B2 is greater than the number of nucleosides in B1.
[0180] Example 105. A guide to any one of Examples 96 to 104, wherein each L1 nucleotide is adenosine.
[0181] Example 106. A wizard of any one of Examples 96 to 105, wherein the wizard binds to a polypeptide of any one of Examples 53 to 77.
[0182] Example 107. The wizard for Example 96 or 97, wherein T1 and X do not exist; T2 is absent or consists of 2 to 6 linked uridines; S 1a and S 1a Each is composed of 7 linked nucleosides; S 1b and S 1b Each is composed of 3 to 9 linked nucleosides; S2 and S2' are each composed of three linked nucleosides; S3 and S3' are each composed of four linked nucleosides; S4 and S4' are each composed of 5 linked nucleosides; When present, S5 and S5' each consist of 7 to 9 linked nucleosides; n is 1; m is 0 or 1; B1 is composed of one nucleoside; B2 is composed of four linked nucleosides; When present, each of H1, H2, H3, and H4 consists of four linked nucleosides; L1 consists of two linked adenosine nucleotides; L2 consists of 5 linked nucleosides; L3 is composed of 0 nucleosides; and When present, L4 consists of 4 linked nucleosides.
[0183] Example 108. The wizard of Example 107, wherein S 1b and S 1b Each is composed of 7 linked nucleosides.
[0184] Example 109. The guide to Example 107 or 108, wherein S5 and S5' are each composed of 7 linked nucleosides.
[0185] Example 110. A wizard for any one of Examples 107 to 109, wherein the nucleobase sequence of B2 is AAAG.
[0186] Example 111. A wizard for any one of Examples 107 to 108, wherein the nucleobase sequence of L2 is AAACU.
[0187] Example 112. A wizard for any one of Examples 107 to 111, wherein the nucleobase sequence of L3 is UUUUAA.
[0188] Example 113. A wizard for any one of Examples 107 to 112, wherein the nucleobase sequence of H2 is GUCA.
[0189] Example 114. A wizard for any one of Examples 107 to 113, where n is 1 and m is 0.
[0190] Example 115. A wizard for any of Examples 107 to 114, where both n and m are 1.
[0191] Example 116. The wizard for Example 107, wherein S 1b and S 1b Each is composed of 7 linked nucleosides; n is 1 and m is 0; the nucleobase sequence of B2 is AAAG; the nucleobase sequence of L2 is AAACU; the nucleobase sequence of L3 is UUUUAA, and the nucleobase sequence of H2 is GUCA.
[0192] Example 117. A guide of any one of Examples 107 to 116, wherein the guide comprises a region having at least 40, at least 50, at least 60, or at least 70 nucleosides, wherein the nucleobase sequence of the region has at least 90%, 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with an isoplethora of any one of SEQ ID NO: 131, 961 to 964, 1016 to 1020, 1052 to 1070, 1313 to 1314, or 2825.
[0193] Example 118. The guide to Example 117, wherein the nucleobase sequence of the region has 100% identity with the isolength region of any of SEQ ID NO: 131, 961 to 964, 1016 to 1020, 1052 to 1070, 1313 to 1314 or 2825.
[0194] Example 119. A guide of any one of Examples 107 to 118, wherein D consists of 20 to 23 nucleosides having a nucleobase sequence that is 100% complementary to the nucleobase sequence of the target DNA.
[0195] Example 120. The wizard for Example 96 or 97, wherein T1 and X do not exist; T2 is absent or consists of 6 linked uridines; S 1a and S 1a Each is composed of 9 linked nucleosides; S 1b and S 1b Each is composed of 2 to 8 linked nucleosides; S2 and S2' are each composed of 6 linked nucleosides; S3 and S3' are each composed of 7 linked nucleosides; n and m are 0; B1 is composed of 3 nucleosides; B2 is composed of four linked nucleosides; H1 consists of 4 linked nucleosides; H2 is composed of 5 linked nucleosides; H3 is composed of 4 linked nucleosides; L1 consists of two linked adenosine nucleotides; and L2 consists of 0 linked nucleosides.
[0196] Example 121. The wizard for Example 120, wherein S 1b and S 1b Each is composed of 4 linked nucleosides.
[0197] Example 122. A guide to any of Examples 120 to 121, wherein L2 consists of 0 linked nucleosides and T2 consists of 6 linked uridines.
[0198] Example 123. A wizard for any one of Examples 120 to 122, wherein the nucleobase sequence of B2 is UAAC.
[0199] Example 124. A wizard for any one of Examples 120 to 123, wherein the nucleobase sequence of L2 is GAACUC.
[0200] Example 125. A wizard for any one of Examples 120 to 124, wherein the nucleobase sequence of H2 is UUUAU.
[0201] Example 126. The wizard for Example 120, wherein S 1b and S 1b Each is composed of 4 linked nucleosides; n and m are 0; the nucleobase sequence of B2 is UACC; the nucleobase sequence of L2 is GAACUC; and the nucleobase sequence of H2 is UUUAU.
[0202] Example 127. A wizard of any one of Examples 120 to 126, wherein the wizard comprises a region having a nucleobase sequence having at least 90%, 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any one of SEQ ID NO: 520, 965 to 968, 1071 to 1094, or 1315 to 1318, wherein the region comprises at least 40, at least 50, at least 60, or at least 70 nucleosides.
[0203] Example 128. The guide to Example 127, wherein the nucleobase sequence of the region has 100% identity with the iso-length region of any one of SEQ ID NO: 520, 965 to 968, 1071 to 1094 or 1315 to 1318.
[0204] Example 129. A guide of any one of Examples 120 to 127, wherein D consists of 20 to 23 nucleosides having a nucleobase sequence that is 100% complementary to the nucleobase sequence of the target DNA.
[0205] Example 130. The wizard for Example 96 or 97, wherein T1 and X do not exist; T2 is absent or consists of 1 to 7 linked nucleosides; S 1a and S 1a Each is composed of 7 linked nucleosides; S 1b and S 1b Each nucleotide is composed of 4 to 12 linked nucleosides; S2 and S2' are each composed of 5 linked nucleosides; S3 and S3' are each composed of 6 linked nucleosides; S4 and S4' are each composed of 6 linked nucleotides; When present, S5 and S5' are each composed of 7 linked nucleosides; n is 1; m is 0 or 1; B1 is composed of one nucleoside; B2 is composed of three linked nucleosides; Each of H1, H2 and H3 consists of 4 linked nucleosides; When present, H4 and H5 consist of 3 linked nucleosides; L1 consists of three linked adenosine nucleotides; L2 and L3 consist of 0 linked nucleosides; When present, L4 consists of 1 to 7 linked nucleosides.
[0206] Example 131. The wizard of Example 130, wherein S 1b and S 1b Each is composed of 6 linked nucleosides.
[0207] Example 132. A wizard for any one of Examples 130 to 131, wherein the nucleobase sequence of B2 is GAG.
[0208] Example 133. A wizard for any one of Examples 130 to 132, wherein the nucleobase sequence of H2 is AUCC.
[0209] Example 134. The wizard for Example 130, wherein S 1b and S 1b Each is composed of 6 linked nucleosides; n is 1 and m is 0; the nucleobase sequence of B2 is GAG; and the nucleobase sequence of H2 is AUCC.
[0210] Example 135. A wizard of any one of Examples 130 to 134, wherein the wizard comprises a region having a nucleobase sequence having at least 90%, 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any one of SEQ ID NO: 546, 973 to 976, 1021 to 1023, 1095 to 1126, or 1319 to 1342, wherein the region comprises at least 40, at least 50, at least 60, or at least 70 nucleosides.
[0211] Example 136. The guide to Example 135, wherein the nucleobase sequence of the region has 100% identity with the isolength region of any of SEQ ID NO: 546, 973 to 976, 1021 to 1023, 1095 to 1126 or 1319 to 1342.
[0212] Example 137. A guide of any one of Examples 130 to 136, wherein D consists of 20 to 24 nucleosides having a nucleobase sequence that is 100% complementary to the nucleobase sequence of the target DNA.
[0213] Example 138. The wizard for Example 96 or 97, wherein T1 and X do not exist; T2 is absent or consists of 1 to 6 linked nucleosides; S 1a and S 1a Each is composed of 7 linked nucleosides; S 1b and S 1b Each nucleotide is composed of 4 to 12 linked nucleosides; S2 and S2' are each composed of four linked nucleosides; S3 and S3' are each composed of four linked nucleosides; S4 and S4' are each composed of 5 linked nucleosides; When present, S5 and S5' are each composed of 7 linked nucleosides; n is 1; m is 0 or 1; B1 is composed of one nucleoside; B2 is composed of 3 nucleosides; H1 consists of 4 linked nucleosides; H2 is composed of three linked nucleosides; H3 consists of 6 linked nucleosides; H4 consists of four linked nucleosides; When present, H5 consists of 3 linked nucleosides; L1 consists of two linked adenosine nucleotides; L2 consists of two linked nucleosides; L3 consists of 0 linked nucleosides; When present, L4 consists of 7 linked nucleosides.
[0214] Example 139. The wizard of Example 138, wherein S 1b and S 1b It consists of 10 linked nucleosides each.
[0215] Example 140. A wizard for any of Examples 138 to 139, where n is 1 and m is 0.
[0216] Example 141. A wizard for any one of Examples 138 to 140, wherein the nucleobase sequence of B2 is GAG.
[0217] Example 142. A wizard for any of Examples 138 to 141, wherein the nucleobase sequence of L2 is AA.
[0218] Example 143. A wizard for any of Examples 138 to 142, wherein the nucleobase sequence of H2 is UAA.
[0219] Example 144. The wizard of Example 138, wherein S 1b and S 1b Each is composed of 10 linked nucleosides; n is 1 and m is 0; the nucleobase sequence of B2 is GAG; the nucleobase sequence of L2 is AA; and the nucleobase sequence of H2 is UAA.
[0220] Example 145. A guide of any one of Examples 138 to 144, wherein the guide comprises a region having a nucleobase sequence having at least 90%, 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any one of SEQ ID NO: 562, 985 to 989, 1027 to 1029, 1153 to 1182, or 1362 to 1390, wherein the region comprises at least 40, at least 50, at least 60, or at least 70 nucleosides.
[0221] Example 146. The guide to Example 145, wherein the nucleobase sequence of the region has 100% identity with the isolength region of any of SEQ ID NO: 562, 985 to 989, 1027 to 1029, 1153 to 1182 or 1362 to 1390.
[0222] Example 147. A guide of any one of Examples 138 to 146, wherein D consists of 20 to 24 nucleosides having a nucleobase sequence that is 100% complementary to the nucleobase sequence of the target DNA.
[0223] Example 148. The wizard for Example 96 or 97, wherein T1 and X do not exist; T2 is absent or consists of 2 to 5 linked nucleosides; S 1a and S 1a Each is composed of 9 linked nucleosides; S 1b and S 1b Each is composed of 2 to 10 linked nucleosides; S2 and S2' are each composed of four linked nucleosides; S3 and S3' are each composed of 9 linked nucleosides; n and m are 0; B1 is composed of one nucleoside; B1 is composed of 3 nucleosides; H1 consists of 4 linked nucleosides; H2 is composed of three linked nucleosides; H3 is composed of 4 linked nucleosides; L1 consists of two linked adenosine nucleotides; L2 consists of 7 linked nucleosides.
[0224] Example 149. The wizard of Example 148, wherein S 1b and S 1b Each is composed of 6 linked nucleosides.
[0225] Example 150. A wizard for any one of Examples 148 to 149, wherein the nucleobase sequence of B2 is CUA.
[0226] Example 151. A wizard for any one of Examples 148 to 150, wherein the nucleobase sequence of L2 is GUGUUUA.
[0227] Example 152. A wizard for any of Examples 148 to 151, wherein the nucleobase sequence of H2 is AAA.
[0228] Example 153. The wizard of Example 148, wherein S 1b and S 1b Each is composed of 10 linked nucleosides; n and m are 0; the nucleobase sequence of B2 is CUA; the nucleobase sequence of L2 is GUGUUUA; and the nucleobase sequence of H2 is AAA.
[0229] Example 154. A guide of any one of Examples 148 to 153, wherein the guide comprises a region having a nucleobase sequence having at least 90%, 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any one of SEQ ID NO: 593, 1006 to 1010, 1272 to 1292, 1394 to 1408, 2827 to 2833, wherein the region comprises at least 40, at least 50, at least 60, or at least 70 nucleosides.
[0230] Example 155. The guide to Example 160, wherein the nucleobase sequence of the region has 100% identity with the isolength region of any one of SEQ ID NO: 593, 1006 to 1010, 1272 to 1292, 1394 to 1408, 2827 to 2833.
[0231] Example 156. A guide of any one of Examples 148 to 155, wherein D consists of 20 to 24 nucleosides having a nucleobase sequence that is 100% complementary to the nucleobase sequence of the target DNA.
[0232] Example 157. The wizard for Example 96 or 97, wherein T1 and X do not exist; T2 either does not exist or is composed of 3 linked nucleosides; S 1a and S 1b Together, and S 1a 'and S 1b Together, they are composed of 12 to 20 linked nucleosides; S2 and S2' are each composed of three linked nucleosides; S3 and S3' are each composed of three linked nucleosides; S3 and S3' are each composed of 10 to 13 linked nucleosides; n is 1 and m is 0; B1 consists of 0 nucleosides; B2 consists of 0 nucleosides; H1, H2, H3, and H4 are composed of four linked nucleosides; L1 consists of two linked adenosine nucleotides; L2 consists of two linked nucleosides; and L3 consists of 6 linked nucleosides.
[0233] Example 158. The wizard for Example 157, wherein S 1a and S 1b Together, and S 1a 'and S 1b Together, they are composed of 15 linked nucleosides.
[0234] Example 159. A guide to any of Examples 157 to 158, wherein L3 consists of 3 linked nucleosides.
[0235] Example 160. A wizard of any one of Examples 157 to 159, wherein S3 and S3' are each composed of 13 linked nucleosides.
[0236] Example 161. A wizard for any one of Examples 157 to 159, wherein the nucleobase sequence of L2 is GU.
[0237] Example 162. A wizard for any of Examples 160 to 161, wherein the nucleobase sequence of H2 is GAAA.
[0238] Example 163. The wizard for Example 157, wherein S 1a and S 1b Together, and S 1a 'and S 1b Together, each is composed of 15 linked nucleosides; n is 1 and m is 0; the nucleobase sequence of L2 is GU; and the nucleobase sequence of H2 is GAAA.
[0239] Example 164. A guide of any of Examples 157 to 163, wherein the guide comprises a region having a nucleobase sequence having at least 90%, 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any of SEQ ID NO: 568, 990 to 993, 1030 to 1034, 1183 to 1206, or 1391 to 1393, wherein the region comprises at least 40, at least 50, at least 60, or at least 70 nucleosides.
[0240] Example 165. The guide to Example 164, wherein the nucleobase sequence of the region is 100% identical to the isolength region of any one of SEQ ID NO: 568, 990 to 993, 1030 to 1034, 1183 to 1206 or 1391 to 1393.
[0241] Example 166. A guide of any one of Examples 158 to 165, wherein D consists of 20 to 24 nucleosides having a nucleobase sequence that is 100% complementary to the nucleobase sequence of the target DNA.
[0242] Example 167. A guide composed of oligonucleotides according to the following formula: T1-DXS 1a -B1-S 1b -H1-S 1b '-B2-S 1a '-L1-S2-H2-S2'-L2-S3-H3-S3'-(L3-S4-H4-S4') n -T2; in: D consists of 17 to 25 linked nucleosides, wherein the nucleobase sequence of D is complementary to the nucleobase sequence of the target DNA; X is absent, or consists of one or two linked nucleosides that are not complementary to the nucleobase sequence of the target DNA; Each of T1 and T2 either lacks or is independently composed of 1 to 30 linked nucleosides; S 1a and S 1a Each is composed of 7 linked nucleosides, of which S 1a The nucleobase sequence and S 1a The nucleobase sequences of ' are 100% complementary; S 1b and S 1b Each is composed of 4 to 12 linked nucleosides, of which S 1b The nucleobase sequence and S 1b The nucleobase sequences of ' are 100% complementary; S2 and S2' are each composed of 6 linked nucleosides, and the nucleobase sequence of S2 is 100% complementary to the nucleobase sequence of S2'. S3 and S3' are each composed of 5 or 6 linked nucleosides, wherein the nucleobase sequence of S3 is 100% complementary to the nucleobase sequence of S3'. When present, S4 and S4' each consist of 6 linked nucleosides, wherein the nucleobase sequence of S4 is 100% complementary to the nucleobase sequence of S4'. n is 0 or 1; B1 is composed of one nucleoside; B2 is composed of three linked nucleosides; Each of H1, H2, H3, and H4 is independently composed of 3 to 4 linked nucleosides; L1 consists of 17 linked nucleosides; L2 consists of 0 linked nucleosides; and L3 consists of 8 linked nucleosides.
[0243] Example 168. The wizard for Example 167, wherein S 1b and S 1b Each is composed of 8 linked nucleosides.
[0244] Example 169. A wizard for any one of Examples 167 to 168, wherein the nucleobase sequence of B2 is GAA.
[0245] Example 170. A guide to any of Examples 167 to 169, wherein the nucleobase sequence of L2 is AAAAAUUUAUUCAAAAC (SEQ ID NO: 78).
[0246] Example 171. A wizard for any one of Examples 167 to 170, wherein the nucleobase sequence of H2 is GAAA.
[0247] Example 172. The wizard of Example 171, wherein S 1b and S 1b Each is composed of 8 linked nucleosides; n and m are 0; the nucleobase sequence of B2 is GAA; the nucleobase sequence of L2 is AAAAAUUUAUUCAAAAC (SEQ ID NO: 78); and the nucleobase sequence of H2 is GAAA.
[0248] Example 173. A wizard of any one of Examples 167 to 172, wherein the wizard comprises a region having a nucleobase sequence having at least 90%, 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any one of SEQ ID NO: 554, 981 to 984, 1024 to 1026, 1127 to 1152, or 1343 to 1361, wherein the region comprises at least 40, at least 50, at least 60, or at least 70 nucleosides.
[0249] Example 174. The guide to Example 173, wherein the nucleobase sequence of the region has 100% identity with the isolength region of any of SEQ ID NO: 554, 981 to 984, 1024 to 1026, 1127 to 1152 or 1343 to 1361.
[0250] Example 175. A guide of any one of Examples 167 to 174, wherein D consists of 20 to 24 nucleosides having a nucleobase sequence that is 100% complementary to the nucleobase sequence of the target DNA.
[0251] Example 176. A guide of any one of Examples 93 to 95, wherein the guide comprises a region having a nucleobase sequence having at least 90%, 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any one of SEQ ID NO: 121 to 141, 510 to 602, 961 to 1430, 1436 to 2833, wherein the region comprises at least 40, at least 50, at least 60, at least 70, at least 80, or at least 90 nucleosides.
[0252] Example 177. A guide of any one of Examples 93 to 95, wherein the guide comprises a protein recognition region having a nucleobase sequence having at least 90%, 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any one of SEQ ID NO: 121 to 141, 510 to 602, 961 to 1430, 1436 to 2833.
[0253] Example 178. A wizard of any one of Examples 93 to 177, wherein the wizard comprises a modified oligonucleotide.
[0254] Example 179. A guide of any one of Examples 93 to 178, wherein the guide comprises at least one modified sugar portion.
[0255] Example 180. The guide to Example 179, wherein the modified sugar moiety is selected from 2'-OMe and 2'-F.
[0256] Example 181. The guide of Example 179, wherein the modified oligonucleotide contains a 2'-OMe or 2'-F sugar moiety in the first five nucleotides at the 5' end of the guide or in the last five nucleotides at the 3' end of the guide.
[0257] Example 182. The guide of Examples 179 to 181, wherein the guide contains 2 to 3 2'-OMe sugar portions at the 5' end of the guide and 2 to 3 2'-OMe sugar portions at the 3' end of the guide.
[0258] Example 183. The wizard of Example 182, wherein each non-2'-OMe sugar portion of the wizard contains an unmodified RNA sugar portion.
[0259] Example 184. A wizard of any one of Examples 179 to 183, wherein the wizard contains a modified nucleoside inter-bond.
[0260] Example 185. The guide to Example 184, wherein the modified nucleoside inter-bond is a thiophosphate nucleoside inter-bond.
[0261] Example 186. The guide of Example 184 or 185, wherein the guide contains at least one modified nucleotide bond within the first five nucleotides at the 5' end of the guide or within the last five nucleotides at the 3' end of the guide.
[0262] Example 187. The guide of Example 186, wherein the guide contains 2 to 3 thiophosphate nucleoside inter-bonds at the 5' end of the guide and 2 to 3 thiophosphate nucleoside inter-bonds at the 3' end of the guide.
[0263] Example 188. The wizard of Example 187, wherein each remaining nucleoside inter-bond is an unmodified phosphodiester nucleoside inter-bond.
[0264] Example 189. A wizard of any one of Examples 184 to 188, wherein each nucleoside inter-bond of the wizard is a standard-length nucleoside inter-bond.
[0265] Example 190. A wizard of any one of Examples 93 to 177, wherein the wizard contains unmodified oligonucleotides.
[0266] Example 191. A wizard of any one of Examples 93 to 190, wherein the wizard includes a target recognition region that is at least 90%, at least 95%, or 100% complementary to the target DNA sequence.
[0267] Example 192. The wizard of Example 191, wherein the target recognition region contains 17 to 25, 19 to 24, 20 to 24, 20 to 23, 20 to 22, 21 to 22, 21 to 23, 21 to 24, 20, 21, 22, 23 or 24 linked nucleosides.
[0268] Example 193. An editing system comprising a polypeptide of any one of Examples 53 to 77 and a wizard of any one of Examples 93 to 192.
[0269] Example 194. An editing system comprising: Nucleic acid of any one of Examples 78 to 92; and The wizard for any one of Examples 93 to 192.
[0270] Example 195. An editing system comprising: Nucleic acid of any one of Examples 78 to 92; and A nucleic acid that encodes a guide for any one of Examples 93 to 177 or 190.
[0271] Example 196. The editing system of Example 194 or 195, wherein the nucleic acid encoding the polypeptide is exogenous mRNA.
[0272] Example 197. The editing system of Example 194 or 195, wherein the nucleic acid encoding the polypeptide is DNA.
[0273] Example 198. An editing system of any one of Examples 195 to 197, wherein the nucleic acid of the coding wizard is DNA.
[0274] Example 199. A composition comprising the editing system and LNP of any one of Examples 193 to 198.
[0275] Example 200. A viral vector comprising the editing system of Example 195.
[0276] Example 201. A method for editing target nucleic acid, the method comprising administering to a subject the composition of Example 199 or the viral vector of Example 200.
[0277] Example 202. A method for editing target nucleic acid, the method comprising contacting a cell with the composition of Example 199 or the viral vector of Example 200.
[0278] Example 203. A method for generating double-strand breaks or nicks in a target nucleic acid, the method comprising contacting the composition of Example 199 or the viral vector of Example 200 with a cell.
[0279] Example 204. The composition of Example 199 or the viral vector of Example 200 is used for therapy.
[0280] The embodiments provided herein are directed to guided nucleic acid binding agents, editing systems comprising the guided nucleic acid binding agents, and methods of using the same. In some embodiments, the guided nucleic acid binding agent comprises a Cas protein. In some embodiments, the guided nucleic acid binding agent comprises a Cas enzyme. In some embodiments, the guided nucleic acid binding agent comprises a Cas protein containing a RuvC domain and an HNH domain. In some embodiments, the guided nucleic acid binding agent comprises a type II Cas protein. In some embodiments, the guided nucleic acid binding agent comprises a type II Cas endonuclease. In some embodiments, the guided nucleic acid binding agent comprises a Cas protein containing a RuvC domain and an HNH domain. In some embodiments, the guided nucleic acid binding agent comprises a modified Cas protein (“Cas cleavage enzyme”) having an inactivating mutation in one of the nuclease active sites. In some embodiments, the guided nucleic acid binding agent comprises a modified Cas protein (“inactivated Cas” or “dCas”) having an inactivating mutation in both nuclease active sites. In some embodiments, the protein is isolated. In some embodiments, the protein is engineered.
[0281] In some embodiments, the editing system comprises a guided nucleic acid binder and a guide. In some embodiments, the guide consists of a single oligonucleotide and is a single guide, comprising a target recognition region and a protein recognition region. In some embodiments, the guide is a dual guide, wherein a first oligonucleotide comprises a first portion of the target recognition region and a first portion of the protein recognition region, and a second oligonucleotide comprises a second portion of the protein recognition region. In some such embodiments, the first oligonucleotide is “crRNA” and the second oligonucleotide is “tracrRNA”. The target recognition region is a region of the oligonucleotide (single guide or crRNA) having a nucleobase sequence complementary to an equal-length portion of the target sequence within the complementary strand of the target DNA.
[0282] I. Some nucleic acid binding agents A. Certain polypeptides In some embodiments, the guided nucleic acid binder is an isolated or engineered polypeptide. In some embodiments, the polypeptide is composed of a Cas protein or a fusion protein comprising a Cas protein. In some embodiments, the Cas protein is a Cas enzyme. In alternative embodiments, the Cas protein is an inactivated Cas protein. In some embodiments, the polypeptide constituting the Cas protein has one or more domains. In some embodiments, the domain of the polypeptide is a RuvC domain, which is composed of one or more subdomains (such as RuvCI, RuvCII, and / or RuvCIII subdomains), as described herein. In some embodiments, the domain of the polypeptide is an HNH domain, as described herein. In some embodiments, the domain of the polypeptide is a PAM interaction (PI) domain, as described herein. In some embodiments, the polypeptide comprises a RuvC domain and an HNH domain. In some embodiments, the polypeptide comprises a RuvC domain, an HNH domain, and a PI domain. In some embodiments, the polypeptide comprises a RuvC domain, an HNH domain, a PI domain, and a nuclear localization sequence (NLS). In some embodiments, the RuvC domain is divided into two or three discontinuous subdomains, referred to as RuvCI, RuvCII, and RuvCIII. In some embodiments, the HNH domain is part of the NUC leaf. In some embodiments, the polypeptide further comprises a REC leaf. In some embodiments, the REC leaf is composed of subdomains corresponding to the three REC domains RECI, RECII, and RECIII. In some embodiments, RECI is divided into two discontinuous subdomains, RECIa and RECIb. In some embodiments, the polypeptide comprises a bridging helical domain. Representative sequences corresponding to these domains are described in the table below.
[0283] Table 1 Predicted domain and subdomain structures of the novel ION-Cas protein *The start and end sites correspond to SEQ ID NO: 12 of IONCas008, SEQ ID NO: 13 of IONCas009, and SEQ ID NO: 20 of IONCas016.
[0284] In some embodiments, the polypeptide includes a region having at least 20, 25, or 30 amino acids having a sequence having at least 70% sequence identity with the equal-length portion of SEQ ID NO: 85. In some embodiments, the polypeptide includes a region having at least 20, 25, or 30 amino acids having a sequence having at least 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the equal-length portion of SEQ ID NO: 85.
[0285] In some embodiments, the polypeptide includes a region having at least 30, 40, 50, or 55 amino acids having a sequence identity of at least 70% with the equal-length portion of SEQ ID NO: 92. In some embodiments, the polypeptide includes a region having at least 30, 40, 50, or 55 amino acids having a sequence identity of at least 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% with the equal-length portion of SEQ ID NO: 92.
[0286] In some embodiments, the polypeptide includes a region having at least 70, 80, 90, or 100 amino acids having a sequence identity of at least 70% with the equal-length portion of SEQ ID NO: 94. In some embodiments, the polypeptide includes a region having at least 70, 80, 90, or 100 amino acids having a sequence identity of at least 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% with the equal-length portion of SEQ ID NO: 94.
[0287] In some embodiments, the polypeptide includes a region having at least 15, 16, 17, or 18 amino acids having a sequence identity of at least 70% with the equal-length portion of SEQ ID NO: 97. In some embodiments, the polypeptide includes a region having at least 15, 16, 17, or 18 amino acids having a sequence identity of at least 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% with the equal-length portion of SEQ ID NO: 97.
[0288] In some embodiments, the polypeptide includes a region having at least 30, 40, 45, or 50 amino acids having a sequence having at least 70% sequence identity with the equal-length portion of SEQ ID NO: 104. In some embodiments, the polypeptide includes a region having at least 30, 40, 45, or 50 amino acids having a sequence having at least 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the equal-length portion of SEQ ID NO: 104.
[0289] In some embodiments, the polypeptide includes a region having at least 70, 80, 90, or 100 amino acids having a sequence having at least 70% sequence identity with the equal-length portion of SEQ ID NO: 106. In some embodiments, the polypeptide includes a region having at least 70, 80, 90, or 100 amino acids having a sequence having at least 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the equal-length portion of SEQ ID NO: 106.
[0290] In some embodiments, the polypeptide includes a region having at least 35, 40, 45, or 48 amino acids having a sequence having at least 70% sequence identity with the equal-length portion of SEQ ID NO: 109. In some embodiments, the polypeptide includes a region having at least 35, 40, 45, or 48 amino acids having a sequence having at least 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the equal-length portion of SEQ ID NO: 109.
[0291] In some embodiments, the polypeptide includes a region having at least 35, 40, 45, or 47 amino acids having a sequence having at least 70% sequence identity with the equal-length portion of SEQ ID NO: 116. In some embodiments, the polypeptide includes a region having at least 35, 40, 45, or 47 amino acids having a sequence having at least 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the equal-length portion of SEQ ID NO: 116.
[0292] In some embodiments, the polypeptide includes a region having at least 70, 80, 90, or 100 amino acids having a sequence having at least 70% sequence identity with the equal-length portion of SEQ ID NO: 118. In some embodiments, the polypeptide includes a region having at least 70, 80, 90, or 100 amino acids having a sequence having at least 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the equal-length portion of SEQ ID NO: 118.
[0293] In some embodiments, the polypeptide comprises a RuvCI domain having at least 70% sequence identity with SEQ ID NO: 85, SEQ ID NO: 97, or SEQ ID NO: 109. In some embodiments, the polypeptide comprises a RuvCI domain having at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 85, SEQ ID NO: 97, or SEQ ID NO: 109.
[0294] In some embodiments, the polypeptide comprises a RuvCII domain having at least 70% sequence identity with SEQ ID NO: 92, SEQ ID NO: 104, or SEQ ID NO: 116. In some embodiments, the polypeptide comprises a RuvCII domain having at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 92, SEQ ID NO: 104, or SEQ ID NO: 116.
[0295] In some embodiments, the polypeptide comprises a RuvCIII domain having at least 70% sequence identity with SEQ ID NO: 94, SEQ ID NO: 106, or SEQ ID NO: 118. In some embodiments, the polypeptide comprises a RuvCIII domain having at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 94, SEQ ID NO: 106, or SEQ ID NO: 118.
[0296] In some embodiments, the polypeptide comprises a RuvC domain consisting of subdomains having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 85, 92, and 94; or SEQ ID NO: 97, 104, and 106; or SEQ ID NO: 109, 116, and 118.
[0297] In some embodiments, the polypeptide includes a region having at least 70, 80, 90, or 100 amino acids having a sequence identity of at least 70% with the equal-length portion of SEQ ID NO: 93. In some embodiments, the polypeptide includes a region having at least 70, 80, 90, or 100 amino acids having a sequence identity of at least 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% with the equal-length portion of SEQ ID NO: 93.
[0298] In some embodiments, the polypeptide includes a region having at least 70, 80, 90, or 100 amino acids having a sequence having at least 70% sequence identity with the equal-length portion of SEQ ID NO: 105. In some embodiments, the polypeptide includes a region having at least 70, 80, 90, or 100 amino acids having a sequence having at least 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the equal-length portion of SEQ ID NO: 105.
[0299] In some embodiments, the polypeptide includes a region having at least 70, 80, 90, or 100 amino acids having a sequence having at least 70% sequence identity with the equal-length portion of SEQ ID NO: 117. In some embodiments, the polypeptide includes a region having at least 70, 80, 90, or 100 amino acids having a sequence having at least 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the equal-length portion of SEQ ID NO: 117.
[0300] In some embodiments, the polypeptide comprises an HNH domain having at least 70% sequence identity with any of SEQ ID NO: 93, 105, or 117. In some embodiments, the polypeptide comprises an HNH domain having at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any of SEQ ID NO: 93, 105, or 117.
[0301] In some embodiments, the polypeptide includes a region having at least 70, at least 75, or at least 80 amino acids having a sequence identity of at least 70% with the equal-length portion of SEQ ID NO: 87. In some embodiments, the polypeptide includes a region having at least 70, at least 75, or at least 80 amino acids having a sequence identity of at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% with the equal-length portion of SEQ ID NO: 87.
[0302] In some embodiments, the polypeptide includes a region having at least 70, at least 75, or at least 80 amino acids having a sequence having at least 70% sequence identity with the equal-length portion of SEQ ID NO: 99. In some embodiments, the polypeptide includes a region having at least 70, at least 75, or at least 80 amino acids having a sequence having at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with the equal-length portion of SEQ ID NO: 99.
[0303] In some embodiments, the polypeptide includes a region having at least 70, 75, or 77 amino acids having a sequence having at least 70% sequence identity with the equal-length portion of SEQ ID NO: 111. In some embodiments, the polypeptide includes a region having at least 70, 75, or 77 amino acids having a sequence having at least 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the equal-length portion of SEQ ID NO: 111.
[0304] In some embodiments, the polypeptide includes a region having at least 150, at least 175, or at least 200 amino acids having a sequence identity of at least 70% with the equal-length portion of SEQ ID NO: 89. In some embodiments, the polypeptide includes a region having at least 150, at least 175, or at least 200 amino acids having a sequence identity of at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% with the equal-length portion of SEQ ID NO: 89.
[0305] In some embodiments, the polypeptide includes a region having at least 150, at least 175, or at least 200 amino acids having a sequence identity of at least 70% with the equal-length portion of SEQ ID NO: 101. In some embodiments, the polypeptide includes a region having at least 150, at least 175, or at least 200 amino acids having a sequence identity of at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% with the equal-length portion of SEQ ID NO: 101.
[0306] In some embodiments, the polypeptide includes a region having at least 150, at least 175, or at least 200 amino acids having a sequence identity of at least 70% with the equal-length portion of SEQ ID NO: 113. In some embodiments, the polypeptide includes a region having at least 150, at least 175, or at least 200 amino acids having a sequence identity of at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% with the equal-length portion of SEQ ID NO: 113.
[0307] In some embodiments, the polypeptide includes a region having at least 125, at least 150, or at least 175 amino acids having a sequence identity of at least 70% with the equal-length portion of SEQ ID NO: 88. In some embodiments, the polypeptide includes a region having at least 125, at least 150, or at least 175 amino acids having a sequence identity of at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% with the equal-length portion of SEQ ID NO: 88.
[0308] In some embodiments, the polypeptide includes a region having at least 125, at least 150, or at least 175 amino acids having a sequence having at least 70% sequence identity with the equal-length portion of SEQ ID NO: 100. In some embodiments, the polypeptide includes a region having at least 125, at least 150, or at least 175 amino acids having a sequence having at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with the equal-length portion of SEQ ID NO: 100.
[0309] In some embodiments, the polypeptide includes a region having at least 50, 55, or 60 amino acids having a sequence having at least 70% sequence identity with the equal-length portion of SEQ ID NO: 112. In some embodiments, the polypeptide includes a region having at least 50, 55, or 60 amino acids having a sequence having at least 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the equal-length portion of SEQ ID NO: 112.
[0310] In some embodiments, the polypeptide includes a region having at least 125, at least 150, or at least 175 amino acids having a sequence identity of at least 70% with the equal-length portion of SEQ ID NO: 90. In some embodiments, the polypeptide includes a region having at least 125, at least 150, or at least 175 amino acids having a sequence identity of at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% with the equal-length portion of SEQ ID NO: 90.
[0311] In some embodiments, the polypeptide includes a region having at least 125, at least 150, or at least 175 amino acids having a sequence having at least 70% sequence identity with the equal-length portion of SEQ ID NO: 102. In some embodiments, the polypeptide includes a region having at least 125, at least 150, or at least 175 amino acids having a sequence having at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with the equal-length portion of SEQ ID NO: 102.
[0312] In some embodiments, the polypeptide includes a region having at least 125, at least 150, or at least 175 amino acids having a sequence having at least 70% sequence identity with the equal-length portion of SEQ ID NO: 114. In some embodiments, the polypeptide includes a region having at least 125, at least 150, or at least 175 amino acids having a sequence having at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with the equal-length portion of SEQ ID NO: 114.
[0313] In some embodiments, the polypeptide comprises a RECIa domain having at least 70% sequence identity with SEQ ID NO: 87, SEQ ID NO: 99, or SEQ ID NO: 111. In some embodiments, the polypeptide comprises a RECIa domain having at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 87, SEQ ID NO: 99, or SEQ ID NO: 111.
[0314] In some embodiments, the polypeptide comprises a RECIb domain having at least 70% sequence identity with SEQ ID NO: 89, SEQ ID NO: 101, or SEQ ID NO: 113. In some embodiments, the polypeptide comprises a RECIb domain having at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 89, SEQ ID NO: 101, or SEQ ID NO: 113.
[0315] In some embodiments, the polypeptide comprises a RECII domain having at least 70% sequence identity with SEQ ID NO: 88, SEQ ID NO: 100, or SEQ ID NO: 112. In some embodiments, the polypeptide comprises a RECII domain having at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 88, SEQ ID NO: 100, or SEQ ID NO: 112.
[0316] In some embodiments, the polypeptide comprises a RECIII domain having at least 70% sequence identity with SEQ ID NO: 90, SEQ ID NO: 102, or SEQ ID NO: 114. In some embodiments, the polypeptide comprises a RECIII domain having at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 90, SEQ ID NO: 102, or SEQ ID NO: 114.
[0317] In some embodiments, the polypeptide includes a region having at least 500, 550, 600, or 648 amino acids having a sequence identity of at least 70% with the equal-length portion of SEQ ID NO: 91. In some embodiments, the polypeptide includes a region having at least 500, 550, 600, or 648 amino acids having a sequence identity of at least 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% with the equal-length portion of SEQ ID NO: 91.
[0318] In some embodiments, the polypeptide includes a region having at least 500, 550, 600, or 650 amino acids having a sequence identity of at least 70% with the equal-length portion of SEQ ID NO: 103. In some embodiments, the polypeptide includes a region having at least 500, 550, 600, or 650 amino acids having a sequence identity of at least 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% with the equal-length portion of SEQ ID NO: 103.
[0319] In some embodiments, the polypeptide includes a region having at least 500, 550, 600, 650, or 656 amino acids having a sequence identity of at least 70% with the equal-length portion of SEQ ID NO: 115. In some embodiments, the polypeptide includes a region having at least 500, 550, 600, 650, or 656 amino acids having a sequence identity of at least 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% with the equal-length portion of SEQ ID NO: or 115.
[0320] In some embodiments, the polypeptide comprises a REC leaf having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 91, 103, or 115.
[0321] In some embodiments, the polypeptide includes a region having at least 500, 550, 600, or 616 amino acids having a sequence having at least 70% sequence identity with the equal-length portion of SEQ ID NO: 96. In some embodiments, the polypeptide includes a region having at least 500, 550, 600, or 616 amino acids having a sequence having at least 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the equal-length portion of SEQ ID NO: 96.
[0322] In some embodiments, the polypeptide includes a region having at least 500, 550, 600, or 648 amino acids having a sequence having at least 70% sequence identity with the equal-length portion of SEQ ID NO: 108. In some embodiments, the polypeptide includes a region having at least 500, 550, 600, or 648 amino acids having a sequence having at least 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the equal-length portion of SEQ ID NO: 108.
[0323] In some embodiments, the polypeptide includes a region having at least 500, 550, 600, or 650 amino acids having a sequence identity of at least 70% with the equal-length portion of SEQ ID NO: 120. In some embodiments, the polypeptide includes a region having at least 500, 550, 600, or 648 amino acids having a sequence identity of at least 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% with the equal-length portion of SEQ ID NO: 120.
[0324] In some embodiments, the polypeptide comprises a NUC leaf having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 96, 108, or 120.
[0325] In some embodiments, the polypeptide comprises a region having at least 175, at least 200, at least 225, or 244 amino acids having a sequence identity of at least 70% with the equal-length portion of SEQ ID NO: 95. In some embodiments, the polypeptide comprises at least 175, at least 200, at least 225, or 244 amino acids having a sequence identity of at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% with the equal-length portion of SEQ ID NO: 95.
[0326] In some embodiments, the polypeptide comprises a region having at least 175, at least 200, at least 225, or 242 amino acids having a sequence identity of at least 70% with the equal-length portion of SEQ ID NO: 107. In some embodiments, the polypeptide comprises at least 175, at least 200, at least 225, or 242 amino acids having a sequence identity of at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% with the equal-length portion of SEQ ID NO: 107.
[0327] In some embodiments, the polypeptide comprises a region having at least 175, at least 200, at least 225, at least 250, or 278 amino acids having a sequence identity of at least 70% with the equal-length portion of SEQ ID NO: 119. In some embodiments, the polypeptide comprises at least 175, at least 200, at least 225, at least 250, or 278 amino acids having a sequence identity of at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% with the equal-length portion of SEQ ID NO: 119.
[0328] In some embodiments, the polypeptide comprises a PI domain having at least 70% sequence identity with any of SEQ ID NO: 95, 107, or 119. In some embodiments, the polypeptide comprises a PI domain having at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any of SEQ ID NO: 95, 107, or 119.
[0329] In some embodiments, the polypeptide comprises a REC leaf and a NUC leaf, the REC leaf having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 91, and the NUC leaf having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 96.
[0330] In some embodiments, the polypeptide comprises a REC leaf and a NUC leaf, the REC leaf having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 103, and the NUC leaf having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 108.
[0331] In some embodiments, the polypeptide comprises a REC leaf and a NUC leaf, the REC leaf having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 115, and the NUC leaf having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 120.
[0332] In some embodiments, the polypeptide comprises a REC leaf having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 91; a NUC leaf and a PI domain having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 96; and the PI domain having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 95.
[0333] In some embodiments, the polypeptide comprises a REC leaf having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 103; a NUC leaf and a PI domain having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 108; and the PI domain having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 107.
[0334] In some embodiments, the polypeptide comprises a REC leaf having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 115; a NUC leaf and a PI domain having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 120; and the PI domain having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 119.
[0335] In some embodiments, the polypeptide comprises: RuvCI-bridging helix-RECIa-RECII-RECIb-RECIII-RuvCII-HNH-RuvCIII-PI, wherein each (sub)domain has at least 85%, at least 90%, at least 95%, at least 99%, or 100% identity with the corresponding sequence in Table 1. In some embodiments, the polypeptide comprises: REC leaf-NUC leaf, wherein each (sub)domain has at least 85%, at least 90%, at least 95%, at least 99%, or 100% identity with the corresponding sequence in Table 1.
[0336] Table 2 Predicted non-PI region and PI domain of novel ION-Cas protein In some embodiments, the polypeptide comprises a non-PI region having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3186. In some embodiments, the polypeptide comprises a non-PI region and a PI domain, the non-PI region having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3186, and the PI domain having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3187. In some embodiments, the polypeptide comprises a non-PI region and a PI domain, the non-PI region having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3186, and the PI domain having less than 75%, less than 60%, less than 50%, less than 40%, or less than 30% identity with SEQ ID NO: 3187.
[0337] In some embodiments, the polypeptide comprises a non-PI region having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3188. In some embodiments, the polypeptide comprises a non-PI region and a PI domain, the non-PI region having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3188, and the PI domain having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3189. In some embodiments, the polypeptide comprises a non-PI region and a PI domain, the non-PI region having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3188, and the PI domain having less than 75%, less than 60%, less than 50%, less than 40%, or less than 30% identity with SEQ ID NO: 3189.
[0338] In some embodiments, the polypeptide comprises a non-PI region having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3190. In some embodiments, the polypeptide comprises a non-PI region and a PI domain, the non-PI region having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3190, and the PI domain having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3191. In some embodiments, the polypeptide comprises a non-PI region and a PI domain, the non-PI region having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3190, and the PI domain having less than 75%, less than 60%, less than 50%, less than 40%, or less than 30% identity with SEQ ID NO: 3191. In some embodiments, the polypeptide comprises a non-PI region having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3192. In some embodiments, the polypeptide comprises a non-PI region and a PI domain, the non-PI region having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3192, and the PI domain having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3193. In some embodiments, the polypeptide comprises a non-PI region and a PI domain, the non-PI region having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3192, and the PI domain having less than 75%, less than 60%, less than 50%, less than 40%, or less than 30% identity with SEQ ID NO: 3193. In some embodiments, the polypeptide comprises a non-PI region that has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO:3194.In some embodiments, the polypeptide comprises a non-PI region and a PI domain, the non-PI region having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3194, and the PI domain having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3195. In some embodiments, the polypeptide comprises a non-PI region and a PI domain, the non-PI region having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3194, and the PI domain having less than 75%, less than 60%, less than 50%, less than 40%, or less than 30% identity with SEQ ID NO: 3195. In some embodiments, the polypeptide comprises a non-PI region having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3196. In some embodiments, the polypeptide comprises a non-PI region and a PI domain, the non-PI region having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3196, and the PI domain having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3197. In some embodiments, the polypeptide comprises a non-PI region and a PI domain, the non-PI region having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3196, and the PI domain having less than 75%, less than 60%, less than 50%, less than 40%, or less than 30% identity with SEQ ID NO: 3197.
[0339] In some embodiments, the polypeptide comprises a non-PI region having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3198. In some embodiments, the polypeptide comprises a non-PI region and a PI domain, the non-PI region having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3198, and the PI domain having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3199. In some embodiments, the polypeptide comprises a non-PI region and a PI domain, the non-PI region having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3198, and the PI domain having less than 75%, less than 60%, less than 50%, less than 40%, or less than 30% identity with SEQ ID NO: 3199.
[0340] In some embodiments, the polypeptide described above is a Cas protein. In some embodiments, the polypeptide described above is a Cas enzyme. In some embodiments, the polypeptide described above is an inactivated Cas protein. In some embodiments, the polypeptide described above is a Cas fusion protein.
[0341] B. Protein Engineering Structure-guided protein engineering methods and protein engineering via directed evolution have been applied to various Cas proteins (see Liu et al.). Trends in Biotechnology 39(3):262-273, 2021).
[0342] 1. Certain heterogeneous structural domains In some embodiments, heterologous domains from non-Cas proteins may be attached to the Cas protein to enhance functionality (see, for example, WO2014 / 152432; WO2016 / 063264; WO2022 / 162247; WO2022 / 140577; US2019 / 0233805), thereby causing a Cas fusion protein. In some embodiments, the Cas moiety of the fusion protein is used to target the heterologous domain to a DNA target of interest. In some embodiments, the Cas moiety of the fusion protein is an inactivated Cas protein.
[0343] In some such embodiments, the heterodomain is a transcription activator or a transcription repressor, or a domain thereof. For example, in some cases, the heterodomain represses transcription (e.g., the heterodomain is a transcription repressor, a polypeptide that acts by recruiting a transcription repressor protein, a polypeptide that causes target DNA modification such as methylation, a polypeptide that causes recruitment of DNA modifiers, a polypeptide that regulates histones associated with the target DNA, or a polypeptide that recruits histone modifiers such as those that modify histone acetylation and / or methylation). In other cases, the heterodomain increases transcription (e.g., the heterodomain is a transcription activator, a polypeptide that acts by recruiting a transcription activator protein, a polypeptide that causes target DNA modification such as demethylation, a polypeptide that causes recruitment of DNA modifiers, a polypeptide that regulates histones associated with the target DNA, or a polypeptide that recruits histone modifiers such as those that modify histone acetylation and / or methylation).
[0344] In some embodiments, the Cas fusion protein includes a heterologous domain having enzymatic activity that modifies the target nucleic acid (e.g., nuclease activity, methyltransferase activity, demethylase activity, DNA repair activity, DNA damage activity, deaminase activity, superoxide dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer formation activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity, or glycosylation activity).
[0345] In some embodiments, a heterologous domain modifies a second polypeptide (e.g., a histone) associated with the target nucleic acid (e.g., a polypeptide having methyltransferase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylation activity, deSUMOylation activity, ribosylation activity, deribosylation activity, myristoylation activity, or demyristoylation activity).
[0346] Examples of peptides that can be used to enhance transcription include, but are not limited to, transcription activators such as VP16, VP64, VP48, VP160, MyoD1, HSF1, RTA, SET7 / 9, or their domains; p65 subdomains (e.g., from NFkB), and activation domains of EDLL and / or TAL (e.g., for activity in plants); histone lysine methyltransferases such as SET1A, SET1B, MLL1 to 5, ASH1, SYMD2, NSD1, or their domains; histone lysine demethylases such as JHDM2a / b, UTX, JMJD3, or their domains; histone acetyltransferases such as GCN5, PCAF, CBP, p300, TAF1, TIP60 / PLIP, MOZ / MYST3, MORF / MYST4, SRC1, ACTR, P160, CLOCK, or their domains; and DNA demethylases such as undecyltransfer (TET) dioxygenase 1. (TET1CD), TET1, DME, DML1, DML2, ROS1 or their domains.
[0347] Examples of peptides that can be used to reduce transcription include, but are not limited to, transcriptional repressor peptides, such as Krüppel-associated box (KRAB or SKD) domains; KOX1 repressor domains; and Mad... mSIN3 interaction domain (SID); ERF repressor domain (ERD), SRDX repressor domain (e.g., for inhibition in plants); histone lysine methyltransferases, such as Pr-SET7 / 8, SUV4-20H1, RIZ1, or their domains; histone lysine demethylases, such as JMJD2A / JHDM3A, JMJD2B, JMJD2C / GASCI, JMJD2D, JARID1A / RBP2, JARID1B / PLU-I, JARID1C / SMCX, JARID1D / SMCY; histone lysine deacetylases, such as HDAC1, HDAC2, HDAC3, HDAC8, HDAC4, HDAC5, HDAC7, HDAC9, SIRTI, SIRT2, HDAC11, or their domains; DNA methyltransferases, such as HhaI DNA m5c-methyltransferase (M.HhaI), DNA methyltransferase 1 (DNMT1), DNA methyltransferase 3a (DNMT3a), DNA methyltransferase 3b (DNMT3b), METI, DRM3 (plant), ZMET2, CMTI, CMT2 (plant), or their domains; and peripheral recruitment elements, such as lamin A and lamin B, or their domains.
[0348] In some embodiments, the heterologous domain has enzymatic activity that modifies the target nucleic acid. Examples of enzymatic activity that can be provided by the heterologous domain include, but are not limited to: nuclease activity, such as that provided by restriction proteins (e.g., Fokl nuclease); methyltransferase activity, such as that provided by methyltransferases (e.g., Hhal DNA m5c-methyltransferase (M.Hhal), DNA methyltransferase 1 (DNMT1), DNA methyltransferase 3a (DNMT3a), DNA methyltransferase 3b (DNMT3b), METI, DRM3 (plant), ZMET2, CMT1, CMT2 (plant)); and demethylase activity, such as that provided by demethylases (e.g., undecyltransfer (TET) dioxygenase 1). (TET1CD), TET1, DME, DML1, DML2, ROSI, etc.) provide demethylase activity; DNA repair activity; DNA damage activity; deaminase activity, such as deaminase activity provided by deaminases (e.g., cytosine deaminase, such as rat APOBEC1); dismutase integrase activity, such as activity provided by integrase and / or dissociation enzymes (e.g., Gin convertase, such as the highly active mutant GinH106Y of Gin convertase; human immunodeficiency virus type 1 integrase (IN); Tn3 dissociation enzyme), transposase activity, recombinase activity, such as activity provided by recombinase (e.g., the catalytic domain of Gin recombinase), polymerase activity, ligase activity, helicase activity, photolyase activity, and glycosylation enzyme activity).
[0349] In some embodiments, the heterodomain has enzymatic activity that modifies proteins associated with the target nucleic acid (e.g., histones, RNA-binding proteins, DNA-binding proteins). Examples of enzymatic activity (modifying proteins associated with the target nucleic acid) provided by the heterodomain include, but are not limited to, methyltransferase activity, such as that provided by histone methyltransferases (HMTs) (e.g., mottled inhibitor 3-9 homolog 1 (SUV39H1, also known as KMT1A), euchromatin histone lysine methyltransferase 2 (G9A, also known as KMT1C and EHMT2), SUV39H2, ESET / SETDBI, SET1A, SET1B, MLL1 to 5, ASH1, SYMD2, etc.). Methyltransferase activity provided by NSD1, DOT1L, Pr-SET7 / 8, SUV4-20H1, EZH2, RIZ1; demethylase activity, such as that provided by histone demethylases (e.g., lysine demethylase 1A (KDM1A, also known as LSD1), JHDM2a / b, JMJD2A / JHDM3A, JMJD2B, JMJD2C / GASC1, JMJD2D, JARID1A / RBP2, JARID1B / PLU-1, JARID1C / SM). Demethylase activity provided by CX, JARID1D / SMCY, UTX, JMJD3; acetyltransferase activity, such as acetyltransferase activity provided by histone acetyltransferases (e.g., human acetyltransferase p300, GCN5, PCAF, CBP, TAF1, TIP60 / PLIP, MOZ / MYST3, MORF / MYST4, HBO1 / MYST2, HMOF / MYST1, SRC1, ACTR, P160, CLOCK catalytic core / fragment); deacetylation Enzyme activities, such as deacetylase activity provided by histone deacetylases (e.g., HDAC1, HDAC2, HDAC3, HDAC8, HDAC4, HDAC5, HDAC7, HDAC9, SIRTI, SIRT2, HDAC11); kinase activity; phosphatase activity; ubiquitin ligase activity; deubiquitination activity; adenylation activity; deadenylation activity; SUMOylation activity; deSUMOylation activity; ribosylation activity; deribosylation activity; myristylation activity; and demyristylation activity.
[0350] Domains that have been attached to Cas proteins include FokI (see, e.g., WO2014 / 152432) for highly specific genome editing, and transcriptional repressor domains (e.g., KRAB; see, e.g., Alerasool et al.). Nature Methods, 1093 to 1096, 2020) and transcriptional activation domains (e.g., VP64—see, e.g., WO2014 / 197748; p65; and RTA) for the regulation of gene expression; histone modifying enzymes, such as methyltransferases for epigenetic editing (see, e.g., WO2016 / 103233); and fluorescent proteins for genome imaging (see, e.g. Trends in Biotechnology , 39(3):262-273, 2021). One or more of these polypeptides may be attached to the Cas protein at the N-terminus or C-terminus. In some embodiments, multiple domains are attached to the C-terminus, N-terminus, or both of the Cas protein (see, for example, WO2019 / 204766). Additional repressive domains that may be attached to the Cas protein are provided, for example, in WO2022 / 140577; Alerasool et al., Nature Methods , 1093-1096, 2020.
[0351] For some examples of the above-mentioned peptides used in the context of fusions with Cas9, zinc finger, and / or TALE proteins (for site-specific target nucleic acid modification, transcriptional regulation, and / or target protein modification, such as histone modification), see, for example: Nomura et al., J Am Chem Soc. 2007 Jul 18;129(28):8676-7; Rivenbark et al., Epigenetics. 2012 Apr;7(4):350-60; Nucleic Acids Res. 2016 Jul 8;44(12):5615-28; Gilbert et al., Cell. 2013 Jul 18;154(2):442-51; Kearns et al., NatMethods. 2015 May;12(5):401-3; Mendenhall et al., Nat Biotechnol. Dec 2013;31(12):1133-6; Hilton et al., Nat Biotechnol. May 2015;33(5):510-7; Gordley et al., Proc Natl Acad Sci US A. March 31, 2009;106(13):5053-8; Akopian et al., Proc Natl Acad Sci US A. July 22, 2003;100(15):8688-91; Tan et al., J Virol. Feb 2006;80(4):1939-48; Tan et al., Proc Natl Acad Sci US A. October 14, 2003;100(21):11997-2002; Papworth et al., Proc Natl Acad Sci US A. February 18, 2003; 100(4):1621-6; Sanjana et al., Nat Protoc. January 5, 2012; 7(1):171-92; Beerli et al., Proc Natl AcadSci USA. December 8, 1998; 95(25):14628-33; Snowden et al., Curr Biol. December 23, 2002; 12(24):2159-66; Xu et al., Xu et al., Cell Discov. May 3, 2016; 2:16009; Komor et al., Nature. April 20, 2016; 533(7603):420-4; Chaikind et al., Nucleic Acids Res. August 11, 2016; Choudhury et al., Oncotarget.June 23, 2016; Du et al., Cold Spring HarbProtoc. January 4, 2016; Pham et al., Methods Mol Biol. 2016; 1358:43-57; Balboa et al., Stem Cell Reports. September 8, 2015; 5(3):448-59; Hara et al., Sci Rep. June 9, 2015; 5:11221; Piatek et al., Plant Biotechnol J. May 2015; 13(4):578-89; Hu et al., Nucleic Acids Res. April 2014; 42(7):4375-90; Cheng et al., Cell Res. October 2013; 23(10):1163-71; Cheng et al., Cell Res. Oct 2013;23(10):1163-71; and Maeder et al., NatMethods. Oct 2013;10(10):977-9.
[0352] 2. Certain nuclear localization sequences In some embodiments, the Cas protein may include one or more nuclear localization sequences. Nuclear localization sequences are short amino acid sequences that facilitate protein uptake into the cell nucleus. Various eukaryotic nuclear localization signal (NLS) sequences have been used to improve protein nuclear uptake (see, for example, Lu et al.). Cell Comm. and Signalling, (2021). NLS can be located at the N-terminus and / or C-terminus of a polypeptide, or it can be located at any point within the polypeptide sequence. A single-part classical NLS has the sequence K(K / R)X(K / R), where K is a lysine and X can be any amino acid. A bipartite NLS comprises two clusters of 2 to 3 positively charged amino acids separated by a 9 to 12-amino acid group containing several prolines, having R / K(X). 10-12 The common sequence of KRXK, where R is arginine, K is lysine, and X is any amino acid.
[0353] In some embodiments, the Cas protein may comprise one or more nuclear localization sequences or variants thereof, selected from the table below.
[0354] Table 3 nuclear localization sequence 3. Engineered Cas variants with modified PAM specificity The Cas protein binds to the target DNA at a sequence defined by the complementary region between the guide's target recognition region and the target DNA. Site-specific binding (and / or cleavage) of the Cas protein to the double-stranded target DNA typically occurs at a site determined by both: (i) hybridization of the guide's target recognition region to the target DNA, and (ii) the presence of a short motif (referred to as a PAM, or protospacer adjacent motif) within the non-complementary strand of the DNA. In some embodiments, the Cas protein's PAM is adjacent to the 5' of the protospacer sequence. In some embodiments, the Cas protein's PAM is adjacent to the 3' of the protospacer sequence.
[0355] Methods for engineering PAM-interacting domains (PI domains) of Cas proteins (including SpCas9, SaCas9, StlCas9, FnCas9, AsCas12a, etc.) have been previously described (see, for example, WO2016 / 141224; Liu et al., Trends in Biotechnology Table 1 in 39(3):262-273, 2021).
[0356] Bacterial selection systems can be used to optimize Cas protein variants with a non-natural PI domain configured to recognize variant PAMs. Bacteria are transfected using plasmids encoding an inducible virulence gene and plasmids encoding the Cas protein variant. After inducing the virulence gene, the only surviving bacteria are those in which the Cas protein variant cleaves the virulence gene. Using this method, large Cas protein libraries with randomized mutations within the PI domain can be screened against large randomized PAM libraries. For surviving clones, the PAM specificity of evolving Cas proteins can then be analyzed using a bacterial site depletion assay. These methods have previously been used to identify PI domains (Kleinstiver, Kleinstiver, Kleinstiver) with novel PAM recognition sequences targeting SpCas9, SaCas9, and StlCas9. Nature , 523(7561): 481-485, 2015); SaCas9 (Kleinstiver, Nature Biotech. 33(12):1293-1298, 2015). Similarly, the phage-assisted continuous evolution (PACE) selection system has been used to identify Cas9 variants with extended PAM recognition (Hu et al., Nature 556: 57-63, 2918; Miller et al., Nature Biotech. 38: 471-481, 2020).
[0357] Another approach previously used to engineer alternative PAM-interacting domains for identifying variant PAMs is structure-guided protein engineering. Using this method, amino acid residues that form contact with the target DNA are identified via crystal structure, and the activity of mutated Cas protein variants containing these amino acids is subsequently tested. Variants of saCas9 with reduced off-target editing were identified using these methods (Tan et al., PNAS, 116(42):20969 to 20976, 2019), and the same is true for the AsCas12a variant (Kleinstiver et al., Nature Biotech. 37(3): 276-282, 2019). This method has also been used to identify variants of FnCas9 with broader PAM recognition sequences (Hirano et al., Cell , 164:950-961).
[0358] Another approach used to generate Cas protein variants that recognize PAM variants is to use direct substitution of the PI domain of a Cas protein with the PI domain of a closely related orthologous Cas protein to produce domain-exchanged chimeras. Previous work showed that exchanging the PI domain of SpCas9 with the PI domain of St3Cas9 produced a variant of St3Cas9 that cleaved target DNA with typical SpCas9 5'-NGG-3' PAM (Nishimasu et al., Cell, 156(5):935-949, 2014). Similarly, several Cas9 variants that recognize PAM variants have been identified by exchanging the PI domain with closely related Cas proteins (Ma et al., Nature Comm ., 10:560, 2019).
[0359] C. Nucleic acids encoding certain polypeptides 1. Certain nucleic acids In some embodiments, this document provides a nucleic acid encoding a polypeptide. In some embodiments, the polypeptide is a guided nucleic acid binder or a portion thereof. In some embodiments, the polypeptide comprises or is composed of a Cas protein. In some embodiments, the polypeptide comprises or is composed of an ION-Cas protein.
[0360] In some embodiments, the nucleic acid is DNA. In some embodiments, the nucleic acid is RNA. In some embodiments, the nucleic acid is exogenous mRNA.
[0361] In some embodiments, the exogenous mRNA comprises a coding region (e.g., an open reading frame (ORF)) encoding a polypeptide sequence, a cap, and one or more non-coding regions. In some embodiments, the exogenous mRNA comprises a cap, a 5' UTR, a 3' UTR, a coding region, a poly(A) tail, and optionally one or more introns. The mRNA may comprise nucleotides selected from adenosine, guanosine, cytosine, uridine, and N1-methylpseudouridine, wherein the N1-methylpseudouridine is optionally selected from adenosine, guanosine, cytosine, and uridine.
[0362] Naturally occurring eukaryotic mRNA molecules may contain non-coding regions, including but not limited to untranslated regions (UTRs) at their 5' end (5' UTR) and / or at their 3' end (3' UTR), 5' cap structures, and 3'-poly(A) tails. Exogenous mRNAs may be configured to include such regions, which facilitate cellular processing and translation. In some embodiments, formulations (e.g., LNPs) may be configured to deliver exogenous mRNA having an open reading frame encoding a polypeptide. In some embodiments, methods in which the exogenous mRNA is translated in cells are provided.
[0363] Exogenous mRNA may contain repeating regions of consecutive nucleotides, such as adenosine nucleotides, for example, poly(A) tails. A poly(A) tail is a region located at the 3' end of a 3' UTR containing multiple consecutive adenosine monophosphates. It is believed that the poly(A) tail protects mRNA from enzymatic degradation (e.g., in the cytoplasm) and facilitates transcription termination, and / or mRNA export from the nucleus and translation. A poly(A) tail may contain 10 to 300 consecutive adenosine nucleotides. For example, exogenous mRNA may contain repeating regions of 50 to 100, 50 to 150, 50 to 200, 100 to 250, 120 to 160, or 200 to 300 adenosine monophosphates at its 3' end.
[0364] 2. Capped group Eukaryotic mRNA, and RNA virus genomes, include a 7-methylguanosine (m7G) cap at the 5' end of the mRNA sequence. This 7-methylguanosine (m7G) cap is linked to the last 5' mRNA nucleotide via a 5',5'-triphosphate bridge (ppp) during in vitro transcription of mRNA. The cap structure plays a fundamental role in mRNA translation by recruiting translation initiation factors, and different 5' caps can be incorporated into naturally occurring mRNA. Cap0 protects endogenous mRNA from nuclease attack and is also involved in nuclear export and translation initiation. Cap1 and Cap2 are both types of 5' caps containing an additional methyl group on the second or third ribonucleotide. Compared to Cap0, the additional modification of Cap1 and Cap2 is considered to reduce immunogenicity.
[0365] Exogenous mRNA may contain specific capping groups as described herein or as known in the art. Examples of 5' capping groups include: “Cap 0”: m7G(5')ppp(5')N; “Cap 1”: m7G(5')ppp(5')(2'OMeN); and “Cap 2”: m7G(5')ppp(5')(2'OMeN)(2'OMeN); wherein m7G represents a guanosine nucleoside methylated at position 7 and having a free 3'-OH, (5') represents a 5' attachment site, p is a phosphate bond, each N is independently a nucleoside, such as guanosine or adenosine, and (2'OMeN) is independently a 2'-O-methyl nucleoside, such as 2'-O-methylguanosine or 2'-O-methyladenosine. The first nucleoside adenosine following the cap may also be methylated at its N6 position.
[0366] Specific capping groups include (m7(3'OMeG)(5')ppp(5')(2'OMeA)pG; 3'-O-Me-m7G(5')ppp(5')G (“ARCA” cap); G(5')ppp(5')A; G(5')ppp(5')G; m7G(5')ppp(5')A; and m7G(5')ppp(5')(2'OMeN)pG (“CleanCap™”), where m7(3'OMeG) represents a guanosine nucleoside methylated at position 7 with a 3'-O-methyl group. Some groups are available from commercial sources (e.g., New England BioLabs, Ipswich, MA, and TriLink Biotechnologies, San Diego, CA). 5'-capping of exogenous mRNA can be performed post-transcriptionally, for example, using a vaccinia virus capping enzyme. The Cap 1 structure can be generated using both a vaccinia virus capping enzyme and a 2'-O methyltransferase. The Cap 2 structure can be generated from the Cap 1 structure, followed by 2'-O-methylation of the third nucleotide at the 5' end using a 2'-O methyltransferase. The Cap 3 structure can be generated from the Cap 2 structure, followed by 2'-O-methylation of the fourth nucleotide at the 5' end using a 2'-O methyltransferase. The enzymes can be derived, for example, from recombinant sources.
[0367] 3. Aclustered (A) tail The 3'-poly(A) tail is a region of continuous adenine nucleotides at the 3' end of the transcribed mRNA. In some embodiments, the 3'-poly(A) tail comprises one to 400 adenine nucleotides. In some embodiments, the 3'-poly(A) tail comprises unmodified adenine nucleotides linked by phosphodiester nucleotide bonds.
[0368] In some embodiments, the exogenous mRNA contains a stabilizing element. This stabilizing element may be a histone stem-loop. Stem-loop-binding protein (SLBP), a 32 kDa protein, has been identified. It associates with the histone stem-loop at the 3' end of histone messengers in both the nucleus and cytoplasm and is believed to promote efficient 3' end processing and translation stimulation of histone precursor mRNA. In some embodiments, the histone stem-loop sequence comprises 15 to 45 nucleotides in length.
[0369] In some embodiments, the exogenous mRNA includes a coding region, at least one histone stem-loop, and optionally a poly(A) sequence or polyadenylation signal. The poly(A) sequence or polyadenylation signal is thought to increase the expression of the encoded protein. In some embodiments, the encoded protein is not a histone, a reporter protein (e.g., luciferase, GFP, EGFP, β-galactosidase, EGFP), or a tagging or selecting protein (e.g., α-globin, galactokinase, and xanthine:guanine phosphoribosyltransferase (GPT)).
[0370] 4. Coding region and sequence optimization The coding region may contain a continuous sequence that begins with a start codon (e.g., methionine (AUG)) and ends with a stop codon (e.g., UAA, UAG, or UGA), and contains one or more nucleotide sequences encoding amino acids. Typically, the coding region encodes a polypeptide that forms a protein in vivo when administered to a subject.
[0371] In some embodiments, the coding region contains selected nucleotide codons to improve one or more properties of the foreign mRNA or the encoded protein, and may contain a codon-optimized ORF. For example, the coding region of a foreign mRNA sequence may be codon-optimized. In some embodiments, codon optimization can be used to match codon frequencies in the target and host organisms to ensure correct folding; bias GC content to increase mRNA stability or reduce secondary structures; minimize tandem repeat codons or base chains that may impair gene construction or expression; customize transcription and translation control regions; insert or remove protein transport sequences; add or remove post-translational modification sites (e.g., glycosylation sites) in the encoded protein; add, remove, or reorganize protein domains; insert or delete restriction sites; modify ribosome binding sites and mRNA degradation sites; modulate translation rates to allow correct folding of various protein domains; or generally reduce or eliminate problematic secondary structures within polynucleotides. Codon optimization tools, algorithms, and services are known in the art—non-limiting examples include services and / or proprietary methods from GeneArt (Life Technologies), DNA2.0 (Menlo Park, CA). In some embodiments, the coding region sequence is optimized using optimization algorithms as provided herein or as known in the art.
[0372] In some embodiments, the coding region shares less than 95%, less than 90%, less than 80%, or less than 70% sequence identity with a naturally occurring or wild-type sequence ORF (e.g., a naturally occurring or wild-type mRNA sequence encoding a protein), or a range of values between those values.
[0373] In some embodiments, the exogenous mRNA may be one in which the G / C level is enhanced. The G / C content of a nucleic acid molecule (e.g., mRNA) can affect RNA stability. RNA with increased amounts of guanine (G) and / or cytosine (C) residues may be more stable than RNA containing large amounts of adenine (A) and thymine (T) or uracil (U) nucleotides. See, for example, WO02 / 098443.
[0374] The coding region contains at least one start codon (AUG nucleotide triplet) at its 5' last nucleotide and at least one stop codon (UAG, UAA, or UGA nucleotide triplet) at its 3' last nucleotide. In some embodiments, the exogenous mRNA includes two or more stop codons, for example, two to ten stop codons.
[0375] Exemplary codons for a particular amino acid are known in the art.
[0376] 5. Chemical modification for exogenous mRNA In some embodiments, the compositions disclosed herein comprise exogenous mRNA containing modified nucleotides or nucleosides (other than those A, C, G, and U linked only by phosphodiester bonds; excluding the 5' cap). Such modified nucleotides and nucleosides may be naturally occurring or non-naturally occurring. Such modifications may include those described herein with respect to oligonucleotides.
[0377] In some embodiments, the naturally occurring modified nucleotides or nucleotides of this disclosure are as commonly known or recognized in the art. Non-limiting examples of such naturally occurring modified nucleotides and nucleotides can be found, in particular, in the widely recognized MODOMICS database. In some embodiments, the exogenous mRNA comprises a modified nucleoside described in one or more of the following U.S. applications: PCT / US2012 / 058519; PCT / US2013 / 075177; PCT / US2014 / 058897; PCT / US2014 / 058891; PCT / US2014 / 070413; PCT / US2015 / 36773; PCT / US2015 / 36759; PCT / US2015 / 36771; or PCT / IB 2017 / 051367, the entire contents of which are incorporated herein by reference. Therefore, exogenous mRNA may contain unmodified nucleosides, modified nucleosides, or combinations thereof. In some embodiments, each nucleoside in the exogenous mRNA is modified in a manner similar to other nucleosides of the same type; for example, all uridine nucleosides are present as N1-methylpseudouridine nucleosides. Compared to RNA containing only unmodified nucleosides, modification can provide reduced degradation and / or reduced immunogenicity.
[0378] In some embodiments, the exogenous mRNA comprises a modified nucleoside selected from 1-methylpseuuridine, 1-ethylpseuuridine, 5-methoxyuridine, 5-methylcytidine, and pseudouridine. In some embodiments, the exogenous mRNA comprises a modified nucleoside selected from 5-methoxymethyluridine, 5-methylthiouridine, 1-methoxymethylpseuuridine, 5-methylcytidine, and 5-methoxycytidine. In some embodiments, the exogenous mRNA comprises a modified nucleoside comprising a combination of two or more (e.g., 2, 3, or 4) of any modified nucleobases selected from this paragraph.
[0379] In some embodiments, the exogenous mRNA contains 1-methylpseuuridine at one or more, for example, all uridine nucleotides. In some embodiments, the exogenous mRNA contains 1-methylpseuuridine at one or more, for example, all uridine sites, and 5-methylcytidine at one or more, for example, all cytidine sites.
[0380] In some embodiments, the exogenous mRNA contains pseudouridine at one or more, e.g., all uridine nucleosides. In some embodiments, the exogenous mRNA contains pseudouridine at one or more, e.g., all uridine sites, and contains 5-methylcytidine at one or more, e.g., all cytidine sites.
[0381] In some embodiments, the exogenous mRNA contains unmodified uridine at one or more, such as all, uridine sites in the nucleic acid.
[0382] In some embodiments, all nucleotides of a specific type in the exogenous mRNA (or its sequence region) are modified nucleotides, wherein the modified type can be any of nucleotides A, G, U, C, or a combination of any of A+G, A+U, A+C, G+U, G+C, U+C, A+G+U, A+G+C, G+U+C, or A+G+C.
[0383] 6. Untranslated Region (UTR) In wild-type mRNA, certain regions of the nucleic acid can be transcribed into RNA but not translated. In the exogenous mRNA described herein, the 5' UTR may begin at the transcription start site and continue to the start codon at the beginning of the coding region, but does not include that start codon. The 3' UTR follows the stop codon and may include a transcription termination signal. Both the 5' UTR and 3' UTR do not encode proteins (they are non-coding regions).
[0384] In some embodiments, the presence of a UTR enhances the stability of the foreign mRNA or induces its downregulation at undesired sites or tissues. Various 5' UTR and 3' UTR sequences are known and available in the art.
[0385] The 5' UTR is a region of mRNA located directly upstream (5') of the start codon (the first codon of the mRNA transcript translated by ribosomes). In some embodiments, the 5' UTR may provide translation initiation and / or form secondary structures involving elongation factor binding. The 5' UTR may include a start sequence, such as a Kozak sequence. The Kozak sequence, also known as the Kozak concordant sequence, is considered to be involved in the ribosomal initiation of translation. In some embodiments, the Kozak sequence contains AUGG. In some embodiments, the Kozak sequence is GCCRCCAUGG (SEQ ID NO: 73) or CCRCCAUGG, where R is a purine nucleoside (adenosine or guanosine).
[0386] In some embodiments of this disclosure, the 5' UTR is an unmodified UTR, i.e., a UTR found in nature. In another embodiment, the 5' UTR is a modified UTR, i.e., not found in nature. In some embodiments, the modified UTR increases gene expression relative to its unmodified counterpart. Exemplary 5' UTRs include Xenopus laevis or human-derived α-globin or β-globin (e.g., U.S. Patent No. 9,012,219), human cytochrome b-245 α polypeptide and hydroxysteroid (17β) dehydrogenase, and tobacco etch virus (e.g., US9,012,219), CMV Immediate Early 1 (IE1) gene (US2014 / 0206753, WO2013 / 185069), sequence GGGAUCCUACC (SEQ ID NO: 74) (WO2014 / 144196), and the 5' UTR described in U.S. Patent Application Publication No. 2010 / 0293625 or WO2015085318. In another embodiment, the 5' UTR of the TOP gene is the 5' UTR of a TOP gene lacking a 5' TOP motif (oligopyrimidine bundle) (e.g., WO / 2015 / 101414, WO2015 / 101415, WO / 2015 / 062738, WO2015 / 024667, WO2015 / 024667); a 5' UTR element derived from the ribosomal protein large 32 (L32) gene (WO / 2015 / 101414, WO2015 / 101415, WO / 2015 / 062738), a 5' UTR element derived from the hydroxysteroid (17-β) dehydrogenase 4 gene (HSD17B4) (WO2015 / 024667), or a 5' UTR element derived from ATP5A1 (WO2015 / 024667) can be used. The 5' UTR element of the UTR. In some embodiments, the internal ribosome entry site (IRES) is used instead of the 5' UTR.
[0387] Typically, exogenous mRNA can include the UTR from any suitable gene. Naturally occurring UTRs, or portions thereof, may be oriented in the same direction as in the transcript from which they are derived, or their orientation or position may be altered. Thus, 5' or 3' UTRs can be reversed, shortened, or lengthened. Reference UTRs, such as naturally occurring UTRs, can be altered by including additional nucleotides, deleted nucleotides, exchanged or transposable nucleotides, or their sequences.
[0388] In some embodiments, the 3' UTR sequence may include repeating sequences containing adenosine and uridine. Such AU-rich sequences are considered to provide high turnover rates. Based on their sequence characteristics and functional properties, AU-rich elements (AREs) can be classified into three classes (Chen et al., 1995): Class I AREs contain several dispersed copies of the AUUUA motif within the U-rich region. C-Myc and MyoD contain Class I AREs. Class II AREs have two or more overlapping UUAUUUASS nonamers, where S is adenosine or uridine. Molecules containing this type of ARE include GM-CSF and TNF-α. Class III AREs are less clearly defined. These U-rich regions do not contain the AUUUA motif. c-Jun and Myogenin are two well-studied examples in this category. It is believed that including a HuR-specific binding site in the 3' UTR will provide stabilization of the messenger in vivo. In some embodiments, the 3' UTR includes repeating AREs.
[0389] Unmodified (natural) and modified (non-natural) 3' UTR sequences are known in the art. Known UTRs in this art include globin UTRs, including Xenopus laevis β-globin UTR and human β-globin UTR (9012219, US2011 / 0086907), modified β-globin constructs (US2012 / 0195936, WO2014 / 071963), α2-globin, α1-globin, UTRs (WO2015 / 101415, WO2015 / 024667), CYBA (Ferizi et al., 2015) and albumin (Thess et al., 2015), bovine or human growth hormone (wild-type or modified) (WO2013 / 185069, US2014 / 0206753, WO2014152774), rabbit β-globin and hepatitis B virus (HBV), α-globin 3' UTR and viral VEEV3'. UTR sequence, sequence UUUGAAUU (WO2014 / 144196), human and mouse ribosomal protein, rps93'UTR (WO2015 / 101414), FIG4 (WO2015 / 101415) and human albumin 7 (WO2015 / 101415).
[0390] The untranslated region may also include up-modulated motifs, such as translation enhancer elements (TEEs). As a non-limiting example, TEE may include those described in WO1999 / 024595, WO2012 / 009644, WO2009 / 075886, WO2007 / 025008, WO1999 / 024595, European Patent Publications EP2610341A1 and EP2610340A1, U.S. Patents US6310197, US6849405, US7456273, US7183395, U.S. Patents US2009 / 0226470, US2011 / 0124100, US2007 / 0048776, US2009 / 0093049, or U.S. Patent Publication US2013 / 0177581, each of which is incorporated herein by reference in its entirety.
[0391] In some embodiments, the 5' UTR or 3' UTR may include one or more additional functional regions selected from the upregulated region or the ribosome-binding region. In some embodiments, the functional region is an upregulated region (e.g., a TEE). In some embodiments, the TEE is a TEE known in the art, such as that in U.S. Application No. 2009 / 0226470. In some embodiments, the exogenous mRNA contains multiple upregulated motifs, which may be the same as or different from each other, and the number of them is, for example, 2, 3, 4, 5 or more.
[0392] In some embodiments, the exogenous mRNA includes a ribosome-binding region, such as an internal ribosome entry site (IRES). An IRES may be, for example, the IRES described in U.S. Patent No. US7468275 and International Patent Publication No. WO2001 / 055369, each of which is incorporated herein by reference in its entirety. In some embodiments, the IRES is an IRES known in the art, such as in WO2014 / 081507.
[0393] In some embodiments, the exogenous mRNA may contain dual, triple, or quadruple UTRs, such as 5' UTRs or 3' UTRs. As used herein, a "dual" UTR is a UTR in which two copies of the same UTR are continuously or substantially continuously included. For example, the exogenous mRNA may contain a dual β-globin 3' UTR as described in U.S. Patent Publication No. 2010 / 0129877, which is incorporated herein by reference in its entirety.
[0394] In some embodiments, the 5' UTR or 3' UTR may contain repeating groups of functional sequences, such as AA, ABAB, AABB, or ABCABC or variations thereof, where each letter A, B, and C represents a different functional area. The pattern may be repeated once, twice, or three or more times.
[0395] Typically, any 5' UTR sequence and any 3' UTR sequence can be combined in a specific exogenous mRNA.
[0396] Other non-coding sequences may also be used as regions or subregions within exogenous mRNA. For example, introns or portions of intron sequences may be incorporated into regions of the nucleic acids disclosed herein. In some embodiments, including intron sequences may increase protein production and nucleic acid levels.
[0397] 7. Length In some embodiments, the exogenous mRNA comprises 200 to 3,000 nucleotides. For example, the exogenous mRNA may comprise 200 to 500, 200 to 1000, 200 to 1500, 200 to 3000, 500 to 1000, 500 to 1500, 500 to 2000, 500 to 3000, 1000 to 1500, 1000 to 2000, 1000 to 3000, 1500 to 3000, or 2000 to 3000 nucleotides. Unless otherwise stated, the exogenous mRNA comprises a continuous sequence of nucleosides.
[0398] 8. Preparation In some embodiments, the RNA transcript is produced in an in vitro transcription reaction using a non-amplified, linear DNA template. In some embodiments, the template DNA is isolated DNA. In some embodiments, the template DNA is cDNA. In some embodiments, the cDNA is formed by reverse transcription of RNA polynucleotides.
[0399] In vitro transcription systems typically contain transcription buffer, nucleoside triphosphates (NTPs), RNase inhibitors, and polymerases. NTPs can be purchased from suppliers or synthesized according to methods known in the art as described herein.
[0400] In vitro transcription of RNA is known in the art and described in international publication WO 2014 / 152027, which is incorporated herein by reference in its entirety. In some embodiments, the exogenous mRNA is prepared according to any one or more of the methods described in WO 2018 / 053209 and WO 2019 / 036682, each of which is incorporated herein by reference.
[0401] 9. Purification The purification of exogenous mRNA described herein may include steps including nucleic acid purification, quality assurance, and quality control. Purification may be performed using methods known in the art, such as, but not limited to, AGENCOURT® magnetic beads (Beckman-Coulter Genomics, Danvers, MA), polythymidine magnetic beads, LNATM oligo-T capture probes (EXIQON® Inc, Vedbaek, Denmark), or HPLC-based purification methods, such as, but not limited to, strong anion exchange HPLC, weak anion exchange HPLC, reversed-phase HPLC (RP-HPLC), and hydrophobic interaction HPLC (HIC-HPLC). The term “purified” when used for exogenous mRNA refers to exogenous mRNA isolated from at least one contaminant. A “contaminant” is any substance that renders another substance unsuitable, impure, or inferior. Therefore, purified nucleic acids (e.g., DNA and RNA) exist in a form or environment different from when they are found in nature, or in a form or environment different from when they exist prior to undergoing treatment or purification methods.
[0402] Quality assurance and / or quality control checks may be performed using methods such as, but not limited to, gel electrophoresis, ultraviolet absorbance, or analytical HPLC.
[0403] In some embodiments, nucleic acids may be sequenced by methods including but not limited to reverse transcriptase-PCR.
[0404] II. certain oligonucleotides In some embodiments, this document provides oligonucleotides composed of linked nucleosides. In some embodiments, a unidirectional guide comprises an oligonucleotide. In some embodiments, a bidirectional guide comprises two oligonucleotides. The oligonucleotide may be an unmodified oligonucleotide or a modified oligonucleotide. A modified oligonucleotide comprises at least one modification relative to the unmodified nucleic acid. That is, a modified oligonucleotide comprises at least one modified nucleoside (comprising a modified sugar moiety and / or a modified nucleobase) and / or at least one modified internucleotide bond.
[0405] A. Certain modified nucleosides Modified nucleosides contain either a modified sugar moiety or a modified nucleobase, or both.
[0406] 1. Certain sugar components In some embodiments, the modified sugar moiety is a non-bicyclic modified sugar moiety. In some embodiments, the modified sugar moiety is a bicyclic or tricyclic sugar moiety. In some embodiments, the modified sugar moiety is a sugar substitute. Such sugar substitutes may contain one or more substitutions corresponding to substitutions for other types of modified sugar moiety.
[0407] In some embodiments, the modified sugar moiety is a non-bicyclic modified furanyl sugar moiety comprising one or more substituents, including but not limited to substituents at the 2', 3', 4' and / or 5' positions, as indicated by the following numbers: In some embodiments, the modified furanyl sugar moiety is not the unmodified sugar moiety ( Right now The ribosyl sugar moiety of the unmodified RNA or unmodified DNA portion. In some embodiments, the modified furanosyl sugar moiety is a xylose, lythose, or arabinose sugar moiety.
[0408] In some embodiments, the non-bicyclic modified sugar moiety is a 2'-substituted sugar moiety and includes a substituent at the 2' position. In some embodiments, one or more non-bridging substituents of the non-bicyclic modified sugar moiety are branched. Examples of suitable substituents for the 2' position of the modified sugar moiety include, but are not limited to: 2'-F, 2'-OCH3 (“OMe” or “O-methyl”), and 2'-O(CH2)2OCH3 (“MOE” or “O-methoxyethyl”). In some embodiments, the 2'-substituted group is selected from: halogen, allyl, amino, azide, SH, CN, OCN, CF3, OCF3, C1-C 10 Alkoxy, C1-C 10 Substituted alkoxy groups, C1-C 10 Alkyl, C1-C 10 Substituted alkyl, S-alkyl, N(R) m )-alkyl, O-alkenyl, S-alkenyl, N(R m )-Alkenyl, O-alkynyl, S-alkynyl, N(R m )-Alkyne, O-alkylene-O-alkyl, Alkyne, Alkylaryl, Arylalkyl, O-Alkylaryl, O-Arylalkyl, O(CH2)2SCH3, O(CH2)2ON(R m (R) n ) or OCH2C(=O)-N(R m (R) n ), where each R m and R n Independently, it is H, an amino protecting group, or a substituted or unsubstituted C1-C. 10Alkyl groups, O(CH2)2ON(CH3)2 (“DMAOE”), or 2'-O(CH2)2O(CH2)2N(CH3)2 (“DMAEOE”). The synthetic methods for some of these 2'-substituents can be found in... For example Cook et al. US6,531,584; Cook et al. US 5,859,221; and Cook et al. This was found in US 6,005,087. Some embodiments of these 2'-substituents may be further substituted with one or more substituents, which are independently selected from: halogens, cyano groups, OR... a2 NO2, NH2, NHR a2 、N(R a2 2. C1-C6 alkyl, C1-C6 haloalkyl, C2-C6 alkenyl, C2-C6 ynyl, C3-C 10 cycloalkyl, C6-C 10 Aryl, heteroaryl, heterocyclic, C1-C6 alkylene-NH2, C1-C6 alkylene-NHR a2 C1-C6 alkylene-N(R) a2 2. C(O)R a3 C(O)OR a3 C(O)NHR a3 C(O)N(C1-C4 alkyl)R a3 SR a3 S(O) )2 R a3 S(O)R a3 ,NHC(O)R a3 N(C1-C4 alkyl)C(O)R a3 NHS(O)R a3 N(C1-C4 alkyl)S(O)R a3 NHS(O)2R a3 and N(C1-C4 alkyl)S(O)2R a3 ; Each R a2 Independently selected from C2-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C3-C 10 cycloalkyl, C6-C 10 Aryl, heteroaryl, and heterocyclic; each R a3 Independently, it is hydrogen, OH, C1-C6 alkyl, C1-C6 haloalkyl, C3-C 10 cycloalkyl, C6-C 10Aryl, heteroaryl, or heterocyclic groups. In some embodiments, the sugar moiety contains two of the above-described substituents at the 2' position. In some embodiments, the sugar moiety contains a 2'-fluorine and a second 2'-substituent.
[0409] In some embodiments, the 2'-substituted sugar moiety comprises a non-bridging 2'-substituent group selected from: F, NH2, N3, OCF3, OCH3, O(CH2)3NH2, CH2CH=CH2, OCH2CH=CH2, OCH2CH2OCH3, O(CH2)2SCH3, O(CH2)2ON(R) m (R) n O(CH2)2O(CH2)2N(CH3)2 and N-substituted acetamides (OCH2C(=O)-N(R) m (R) n ), where each R m and R n Independently, it is H, an amino protecting group, or a substituted or unsubstituted C1-C. 10 Alkyl group. In some embodiments, the 2'-substituted sugar moiety comprises a non-bridging 2'-substituted group selected from the following: F, OCF3, OCH3, OCH2CH2OCH3, O(CH2)2SCH3, O(CH2)2ON(CH3)2, O(CH2)2O(CH2)2N(CH3)2, O(CH2)2ON(CH3)2 (“DMAOE”), O(CH2)2O(CH2)2N(CH3)2 (“DMAEOE”), and OCH2C(=O)-N(H)CH3 (“NMA”).
[0410] In some embodiments, the 2'-substituted sugar moiety comprises a 2'-substituent selected from the following: F, OCH3, and OCH2CH2OCH3.
[0411] In some embodiments, the modified furanyl sugar moiety and the nucleotide incorporated into such modified furanyl sugar moiety are further defined by stereochemical configuration. For example, the 2'-deoxyfuranyl sugar moiety ( Right now The 2'-(H)H furanylose moiety can exist in seven isomers other than the naturally occurring β-D-deoxyribosyl configuration. These modified sugar moieties are described in... For example This document is incorporated herein by reference in WO 2020 / 072991. The 2'-modified sugar moiety has an additional stereocenter at the 2' position relative to the 2'-deoxyfuranosyl sugar moiety; thus, such sugar moieties have a total of sixteen possible stereochemical configurations. Unless otherwise stated, the modified furanosyl sugar moieties described herein are in the β-D-ribosyl stereochemical configuration.
[0412] In some embodiments, the non-bicyclic modified sugar moiety includes a substituent at the 4' position. Examples of suitable substituents at the 4' position of the modified sugar moiety include, but are not limited to, alkoxy groups. For example , methoxy), alkyl and in Manoharan et al. Those described in WO 2015 / 106128.
[0413] In some embodiments, the non-bicyclic modified sugar moiety includes a substituent at the 3' position. Examples of suitable substituents at the 3' position of the modified sugar moiety include, but are not limited to, alkoxy groups. For example , methoxy), alkyl ( For example (methyl, ethyl).
[0414] In some embodiments, the non-bicyclic modified sugar moiety includes a substituent at the 5' position. Examples of suitable substituents at the 5' position of the modified sugar moiety include, but are not limited to, vinyl, alkoxy ( For example , methoxy), alkynyl, allyl and alkyl ( For example ,methyl( R or S ), ethyl ( R or S )).
[0415] In some instances, the non-bicyclic modified sugar moiety contains more than one non-bridging sugar substituent, such as the 2'-F-5'-methyl sugar moiety, as in Migawa. et al. As described in US 2010 / 0190837, which is incorporated herein by reference, or as Rajeev... et al. The alternative sugar moieties modified with 2'- and 5'- as described in US 2013 / 0203836.
[0416] Some modified sugar moieties are bicyclic sugar moieties and contain substituents that bridge the two atoms of the furanyl ring to form a second ring. In some embodiments, the bicyclic sugar moieties contain a bridge between the 4' and 2' furanyl ring atoms. Examples of such 4' to 2' bridging sugar substituents include, but are not limited to: 4'-CH-2-2', 4'-(CH2)2-2', 4'-(CH2)3-2', 4'-CH2-O-2' (“LNA”), 4'-CH2-S-2', 4'-(CH2)2-O-2' (“ENA”), 4'-CH(CH3)-O-2' (when present as… SWhen configured, they are referred to as "restricted ethyl" or "cEt"), 4'-CH2-O-CH2-2', 4'-CH2-N(R)-2', 4'-CH(CH2OCH3)-O-2' ("restricted MOE" or "cMOE") and their analogues, 4'-C(CH3)(CH3)-O-2' and their analogues, 4'-CH2-N(OCH3)-2' and their analogues, 4'-CH2-ON(CH3)-2', 4'-CH2-C(H)(CH3)-2', 4'-CH2-C(=CH2)-2' and their analogues, 4'-C(R) a R b )-N(R)-O-2'、4'-C(R a R b )-ON(R)-2', 4'-CH2-ON(R)-2' and 4'-CH2-N(R)-O-2', wherein each R, R a and R b Independently, it is H, a protecting group, or C1-C. 12 Alkyl groups. Representative U.S. patents teaching the preparation of such bicyclic sugar moieties include, but are not limited to: Imanishi et al. US 7,427,672; Swayze et al. US 7,741,457; Swayze et al. US 8,022,193; Seth et al. US 8,278,283; Prakash et al. US 8,278,425; and Seth et al. US 8,278,426.
[0417] 2. Certain modified nucleobases In some embodiments, the modified oligonucleotide comprises one or more nucleosides containing unmodified nucleosides. In some embodiments, the modified oligonucleotide comprises one or more nucleosides containing modified nucleosides. In some embodiments, the modified oligonucleotide comprises one or more nucleosides that do not contain nucleosides, referred to as debased nucleosides. In some embodiments, the modified oligonucleotide comprises one or more inosine nucleosides (i.e., nucleosides containing hypoxanthine nucleosides). An "unmodified nucleobase" is adenine (A), thymine (T), cytosine (C), uracil (U), or guanine (G). A modified nucleobase is a group of atoms other than the unmodified A, T, C, U, or G that can pair with at least one other nucleobase. 5-Methylcytosine is an example of a modified nucleobase. A universal base is a modified nucleobase that can pair with any one of the five unmodified nucleosides.
[0418] In some embodiments, the modified nucleobase is selected from: 5-substituted pyrimidines, 6-azapyrimidines, alkyl or alkynyl-substituted pyrimidines, alkyl-substituted purines, and N-2, N-6, and O-6-substituted purines. In some embodiments, the modified nucleobases are selected from: 5-methylcytosine, 2-aminopropyladenine, 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-N-methylguanine, 6-N-methyladenine, 2-propyladenine, 2-thiouracil, 2-thiothymidine and 2-thiocytosine, 5-propynyl(C≡C-CH3)uracil, 5-propynylcytosine, 6-azauracil, 6-azacytosine, 6-azathymidine, 5-ribosyluracil (pseudoruracil), N1-methylpseudoruracil, 4-thiouracil, 8-halogenated, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxy 8-Aza and other 8-substituted purines, 5-halogenated (especially 5-bromo), 5-trifluoromethyl, 5-halogenated uracil and 5-halogenated cytosine, 7-methylguanine, 7-methyladenine, 2-F-adenine, 2-aminoadenine, 7-deadenine, 7-deadenine, 3-deadenine, 3-deadenine, 6-N-benzoyladenine, 2-N-isobutyrylguanine, 4-N-benzoylcytosine, 4-N-benzoyluracil, 5-methyl4-N-benzoylcytosine, 5-methyl4-N-benzoyluracil, universal bases, hydrophobic bases, hybrid bases, size-enlarged bases and fluorinated bases. Other modified nucleobases include tricyclic pyrimidines, such as 1,3-diazaphenoxazin-2-one, 1,3-diazaphenthiazin-2-one, and 9-(2-aminoethoxy)-1,3-diazaphenoxazin-2-one (G-clamp). Modified nucleobases may also include nucleobases in which purine or pyrimidine bases are replaced by other heterocycles (e.g., 7-deadenine, 7-deadenine, 2-aminopyridine, and 2-pyridone). Other nucleobases include those disclosed in the following literature: Englisch et al., Angewandte Chemie, International Edition, 1991, 30, 613; Sanghvi, YS, Chapter 15, Antisense Research and Applications, Crooke, ST and Lebleu, B., eds., CRC Press, 1993, 273-288; and those disclosed in Chapters 6 and 15 of Antisense Drug Technology, Crooke ST, ed., CRC Press, 2008, 163-166 and 442-443.
[0419] Publications teaching the preparation of certain modified nucleosides and other modified nucleosides are, but are not limited to, Rogers et al., US 5,134,066; Benner et al., US 5,432,272; Matteucci et al., US 5,502,177; Froehler et al., US 5,594,121; and Cook et al., US 5,681,941.
[0420] B. Certain modified nucleoside bonds Naturally occurring RNA and DNA nucleoside linkages are 3' to 5' phosphodiester bonds. In some embodiments, the nucleosides of modified oligonucleotides may be linked together using one or more modified nucleoside linkages. Two main classes of internucleotide linkage groups are defined by the presence or absence of a phosphorus atom. Representative phosphorus-containing nucleoside linkages include, but are not limited to, phosphodiester bonds containing phosphodiester bonds (also referred to as unmodified or naturally occurring bonds), phosphodiester bonds, methylphosphonates, phosphoramides, thiophosphates, phosphonoacetates (“PACE”), thiophosphonoacetates (“Thio-PACE”), and dithiophosphates. Representative phosphorus-free nucleoside linkage groups include, but are not limited to: methylene methylimino (-CH2-N(CH3)-O-CH2-), thiodiester bonds, thiocarbamates (-OC(=O)(NH)-S-); siloxanes (-O-SiH2-O-); and N,N'-dimethylhydrazine (-CH2-N(CH3)-N(CH3)-). Modified nucleoside bonds, compared to naturally occurring phosphodiester nucleoside bonds, can be used to alter, and typically increase, the nuclease resistance of oligonucleotides. In some embodiments, nucleoside bonds with chiral atoms can be prepared as racemic mixtures or as individual enantiomers. Methods for preparing phosphorus-containing and phosphorus-free nucleoside bonds are well known to those skilled in the art.
[0421] In some embodiments, the modified nucleoside interbond is any modified nucleoside interbond described in WO2021 / 030778, which is incorporated herein by reference. In some embodiments, the modified nucleoside interbond comprises the following: Independently, for each of these nucleoside linking groups of the modified oligonucleotide: X is selected from O or S; R1 is selected from H, C1-C6 alkyl groups, and substituted C1-C6 alkyl groups; and T is selected from SO2R2, C(=O)R3 and P(=O)R4R5, where: R2 is selected from aryl, substituted aryl, heterocycle, substituted heterocycle, aromatic heterocycle, substituted aromatic heterocycle, diazole, substituted diazole, C1-C6 alkoxy, C1-C6 alkyl, C1-C6 alkenyl, C1-C6 alkynyl, substituted C1-C6 alkyl, substituted C1-C6 alkenyl, substituted C1-C6 alkynyl, and linker; R3 is selected from aryl, substituted aryl, CH3, N(CH3)2, OCH3 and linker; R4 is selected from OCH3, OH, C1-C6 alkyl, substituted C1-C6 alkyl, and linker; and R5 is selected from OCH3, OH, C1-C6 alkyl, and substituted C1-C6 alkyl.
[0422] In some embodiments, the modified nucleoside interbond comprises a methanesulfonylaminophosphate linker group having the following formula: .
[0423] In some embodiments, the modified nucleoside interbonds are neutral nucleoside interbonds. Neutral nucleoside interbonds include, but are not limited to, triphosphates, methylphosphonates, MMI (3'-CH2-N(CH3)-O-5'), amide-3 (3'-CH2-C(=O)-N(H)-5'), amide-4 (3'-CH2-N(H)-C(=O)-5'), formaldehyde (3'-O-CH2-O-5'), methoxypropyl (MOP), and thioformaldehyde (3'-S-CH2-O-5'). Other neutral nucleoside interbonds include nonionic bonds, including siloxanes (dialkylsiloxanes), carboxylic esters, formamides, sulfides, sulfonates, and amides (see, for example: Carbohydrate Modifications in Antisense Research ;YSSanghvi and PDCook, eds., ACS Symposium Series 580; Chapters 3 and 4, 40–65). Other neutral nucleoside interbonds include nonionic bonds, comprising a mixture of N, O, S, and CH2 components.
[0424] In some embodiments, the modified nucleoside interbonds are standard-length nucleoside interbonds and are represented by the following formula Z1, Z2, or Z3. As used herein, a "standard-length nucleoside interbond" refers to a nucleoside interbond having a structure represented by the formula Z1, Z2, or Z3: Wherein, for each nucleoside linker of formula Z1, Z2, or Z3, the following is independent: Each X 1 Independently selected from O and S; X 2 Selected from O, NR1 CH2 and S; X 3 Selected from O, NR 1 CH2 and S; L does not exist, NR 1 、N(R 1 SO2, -N=, O, C1-C6 alkylene or C1-C6 heteroalkylene; Each R 1 Independently selected from H, C1-C6 alkyl and substituted C1-C6 alkyl, or two R on the same atom 1 Together they form = O; and R 2 Selected from -OH, -SH, Cl-C 22 Alkyl, substituted C1-C 22 Alkyl, C2-C 22 Alkenyl, substituted C2-C 22 Alkenyl, cycloalkyl, substituted cycloalkyl, heterocyclic, substituted heterocyclic, heteroaryl, substituted heteroaryl, aryl and substituted aryl; When the group is substituted, it contains one or more elements selected from halogens, -OH, -N(R) 1 )2、-O-C1-C6 alkyl、C1-C 22 Alkyl, C2-C 22 Substituents of alkenyl, cycloalkyl, heterocyclic, heteroaryl, and aryl groups.
[0425] In some embodiments, the extended internucleotide bond is an extended length of internucleotide bond. For example, in some instances, an extended internucleotide bond is formed when two functional groups on two oligonucleotides react to link the two oligonucleotides together to form a single oligonucleotide containing an extended internucleotide bond. One such reaction is the click reaction between bicyclic [6.1.0]nonyne and an azide. Additional linkers suitable for linking two oligonucleotides via click chemistry are described in "Click Chemistry for Biotechnology and Materials Science" edited by Joerg Lahann, Wiley 2009. Other examples of linker chemistry include the reverse electron-demanding Diels-Alder reaction. For example For example in Argamunt et al. , J. Org. Chem. 2020, 85 , 10, 6593–6604, Sarrett et al. , Nat. Protocols 2021, 16, 3348–3381; Handula et al. , Molecule s, 2021, 26 (15), 4640, Wiessler et al. , Int. J. Med.Sci. 2010, 7 (1), 19–28; Copper-catalyzed azide-alkyne cycloaddition (CuAAC) see, for example, SI Presolski et al. , J. Am. Chem. Soc. 2010, 132, 14570–14576; D.Soriano Del Amo et al. , J. Am. Chem. Soc. , 2010, 132 , 16893–16899; Staudinger reaction, see For example Saxon and CR Bertozzi, Science , 2000, 287 , 2007–2010; BL Nilsson et al. , Org.Lett. , 2000, 2 E. Saxon, 1939–1941 et al. , Org.Lett. , 2000, 2 ,2141–2143; For the formation of hydrazones and oximes, see 2141–2143. For example JY Axup et al. , Proc.Natl.Acad.Sci.USA ,2012, 109 , 16101–16106; photoclick reaction, see For example W. Song et al. , Angew.Chem., Int. Ed. ,2008, 47 , 2832–2835, A. Herner and Q. Lin, Top.Curr.Chem. , 2016, 374 1; Strain-promoted alkyne-nitroketone cycloaddition (SPANC) reaction, see [reference needed]. For example DA MacKenzie et al. , Curr.Opin.Chem.Biol. , 2014, 21 , 81–88; Transition metal-catalyzed cross-coupling reactions, see For example M. Chalker et al. , J. Am. Chem. Soc. , 2009, 131, 16346–16347; Nucleophilic addition reactions, particularly the addition of thiols to maleimides, see [reference needed]. For example Kang et al. , Chem.Sci. , 2021, 12 , 13613-13647, Bernardim et al. , Nat. Comm. 2016, 7 , 13128, Jain et al. , Pharm . Res. 2015, 32 (11), 3526-3540.
[0426] C. Certain motifs In some embodiments, the guide (modified oligonucleotide) comprises one or more modified nucleosides comprising a modified sugar moiety. In some embodiments, the modified oligonucleotide comprises one or more modified nucleosides comprising a modified nucleotide. In some embodiments, the modified oligonucleotide comprises one or more modified nucleotide interbonds. In such embodiments, the modified, unmodified, and differently modified sugar moiety, nucleotide, and / or nucleotide interbond of the modified oligonucleotide define a pattern or motif. In some embodiments, the patterns of the sugar moiety, nucleotide, and nucleotide interbond are each independent of each other. Therefore, the modified oligonucleotide can be described by its sugar motif, nucleotide motif, and / or nucleotide interbond motif (as used herein, the nucleotide motif describes modifications to the nucleotide sequence independent of the nucleotide sequence).
[0427] 1. Certain glycosylations In some embodiments, the oligonucleotide comprises one or more types of modified and / or unmodified sugar motifs arranged along the oligonucleotide or its regions in a defined pattern or glycomolecular motif. In some cases, such glycomolecular motifs include, but are not limited to, any of the sugar modifications discussed herein. Typically, the guide comprises an unmodified RNA nucleoside and optionally one or more modified nucleosides as described herein. In some embodiments, the guide comprises or consists of modified oligonucleotides.
[0428] In some embodiments, the 5' nucleotide of the modified oligonucleotide comprises a modified sugar moiety. In some embodiments, the 3' nucleotide of the modified oligonucleotide comprises a modified sugar moiety. In some embodiments, at least 1, 2, 3, 4, or 5 of the 5' nucleotides of the modified oligonucleotide comprise a modified sugar moiety. In some embodiments, at least 1, 2, 3, 4, or 5 of the 3' nucleotides of the modified oligonucleotide comprise a modified sugar moiety.
[0429] In some embodiments, the above modifications (sugars, nucleotides, nucleoside bonds) are incorporated into the modified oligonucleotide. In some embodiments, the modified oligonucleotide is characterized as described herein with its modified motif and total length, and such parameters are independent of each other. Unless otherwise stated, all modifications are independent of the nucleotide sequence. In some embodiments, the modified oligonucleotide is a guide. In some embodiments, the guide has a sugar motif, nucleobase, and / or nucleoside intergrowth motif as described in any of the following references, each of which is incorporated herein by reference: WO 2014 / 144761, WO 2015 / 026885, WO 2016 / 089433, WO 2016 / 100951, WO 2016 / 123230, WO 2016 / 164356, WO2017 / 004261, WO 2017 / 004279, WO 2017 / 068377, WO 2017 / 136794, WO 2017 / 181107, WO2017 / 214460, WO 2018 / 009822, WO 2018 / 057946, WO 2018 / 098383、WO 2018 / 107028、WO2018 / 125964、WO 2019 / 084664、WO 2019 / 147275、WO 2019 / 147743、WO 2019 / 183000、WO2019 / 237069、WO 2021 / 119006, WO 2021 / 119275, WO 2021 / 125840, WO 2021 / 207651, WO2021 / 207711, WO 2022 / 086846.
[0430] 2. Certain nucleobase motifs In some embodiments, the oligonucleotide comprises modified and / or unmodified nucleobases arranged in a defined pattern or motif along the oligonucleotide or its region. In some embodiments, each nucleobase is modified. In some embodiments, none of the nucleobases are modified. In some embodiments, each purine or each pyrimidine is modified. In some embodiments, each adenine is modified. In some embodiments, each guanine is modified. In some embodiments, each thymine is modified. In some embodiments, each uracil is modified. In some embodiments, each cytosine is modified. In some embodiments, some or all of the cytosine nucleobases in the modified oligonucleotide are 5-methylcytosine. In some embodiments, all of the cytosine nucleobases are 5-methylcytosine, and all other nucleobases in the modified oligonucleotide are unmodified nucleobases.
[0431] In some embodiments, the modified oligonucleotide comprises a block of modified nucleobases. In some such embodiments, the block is located at the 3' end of the oligonucleotide. In some embodiments, the block is located within three nucleosides at the 3' end of the oligonucleotide. In some embodiments, the block is located at the 5' end of the oligonucleotide. In some embodiments, the block is located within three nucleosides at the 5' end of the oligonucleotide.
[0432] In some embodiments, the modified nucleobases are selected from pseudouracil, 2-6-diaminopurine, 2-thiouracil, 4-thiouracil, 2-aminoadenine, 6-methyladenine, hypoxanthine, or 5-methylcytosine.
[0433] 3. Certain nucleoside internucleotide motifs In some embodiments, the oligonucleotide comprises modified and / or unmodified internucleotide links arranged in a defined pattern or motif along the oligonucleotide or its regions. In some embodiments, each internucleotide linking group is a phosphodiester internucleotide linker (P=O). In some embodiments, each internucleotide linking group of the modified oligonucleotide is a phosphate thioester internucleotide linker (P=S). In some embodiments, each internucleotide linker of the modified oligonucleotide is independently selected from phosphate thioester internucleotide links and phosphodiester internucleotide links. In some embodiments, each internucleotide linker of the modified oligonucleotide is independently selected from phosphate thioester internucleotide links, phosphodiester internucleotide links, and phosphonoacetate internucleotide links.
[0434] In some embodiments, the 5' internucleotide bonds of the modified oligonucleotide are modified internucleotide bonds. In some embodiments, the 3' internucleotide bonds of the modified oligonucleotide are modified internucleotide bonds. In some embodiments, at least 1, 2, 3, 4, or 5 of the 5' internucleotide bonds of the modified oligonucleotide are modified internucleotide bonds. In some embodiments, at least 1, 2, 3, 4, or 5 of the 3' internucleotide bonds of the modified oligonucleotide are modified internucleotide bonds. In some embodiments, the modified internucleotide bonds are phosphate thioester internucleotide bonds.
[0435] 4. Certain guide chemical modification motifs The guide may contain any modifications commonly applied to the modified oligonucleotides described herein, including modified sugar moieties, modified nucleotide internucleotides, and modified nucleobases. In some embodiments, the guide is modified at the 3' end, the 5' end, or both the 3' and 5' ends. In some embodiments, the three nucleotides at the 3' end and the three nucleotides at the 5' end are 2'-OMe nucleotides, and the remainder of the guide's nucleotides is unmodified RNA nucleotide. In some embodiments, each nucleotide within the 5' stem-loop is an unmodified RNA nucleotide. In some embodiments, In some embodiments, each internucleotide bond in the guide is an unmodified phosphodiester bond. In some embodiments, the guide contains one or more modified internucleotide bonds. In some embodiments, the guide contains one or more thiophosphate internucleotide bonds. In some embodiments, one, two, three, four, or five internucleotide bonds at point 3' are modified internucleotide bonds. In some embodiments, one, two, three, four, or five internucleotide bonds at point 5' are modified internucleotide bonds. In some embodiments, one, two, three, four, or five internucleotide bonds at points 3' and 5' are modified internucleotide bonds, and the remainder of these bonds is an unmodified phosphodiester internucleotide bond. In some embodiments, one, two, three, four, or five internucleotide bonds at point 3' are thiophosphate internucleotide bonds. In some embodiments, one, two, three, four, or five internucleotide bonds at point 5' are thiophosphate internucleotide bonds. In some embodiments, one, two, three, four, or five nucleoside interbonds at the 3' and 5' ends are phosphate thioester nucleoside interbonds, and the remainder of these bonds are unmodified phosphodiester nucleoside interbonds. In some embodiments, three nucleoside interbonds at the 3' end and three nucleoside interbonds at the 5' end are phosphate thioester nucleoside interbonds, and the remainder of these bonds are unmodified phosphodiester nucleoside interbonds.
[0436] D. certain length In some embodiments, the oligonucleotide (including the guide) may have any of a variety of length ranges. In some embodiments, the oligonucleotide consists of X to Y linked nucleosides, where X represents the minimum number of nucleosides in the range and Y represents the maximum number of nucleosides in the range. In some such embodiments, X and Y are each independently selected from 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69. 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 11 2, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 14 7, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180; provided that X≤Y. For example, in some instances, the oligonucleotide consists of 50 to 180, 80 to 180, 90 to 180, 100 to 180, 110 to 180, 120 to 180, 130 to 180, 140 to 180, 150 to 180, 160 to 180, 160 to 170, 90 to 100, 100 to 110, 110 to 120, 120 to 130, 130 to 140, 140 to 150, 100 to 140, 110 to 140, 110 to 130, 120 to 140, or 120 to 150 linked nucleosides.
[0437] In some embodiments, the unidirectional oligonucleotides consist of 50 to 180, 80 to 180, 90 to 180, 100 to 180, 110 to 180, 120 to 180, 130 to 180, 140 to 180, 150 to 180, 160 to 180, 160 to 170, 90 to 100, 100 to 110, 110 to 120, 120 to 130, 130 to 140, 140 to 150, 100 to 140, 110 to 140, 110 to 130, 120 to 140, or 120 to 150 linked nucleosides. In some embodiments, the target recognition region of the guide is composed of 15 to 30, 18 to 26, 18 to 24, 18 to 22, 20 to 26, 20 to 24, 20 to 22, 21 to 23, 22 to 23, 20, 21, 22, or 23 linked nucleosides.
[0438] In some embodiments, the protein recognition region of the guide comprises 60 to 130, 60 to 120, 60 to 110, 60 to 100, 70 to 130, 70 to 120, 70 to 110, 70 to 100, 70 to 90, 70 to 80, 80 to 130, 80 to 120, 80 to 110, 80 to 100, 80 to 90, 90 to 130, 90 to 120, 90 to 110, 90 to 100, 90 to 100, It consists of 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, or 110 linked nucleosides. In some embodiments, the protein recognition region of the guide represents a truncated form of the native guide corresponding to the Cas protein. In some embodiments, the guide is truncated at the 3' end. In some embodiments, the guide is shortened by removing nucleotides from the stem-loop and / or hairpin structure, such as removing a base pair from the hairpin structure to shorten the hairpin structure. In some embodiments, the truncated version is at least 5, 10, 15, or 20 nucleotides shorter than the natural guide. In some embodiments, the protein recognition region of the guide represents a longer version of the natural guide corresponding to the Cas protein. In some embodiments, an additional nucleotide is added at the 3' end. In some embodiments, an additional nucleotide is inserted within the loop and / or hairpin structure of the natural guide.
[0439] E. Nucleobase sequence In some embodiments, the oligonucleotide (unmodified or modified) is further described by its nucleotide sequence. In some embodiments, the oligonucleotide has a nucleotide sequence complementary to a second oligonucleotide or an identified reference nucleic acid (such as a target nucleic acid). In some such embodiments, a region of the oligonucleotide has a nucleotide sequence complementary to a second oligonucleotide or an identified reference nucleic acid (such as a target nucleic acid). In some embodiments, the nucleotide sequence of a region or the entire length of the oligonucleotide is at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% complementary to a second oligonucleotide or nucleic acid (such as a target nucleic acid).
[0440] F. Secondary structure In some embodiments, the oligonucleotide or a portion thereof adopts a defined secondary structure. In some embodiments, the secondary structure is determined by Watson-Crick base pairing, Hoogsteen base pairing, and / or non-canonical base pairing interactions. The secondary structure of the oligonucleotide sequence can be predicted using standard software, such as the Vienna RNA package RNAfold (http: / / rna.tbi.univie.ac.at / cgi-bin / RNAWebSuite / RNAfold.cgi; see Lorenz et al.). Algorithms for Molecular Biology (6:1 26, 2011). In some embodiments, the oligonucleotide may have more than one predicted secondary structure.
[0441] Guide oligonucleotides possess secondary structures, at least a portion of which is important for forming a contact with their corresponding Cas proteins. For example, the secondary structure features of guide oligonucleotides for spCas9 binding have been established based on empirical evidence (see, for example, Zhang et al.). ChemPlusChem , 2021 and Dong et al., Curr. Opinion in Biotech. (2022) and the crystal structure of the guide-SpCas9-target DNA complex (see Nishimasu et al., 2022). Cell Both (2014) provide detailed descriptions. By adding a four-loop to the double-stranded structure formed between the native crRNA and native tracrRNA of SpCas9, the functions of crRNA and tracrRNA are combined into a single-guide RNA. The same four-loop can be added to crRNA and tracrRNA linking novel type II Cas proteins. The single-guide RNA for SpCas9 has conserved structural features within its protein-binding region, including a lower stem, an upper stem, a four-loop, a protrusion, a linker region (a hairpin structure with 5 nucleotide loops), and two hairpin structures at the 3' end. Figure 6Guide RNAs of certain novel type II Cas proteins have been shown to possess unique secondary structures, sharing common features in their protein-binding regions, including double-stranded regions (double guides) or stem-loops (single guides) and one or more hairpin structures at the 3' end. The 5' stem-loop feature of the guide RNA is also known as the repeat-anti-repeat region (R-AR).
[0442] In some embodiments, the guide is composed of oligonucleotides that can be represented by the following formula: T1-DXS 1a -B1-S 1b -H1-S 1b '-B2-S 1a '-L1-S2-H2-S2'-L2-S3-H3-S3'-(L3-S4-H4-S4') n -(L4-S5-H5-S5') m -T2; in: D consists of 15 to 25 linked nucleosides, wherein the nucleobase sequence of D is complementary to the nucleobase sequence of the target DNA; X is absent, or consists of one or two linked nucleosides that are not complementary to the sequence of the target DNA; Each of T1 and T2 either lacks or is independently composed of 1 to 30 linked nucleosides; S 1a and S 1a Each is composed of 6 to 12 linked nucleosides, of which S 1a The nucleobase sequence and S 1a The nucleobase sequences of ' are 100% complementary; S 1b and S 1b Each is composed of 2 to 14 linked nucleosides, of which S 1b The nucleobase sequence and S 1b The nucleobase sequences of ' are 100% complementary; S2 and S2' are each composed of 2 to 6 linked nucleosides, wherein the nucleobase sequence of S2 is 100% complementary to the nucleobase sequence of S2'; S3 and S3' are each composed of 2 to 12 linked nucleosides, wherein the nucleobase sequence of S3 is 100% complementary to the nucleobase sequence of S3'; S4 and S4' are each composed of 4 to 14 linked nucleosides, wherein the nucleobase sequence of S4 is 100% complementary to the nucleobase sequence of S4'. S5 and S5' are each composed of 4 to 11 linked nucleosides, wherein the nucleobase sequence of S5 is 100% complementary to the nucleobase sequence of S5'. n is 0 or 1; m is 0 or 1; B1 consists of 0 to 3 linked nucleosides; B2 consists of 0 to 4 linked nucleosides; B1 is not equal to B2 unless both are composed of 0 linked nucleosides; Each of H1, H2, H3, H4, and H5 is independently composed of 3 to 6 linked nucleosides; L1 consists of 2 to 3 linked nucleosides; Each of L2, L3, and L4 is independently composed of 0 to 10 linked nucleosides.
[0443] In some embodiments, a guide specifically targeting the ION-CAS protein is represented by the following formula: T1-DXS 1a -B1-S 1b -H1-S 1b '-B2-S 1a '-L1-S2-H2-S2'-L2-S3-H3-S3'-(L3-S4-H4-S4') n -(L4-S5-H5-S5') m -T2; Each of these variables is defined in the table below. "x" indicates that a specific nucleoside or region does not exist.
[0444] Secondary structural features of the guide responsible for binding to the Cas enzyme include at least one double-stranded region (or one or more stem loops for single guides) and two or more additional stem loops and / or hairpin regions at the 5' end of the protein recognition region. The protrusion in the 5' stem loop is conserved in guides of most type II Cas proteins; however, guides with a variety of predicted secondary structures have also been reported (see, for example, Alexander et al.). CRISPR J., 6(3): 261-277, 2023).
[0445] In some embodiments, the activity of the guide is not driven by a specific nucleobase sequence, but by the formation of a specific RNA secondary structure. The sequence that produces this secondary structure can be altered. A common method of altering the sequence while maintaining the three-dimensional structure required for the guide to bind to its corresponding Cas is to flip base pairs within the stem or hairpin structure, such as an "A to U" flip. In such embodiments, the native A to U base pair is replaced with a U to A base pair, swapping the positions of "A" and "U" in the guide sequence. This can be particularly effective when the original guide contains four consecutive Us (the presumed pol-III terminator) (see, for example, Chen et al. Cell. (2013); 155(7): 1479-1491). Another method of altering the sequence within the stem or hairpin structure involves replacing less stable A to U or U to A base pairs with more stable C to G or G to C base pairs to increase the overall hairpin structure stability (see, for example, Riesenberg et al. Nat. Commun. , 2022).
[0446] In some embodiments, the natural hairpin or stem-loop structure of the guide molecule is elongated or replaced with an elongated stem-loop. In some cases, stem elongation has been shown to enhance the assembly of the guide with CRISPR-Cas proteins (Chen et al., Cell. (2013); 155(7): 1479-1491). In some embodiments, the stem of the stem-loop is elongated by at least 1, 2, 3, 4, 5, or more complementary base pairs (i.e., corresponding to the addition of 2, 4, 6, 8, 10, or more nucleotides to the guide). In some embodiments, these are located at the end of the stem, adjacent to the stem-loop. In some embodiments, the natural hairpin or stem-loop structure of the guide is shortened. In some such embodiments, the stem of the stem-loop is shortened by 1, 2, or 3 complementary base pairs (i.e., corresponding to the removal of 2, 4, or 6 nucleotides from the guide).
[0447] It has also been demonstrated that, for the SpCas9 guide, the loop region of the upper stem or hairpin structure 1 can be replaced with a longer loop containing the aptamer. In some embodiments, the aptamer binds to effector proteins. In some embodiments, the aptamer binds to MS2, PP7, com, or boxB. In some embodiments, the aptamer can be attached to the 3' end of the guide, or can be inserted into the guide as a new hairpin structure outside the 3' hairpin structure of the natural guide (see, for example, Zalatan et al.). Cell (2015).
[0448] In some embodiments, the guide may be modified so that it is activated only by an external chemical signal. In one such embodiment, the guide includes a 5' extension attached to the 5' end of the target recognition region of the guide, forming a hairpin double strand with the target recognition region. This 5' extension also includes a theophylline aptamer. In this system, the addition of theophylline causes the cleavage of the 5' extension and exposes the target recognition region of the guide. In some embodiments, the guide may be modified so that it can be inactivated by an external chemical signal. In one such embodiment, the guide includes a theophylline aptamer incorporated into the ring region of a first stem ring and a second theophylline aptamer incorporated into the ring of a second hairpin structure. The guide is adapted for Cas9 cleavage of the target, but is cleaved when theophylline is added to the system. Various other external control systems for the guide have been described (see, for example, Zhang et al.). ChemPlusChem , 2021).
[0449] C. Guided chemical modification The guide may contain any modifications commonly applied to the modified oligonucleotides described herein, including modified sugar moieties, modified nucleotide internucleotides, and modified nucleobases. In some embodiments, the guide is modified at the 3' end, the 5' end, or both the 3' and 5' ends. In some embodiments, the three nucleotides at the 3' end and the three nucleotides at the 5' end are 2'-OMe nucleotides, and the remainder of the guide's nucleotides is unmodified RNA nucleotide. In some embodiments, each nucleotide within the 5' stem-loop is an unmodified RNA nucleotide. In some embodiments, In some embodiments, each internucleotide bond in the guide is an unmodified phosphodiester bond. In some embodiments, the guide contains one or more modified internucleotide bonds. In some embodiments, the guide contains one or more thiophosphate internucleotide bonds. In some embodiments, one, two, three, four, or five internucleotide bonds at point 3' are modified internucleotide bonds. In some embodiments, one, two, three, four, or five internucleotide bonds at point 5' are modified internucleotide bonds. In some embodiments, one, two, three, four, or five internucleotide bonds at points 3' and 5' are modified internucleotide bonds, and the remainder of these bonds is an unmodified phosphodiester internucleotide bond. In some embodiments, one, two, three, four, or five internucleotide bonds at point 3' are thiophosphate internucleotide bonds. In some embodiments, one, two, three, four, or five internucleotide bonds at point 5' are thiophosphate internucleotide bonds. In some embodiments, one, two, three, four, or five nucleoside interbonds at the 3' and 5' ends are phosphate thioester nucleoside interbonds, and the remainder of these bonds are unmodified phosphodiester nucleoside interbonds. In some embodiments, three nucleoside interbonds at the 3' end and three nucleoside interbonds at the 5' end are phosphate thioester nucleoside interbonds, and the remainder of these bonds are unmodified phosphodiester nucleoside interbonds.
[0450] III. Some editing systems In some embodiments, the editing system comprises a guided nucleic acid binding agent and a guide. In some embodiments, the guide consists of a single oligonucleotide and is a single guide. In some embodiments, the guide consists of two oligonucleotides and is a dual guide. In some embodiments, the guide comprises a target recognition region and a protein recognition region. In some embodiments, the guide is a single guide, consisting of an oligonucleotide comprising a target recognition region and a protein recognition region. In some embodiments, the guide is a dual guide, wherein a first oligonucleotide comprises a first component of the target recognition region and the protein recognition region, and a second oligonucleotide comprises a second component of the protein recognition region. In some such embodiments, the first oligonucleotide is “crRNA”, and the second oligonucleotide is “tracrRNA”. In some embodiments, the target recognition region is a region of an oligonucleotide having a sequence complementary to an equal-length portion of the target sequence within the complementary strand of the target DNA. In some embodiments, the editing system comprises a guided nucleic acid binder and a guide. In some embodiments, the guided nucleic acid binder comprises a Cas protein. In some embodiments, the guided nucleic acid binder is composed of a Cas protein. In some embodiments, the Cas protein has an amino acid sequence of any one of SEQ ID NO: 4 to 23 or 231 to 323. In some embodiments, the Cas protein is encoded by DNA having a nucleobase sequence of any one of SEQ ID NO: 24 to 43 or 324 to 416. In some embodiments, the guide is an oligonucleotide. In some embodiments, the guide is composed of two oligonucleotides. In some embodiments, the guide comprises a protein recognition region having a nucleobase sequence selected from any one of SEQ ID NO: 121 to 141, 510 to 602, 961 to 1408, 2825, 2827 to 2833. In some embodiments, the editing system comprises a guided nucleic acid agent comprising a guided nucleic acid protein and a guide, wherein the guide's protein recognition region has a nucleobase sequence that binds to the guided nucleic acid protein. In some embodiments, the guided nucleic acid agent comprises an ION-Cas protein. In some embodiments, the editing system comprises an ION-Cas protein and a guide, the ION-Cas protein having the amino acid sequence of SEQ ID NO: 12, and the guide comprising a protein recognition region having a nucleobase sequence of SEQ ID NO: 129 or 130. In some embodiments, the editing system comprises an ION-Cas protein and a guide, the ION-Cas protein having the amino acid sequence of SEQ ID NO: 13, and the guide comprising a protein recognition region having a nucleobase sequence of SEQ ID NO: 131, 961 to 964, 1016 to 1020, 1052 to 1070, 1313 to 1314, or 2825. In some embodiments, the editing system comprises an ION-Cas protein and a guide, the ION-Cas protein having the amino acid sequence of SEQ ID NO: 20, and the guide comprising a protein recognition region having the nucleotide sequence of SEQ ID NO: 139. In some embodiments, the editing system comprises an ION-Cas protein and a guide, the ION-Cas protein having the amino acid sequence of SEQ ID NO: 241, and the guide comprising a protein recognition region having the nucleotide sequence of SEQ ID NO: 520, 965 to 968, 1071 to 1094, or 1315 to 1318.In some embodiments, the editing system comprises an ION-Cas protein and a guide, the ION-Cas protein having the amino acid sequence of SEQ ID NO: 267, and the guide comprising a protein recognition region having a nucleotide sequence of SEQ ID NO: 546, 973 to 976, 1021 to 1023, 1095 to 1126, or 1319 to 1342. In some embodiments, the editing system comprises an ION-Cas protein and a guide, the ION-Cas protein having the amino acid sequence of SEQ ID NO: 275, and the guide comprising a protein recognition region having a nucleotide sequence of SEQ ID NO: 554, 981 to 984, 1024 to 1026, 1127 to 1152, or 1343 to 1361. In some embodiments, the editing system comprises an ION-Cas protein and a guide, the ION-Cas protein having the amino acid sequence of SEQ ID NO: 284, and the guide comprising a protein recognition region having a nucleotide sequence of SEQ ID NO: 562, 985 to 989, 1027 to 1029, 1153 to 1182, or 1362 to 1390. In some embodiments, the editing system comprises an ION-Cas protein and a guide, the ION-Cas protein having the amino acid sequence of SEQ ID NO: 289, and the guide comprising a protein recognition region having a nucleotide sequence of SEQ ID NO: 568, 990 to 993, 1030 to 1034, 1183 to 1206, or 1392 to 1393. In some embodiments, the editing system includes an ION-Cas protein and a guide, the ION-Cas protein having the amino acid sequence of SEQ ID NO: 314, and the guide including a protein recognition region having the nucleobase sequence of SEQ ID NO: 593, 1006 to 1010, 1272 to 1292, 1394 to 1408, 2827 to 2833.
[0451] In some embodiments, the editing system includes an ION-Cas protein and a guide, the ION-Cas protein having an amino acid sequence that is at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 12, and the guide including a protein recognition region having a nucleobase sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 129 or 130. In some embodiments, the editing system includes an ION-Cas protein and a guide, the ION-Cas protein having an amino acid sequence that is at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 13, and the guide including a protein recognition region having a nucleobase sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 131, 961 to 964, 1016 to 1020, 1052 to 1070, 1313 to 1314, or 2825. In some embodiments, the editing system includes an ION-Cas protein and a guide, the ION-Cas protein having an amino acid sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO: 20, and the guide including a protein recognition region having a nucleobase sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO: 141.
[0452] In some embodiments, the editing system includes an ION-Cas protein and a guide, the ION-Cas protein having an amino acid sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO: 241, and the guide including a protein recognition region having a nucleobase sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with any of SEQ ID NO: 520, 965 to 968, 1071 to 1094, or 1315 to 1318. In some embodiments, the editing system includes an ION-Cas protein and a guide, the ION-Cas protein having an amino acid sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO: 267, and the guide including a protein recognition region having a nucleobase sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with any of SEQ ID NO: 546, 973 to 976, 1021 to 1023, 1095 to 1126, or 1319 to 1342. In some embodiments, the editing system includes an ION-Cas protein and a guide, the ION-Cas protein having an amino acid sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO: 275, and the guide including a protein recognition region having a nucleobase sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with any of SEQ ID NO: 554, 981 to 984, 1024 to 1026, 1127 to 1152, or 1343 to 1361.In some embodiments, the editing system includes an ION-Cas protein and a guide, the ION-Cas protein having an amino acid sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO: 284, and the guide including a protein recognition region having a nucleobase sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with any of SEQ ID NO: 562, 985 to 989, 1027 to 1029, 1153 to 1182, or 1362 to 1390. In some embodiments, the editing system includes an ION-Cas protein and a guide, the ION-Cas protein having an amino acid sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with SEQ ID NO: 289, and the guide including a protein recognition region having a nucleobase sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with any of SEQ ID NO: 568, 990 to 993, 1030 to 1034, 1183 to 1206, or 1392 to 1393. In some embodiments, the editing system includes an ION-Cas protein and a guide, the ION-Cas protein having an amino acid sequence that is at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 314, and the guide including a protein recognition region having a nucleobase sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to any one of SEQ ID NO: 593, 1006 to 1010, 1272 to 1292, 1394 to 1408, 2827 to 2833.
[0453] IV. Some delivery carriers A. Lipid nanoparticles ("LNP") 1. Composition In some embodiments, the LNP comprises a protonable or cationic lipid, a neutral or zwitterionic lipid, a polymeric lipid, cholesterol, or a derivative thereof, and optionally one or more excipients. The LNP at least partially encapsulates the load. In some embodiments, the load comprises a nucleic acid, such as exogenous mRNA. In some embodiments, the load comprises a guide, such as an oligomer. In some embodiments, the load comprises exogenous mRNA and a guide. In some embodiments, the LNP is provided as a suspension in an aqueous medium. In some embodiments, the pharmaceutical composition comprises an LNP at least partially encapsulated in an aqueous medium.
[0454] In some embodiments, the LNP comprises 30 to 70 mol% of protonable or cationic lipids; 5 to 30 mol% of neutral or zwitterionic lipids; 20 to 50 mol% of cholesterol or a derivative thereof; and 1 to 10 mol% of polymer-lipids. In some embodiments, the LNP comprises 40 to 60 mol% of protonable or cationic lipids; 5 to 15 mol% of neutral or zwitterionic lipids; 30 to 50 mol% of cholesterol or a derivative thereof; and 2 to 4 mol% of polymer-lipids. The molar percentage of each of the protonable or cationic lipids, neutral or zwitterionic lipids, polymer-lipids, cholesterol or a derivative thereof is determined together (“total lipids”), regardless of the loading or excipients.
[0455] LNPs can be characterized as the molar or mass ratio of payload (such as exogenous mRNA or oligomers) to total lipids. In some embodiments, the lipid to payload ratio is 20:1 to 1:1, 10:1 to 1:1, or 5:1 to 1:1 (mass:mass).
[0456] LNP compositions may contain one or more additional excipients, such as triglycerides, surfactants, and / or hydrophobic excipients such as waxes. Triglycerides include triglycerides of myristate (Dynasan 114), triglycerides of palmitate (Dynasan 116), or triglycerides of stearate (Dynasan 118), Witeposol matrix, glyceryl stearate (Imwitor 900), glyceryl behenate (Compritol 888 ATO), and glyceryl palmitate (Precirol ATO 5). Nonionic surfactants may include fractions such as ethylene glycol esters, propylene glycol esters, glyceryl esters, polyglycerol esters, sorbitan esters, sucrose esters, and ethoxylated esters, nonionic alkanolamides, and ethers such as fatty alcohol ethoxylates, propoxylated alcohols, and ethoxylated / propoxylated block polymers. Anionic surfactants include carboxylates, acyl lactyl lactates, acyl amides of amino acids, sulfates such as alkyl sulfates and ethoxylated alkyl sulfates, sulfonates such as alkylbenzene sulfonates, acyl hydroxyethyl sulfonates, acyl taurate and sulfosuccinate, and phosphates. Specific surfactants include lecithin, poloxamer 188, poloxamer 407, tyloxacillin, polysorbate 20, polysorbate 60, polysorbate 80, sodium cholate, sodium glycocholate, sodium taurine deoxycholate, butanol, butyric acid, cetylpyridinium chloride, sodium lauryl sulfate, sodium oleate, polyvinyl alcohol, or Cremophor EL. Waxes include beeswax and cetyl palmitate. Other hydrophobic excipients include stearic acid, palmitic acid, behenic acid, Miglyol 812, and paraffin.
[0457] a. Protonable and cationic lipids In some embodiments, LNPs comprising protonable and / or cationic lipids are provided. As used herein, a “protonable lipid” is a lipid-containing compound that is substantially protonated at or below physiological pH (e.g., pH 7–7.5 or pH 7.4). As used herein, a “cationic lipid” is a lipid-containing compound carrying a non-exchangeable net positive charge. Lipids are typically hydrophobic moieties and may contain hydrocarbon chains.
[0458] It is believed that the design of protonable and / or cationic lipids can enhance endosomal drug escape. Therefore, it is thought that in vivo, the delivery capability of LNPs can benefit from electrostatic charge interactions with endosomals and subsequent LNP destabilization to release the payload. Protonable lipids are thought to facilitate delivery by either (i) protonation in a weakly acidic environment to acquire a positive charge, and / or (ii) incorporation of pH-indestabilizing groups that cleave or alter conformation in low pH environments. Accordingly, such pH-sensitive lipids can help enhance charge-based uptake by liposomes in target cells or trigger payload release by destabilizing the liposome membrane. Cationic lipids can be characterized in anticubic or hexagonal liquid crystal phases.
[0459] Fusion property describes the ability of a lipid or multi-lipid construct to bind to a lipid layer. See, for example, Hashiba et al., Small Science, 3(1), 2023, 2200071; V. Gyanani and R. Goswami, Pharmaceutics, 2023, 15(4): 1184. Lipid designs considered to confer fusion properties include (i) unsaturation in the lipid tail, (ii) branching in the lipid carbon chain, (iii) multi-tailed lipids, such as 7C1 and G0-C14 lipids, see, for example, Dahlman et al., Nat Nanotechnol. Aug. 2014; 9(8): 648–655, and (iv) polymer-lipid incorporation.
[0460] It is believed that protonable or cationic lipids containing unsaturated hydrocarbon chains can provide LNPs with higher fluidity. Exemplary publications describing protonable and / or cationic lipids include U.S. Patent Publications Nos. 2006 / 0083780 and 2006 / 0240554; U.S. Patent Nos. 5,208,036; 5,264,618; 5,279,833; 5,283,185; 5,753,613; and 5,785,992; and PCT Publication No. WO 96 / 10390, the contents of which are incorporated herein by reference in their entirety.
[0461] Exemplary protonable and / or cationic lipids include a protonable or cationic head group, optionally comprising one or more straight-chain hydrocarbon chains of 10 to 20 carbon atoms, and optionally one or more branched, optionally unsaturated hydrocarbon chains of 12 to 30 carbon atoms, each optionally interrupted by one or more functional groups selected from esters, ethers, amines (e.g., tertiary amines), amides, carbonates, carbamates, ureas, and disulfides. Protonable or cationic lipids may contain a branched portion. In some embodiments, the branched portion comprises a tertiary carbon atom, a quaternary carbon atom, a tertiary amine, a vicinal diol, an amide, a carbamate, or an acetal.
[0462] In some instances, protonable or cationic lipids are protonable lipids containing amines (“protonable amino lipids”), such as tertiary amines. It is believed that the proton cycling of the amino group at different pH levels can provide additional stabilization for negatively charged loads (e.g., exogenous mRNA) while promoting load release within low-pH compartments in vivo (e.g., endosomes). In some embodiments, protonable lipids have pKa of protonable (e.g., amino) groups in the range of about 4 to about 7. Such lipids will primarily neutralize the surface at physiological pH. This is believed to reduce the sensitivity to clearance. pKa measurements of lipids within lipid particles can be performed, for example, using the fluorescent probe 2-(p-toluidine)-6-naphthalenesulfonic acid (TNS), using the method described in Cullis et al., (1986) Chem Phys Lipids 40, 127-144.
[0463] In some embodiments, the protonable or cationic lipids are DLin-MC3-DMA, ALC-0315, SM-102, or LP-01: DLin-MC3-DMA ALC-0315 SM-102, LP-01.
[0464] b. Polymer-lipid In some embodiments, a polymer-lipid LNP is provided. The polymer-lipid may be a polymer-functionalized lipid, wherein the polymer and lipid (e.g., a hydrocarbon chain optionally interrupted by one or more inserted functional groups) are linked by covalent bonds having optional inserted atoms. The polymer-lipid optionally includes a branched moiety. The lipid may be, for example, a hydrocarbon chain optionally interrupted by one or more inserted functional groups. The polymer-lipid typically includes an uncharged hydrophilic moiety, which is believed to limit aggregation during the formulation of the LNP, such as PEG, GMI, or ATTA. Therefore, it is believed that when included in the LNP, the polymer-lipid can reduce aggregation. In some embodiments, the amount of this polymer-lipid in the LNP is selected to reduce particle aggregation. Examples of polymer-lipids include polyethylene glycol (PEG)-modified lipids, monosialotetrahexosylganglioside (GMI), and polyamide oligomers (“PAOs”), such as those described in U.S. Patent No. 6,320,017. ATTA-lipids are described, for example, in U.S. Patent No. 6,320,017, and PEG-functionalized lipids are described, for example, in U.S. Patents Nos. 5,820,873, 5,534,499, and 5,885,613. In some embodiments, the polymer-lipid comprises one or more hydrocarbon chains interrupted by biodegradable functional groups (e.g., esters). The polymer-lipid may contain branched moieties. In some embodiments, the branched moieties contain tertiary carbon atoms, quaternary carbon atoms, tertiary amines, vicinal diols, amides, carbamates, or acetals.
[0465] Lipids and optional branched moieties are believed to influence the association strength between the polymer-lipid and the LNP. At least three characteristics are thought to influence the exchange rate: the length of the lipid chain, the rigidity of the lipid chain (e.g., determined by saturation), and the size of the steric barrier head group (see, for example, U.S. Patent No. 5,820,873). For some therapeutic applications, rapid loss of the PEG-functionalized lipid from the LNP in vivo may be preferred, and thus the PEG-lipid will have a relatively short lipid anchor. In other therapeutic applications, longer plasma circulation lifetime of the nucleic acid-lipid particles may be preferred, and thus the PEG-lipid will have a relatively longer lipid anchor. mPEG (mw2000)-distearatephosphatidylethanolamine (PEG-DSPE) is thought to remain associated with the LNP in vivo for several days. Other conjugates, such as PEG-CerC20, have similar retention capabilities. However, PEG-CerC14 is thought to be exchanged more rapidly from the formulation upon exposure to serum.
[0466] In some embodiments, the polymer-lipid is charge-neutral. In some embodiments, the polymer-lipid is zwitterionic.
[0467] The polymer can be a hydrophilic polymer, such as poly(ethylene glycol), also known as poly(ethylene oxide) or PEG (“PEG-lipid”). In some embodiments, the PEG is a linear polymer of optionally substituted ethylene glycol or ethylene oxide. In some embodiments, the PEG portion is substituted, for example, by one or more alkyl, alkoxy, acyl, hydroxyl, or aryl groups. In some embodiments, the PEG portion comprises a PEG copolymer, such as PEG-polyurethane or PEG-polypropylene (see, for example, J. Milton Harris, Poly(ethylene glycol) chemistry: biotechnical and biomedical applications (1992)). In some embodiments, the PEG portion is a PEG homopolymer.
[0468] Typically, the length of the PEG chain can be represented by a number indicating the number of repeating ethylene oxide units, or by the molecular weight (in Daltons) of the PEG moiety of the polymer-lipid. In a specific example, PEG-DMG 2000 has the following structure: Where 2000 is the approximate molecular weight in Daltons for the PEG portion of the molecule. The length of the PEG chain can vary among individual lipids in a sample and can therefore be expressed as the average number of repeating units or the average molecular weight.
[0469] Typically, the length of a PEG chain can be represented by a number indicating the number of repeating oxyethyl units, or by the molecular weight (in Daltons) of the PEG portion of the polymer-lipid.
[0470] c. Neutral and zwitterionic lipids In some embodiments, the LNP comprises (charged) neutral or zwitterionic lipids. Neutral or zwitterionic lipids can generally be any type of lipid that is uncharged or neutral-zwitterionic at physiological pH. Typically, neutral or zwitterionic lipids will include a polar head group and one or more (e.g., two) hydrophobic tail groups. Exemplary head groups include phosphatidylcholine, phosphatidylethanolamine, phosphatidylserine, and inositol.
[0471] In some embodiments, the neutral lipid comprises two hydrocarbon groups, each optionally interrupted by a biodegradable portion. Neutral or zwitterionic lipids with a variety of acyl chain groups having varying chain lengths and saturation are available or can be isolated or synthesized using well-known techniques. In some embodiments, the neutral or zwitterionic lipid comprises saturated fatty acids, or mono- or di-unsaturated fatty acids. Additionally, lipids having a mixture of saturated and unsaturated fatty acid chains may be used. In some embodiments, the fatty acids are interrupted by a biodegradable portion, such as an ester.
[0472] In some embodiments, the neutral or zwitterionic lipids comprise phosphatidylcholine (PC), phosphatidylethanolamine (PE), glycerophosphatidylcholine, sphingomyelin, sphingolipid, phospholipid, natural lecithin, or hydrogenated phospholipids. Such lipids include, for example, diacylphosphatidylcholine, diacylphosphatidylethanolamine, ceramide, sphingomyelin, dihydrosphingomyelin, cephalin, and cerebrosides.
[0473] The choice between neutral or zwitterionic lipids is usually guided by considerations such as LNP particle size and stability in circulation.
[0474] In some embodiments, the neutral or zwitterionic lipid is a phospholipid. As used herein, "phospholipid" refers to a lipid comprising a hydrophilic phosphate head group and one or more hydrophobic tail groups. In some embodiments, the phospholipid can facilitate fusion with a membrane. For example, cationic phospholipids can interact with one or more negatively charged phospholipids of a membrane (e.g., a cell membrane or intracellular membrane). Fusion of the phospholipid with the membrane can allow one or more components of the LNP (e.g., the load) to be delivered through the membrane, for example, into a cell. In some embodiments, the neutral or zwitterionic lipid is 1,2-distearate-sn-glycerol-3-phosphocholine (DSPC).
[0475] d. Cholesterol and its derivatives Cholesterol is a ubiquitous structural membrane lipid, and its role in lipid compositions depends on the context. When bound to phospholipids with low gel-liquid crystal phase transition (Tm), cholesterol is thought to facilitate the formation of an ordered liquid phase, characterized by increased bilayer thickness and membrane rigidity. It is believed that cholesterol binds to other lipids containing straight-chain and / or branched hydrocarbons (e.g., low-Tm) such that the cross-sectional area of the lipid and cholesterol is less than the sum of their individual cross-sectional areas. However, when bound to high-Tm lipids, cholesterol is thought to increase membrane fluidity and provide a narrower bilayer. In any case, cholesterol is thought to promote an ordered liquid phase. Additionally, cholesterol is thought to reduce the amount of surface-bound proteins and improve circulating half-life. The amount of cholesterol in the LNP can be selected to match in vivo membrane cholesterol levels.
[0476] In some embodiments, LNP containing cholesterol is provided.
[0477] In some embodiments, the cholesterol derivative is selected from β-sitosterol, β-sitosterol acetate, β-sitosterol amino acid conjugates, coccosterol, ergosterol, 9,11-dehydroergosterol, campesterol, stigmasterol, brassicasterol, fucosterol, lycopene, ursolic acid, carotene, cholesterol, 5-heptadecylresorcinol, cholesterol hemisuccinate, 6-keto-5α-hydroxycholesterol, 7α-hydroxycholesterol, 7β-hydroxycholesterol, 7-ketocholesterol, 7β,25-dihydroxycholesterol, 2 7-hydroxycholesterol, 25-hydroxycholesterol, 20α-hydroxycholesterol, 5α-cholesterol, 5β-cosanosterol, cholesterolyl-(2'-hydroxy)-ethyl ether, 6-ketocholesterol, cholesterol-(4'-hydroxy)-butyl ether, 5α-cholesterol, cholesterol ketone, 5β-cholesterol, decanoic acid cholesterol ester, vitamin D3, vitamin D2, calcipotriol, betulin, lupeol, ursolic acid, oleanolic acid, DC-cholesterol, BHEM-cholesterol, oleic acid cholesterol ester, or combinations thereof. See, for example, Paunovska, K. et al., Adv Mater. April 2019; 31(14); Patel et al., Nat. Comm., (2020) 11:983; Ni et al., Nat. Comm. 13, article number: 4766 (2022); Kim et al., ACS Nano 2022, 16(9), 14792–14806.
[0478] 2. LNP formation LNPs can be formed by any method known in the art, including but not limited to continuous mixing or direct dilution processes.
[0479] In some embodiments, a method for preparing LNPs via a continuous mixing process is provided. For example, a process includes providing an aqueous solution comprising nucleic acids in a first reservoir, providing a lipid solution in a second reservoir, and mixing the aqueous solution with the lipid solution to mix an organic lipid solution with the aqueous solution, thereby rapidly producing LNPs encapsulated with a load. The lipid solution contains a lower alcohol, such as ethanol. This process and the apparatus for performing such a process are described in detail in U.S. Patent Publication No. 2004 / 0142025, the disclosure of which is incorporated herein by reference in its entirety. The lipid nanoparticles are produced by mixing an aqueous solution containing a load with an organic lipid solution, which undergoes continuous stepwise dilution in the presence of a buffer solution (i.e., an aqueous solution).
[0480] In some embodiments, a method for preparing LNPs by a direct dilution process is provided, the process comprising forming an LNP solution and directly introducing the LNP solution into a collection container containing a controlled amount of dilution buffer. The collection container may include one or more elements configured to agitate the contents of the collection container to facilitate dilution. In one aspect, the amount of dilution buffer present in the collection container is substantially equal to the volume of the liposome solution introduced therein. As a non-limiting example, introducing a liposome solution in about 45% ethanol into a collection container containing an equal volume of dilution buffer will advantageously produce smaller particles.
[0481] In some embodiments, a method for preparing LNPs via a direct dilution process is provided, wherein a third reservoir containing a dilution buffer is fluidly coupled to a second mixing zone. In this embodiment, the LNP solution formed in a first mixing zone is immediately and directly mixed with the dilution buffer in the second mixing zone. In a preferred aspect, the second mixing zone includes a T-connector arranged such that the LNP solution and the dilution buffer flow meet as opposing 180° flows; however, connectors providing a shallower angle may be used. In one aspect, the flow rate of the dilution buffer supplied to the second mixing zone is controlled to be substantially equal to the flow rate of the liposome solution introduced therein from the first mixing zone. Such control of the dilution buffer flow rate can advantageously allow for the formation of small particle sizes at reduced concentrations. Processes and apparatus for performing the direct dilution process are described in U.S. Patent Publication No. 2007 / 0042031, the disclosure of which is incorporated herein by reference in its entirety.
[0482] B. Viral vector In some embodiments, the coding and editing system, the guided nucleic acid binder, or the guided nucleic acid can be introduced into cells or organisms via a vector. In some embodiments, and in some cases, the vector is a plasmid, microcircle, CELiD, adeno-associated virus (AAV)-derived viral particle, or lentivirus.
[0483] Non-restrictive disclosure and inclusion by reference Each document and patent publication listed in this article is incorporated in its entirety by reference.
[0484] While certain compounds, compositions, and methods described herein have been specifically described with reference to certain embodiments, the following examples are for illustrative purposes only and are not intended to limit the scope of the compounds described herein. Each of the references, GenBank accession numbers, ENSEMBL identifiers, etc., listed in this application is incorporated herein by reference in its entirety.
[0485] The sequence listing submitted with this application identifies each nucleic acid sequence as “RNA” or “DNA” as needed; however, those skilled in the art will readily understand that the naming of “RNA” or “DNA” to describe modified oligonucleotides is arbitrary in some cases. For example, an oligonucleotide containing a nucleoside with a 2'-OH sugar moiety and a thymine base may be described as DNA having a modified sugar (i.e., 2'-OH replacing a 2'-H in DNA) or as RNA having a modified base (i.e., thymine (5-methyluracil) replacing uracil in RNA); and certain nucleic acid compounds described herein contain one or more nucleosides with a modified sugar moiety containing a 2'-substituent that is neither OH nor H. Those skilled in the art will readily understand that labeling such nucleic acid compounds as “RNA” or “DNA” does not alter or limit the description of such nucleic acid compounds.
[0486] In this document, the description of compounds having the nucleobase sequence of SEQ ID NO “ ” describes only the nucleobase sequence. Therefore, unless otherwise described, such a description of the compound by reference to the nucleobase sequence of SEQ ID NO does not limit the presence or absence of sugar or nucleoside inter-bond modifications or additional substituents (such as conjugation groups). Furthermore, unless otherwise described, the nucleobases of compounds having the nucleobase sequence of SEQ ID NO “ ” include modified forms of such compounds having the nucleobases identified as described herein.
[0487] In this document, sugars, nucleotide bonds and nucleobase modifications may be represented within the nucleotide or nucleobase sequence, or may be represented in the accompanying text (e.g., in separate text appearing within, above or below the compound table).
[0488] Although every effort has been made to accurately describe the compounds in the attached sequence listing, in the event of any discrepancy between the description in this specification and the description in the attached sequence listing, the description in the sequence listing shall prevail.
[0489] The compounds described herein include variants in which one or more atoms are replaced by non-radioactive or radioactive isotopes of the element shown. For example, compounds containing hydrogen atoms are covered herein. 1 All possible deuterium substitutions of the H hydrogen atom. Isotopic substitutions covered in the compounds described herein include, but are not limited to: 2 H or 3 H replaces 1 H, 13 C or 14 C replaces 12 C, 15 N replaces 14 N, 17 O or 18O replaces 16 O, and 33 S, 34 S, 35 S or 36 S replaces 32 S. In some embodiments, non-radioactive isotope substitution can endow oligomeric compounds with new properties beneficial for use as therapeutic or research tools. In some embodiments, radioactive isotope substitution can adapt the compounds for research or diagnostic purposes, such as imaging.
[0490] Example Example 1: Computational prediction and characterization of Cas protein in HEK293 cells—expression and localization Metagenomes were assembled using publicly available datasets. CRISPR boxes containing CRISPR arrays and nearby proteins were predicted using CRISPRFinder and CRISPRone and classified into subclasses and isotypes. For type II CRISPR systems, protein domains were identified via a Hidden Markov Model, and sequences were compared with... Streptococcus pyogenes Alignment was performed with Cas9 (SpCas9). Furthermore, the anti-repetitive sequence at the 5' end of the tracrRNA was predicted using a CRISPR array, and the tracrRNA was predicted by searching for rho-independent terminations characterized by GC-enriched palindromic sequences followed by a series of U residues at the 3' end. Nineteen putative Cas proteins were identified. The corresponding Cas protein sequences and the GeneArt (ThermoFisher) codon-optimized human nucleic acid sequences are provided in the sequence listing, as shown in the table below. The Cas protein DNA sequences were cloned by Twist Biosciences into the pTwistCMVpuro vector.
[0491] To facilitate enzyme localization to the cell nucleus, nucleoplasmic protein (SEQ ID NO: 3399) and c-Myc (SEQ ID NO: 3400) nuclear localization signals (NLS) were appended to the N-terminus and C-terminus of each Cas protein, respectively. Additionally, a HiBiT (Promega) tag was added to the C-terminus of the protein for expression detection. These additional elements are linked to the protein via known protein linkers such as GGGS / (GS)3 / (GGGGS)3. The DNA sequences encoding the Cas proteins with NLS and HiBiT tags are provided in the sequence listing, as shown in the DNA sequences SEQ ID NO: 44 to 63 below.
[0492] Table 4 Cas protein and SpCas9 The expression and localization of each Cas protein were then evaluated in HEK293 cells. Using a 4D Nucleofector (Lonza), 400 ng of codon-optimized DNA pairs encoding the estimated Cas proteins were used per well (2 × 10⁻⁶). 5 HEK293 cells were transfected and seeded in 96-well plates after transfection. Cells were harvested after 48 or 72 hours. For each well, HiBit levels, representing Cas protein expression, were measured using the Nano-glo HiBit lysis detection system (Promega) according to the manufacturer's protocol. Cas protein expression was normalized to cell counts as measured by CellTiter-Glo luminescence cell viability assay (Promega). The expression of each Cas protein is presented as the average of four measurements in the table below as HiBit luminescence / CellTiter-Glo luminescence. A negative control "no-program" group, consisting of untransfected cells, was included in the analysis. HiBit expression was detected in all samples compared to the no-program negative control. Some samples showed high protein expression comparable to SpCas9, while others showed relatively low expression.
[0493] Table 5 Expression of Cas protein and SpCas9 in HEK293 cells (normalized to cell count) The cellular localization of the Cas protein was then examined by immunostaining. HeLa cells were cultured at 1.5 × 10⁶ cells per well. 4 Cells were seeded at a density of 100 g / well in 8-well slides and transfected with 0.4 µL of BioT transfection reagent (Morganville Scientific) at a dose of 100 ng / well of a codon-optimized vector expressing Cas proteins. After 48 hours, cells were fixed in 4% formaldehyde and permeabilized with 0.1% Triton X-100 in PBS. The slides were blocked with 2% BSA in PBS for 1 hour at room temperature, and then incubated overnight at 4°C with mouse anti-HiBit antibody (Promega). The next day, the slides were washed and incubated for 1 hour with donkey anti-mouse secondary antibody Alexa Fluor 555 (1:1000) in combination with DAPI (1:1000). The slides were then washed and mounted with mounting medium. Images were taken using an EVOS cell imaging microscope. The cellular localization of Cas proteins is summarized in the table below. The control SpCas9 group showed nuclear localization as expected. Eight of the Cas proteins showed nuclear localization, while the remainder were predominantly localized to the cytoplasm.
[0494] Table 6 Cellular localization of Cas protein and SpCas9 in HEK293 cells Example 2: Evaluation of Cas protein activity in HEK293T nuclear extract After confirming expression in HEK293T cells, protein recognition region sequences were designed. For each Cas protein, the sequences of protein-specific crRNA and tracrRNA were computationally predicted, and "gaaa" quadruple loops were used to obtain the putative guide protein recognition regions, as shown in the table below. The "gaaa" quadruple loops between the putative crRNA and tracrRNA sequences are represented by lowercase letters.
[0495] Enzymatic activity of the Cas protein described above was screened using a lysis assay in HEK293T cells. The assay measured the lysis of synthesized tool DNA targets in nuclear extracts of HEK293T cells that had been transfected to express the Cas protein and transcribe its corresponding guide. First round of screening ( Figure 1 The second round of screening is performed using a system consisting of a Cas protein and a guide, which contains a universal target recognition region linked to the protein recognition region. Figure 2 The process is performed using a system consisting of a Cas protein and a guide that includes a target recognition region optimized for the Cas protein, which is linked to the protein recognition region.
[0496] Table 7 Sequences targeting the protein recognition regions predicted for the corresponding Cas protein and SpCas9. In the initial experiments, each guide used techniques previously employed by Ran et al. Nature The target recognition region, known as the "Universal Spacer Sequence" (GSS), was disclosed in 2015. The GSS target recognition regions are shown in the table below. In addition, other specific target recognition regions were designed for each Cas protein, as shown in the table below. The DNA sequences used for each target recognition region are included in the sequence listing (SEQ ID NO: 64 to 71).
[0497] Table 8 Sequences of general and Cas-specific target recognition regions used in the wizard Cas protein screening using a universal target recognition region A wizard with a universal target recognition region was designed by appending a universal target recognition region sequence (SEQ ID NO: 944) to the 5' end of a protein recognition sequence (SEQ ID NO: 121 to 141) selected from the table above.
[0498] A guide RNA was constructed. in vitro Synthesized DNA templates. A target recognition sequence is appended to the 5' end of the Cas protein recognition region described above to generate a guide, and then a T7 promoter region (TAATACGACTCACTATA (SEQ ID NO:72)) is appended to the 5' end of the target recognition sequence. Additionally, four AAGC nucleobases are added to the 5' end of this T7 promoter region to improve binding to the T7 promoter. The nucleobase sequences of the DNA templates used for transcription of the guides listed in the table below are provided as SEQ ID NO: 143 to 161, respectively. The HiScribe® T7 High-Yield RNA Synthesis Kit (New England Bioloabs) was used for T7-based in vitro transcription of this template. Each DNA template in the sequence listing represents an on-chain sequence and has the same sequence as the RNA transcript at the end of transcription using the HiScribe® T7 High-Yield RNA Synthesis Kit.
[0499] The guide sequences are listed in the table below. For each guide sequence in the table, the target recognition region is represented by a bold, underlined uppercase letter, and the protein recognition region is represented by a standard uppercase letter, with the four-loop region indicated by a lowercase letter. The Cas protein targeted by each guide is indicated in the table below by its Cas protein ID and its corresponding SEQ ID NO.
[0500] Table 9 Sequence with a guide for universal target recognition region The guides described in the table above are used to test the activity of their respective Cas proteins in a lysis assay. The DNA target used in the lysis assay is a 156 bp double-stranded DNA containing a universal spacer region (as described above) followed by a PAM library. The target DNA sequence (from 5' to 3') is: ACACTCTTTCCCTACACGACGCTCTTCCGATCTCCAACTTCATCCACGTTCAC GGGACTCAACCAAGTCATTC NNNNNNNCGGCAGACTTCTCCTCAGGAGTCAGATGCACCATGGTGTCTGTTAGATCGGAAGAGCACACGTCTGAACTCCAGTC (SEQ ID NO: 230), where the underlined bold nucleobases represent the universal spacer region (target recognition region), and 7N (N = A, C, G, T) represents the PAM library.
[0501] HEK293T cells maintained in DMEM (10% FBS, 10 U / mL penicillin, 100 µg / mL streptomycin) at a growth rate of 1.5 × 10⁻⁶ cells / mL were threshed. 6 Cells were seeded at a density of 10 cells / plate in 10 cm culture dishes (Corning) and incubated for 24 hours. Cells were transfected using 20 µL BioT transfection reagent (Morganville Scientific) at a rate of 5 µg / plate expressing the Cas protein listed in Table 4 above. HEK293T nuclear extract was prepared using NE-PER nuclear and cytoplasmic extraction reagent (Thermo Scientific) according to the manufacturer's protocol. The HiScribe T7 Rapid High-Yield RNA Synthesis Kit (New England Biolabs) was used to extract the corresponding DNA template described above. in vitro Transcription guide RNA.
[0502] For each lysis reaction, 5 µL of HEK293T nuclear extract expressing Cas protein and 200 nM were added. in vitro The transcribed guide RNA was incubated with 3 µL of lysis buffer NEBuffer r3.1 (New England BioLabs) at room temperature for 10 minutes to promote the formation of the Cas-guided ribonucleoprotein (RNP) complex. Subsequently, 10 nM DNA template was added to the reaction and incubated overnight at 37°C, followed by the addition of proteinase K and RNase to clean the reaction. The DNA template was purified using DNA Clean & Concentrator-5 (Zymo Research). Digestion of the DNA template was analyzed and visualized using Agilent TapeStation. Figure 1 As shown, Cas proteins IONCAS009 and IONCAS016 were found to have... body outside Pyrolysis activity.
[0503] To further examine the enzyme activity of relatively high-expressing Cas proteins that did not show activity when using the universal spacer region, additional screening was performed using an enzyme-specific target recognition region instead of the universal spacer region (target recognition region) sequence.
[0504] Screening of Cas proteins using Cas-specific target recognition regions A wizard was designed by appending Cas protein-specific target recognition region sequences (SEQ ID NO: 945 to 951) selected from Table 8 above to the 5' end of protein recognition sequences (SEQ ID NO: 121 to 141) selected from the table above.
[0505] A guide RNA was constructed. in vitro Synthesized DNA template. A Cas-specific target recognition sequence was appended to the 5' end of the Cas protein recognition region described above to generate a guide, and then the T7 promoter region (TAATACGACTCACTATA (SEQ ID NO: 72)) was appended to the 5' end of the target recognition sequence. Furthermore, since T7 strongly prefers G-starting and GG is most effective, GG was added to the 5' end of target recognition regions that begin with a nucleotide other than G. For target recognition regions that begin with G, one G was added to their 5' end. Additionally, four nucleotides AAGC were added to the 5' end of the T7 promoter region to improve binding with T7 polymerase. The nucleotide sequences of the DNA template used for transcription of the guides listed in the table below are provided as SEQ ID NO: 162 to 174, respectively. The HiScribe® T7 High-Yield RNA Synthesis Kit (New England Bioloabs) was used for T7-based in vitro transcription of this template. Each DNA template in the sequence listing represents an on-chain sequence and has the same sequence as the RNA transcript at the end of transcription using the HiScribe® T7 High-Yield RNA Synthesis Kit.
[0506] The guide sequences are listed in the table below. For each guide sequence in the table below, the target recognition region is represented by a bold, underlined uppercase letter, and the protein recognition region is represented by a standard uppercase letter, with the four-loop indicated by a lowercase letter. The Cas protein targeted by each guide is indicated in the table below by its Cas protein ID and its SEQ ID NO.
[0507] Table 10 Sequence of a guide containing a Cas protein-specific target recognition region The lysis assay was repeated as described above, but with the following changes: the guide used consisted of the Cas-specific guides described in the table above; and the DNA target used consisted of a Cas-specific target recognition region attached to a 7-nucleotide degenerate PAM sequence. When using the Cas-specific target recognition region, the Cas proteins IONCAS008 and IONCAS013 showed [further characteristics] in the nuclear extract. in vitro cleavage activity ( Figure 2 ).
[0508] Example 3: Identification of the PAM sequence of the Cas protein The recognition shows in vitro The PAM sequence of each Cas protein with enzyme activity was determined by loss assay (Ran, FA et al.). Nature 2015, 520,186-191). In this loss assay, the uncleaved portion of the DNA template was amplified via PCR using i5 / i7 index primers from Integrated DNA Technologies (IDT). The amplicons were sequenced via Illumina MiSeq. For a given single-end MiSeq sequencing run, the data was demultiplexed using a MiSeq machine or Illumina bcl2fastq conversion software (https: / / support.illumina.com / sequencing / sequencing_software / bcl2fastq-conversion-software / downloads.html) to generate a FASTQ file. A custom program was written to extract the 7-mer PAM sequence located at the 3' end of the known spacer sequence from the FASTQ file. For each extracted PAM sequence in the sample, the count per million (CPM) value was calculated by dividing the number of reads containing the PAM sequence by the total number of reads in the sample. The relative PAM sequence exhaustion frequency was calculated by dividing the CPM value in the Cas protein-treated sample by the corresponding CPM value in the untreated control sample. Then, the Python Logomaker software library (Tareen and Kinney 2019), available at https: / / logomaker.readthedocs.io / en / latest / , was used to generate PAM sequence motifs at different exhaustion frequency thresholds. The common PAM sequences of the Cas proteins IONCAS008, IONCAS009, and IONCAS016 are summarized in the table below and shown in [the table / example]. Figure 3 middle.
[0509] Table 11 Common PAM sequence of Cas proteins In the table above, "N" represents any nucleobase; "M" represents A or C; "H" represents A, C, or T; and "V" represents A, C, or G.
[0510] Example 4: Evaluation of cellular activity of the gene target Cas protein in HEK293 cells The enzymatic activity of the Cas proteins described above was evaluated in HEK293 cells, measuring editing in six genomic regions. For each Cas protein, a wizard was designed containing a target recognition region sequence designed for the genomic region, which was then linked to the 5' end of a protein recognition region sequence selected from the examples above. This was then used... in vitro The lysis assay was evaluated using a wizard in HEK293 cells.
[0511] Six genomic regions widely used in SpCas9 research were selected as targets for assessing cellular Cas protein activity. These regions are located within human genomic loci at AAVS1, EMX1, FANCF, HBB, HEKSite4, and HPRT. For each putative Cas protein and guide system, CHOPCHOP v3 (Labun et al.) was used. Nucleic Acids Res The target recognition regions were designed using the Benchling software (2019) and are listed in the table below. The Cas protein and gene targets targeted by each target recognition region are shown in the table below. The DNA sequences used for each target recognition region are included in the sequence listing (SEQ ID NO: 175 to 199).
[0512] Table 12 Sequences of gene-specific target recognition regions used to assess Cas protein and SpCas9 activity in HEK293 cells. Guides were designed by adding Cas protein-specific target recognition region sequences selected from the table above to the 5' end of protein recognition sequences (SEQ ID NO: 121 to 141) selected from the table above. Guide sequences are listed in the table below. For each guide sequence, the target recognition region is represented by a bold, underlined uppercase letter, and the protein recognition region is represented by a standard uppercase letter, with four-loop sequences indicated by lowercase letters. The Cas protein targeted by each guide is indicated in the table below by its Cas protein ID and its SEQ ID NO. The gene target of each guide is also indicated in the table below. The DNA sequence of each guide is included in the sequence listing (SEQ ID NO: 200 to 229).
[0513] Table 13 Sequences of guides with gene-specific target recognition regions A plasmid expressing the guide was designed and ordered from Twist BioScience by cloning the guide sequence into a vector driven by the U6 promoter for transcription by RNA polymerase III. The U6 promoter sequence (GAGGGCCTATTTCCCATGATTCCTTCATATTTGCATATACGATACAAGGCTGTTAGAGAGATAATTGGAATTAATTTGACTGTAAACACAAAGATATTAGTACAAAATACGTGACGTAGAAAGTAATAATTTCTTGGGTAGTTTGCAGTTTTAAAATTATGTTTTAAAATGGACTATCATATGCTTACCGTAACTTGAAAGTATTTCGATTTCTTGGCTTTATATATCTTGTGGAAAGGACGAAACACCg; SEQ ID NO: 75) was appended to the 5' end of the guide sequence, wherein for guides with 5'-GG, the final g, represented by a lowercase letter, was removed. A transcription termination sequence (GTCTAGAGGTAC; SEQ ID NO: 76) was added to the 3' end.
[0514] Using a 4D Nucleofector (Lonza), 400 ng of codon-optimized DNA encoding the Cas protein was placed per well, along with 400 ng of expression guide plasmid per well, in a 2×10⁻⁶ pair. 5 HEK293 cells were transfected and seeded in 96-well plates after transfection. Two negative controls were included in the analysis: cells transfected with SpCas9 / random spgRNA and untransfected cells. After a 48 or 72-hour incubation period, genomic DNA was extracted from the cells using the DNeasy Blood & Tissue Kit (Qiagen) according to the manufacturer's recommended protocol. Genomic PCR was performed using primers spanning the target region, and libraries were constructed using i5 / i7 index primers from IDT. Amplicon sequencing was performed via MiSeq (Illumina). For a given paired-end MiSeq sequencing run, the data was demultiplexed to generate FASTQ files (as described above). CRISPResso2 software (Clement et al.) was used. Nat Biotechnol,The CRISPRessoBatch module (2019) was used to process and analyze all sample sequencing data. A custom script was written to calculate the percentage of edits based solely on InDel reads from CRISPResso2 analysis results. The gene editing quantification window was set to approximately 40 nt long, centered on a guide RNA sequence of approximately 20 nt, with approximately 10 nt upstream and approximately 10 nt downstream. To enable visualization of low-frequency InDel read alignments, the minimum threshold for the report reads was set to 0.01% (default 0.20%), and the maximum number of report alignments was set to 100 (default 50).
[0515] The results are presented as InDel percentages (%), summarized in the table below. “ND” indicates that no data was obtained.
[0516] Table 14 Percentage of InDel of Cas protein and guide to AAVS1 in HEK293 cells Table 15 Percentage of InDel values of Cas protein and guide to EMX1 in HEK293 cells Table 16 Percentage of InDel values of Cas protein and guide to FANCF in HEK293 cells Table 17 Percentage of InDel of Cas protein and guide to HBB in HEK293 cells Table 18 Percentage of Cas protein and guide protein in HEKSite4 InDel in HEK293 cells Table 19 Percentage of InDel of Cas protein and guide to HPRT in HEK293 cells Example 5: Computational Prediction of Cas Proteins and Related Wizards Metagenomes were assembled using publicly available datasets. CRISPR boxes containing CRISPR arrays and nearby proteins were predicted using CRISPRFinder and CRISPRone and classified into subclasses and isotypes. For type II CRISPR systems, protein domains were identified via a Hidden Markov Model, and sequences were compared with... Streptococcus pyogenesThe Cas9 (SpCas9) sequence was compared. Furthermore, the tracrRNA was predicted by using a CRISPR array to predict the anti-repetitive sequence at the 5' end and by searching for rho-independent terminations in GC-enriched palindromic sequences followed by a series of U residues at the 3' end.
[0517] Following these analyses, 93 putative Cas proteins were identified. The corresponding Cas protein sequences and the GeneArt (ThermoFisher) codon-optimized human nucleic acid sequences are provided in the sequence listing, as shown in the table below. The Cas protein DNA sequences were cloned by Twist Biosciences into the pTwistCMVpuro vector.
[0518] To facilitate enzyme localization to the cell nucleus, nucleoplasmic protein (SEQ ID NO: 3399) and c-Myc (SEQ ID NO: 3400)NLS were appended to the N-terminus and C-terminus of each Cas protein, respectively. Additionally, a HiBiT (Promega) tag was added to the C-terminus of the protein for expression detection. These additional elements are linked to the protein via known protein linkers such as GGGS / (GS)3 / (GGGGS)3. The DNA sequences encoding the Cas proteins with NLS and HiBiT tags are provided in the sequence listing, as shown in the DNA sequences SEQ ID NO: 231 to 509 below.
[0519] Table 20 Cas protein and SpCas9 Example 6: Evaluation of Cas protein activity in HEK293T nuclear extract Protein recognition region sequences for the guide were designed. The sequences of protein-specific crRNA and tracrRNA were computationally predicted for each Cas protein, and the putative guide protein recognition elements were obtained using a "gaaa" four-loop connection, as shown in the table below. The "gaaa" four-loop between the putative crRNA and tracrRNA sequences is represented by lowercase nucleotides.
[0520] Enzymatic activities of the Cas proteins described above were screened using a lysis assay in HEK293T cells (as described above). The assay measured the lysis of synthetic tool DNA targets synthesized in nuclear extracts from HEK293T cells that had been transfected to express Cas proteins and transcribe their corresponding guides. Guides containing target recognition regions linked to the 5' end of the protein recognition region were designed for each Cas protein, as described below. Table 21 Predicted protein recognition region sequence for the corresponding Cas protein Cas protein screening using a universal target recognition region A wizard with a universal target recognition region was designed by adding a universal target recognition region sequence (SEQ ID NO: 944) to the 5' end of the protein recognition region sequences (SEQ ID NO: 510 to 602) selected from the table above.
[0521] A guide RNA was constructed. in vitro Synthesized DNA template. A target recognition sequence was appended to the 5' end of the Cas protein recognition region described above to generate a guide, and then a T7 promoter region (TAATACGACTCACTATA (SEQ ID NO:72)) was appended to the 5' end of the target recognition sequence. Additionally, four nucleotides AAGC were added to the 5' end of the T7 promoter region to improve binding to the T7 promoter. The nucleotide sequences of the DNA templates used for transcription of the guides C20g-C114g listed in the table below are provided as SEQ ID NO: 603 to 695, respectively. The HiScribe® T7 High-Yield RNA Synthesis Kit (NewEngland Biolabs) was used for T7-based transcription of this template. in vitro Transcription. Each DNA template in the sequence listing represents an on-strand sequence and has the same sequence as the RNA transcript at the end of transcription using the HiScribe® T7 High-Yield RNA Synthesis Kit.
[0522] The guide sequences are listed in the table below. For each guide sequence in the table, the target recognition region is represented by a bold, underlined uppercase letter, and the protein recognition region is represented by a standard uppercase letter, with the four-loop region indicated by a lowercase letter. The Cas protein targeted by each guide is indicated in the table below by its Cas protein ID and its corresponding SEQ ID NO.
[0523] Table 22 Sequence with a guide for universal target recognition region exist in vitro The enzymatic activity of the Cas protein described above was screened in a lysis assay, in which the synthesized DNA target was compared with the nuclear extract of HEK293T cells transfected with pTwist-CMV-puro-Cas protein and its corresponding... in vitro The transcribed guide RNA was incubated together. The DNA target was a 156 bp double-stranded DNA containing a universal spacer followed by a PAM library. The DNA sequence of the target is described above (SEQ ID NO: 230).
[0524] HEK293T cells maintained in DMEM (10% FBS, 10 U / mL penicillin, 100 µg / mL streptomycin) at a growth rate of 1.5 × 10⁻⁶ cells / mL were threshed. 6 Cells were seeded at a density of 10 cells / plate in 10 cm culture dishes (Corning) and incubated for 24 hours. Cells were transfected using 20 µL BioT transfection reagent (Morganville Scientific) at a rate of 5 µg / plate expressing the Cas protein listed in Table 2 above. HEK293T nuclear extract was prepared using NE-PER nuclear and cytoplasmic extraction reagent (Thermo Scientific) according to the manufacturer's protocol. The HiScribe T7 Rapid High-Yield RNA Synthesis Kit (New England Biolabs) was used to extract the corresponding DNA template described above. in vitro Transcription guide RNA.
[0525] For each lysis reaction, 5 µL of HEK293T nuclear extract expressing Cas protein and 200 nM were added. in vitroThe transcribed guide RNA was incubated with 3 µL of lysis buffer NEBuffer r3.1 (New England BioLabs) at room temperature for 10 minutes to promote the formation of the Cas-guide ribonucleoprotein (RNP) complex. Subsequently, 10 nM DNA template was added to the reaction and incubated overnight at 37°C, followed by the addition of proteinase K and RNase to clean up the reaction. The DNA template was purified using DNA Clean & Concentrator-5 (Zymo Research). Digestion of the DNA template was analyzed and visualized using Agilent TapeStation.
[0526] As shown in Figures 4a to 4c, the Cas proteins IONCAS030, IONCAS031, IONCAS034, IONCAS036, IONCAS050, IONCAS056, IONCAS057, IONCAS058, IONCAS061, IONCAS064, IONCAS065, IONCAS069, IONCAS073, IONCAS074, IONCAS075, IONCAS079, IONCAS080, IONCAS083, IONCAS088, IONCAS090, IONCAS091, IONCAS092, IONCAS093, IONCAS105, and IONCAS114 were found to possess... in vitro Pyrolysis activity.
[0527] Example 7: Identification of the PAM sequence of the Cas protein show in vitro The PAM sequence of the Cas protein, which provides enzyme activity, was determined by loss assay (Ran, FA, et al.). Nature 2015, 520, 186-191). In this loss assay, the uncleaved portion of the DNA template was amplified via PCR using i5 / i7 index primers from Integrated DNA Technologies (IDT). The amplicons were sequenced via Illumina MiSeq. For a given single-end MiSeq sequencing run, the data was demultiplexed using a MiSeq machine or Illumina bcl2fastq conversion software (https: / / support.illumina.com / sequencing / sequencing_software / bcl2fastq-conversion-software / downloads.html) to generate a FASTQ file. A custom program was written to extract the 7-mer PAM sequence located at the 3' end of the known spacer sequence from the FASTQ file. For each extracted PAM sequence in the sample, the count per million (CPM) value was calculated by dividing the number of reads containing the PAM sequence by the total number of reads in the sample. The relative PAM sequence exhaustion frequency was calculated by dividing the CPM value in the Cas protein-treated sample by the corresponding CPM value in the untreated control sample. Then, the Python Logomaker software library (Tareen and Kinney 2019), available at https: / / logomaker.readthedocs.io / en / latest / , was used to generate PAM sequence motifs at different exhaustion frequency thresholds. The common PAM sequences of the Cas proteins IONCAS030, IONCAS031, IONCAS034, IONCAS036, IONCAS050, IONCAS056, IONCAS057, IONCAS058, IONCAS061, IONCAS064, IONCAS065, IONCAS069, IONCAS073, IONCAS074, IONCAS075, IONCAS079, IONCAS080, IONCAS083, IONCAS088, IONCAS090, IONCAS091, IONCAS092, IONCAS093, IONCAS105, and IONCAS114 are summarized in the table below.
[0528] Table 23 Common PAM sequence of Cas proteins In the table above, "N" represents any nucleobase; "M" represents A or C; "R" represents A or G; "Y" represents C or T; "K" represents G or T; "S" represents G or C; "W" represents A or T; "B" represents G, T, or C; "H" represents A, C, or T; and "V" represents A, C, or G.
[0529] Example 8: Guide design and efficacy in HEK293 cells For the Cas proteins selected from the examples above, a wizard was designed containing a target recognition region sequence linked to the 5' end of the protein recognition region sequence. The wizard was optimized by adjusting various parameters, including the target recognition region length, the length of the first stem-loop in the protein recognition region, the substitution of A-U or U-A base pairs with C-G base pairs within the predicted stem, the substitution of loops and adapters with different sequences, and the total protein recognition region length. Subsequently, [the following was used]... in vitro The guide was evaluated in HEK293 cells by lysis assay.
[0530] A plasmid expressing the guide was designed and ordered from Twist BioScience by cloning the guide sequence into a vector driven by the U6 promoter for transcription by RNA polymerase III. The U6 promoter sequence (SEQ ID NO: 75) was appended to the 5' end of the guide sequence, and for guides already having a 5'-GG, the terminal 3'-G was removed. A transcription termination sequence (SEQ ID NO: 76) was added to the 3' end.
[0531] Using a 4D Nucleofector (Lonza), codon-optimized DNA encoding the Cas protein (shown in the table below) was applied per well at 400 ng per well, along with an expression guide plasmid (shown in the table below) per well, for 2 × 10⁻⁶ cells. 5 HEK293 cells were transfected and seeded in 96-well plates after transfection. After a 48 or 72-hour incubation period, genomic DNA was extracted from the cells using the DNeasy Blood & Tissue Kit (Qiagen) according to the manufacturer's recommended protocol. Genomic PCR was performed using primers spanning the target region for the targets shown in the table below, and libraries were constructed using i5 / i7 index primers from IDT. Amplicon via Sequencing is performed using MiSeq (Illumina). For a given paired-end MiSeq sequencing run, the data is demultiplexed to generate a FASTQ file (as described above). The CRISPResso2 software (Clement...) is then used. Nat et al. Biotechnol The CRISPRessoBatch module (2019) was used to process and analyze all sample sequencing data and perform InDel analysis (as described above).
[0532] Target recognition region length optimization Target recognition region sequences of 20 to 24 nt in length were constructed using PAM sequences determined for each Cas protein and target pair. These target recognition region sequences can be found in the sequence listing, as shown in the table below (SEQ ID NO: 696 to 887). A wizard was designed to ligate the target recognition region sequences of 20 to 24 nt in length to the 5' end of the protein recognition sequences selected from the examples above.
[0533] The wizards described in the table below are constructed from target recognition regions identified by ID and SEQ ID NO in the “Target Recognition Region” column and protein recognition regions identified by ID and SEQ ID NO in the “Protein Recognition Region” column. In the table below, the target for each wizard is indicated in the column labeled “Target”, and the length of the protein recognition region is indicated in the column labeled “Length”. The complete sequence of the wizard can be found in the sequence listing, as shown in the column labeled “Wizard”.
[0534] Cas protein and guide were tested in HEK293 cells. in vitro Activity assays, as described above, and results are presented in the table below. For each Cas protein and guide system, the relative amount of InDel was calculated by dividing the %InDel of the target of a given Cas protein and guide system by the %InDel of the target when the target recognition region length is 20 nt. The results are presented in the table below as “Relative Indel relative to 20 nt”. “ND” indicates no data available.
[0535] Table 24 Percentage of InDel targets containing IONCAS009 protein and guide, and varying target recognition region length. Table 25 Percentage of InDel targets containing IONCAS030 protein and guide, and varying target recognition region length. Table 26 Percentage of InDel targets containing IONCAS050 protein and guide, and varying target recognition region length. Table 27 Percentage of InDel targets containing IONCAS057 protein and guide, and varying target recognition region length. Table 28 Percentage of InDel targets containing IONCAS064 protein and guide, and varying target recognition region length. Table 29 Percentage of InDel targets containing IONCAS065 protein and guide, and varying target recognition region length. Table 30 Percentage of InDel targets containing IONCAS075 protein and guide, and varying target recognition region length. Table 31 Percentage of InDel targets containing IONCAS080 protein and guide, and varying target recognition region length. Table 32 Percentage of InDel targets containing IONCAS083 protein and guide, and varying target recognition region length. Table 33 Percentage of InDel targets containing IONCAS090 protein and guide, and varying target recognition region length. Table 34 Percentage of InDel targets containing IONCAS092 protein and guide, and varying target recognition region length. Table 35 Percentage of InDel targets containing IONCAS105 protein and guide, and varying target recognition region length. Table 36 Percentage of InDel targets containing IONCAS114 protein and guide, and varying target recognition region length. Optimization of the protein recognition region length for the first stem-loop The protein recognition region (R:AR) sequence was modified by varying the repeat:anti-repeat (R:AR) length, which was achieved by changing the length of the upper stem loop in the 5' stem loop of the R:AR, as predicted by the ViennaFold software. The R:AR length is the number of nucleotides at the 5' of the "gaaa" quadruple loop in the R:AR and is indicated in the column labeled "R:AR Length" in the table below. The complete R:AR sequence associated with each modification can be found in the sequence listing (SEQ ID NO: 961 to 1019), as shown in the table below. A wizard was then designed by appending target recognition region sequences selected from the examples above to the 5' end of protein recognition region sequences with different R:AR lengths.
[0536] The wizards described in the table below are constructed from target recognition regions identified by ID and SEQ ID NO in the "Target Recognition Region" column and protein recognition regions identified by ID and SEQ ID NO in the "Protein Recognition Region" column. In the table below, the target for each wizard is indicated in the column labeled "Target," and the R:AR length of the protein recognition region is indicated in the column labeled "R:AR Length." The complete sequences of the wizards can be found in the sequence listing (SEQ ID NO: 1601 to 1766), as shown in the column labeled "Wizard."
[0537] Cas protein and guide were tested in HEK293 cells. in vitro Activity assays, as described above, and results are presented in the table below. For each Cas protein and guide system, the relative amount of InDel was calculated by dividing the %InDel of the target of a given Cas protein and guide system by the %InDel of the target of the Cas protein and guide system having the original R:AR length (15 nt for IONCAS009 and 16 nt for other Cas proteins). The results are presented as “Relative InDel” in the table below.
[0538] Table 37 Percentage of InDel targets containing IONCAS009 protein and guide, and varying protein recognition region R:AR length. Table 38 Percentage of InDel targets containing IONCAS030 protein and guide, and varying protein recognition region R:AR length. Table 39 Percentage of InDel targets containing IONCAS050 protein and guide, and varying protein recognition region R:AR length. Table 40 Percentage of InDel targets containing IONCAS057 protein and guide, and varying protein recognition region R:AR length. Table 41 Percentage of InDel targets containing IONCAS064 protein and guide, and varying protein recognition region R:AR length. Table 42 Percentage of InDel targets containing IONCAS065 protein and guide, and varying protein recognition region R:AR length. Table 43 Percentage of InDel targets containing IONCAS075 protein and guide, and varying protein recognition region R:AR length. Table 44 Percentage of InDel targets containing IONCAS080 protein and guide, and varying protein recognition region R:AR length. Table 45 Percentage of InDel targets containing IONCAS083 protein and guide, and varying protein recognition region R:AR length. Table 46 Percentage of InDel targets containing IONCAS090 protein and guide, and varying protein recognition region R:AR length. Table 47 Percentage of InDel targets containing IONCAS092 protein and guide, and varying protein recognition region R:AR length. Table 48 Percentage of InDel targets containing IONCAS105 protein and guide, and varying protein recognition region R:AR length. Table 49 Percentage of InDel targets containing IONCAS114 protein and guide, and varying protein recognition region R:AR length. Protein recognition region optimization: second stem-loop Additional modifications to the protein recognition region were tested by modifying the sequences of the first linker (L1) and the second stem-loop (S2, S2', and H2) of the protein recognition region of the C9gRNA (SEQ ID NO: 131, as described above). Specific modifications are indicated in the column labeled "Modified" in the table below. The complete protein recognition region sequences associated with each SL-1 modification can be found in the sequence listing (SEQ ID NO: 1016 to 1019), as shown in the table below. A wizard was then designed by appending target recognition region sequences selected from the examples above to the 5' end of the modified protein recognition region sequence.
[0539] The wizards described in the table below are constructed from target recognition regions identified by ID and SEQ ID NO in the "Target Recognition Region" column and protein recognition regions identified by ID and SEQ ID NO in the "Protein Recognition Region" column. In the table below, the target for each wizard is indicated in the column labeled "Target," and modifications in the protein recognition region are indicated in the column labeled "Modification." The complete sequences of the wizards can be found in the sequence listing (SEQ ID NO: 1767 to 1778), as shown in the column labeled "Wizard."
[0540] In HEK293 cells, in vitro activity assays were performed on the Cas proteins and guides (as described above). For each IONAS009 protein and guide system, the percentage of insertions and deletions (InDel percentage) on each target is reported in the table below.
[0541] Table 50 Percentage of InDel targets containing IONCAS009 protein and guide, and variations in protein recognition region sequences. Protein recognition region length optimization The protein recognition region (GCR) sequence is modified to effectively shorten the total sequence length by truncating the 3' stem-loop in the guide or by removing base pairs from the 3' stem-loop. The total length of the GCR sequence is shown in the column labeled "Length" in the table below. The complete GCR sequences associated with each modification can be found in the sequence listing (SEQ ID NO: 1020 to 1051), as shown in the table below.
[0542] An additional target recognition region (PKR) sequence of 20 nt in length was constructed using PAM sequences determined for each Cas protein and target pair. The PKR sequences are available in the sequence listing, as shown in the table below (SEQ ID NO: 888 to 890). A wizard was then designed to append the PKR sequences to the 5' end of protein recognition sequences of varying lengths.
[0543] The wizards described in the table below are constructed from target recognition regions identified by ID and SEQ ID NO in the “Target Recognition Region” column and protein recognition regions identified by ID and SEQ ID NO in the “Protein Recognition Region” column. In the table below, the target for each wizard is indicated in the column labeled “Target”, and the length of the protein recognition region sequence is indicated in the column labeled “Length”. The complete sequences of the wizards can be found in the sequence listing (SEQ ID NO: 1409, 1413, 1417, 1856, 1860, 1868, 1872, 1889 to 1949), as shown in the column labeled “Wizard”.
[0544] In HEK293 cells, in vitro activity assays were performed on the Cas proteins and guides (as described above). For each Cas protein and guide system, the percentage of insertions and deletions (InDel percentage) on each target is reported in the table below.
[0545] Table 51 Percentage of InDel targets containing IONCAS009 protein and guide, and varying protein recognition region length. Table 52 Percentage of InDel targets containing IONCAS057 protein and guide, and varying protein recognition region length. Table 53 Percentage of InDel targets containing IONCAS065 protein and guide, and varying protein recognition region length. Table 54 Percentage of InDel targets containing IONCAS075 protein and guide, and varying protein recognition region length. Table 55 Percentage of InDel targets containing IONCAS080 protein and guide, and varying protein recognition region length. Table 56 Percentage of InDel targets containing IONCAS083 protein and guide, and varying protein recognition region length. Table 57 Percentage of InDel targets containing IONCAS090 protein and guide, and varying protein recognition region length. Table 58 Percentage of InDel targets containing IONCAS092 protein and guide, and varying protein recognition region length. Table 59 Percentage of InDel targets containing IONCAS114 protein and guide, and varying protein recognition region length. Protein recognition region sequence optimization The protein recognition region sequence was modified by interchanging various base pairs in the stem of the predicted stem-loop (using ViennaFold), for example, exchanging A to U for G to C, or exchanging G to U swing pairs for G to C. The complete protein recognition region sequences can be found in the sequence listing (SEQ ID NO: 1053 to 1312), as shown in the table below. A wizard was then designed to ligate the target recognition region sequence to the 5' end of the protein recognition sequence.
[0546] The wizards described in the table below are constructed from target recognition regions identified by ID and SEQ ID NO in the “Target Recognition Region” column and protein recognition regions identified by ID and SEQ ID NO in the “Protein Recognition Region” column. In the table below, the targets targeted by each wizard are indicated in the column labeled “Target”. The complete sequences of the wizards can be found in the sequence listing (SEQ ID NO: 1950 to 2471), as shown in the column labeled “Wizard”.
[0547] In HEK293 cells, in vitro activity assays were performed on Cas proteins and guide systems (as described above). For each Cas protein and guide system, the percentage of insertions and deletions (InDel percentage) on each target is reported in the table below. The relative amount of InDel was also calculated by dividing the %InDel of the target of a given Cas protein and guide system by the %InDel of the target of the Cas protein and guide system having that original protein recognition region. The results are presented as “Relative InDel” in the table below.
[0548] Table 60 Percentage of InDel targets containing IONCAS009 protein and guide, and variations in protein recognition region sequences. Table 61 Percentage of InDel targets containing IONCAS030 protein and its guide, and variations in protein recognition region sequences. Table 62 Percentage of InDel targets containing IONCAS057 protein and guide, and variations in protein recognition region sequences. Table 63 Percentage of InDel targets containing IONCAS065 protein and its guide, and variations in protein recognition region sequences. Table 64 Percentage of InDel targets containing IONCAS075 protein and guide, and variations in protein recognition region sequences. Table 65 Percentage of InDel targets containing IONCAS080 protein and its guide, and variations in protein recognition region sequences. Table 66 Percentage of InDel targets containing IONCAS083 protein and its guide, and variations in protein recognition region sequences. Table 67 Percentage of InDel targets containing IONCAS090 protein and guide, and variations in protein recognition region sequences. Table 68 Percentage of InDel targets containing IONCAS092 protein and guide, and variations in protein recognition region sequences. Table 69 Percentage of InDel targets containing IONCAS105 protein and its guide, and variations in protein recognition region sequences. Table 70 Percentage of InDel targets containing IONCAS114 protein and guide, and variations in protein recognition region sequences. Example 9: Design and efficacy of modified guides in HepG2 cells for LNP delivery For the Cas proteins selected from the examples above, a modified wizard containing a target recognition region sequence linked to the 5' end of the protein recognition region sequence was designed and subsequently purchased from IDT. The wizard was optimized by adjusting the target recognition region length, the protein recognition region length, and the protein recognition region sequence. Then... in vitro The enzyme activity of the screening guide was measured in HepG2 cells incubated with Cas protein and LNP delivery formulation of the guide system.
[0549] Design wizards according to the sequences indicated by their SEQ ID NO. in the table below, and order them from IDT. Each wizard in the table below has the following modifications: the nucleotides at positions 1 and 2 at the 5' end are 2'-OMe sugar moieties, wherein each nucleotide is linked to the next nucleotide via a phosphate thioside internucleotide bond; and the last two nucleotides at the 3' end are 2'-OMe sugar moieties, wherein each nucleotide is linked to the previous nucleotide via a phosphate thioside internucleotide bond.
[0550] Human HepG2 cells were seeded at a density of 20,000 cells per well in 96-well plates and cultured in DMEM medium containing 10% FBS and 1% PenStrep. After 24 hours, the medium was changed. RNA / LNP complexes were prepared by mixing RNA (Cas protein mRNA / guide at a ratio of 1:2) with an LNP formulation (Kazemian et al.). Mol.Pharmaceutics ,2022, 19, 1669-1689; Albertsen et al. Adv.Drug Deliv.Rev. 2022, 118, 114416; Cullis et al. Nat. Rev. Drug Discov. (2024, 23, 709-722.) After incubation for 15 minutes, the mixture was added to the culture medium at the doses indicated in the table below. Cells were harvested after 72 hours for InDel analysis, and InDel analysis was performed (as described above in this document).
[0551] Target recognition region length optimization Target recognition region sequences of 20 to 25 nt in length, selected from the examples above, or designed using PAM sequences determined for each Cas protein and target pair. Target recognition region sequences can be found in the sequence listing, as shown in the table below.
[0552] A protein recognition region sequence was designed for each Cas protein, truncating the UUUUUU (6U) sequence at the 3' end. The protein recognition region sequences can be found in the sequence listing, as shown in the table below.
[0553] A wizard was designed by linking a target recognition region sequence of 20 to 25 nt in length to the 5' end of a protein recognition sequence with or without a 3'-6U sequence.
[0554] The wizards described in the table below are constructed from the target recognition regions identified by ID and SEQ ID NO in the "Target Recognition Region" column and the protein recognition regions identified by ID and SEQ ID NO in the "Protein Recognition Region" column. In the table below, the length of the target recognition region sequence is indicated in the column labeled "Length," and the presence of the 6U at the 3' end of the protein recognition sequence is indicated in the column labeled "3' End." The nucleotide sequence of the wizard and the complete annotation sequence describing the modifications on the wizard can be found in the sequence listing, as shown in the wizard columns labeled "SEQ ID NO." and "Mod. SEQ ID No.", respectively.
[0555] In HepG2 cells, Cas protein and guide were analyzed. in vitro Activity assays (as described above). For each Cas protein and guide system, the percentage of insertions and deletions (InDel percentage) on each target is reported in the table below.
[0556] Table 71 Percentage of InDel with IONCAS009 protein and guide EMX1, varying target recognition region length and 3'-6U Table 72 The percentage of InDel in HEKSite4 containing IONCAS009 protein and its guide, the varying target recognition region length, and the 3'-6U Table 73 The percentage of InDel in FANCF using IONCAS057 protein and guide, variations in target recognition region length, and 3'-6U Table 74 The percentage of InDel in HBB using IONCAS057 protein and guide, variations in target recognition region length, and 3'-6U Table 75 The percentage of InDel in AAVS1 with IONCAS030 protein and guide, the varying target recognition region length and 3'-6U Table 76 The percentage of InDel in AAVS1 with IONCAS065 protein and guide, the varying target recognition region length and 3'-6U Table 77 The percentage of InDel in HEKSite4 containing IONCAS065 protein and its guide, the varying target recognition region length, and the 3'-6U Table 78 The percentage of InDel in AAVS1 with IONCAS075 protein and guide, the varying target recognition region length and 3'-6U Table 79 The percentage of InDel in FANCF using IONCAS080 protein and guide, variations in target recognition region length, and 3'-6U Table 80 The percentage of InDel in HEKSite4 containing IONCAS080 protein and its guide, the varying target recognition region length, and the 3'-6U Table 81 The percentage of InDel in FANCF using IONCAS083 protein and guide, variations in target recognition region length, and 3'-6U Table 82 The percentage of InDel in HPRT using IONCAS090 protein and guide, variations in target recognition region length, and 3'-6U Table 83 The percentage of InDel in AAVS1 with IONCAS105 protein and guide, the varying target recognition region length and 3'-6U Table 84 The percentage of InDel in HBB using IONCAS105 protein and guide, variations in target recognition region length, and 3'-6U Protein recognition region sequence optimization The protein recognition region sequence is modified by interchanging various base pairs in the stem-loop, for example, from A to U to C to G, or from G to U to G to C. The selected protein recognition sequences are shown in the table below. All protein recognition region sequences can be found in the sequence listing, as shown in the table below. A wizard was then designed to ligate the target recognition region sequence to the 5' end of the protein recognition sequence. In the table below, the "gaaa" four-loop is represented by lowercase letters.
[0557] Table 85 Select protein recognition region sequences designed for each Cas protein. The wizards described in the table below are constructed from target recognition regions identified by ID and SEQ ID NO in the "Target Recognition Region" column and protein recognition regions identified by ID and SEQ ID NO in the "Protein Recognition Region" column. In the table below, the target targeted by each wizard is indicated in the column labeled "Target". The nucleotide sequence of the wizard and the complete annotation sequence describing the modifications on the wizard can be found in the sequence listing, as shown in the wizard columns labeled "SEQ ID NO." and "Mod. SEQ ID No.", respectively.
[0558] In HepG2 cells, in vitro activity assays were performed on the Cas proteins and guides (as described above). For each Cas protein and guide system, the percentage of insertions and deletions (InDel percentage) on each target is reported in the table below.
[0559] Table 86 Percentage of InDel targets containing IONCAS009 protein and guide, and variations in protein recognition region sequences. Table 87 Percentage of InDel targets containing IONCAS030 protein and its guide, and variations in protein recognition region sequences. Table 88 Percentage of InDel targets containing IONCAS057 protein and guide, and variations in protein recognition region sequences. Table 89 Percentage of InDel targets containing IONCAS065 protein and its guide, and variations in protein recognition region sequences. Table 90 Percentage of InDel targets containing IONCAS075 protein and guide, and variations in protein recognition region sequences. Table 91 Percentage of InDel targets containing IONCAS080 protein and its guide, and variations in protein recognition region sequences. Table 92 Percentage of InDel targets containing IONCAS105 protein and its guide, and variations in protein recognition region sequences. Example 10: Design and efficacy of Cas fusion protein and modified guide for epigenetic repression of PCSK9 in Hepa1-6 cells, LNP delivery The Cas protein described above was inserted into a Cas fusion protein containing a catalytically inactivated Cas protein and a DNA demethylation domain. A modified guide containing a target recognition region sequence targeting PCSK9, linked to the 5' end of the protein recognition region sequence, was designed for the Cas fusion protein and purchased from IDT. Epigenetic repression of PCSK9 was tested in mouse Hepa1-6 cells using the Cas fusion protein and the modified guide system.
[0560] Design of Cas fusion proteins for DNA methylation A nuclease-inactivating version of the Cas protein described above was designed by replacing the aspartic and histidine residues involved in nuclease activity with alanine. For the Cas protein IONCAS009, this point mutation is D11A and H871A. For the Cas protein IONCAS075, this point mutation is D11A and H868A. To confer DNA methylation activity, the C-terminal domain (3A3L, SEQ ID NO: 958) of Dnmtl3l is fused to the N-terminus of the catalytically inactivating Cas protein via an XTEN80 adapter sequence (SEQ ID NO: 959), the adapter sequence being incorporated to increase flexibility and spacing. The 3A3L domain and the XTEN80 adapter were previously disclosed in International Patent Application WO 2023 / 215711.
[0561] The sequence of the Cas fusion protein is provided in the sequence listing, as shown in the table below. The Cas fusion protein DNA sequence was cloned into the pmRVac vector from VectorBuilder for use in... in vitroTranscription (SEQ ID NO: 77). To facilitate enzyme localization to the nucleus, nucleoplasmic proteins (SEQ ID NO: 3399) and c-Myc (SEQ ID NO: 3400)NLS are attached to the N-terminus and C-terminus of each Cas protein sequence, respectively. Additionally, a HiBiT (Promega) tag is added to the C-terminus of the protein for expression detection. These additional elements are linked to the protein via known protein linkers such as GGGS / (GS)3 / (GGGGS)3. A poly-A tail is added to the 3'-end of the DNA sequence encoding the Cas protein with NLS and the HiBiT tag. These DNA sequences are provided in the sequence listing as shown in SEQ ID NO: 956 to 957 below.
[0562] Table 93 Sequence of Cas fusion protein Design and efficacy of PCSK9-targeting Cas fusion proteins and guide systems Using PAM sequences identified for each Cas fusion protein and PCSK9, target recognition region sequences for PCSK9 were designed. These target recognition region sequences can be found in the sequence listing, as shown in the table below. Subsequently, a wizard was designed by appending the target recognition region sequence to the 5' end of the protein recognition sequence C9g-NU (SEQ ID NO: 1052) for IONCAS009-d or C75g-46 (SEQ ID NO: 1375) for IONCAS075-d. A wizard for the positive control SpCas9 was also designed.
[0563] Each of the guides in the table below has the following modifications: the nucleosides at positions 1 and 2 at the 5' end are 2'-OMe sugar moieties, wherein each nucleoside is linked to the next nucleoside via a thiophosphate nucleoside bond; and the last two nucleosides at the 3' end are 2'-OMe sugar moieties, wherein each nucleoside is linked to the previous nucleoside via a thiophosphate nucleoside bond.
[0564] The wizards described in the table below are constructed from the target recognition regions identified by ID and SEQ ID NO in the "Target Recognition Region" column and the protein recognition regions identified by ID and SEQ ID NO in the "Protein Recognition Region" column. The nucleobase sequence of the wizard and the complete annotation sequence describing the modifications on the wizard can be found in the sequence listing, as shown in the wizard columns labeled "SEQ ID NO." and "Mod. SEQ ID No.", respectively.
[0565] Mouse Hepa1-6 cells were seeded in 96-well plates at a density of 20,000 cells per well 2 to 4 hours prior to transfection. The RNA / LNP complex was prepared by mixing 220 ng of Cas fusion protein mRNA and 440 ng of guide RNA with the LNP formulation. After incubation at room temperature for 10 minutes, the mixture was added to the culture medium.
[0566] Samples were collected 3 days post-transfection. RNA was extracted using RLT buffer (Thermo Fisher) and Pall AcroPrepAdvance 96-well plates. PCSK9 RNA was quantified by quantitative PCR using the TaqMan primer and probe set Mm00463746_m1 (Thermo Fisher). PCSK9 RNA levels were normalized to GAPDH RNA using the TaqMan primer and probe set Mm99999915_g1 (Thermo Fisher) or to ACTB using the TaqMan primer and probe set Mm02619580_g1 (Thermo Fisher). PCSK9 RNA expression is presented in the table below as a percentage of PCSK9 RNA (UTC percentage) relative to the amount of PCSK9 RNA in untreated control cells.
[0567] Table 94 Effects of IONCAS009-d and the guide system on PCKS9 expression in Hepa1-6 cells, 3 days Table 95 Effects of IONCAS075-d and the guide system on PCKS9 expression in Hepa1-6 cells, 3 days Duration of action of PCSK9-targeting Cas fusion protein and guide system The guide RNA was selected from the above-mentioned activity screening, and its activity over time was tested in mouse Hepa1-6 cells. Mouse Hepa1-6 cells were seeded in 96-well plates at a density of 20,000 cells per well 2 to 4 hours before transfection. The RNA / LNP complex was generated by mixing 220 ng of Cas fusion protein mRNA and 440 ng of guide RNA with the LNP formulation. SpCas9 and guide RNA SpPCSKg were included as controls. After incubation at room temperature for 10 minutes, the mixture was added to the culture medium.
[0568] Samples were collected at the time points indicated in the table below. Cells were passaged every 2 to 3 days to maintain optimal cell density. RNA was extracted using RLT buffer (Thermo Fisher) and Pall AcroPrep Advance 96-well plates. PCSK9 RNA was quantified by quantitative PCR using the TaqMan primer and probe set Mm00463746_m1 (Thermo Fisher). PCSK9 RNA levels were normalized to GAPDH RNA using the TaqMan primer and probe set Mm99999915_g1 (Thermo Fisher) or normalized to ACTB using the TaqMan primer and probe set Mm02619580_g1 (Thermo Fisher). PCSK9 RNA expression is presented in the table below as a percentage of PCSK9 RNA (UTC percentage) relative to the amount of PCSK9 RNA in untreated control cells.
[0569] Table 96 Effects of IONCAS009-d and guide on PCKS9 expression in Hepa1-6 cells, 3 to 14 days Table 97 Effects of IONCAS075-d and guide on PCKS9 expression in Hepa1-6 cells, 3 to 14 days
Claims
1. A polypeptide comprising a region having at least 800, at least 900, at least 1000, or at least 1050 linked amino acids, wherein the amino acid sequence of the region has at least 85%, at least 90%, at least 95%, or 100% identity with an isometric portion of any one of SEQ ID NO: 4 to 23 or 231 to 323, optionally, wherein the polypeptide is isolated.
2. A polypeptide comprising a region having at least 900, at least 1000, or at least 1100 linked amino acids, wherein the amino acid sequence of the region has at least 85%, at least 90%, at least 95%, or 100% identity with the isometric portions of SEQ ID NO: 6 to 23, 250 to 308; optionally, wherein the polypeptide is isolated.
3. A polypeptide comprising a region having at least 900, at least 1000, at least 1100, at least 1200, or at least 1300 linked amino acids, wherein the amino acid sequence of the region has at least 85%, at least 90%, at least 95%, or 100% identity with an isometric portion of any one of SEQ ID NO: 10 to 23 or 266 to 308; optionally, wherein the polypeptide is isolated.
4. A polypeptide comprising a region wherein the amino acid sequence of said region has at least 85%, at least 90%, at least 95%, or 100% identity with any one of SEQ ID NO: 4 to 23 or 231 to 323; optionally, said polypeptide is isolated.
5. A polypeptide, wherein the amino acid sequence of the polypeptide has at least 85%, at least 90%, at least 95%, or 100% identity with any one of SEQ ID NO: 4 to 23 or 231 to 323; optionally, wherein the polypeptide is isolated.
6. The polypeptide according to claim 5, wherein the amino acid sequence of the polypeptide has at least 85%, at least 90%, at least 95%, or 100% identity with any one of SEQ ID NO: 12, 13, 20, 241, 242, 245, 247, 260, 266, 267, 268, 271, 274, 275, 279, 282, 283, 284, 288, 289, 292, 297, 299, 300, 301, 302, 314, or 323.
7. The polypeptide according to any one of claims 1 to 4, wherein it comprises the said region.
8. The polypeptide according to any one of claims 1 to 4, comprising a heterologous domain.
9. The polypeptide according to claim 8, wherein the heterologous domain is selected from transcription activators, transcription repressors, methyltransferases, demethylases, deaminases, acetyltransferases, or deacetylases.
10. The polypeptide according to any one of claims 1 to 4 or 7 to 9, wherein the amino acid sequence of said region has at least 85%, at least 90%, at least 95%, or 100% identity with the isometric portion of any one of SEQ ID NO: 13, 241, 267, 275, 284, 289, or 314.
11. The polypeptide according to any one of claims 1 to 10, wherein the polypeptide comprises a non-PI region and a PI domain, wherein the non-PI region of the polypeptide has an amino acid sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3186, 3188, 3190, 3192, 3194, 3196, or 3198.
12. The polypeptide according to any one of claims 1 to 10, wherein the polypeptide comprises a non-PI region and a PI domain, wherein the non-PI region of the polypeptide has an amino acid sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3186, and the PI domain has an amino acid sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3187.
13. The polypeptide according to any one of claims 1 to 10, wherein the polypeptide comprises a non-PI region and a PI domain, wherein the non-PI region of the polypeptide has an amino acid sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3188, and the PI domain has an amino acid sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3189.
14. The polypeptide according to any one of claims 1 to 10, wherein the polypeptide comprises a non-PI region and a PI domain, wherein the non-PI region of the polypeptide has an amino acid sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3190, and the PI domain has an amino acid sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3191.
15. The polypeptide according to any one of claims 1 to 10, wherein the polypeptide comprises a non-PI region and a PI domain, wherein the non-PI region of the polypeptide has an amino acid sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3192, and the PI domain has an amino acid sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3193.
16. The polypeptide according to any one of claims 1 to 10, wherein the polypeptide comprises a non-PI region and a PI domain, wherein the non-PI region of the polypeptide has an amino acid sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3194, and the PI domain has an amino acid sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3195.
17. The polypeptide according to any one of claims 1 to 10, wherein the polypeptide comprises a non-PI region and a PI domain, wherein the non-PI region of the polypeptide has an amino acid sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3196, and the PI domain has an amino acid sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3197.
18. The polypeptide according to any one of claims 1 to 10, wherein the polypeptide comprises a non-PI region and a PI domain, wherein the non-PI region of the polypeptide has an amino acid sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3198, and the PI domain has an amino acid sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3199.
19. The polypeptide according to any one of claims 1 to 18, wherein the polypeptide has exactly two catalytically active nuclease sites.
20. The polypeptide according to any one of claims 1 to 18, wherein the polypeptide has exactly one catalytically active nuclease site.
21. The polypeptide according to any one of claims 1 to 18, wherein the polypeptide has a nuclease active site with zero catalytic activity.
22. The polypeptide according to any one of claims 1 to 21, comprising a nuclear localization signal.
23. The polypeptide of claim 22, wherein the nuclear localization signal is located at the N-terminus of the polypeptide.
24. The polypeptide of claim 22, wherein the nuclear localization signal is located at the C-terminus of the polypeptide.
25. The polypeptide of claim 22, wherein the polypeptide comprises two nuclear localization signals, one at the N-terminus and one at the C-terminus of the polypeptide.
26. A nucleic acid encoding a polypeptide according to any one of claims 1 to 25.
27. The nucleic acid according to claim 26, wherein the nucleic acid is DNA.
28. The nucleic acid of claim 26, wherein the nucleic acid is mRNA comprising a coding region encoding a polypeptide according to any one of claims 1 to 25.
29. The nucleic acid of claim 26, wherein the mRNA is an exogenous mRNA comprising a coding region encoding a polypeptide according to any one of claims 1 to 21.
30. The mRNA according to claim 28 or 29, comprising a 5' cap structure.
31. The mRNA according to any one of claims 28 to 30, comprising a 3'-poly(A) tail.
32. The mRNA according to any one of claims 28 to 31, comprising a 5'-UTR.
33. The mRNA according to any one of claims 28 to 32, comprising a 3'-UTR.
34. The mRNA according to claim 28 or 29, comprising a 5' cap structure and a 3'-poly(A) tail.
35. The mRNA according to any one of claims 28 to 34, wherein the coding region is codon-optimized for expression in eukaryotic cells.
36. The mRNA of claim 35, wherein the coding region has been codon-optimized for expression in mammalian cells.
37. The mRNA of claim 36, wherein the mammalian cell is a human cell.
38. The mRNA according to any one of claims 28 to 37, wherein each nucleobase of the mRNA is selected from adenine, guanine, cytosine, uracil, thymine, and N1-methylpseuuridine.
39. The mRNA according to any one of claims 28 to 37, wherein each nucleobase of the mRNA is selected from adenine, guanine, cytosine, and N1-methylpseuuridine.
40. The mRNA according to any one of claims 28 to 37, wherein each nucleobase of the mRNA is selected from adenine, guanine, cytosine, and uracil.
41. A guide, wherein the guide comprises a protein recognition region, wherein the protein recognition region interacts with a polypeptide according to any one of claims 1 to 25.
42. The guide of claim 41, wherein the guide is a dual guide.
43. The guide according to claim 41, wherein the guide is a single guide.
44. A guide composed of oligonucleotides according to the following formula: T1-D-X-S 1a -B1-S 1b -H1-S 1b '-B2-S 1a '-L1-S2-H2-S2'-L2-S3-H3-S3'-(L3-S4-H4-S4') n -(L4-S5-H5-S5') m -T2; in: D consists of 17 to 25 linked nucleosides, wherein the nucleobase sequence of D is complementary to the nucleobase sequence of the target DNA; X is absent, or consists of one or two linked nucleosides that are not complementary to the sequence of the target DNA; Each of T1 and T2 either lacks or is independently composed of 1 to 30 linked nucleosides; S 1a and S 1a Each is composed of 6 to 10 linked nucleosides, of which S 1a The nucleobase sequence and S 1a The nucleobase sequences of ' are 100% complementary; S 1b and S 1b Each is composed of 2 to 14 linked nucleosides, of which S 1b The nucleobase sequence and S 1b The nucleobase sequences of ' are 100% complementary; S2 and S2' are each composed of 2 to 6 linked nucleosides, wherein the nucleobase sequence of S2 is 100% complementary to the nucleobase sequence of S2'; S3 and S3' are each composed of 2 to 12 linked nucleosides, wherein the nucleobase sequence of S3 is 100% complementary to the nucleobase sequence of S3'; S4 and S4' are each composed of 4 to 14 linked nucleosides, wherein the nucleobase sequence of S4 is 100% complementary to the nucleobase sequence of S4'. S5 and S5' are each composed of 4 to 11 linked nucleosides, wherein the nucleobase sequence of S5 is 100% complementary to the nucleobase sequence of S5'. n is 0 or 1; m is 0 or 1; B1 is absent, or consists of 1 to 3 linked nucleosides; B2 is absent or consists of 1 to 4 linked nucleosides; B1 is not equal to B2 unless neither of them exists; Each of H1, H2, H3, H4, and H5 is independently composed of 3 to 6 linked nucleosides; L1 consists of 2 to 3 linked nucleosides; Each of L2, L3, and L4 is independently composed of 0 to 10 linked nucleosides.
45. The wizard of claim 44, wherein the wizard comprises a protein recognition region incorporating the polypeptide of any one of claims 1 to 25.
46. The guide according to any one of claims 44 to 45, wherein T1 is not present.
47. The guide according to any one of claims 44 to 46, wherein T2 is not present.
48. The guide according to any one of claims 44 to 46, wherein T2 consists of 2 to 6 linked nucleosides comprising uracil nucleobases.
49. The guide according to any one of claims 44 to 48, wherein X does not exist.
50. The guide according to any one of claims 44 to 49, wherein n is 1 and m is 0.
51. The guide according to any one of claims 44 to 50, wherein both n and m are 1.
52. The guide according to any one of claims 44 to 51, wherein the number of nucleosides in B2 is greater than the number of nucleosides in B1.
53. The guide according to any one of claims 44 to 52, wherein each L1 nucleoside is adenosine.
54. The guide according to any one of claims 44 to 53, wherein the guide is combined with the polypeptide according to any one of claims 1 to 25.
55. The wizard according to claim 44 or 45, wherein T1 and X do not exist; T2 is absent or consists of 2 to 6 linked uridines; S 1a and S 1a Each is composed of 7 linked nucleosides; S 1b and S 1b Each is composed of 3 to 9 linked nucleosides; S2 and S2' are each composed of three linked nucleosides; S3 and S3' are each composed of four linked nucleosides; S4 and S4' are each composed of 5 linked nucleosides; When present, S5 and S5' each consist of 7 to 9 linked nucleosides; n is 1; m is 0 or 1; B1 is composed of one nucleoside; B2 is composed of four linked nucleosides; When present, each of H1, H2, H3, and H4 consists of four linked nucleosides; L1 consists of two linked adenosine nucleotides; L2 consists of 5 linked nucleosides; L3 is composed of 0 nucleosides; and When present, L4 consists of 4 linked nucleosides.
56. The wizard of claim 55, wherein S 1b and S 1b Each is composed of 7 linked nucleosides.
57. The guide according to claim 55 or 56, wherein S5 and S5' are each composed of 7 linked nucleosides.
58. The guide according to any one of claims 55 to 57, wherein the nucleobase sequence of B2 is AAAG.
59. The guide according to any one of claims 55 to 58, wherein the nucleobase sequence of L2 is AAACU.
60. The guide according to any one of claims 55 to 59, wherein the nucleobase sequence of L3 is UUUUAA.
61. The guide according to any one of claims 55 to 60, wherein the nucleobase sequence of H2 is GUCA.
62. The guide according to any one of claims 55 to 61, wherein n is 1 and m is 0.
63. The guide according to any one of claims 55 to 62, wherein both n and m are 1.
64. The wizard of claim 55, wherein S 1b and S 1b Each is composed of 7 linked nucleosides; n is 1 and m is 0; the nucleobase sequence of B2 is AAAG; the nucleobase sequence of L2 is AAACU; the nucleobase sequence of L3 is UUUUAA, and the nucleobase sequence of H2 is GUCA.
65. The guide according to any one of claims 55 to 64, wherein the guide comprises a region having at least 40, at least 50, at least 60, or at least 70 nucleosides, wherein the nucleobase sequence of the region has at least 90%, 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with an isolength region of any one of SEQ ID NO: 131, 961 to 964, 1016 to 1020, 1052 to 1070, 1313 to 1314, or 2825.
66. The guide according to claim 65, wherein the nucleobase sequence of the region has 100% identity with the isochronous region of any one of SEQ ID NO: 131, 961 to 964, 1016 to 1020, 1052 to 1070, 1313 to 1314 or 2825.
67. The guide according to any one of claims 55 to 66, wherein D consists of 20 to 23 nucleosides having a nucleobase sequence that is 100% complementary to the nucleobase sequence of the target DNA.
68. The guide according to claim 44 or 45, wherein T1 and X do not exist; T2 is absent or consists of 6 linked uridines; S 1a and S 1a Each is composed of 9 linked nucleosides; S 1b and S 1b Each is composed of 2 to 8 linked nucleosides; S2 and S2' are each composed of 6 linked nucleosides; S3 and S3' are each composed of 7 linked nucleosides; n and m are 0; B1 is composed of 3 nucleosides; B2 is composed of four linked nucleosides; H1 consists of 4 linked nucleosides; H2 is composed of 5 linked nucleosides; H3 is composed of 4 linked nucleosides; L1 consists of two linked adenosine nucleotides; and L2 consists of 0 linked nucleosides.
69. The guide according to claim 68, wherein S 1b and S 1b Each is composed of 4 linked nucleosides.
70. The guide according to any one of claims 68 to 69, wherein L2 consists of 0 linked nucleosides and T2 consists of 6 linked uridines.
71. The guide according to any one of claims 68 to 70, wherein the nucleobase sequence of B2 is UAAC.
72. The guide according to any one of claims 68 to 71, wherein the nucleobase sequence of L2 is GAACUC.
73. The guide according to any one of claims 68 to 72, wherein the nucleobase sequence of H2 is UUUAU.
74. The guide according to claim 68, wherein S 1b and S 1b Each is composed of 4 linked nucleosides; n and m are 0; the nucleobase sequence of B2 is UACC; the nucleobase sequence of L2 is GAACUC; and the nucleobase sequence of H2 is UUUAU.
75. The guide according to any one of claims 68 to 74, wherein the guide comprises a region having a nucleobase sequence having at least 90%, 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any one of SEQ ID NO: 520, 965 to 968, 1071 to 1094, or 1315 to 1318, wherein the region comprises at least 40, at least 50, at least 60, or at least 70 nucleosides.
76. The guide according to claim 75, wherein the nucleobase sequence of the region has 100% identity with the isochronous region of any one of SEQ ID NO: 520, 965 to 968, 1071 to 1094 or 1315 to 1318.
77. The guide according to any one of claims 68 to 75, wherein D consists of 20 to 23 nucleosides having a nucleobase sequence that is 100% complementary to the nucleobase sequence of the target DNA.
78. The guide according to claim 44 or 45, wherein T1 and X do not exist; T2 is absent or consists of 1 to 7 linked nucleosides; S 1a and S 1a Each is composed of 7 linked nucleosides; S 1b and S 1b Each nucleotide is composed of 4 to 12 linked nucleosides; S2 and S2' are each composed of 5 linked nucleosides; S3 and S3' are each composed of 6 linked nucleosides; S4 and S4' are each composed of 6 linked nucleotides; When present, S5 and S5' are each composed of 7 linked nucleosides; n is 1; m is 0 or 1; B1 is composed of one nucleoside; B2 is composed of three linked nucleosides; Each of H1, H2 and H3 consists of 4 linked nucleosides; When present, H4 and H5 consist of 3 linked nucleosides; L1 consists of three linked adenosine nucleotides; L2 and L3 consist of 0 linked nucleosides; When present, L4 consists of 1 to 7 linked nucleosides.
79. The guide according to claim 78, wherein S 1b and S 1b Each is composed of 6 linked nucleosides.
80. The guide according to any one of claims 78 to 79, wherein the nucleobase sequence of B2 is GAG.
81. The guide according to any one of claims 78 to 80, wherein the nucleobase sequence of H2 is AUCC.
82. The guide according to claim 78, wherein S 1b and S 1b Each is composed of 6 linked nucleosides; n is 1 and m is 0; the nucleobase sequence of B2 is GAG; and the nucleobase sequence of H2 is AUCC.
83. The guide according to any one of claims 78 to 82, wherein the guide comprises a region having a nucleobase sequence having at least 90%, 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any one of SEQ ID NO: 546, 973 to 976, 1021 to 1023, 1095 to 1126, or 1319 to 1342, wherein the region comprises at least 40, at least 50, at least 60, or at least 70 nucleosides.
84. The guide according to claim 83, wherein the nucleobase sequence of the region has 100% identity with the isochronous region of any one of SEQ ID NO: 546, 973 to 976, 1021 to 1023, 1095 to 1126 or 1319 to 1342.
85. The guide according to any one of claims 78 to 84, wherein D consists of 20 to 24 nucleosides having a nucleobase sequence that is 100% complementary to the nucleobase sequence of the target DNA.
86. The guide according to claim 44 or 45, wherein T1 and X do not exist; T2 is absent or consists of 1 to 6 linked nucleosides; S 1a and S 1a Each is composed of 7 linked nucleosides; S 1b and S 1b Each nucleotide is composed of 4 to 12 linked nucleosides; S2 and S2' are each composed of four linked nucleosides; S3 and S3' are each composed of four linked nucleosides; S4 and S4' are each composed of 5 linked nucleosides; When present, S5 and S5' are each composed of 7 linked nucleosides; n is 1; m is 0 or 1; B1 is composed of one nucleoside; B2 is composed of 3 nucleosides; H1 consists of 4 linked nucleosides; H2 is composed of three linked nucleosides; H3 consists of 6 linked nucleosides; H4 consists of four linked nucleosides; When present, H5 consists of 3 linked nucleosides; L1 consists of two linked adenosine nucleotides; L2 consists of two linked nucleosides; L3 consists of 0 linked nucleosides; When present, L4 consists of 7 linked nucleosides.
87. The guide according to claim 86, wherein S 1b and S 1b It consists of 10 linked nucleosides each.
88. The guide according to any one of claims 86 to 87, wherein n is 1 and m is 0.
89. The guide according to any one of claims 86 to 88, wherein the nucleobase sequence of B2 is GAG.
90. The guide according to any one of claims 86 to 89, wherein the nucleobase sequence of L2 is AA.
91. The guide according to any one of claims 86 to 90, wherein the nucleobase sequence of H2 is UAA.
92. The guide according to claim 86, wherein S 1b and S 1b Each is composed of 10 linked nucleosides; n is 1 and m is 0; the nucleobase sequence of B2 is GAG; the nucleobase sequence of L2 is AA; and the nucleobase sequence of H2 is UAA.
93. The guide according to any one of claims 86 to 92, wherein the guide comprises a region having a nucleobase sequence having at least 90%, 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any one of SEQ ID NO: 562, 985 to 989, 1027 to 1029, 1153 to 1182, or 1362 to 1390, wherein the region comprises at least 40, at least 50, at least 60, or at least 70 nucleosides.
94. The guide according to claim 93, wherein the nucleobase sequence of the region has 100% identity with the isochronous region of any one of SEQ ID NO: 562, 985 to 989, 1027 to 1029, 1153 to 1182 or 1362 to 1390.
95. The guide according to any one of claims 86 to 94, wherein D consists of 20 to 24 nucleosides having a nucleobase sequence that is 100% complementary to the nucleobase sequence of the target DNA.
96. The guide according to claim 44 or 45, wherein T1 and X do not exist; T2 is absent or consists of 2 to 5 linked nucleosides; S 1a and S 1a Each is composed of 9 linked nucleosides; S 1b and S 1b Each is composed of 2 to 10 linked nucleosides; S2 and S2' are each composed of four linked nucleosides; S3 and S3' are each composed of 9 linked nucleosides; n and m are 0; B1 is composed of one nucleoside; B1 is composed of 3 nucleosides; H1 consists of 4 linked nucleosides; H2 is composed of three linked nucleosides; H3 is composed of 4 linked nucleosides; L1 consists of two linked adenosine nucleotides; L2 consists of 7 linked nucleosides.
97. The guide according to claim 96, wherein S 1b and S 1b Each is composed of 6 linked nucleosides.
98. The guide according to any one of claims 96 to 97, wherein the nucleobase sequence of B2 is CUA.
99. The guide according to any one of claims 96 to 98, wherein the nucleobase sequence of L2 is GUGUUUA.
100. The guide according to any one of claims 96 to 99, wherein the nucleobase sequence of H2 is AAA.
101. The guide according to claim 96, wherein S 1b and S 1b Each is composed of 10 linked nucleosides; n and m are 0; the nucleobase sequence of B2 is CUA; the nucleobase sequence of L2 is GUGUUUA; and the nucleobase sequence of H2 is AAA.
102. The guide according to any one of claims 96 to 101, wherein the guide comprises a region having a nucleobase sequence having at least 90%, 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any one of SEQ ID NO: 593, 1006 to 1010, 1272 to 1292, 1394 to 1408, 2827 to 2833, wherein the region comprises at least 40, at least 50, at least 60, or at least 70 nucleosides.
103. The guide according to claim 102, wherein the nucleobase sequence of the region has 100% identity with the isochronous region of any one of SEQ ID NO: 593, 1006 to 1010, 1272 to 1292, 1394 to 1408, 2827 to 2833.
104. The guide according to any one of claims 96 to 103, wherein D consists of 20 to 24 nucleosides having a nucleobase sequence that is 100% complementary to the nucleobase sequence of the target DNA.
105. The guide according to claim 44 or 45, wherein T1 and X do not exist; T2 either does not exist or is composed of 3 linked nucleosides; S 1a and S 1b Together, and S 1a 'and S 1b Together, they are composed of 12 to 20 linked nucleosides; S2 and S2' are each composed of three linked nucleosides; S3 and S3' are each composed of three linked nucleosides; S3 and S3' are each composed of 10 to 13 linked nucleosides; n is 1 and m is 0; B1 consists of 0 nucleosides; B2 consists of 0 nucleosides; H1, H2, H3, and H4 are composed of four linked nucleosides; L1 consists of two linked adenosine nucleotides; L2 consists of two linked nucleosides; and L3 consists of 6 linked nucleosides.
106. The wizard of claim 105, wherein S 1a and S 1b Together, and S 1a 'and S 1b Together, they are composed of 15 linked nucleosides.
107. The guide according to any one of claims 105 to 106, wherein L3 consists of 3 linked nucleosides.
108. The guide according to any one of claims 105 to 107, wherein S3 and S3' are each composed of 13 linked nucleosides.
109. The guide according to any one of claims 105 to 107, wherein the nucleobase sequence of L2 is GU.
110. The guide according to any one of claims 108 to 109, wherein the nucleobase sequence of H2 is GAAA.
111. The guide according to claim 105, wherein S 1a and S 1b Together, and S 1a 'and S 1b Together, each is composed of 15 linked nucleosides; n is 1 and m is 0; the nucleobase sequence of L2 is GU; and the nucleobase sequence of H2 is GAAA.
112. The guide according to any one of claims 105 to 111, wherein the guide comprises a region having a nucleobase sequence having at least 90%, 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any one of SEQ ID NO: 568, 990 to 993, 1030 to 1034, 1183 to 1206, or 1391 to 1393, wherein the region comprises at least 40, at least 50, at least 60, or at least 70 nucleosides.
113. The guide according to claim 112, wherein the nucleobase sequence of the region has 100% identity with the isochronous region of any one of SEQ ID NO: 568, 990 to 993, 1030 to 1034, 1183 to 1206 or 1391 to 1393.
114. The guide according to any one of claims 106 to 113, wherein D consists of 20 to 24 nucleosides having a nucleobase sequence that is 100% complementary to the nucleobase sequence of the target DNA.
115. A guide comprising oligonucleotides according to the following formula: T1-DXS 1a -B1-S 1b -H1-S 1b '-B2-S 1a '-L1-S2-H2-S2'-L2-S3-H3-S3'-(L3-S4-H4-S4') n -T2; in: D consists of 17 to 25 linked nucleosides, wherein the nucleobase sequence of D is complementary to the nucleobase sequence of the target DNA; X is absent, or consists of one or two linked nucleosides that are not complementary to the nucleobase sequence of the target DNA; Each of T1 and T2 either lacks or is independently composed of 1 to 30 linked nucleosides; S 1a and S 1a Each is composed of 7 linked nucleosides, of which S 1a The nucleobase sequence and S 1a The nucleobase sequences of ' are 100% complementary; S 1b and S 1b Each is composed of 4 to 12 linked nucleosides, of which S 1b The nucleobase sequence and S 1b The nucleobase sequences of ' are 100% complementary; S2 and S2' are each composed of 6 linked nucleosides, and the nucleobase sequence of S2 is 100% complementary to the nucleobase sequence of S2'. S3 and S3' are each composed of 5 or 6 linked nucleosides, wherein the nucleobase sequence of S3 is 100% complementary to the nucleobase sequence of S3'. When present, S4 and S4' each consist of 6 linked nucleosides, wherein the nucleobase sequence of S4 is 100% complementary to the nucleobase sequence of S4'. n is 0 or 1; B1 is composed of one nucleoside; B2 is composed of three linked nucleosides; Each of H1, H2, H3, and H4 is independently composed of 3 to 4 linked nucleosides; L1 consists of 17 linked nucleosides; L2 consists of 0 linked nucleosides; and L3 consists of 8 linked nucleosides.
116. The guide according to claim 115, wherein S 1b and S 1b Each is composed of 8 linked nucleosides.
117. The guide according to any one of claims 115 to 116, wherein the nucleobase sequence of B2 is GAA.
118. The guide according to any one of claims 115 to 117, wherein the nucleobase sequence of L2 is AAAAAUUUAUUCAAAAC (SEQ ID NO: 78).
119. The guide according to any one of claims 115 to 118, wherein the nucleobase sequence of H2 is GAAA.
120. The guide according to claim 119, wherein S 1b and S 1b Each is composed of 8 linked nucleosides; n and m are 0; the nucleobase sequence of B2 is GAA; the nucleobase sequence of L2 is AAAAAUUUAUUCAAAAC (SEQ ID NO: 78); and the nucleobase sequence of H2 is GAAA.
121. The guide according to any one of claims 115 to 120, wherein the guide comprises a region having a nucleobase sequence having at least 90%, 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any one of SEQ ID NO: 554, 981 to 984, 1024 to 1026, 1127 to 1152, or 1343 to 1361, wherein the region comprises at least 40, at least 50, at least 60, or at least 70 nucleosides.
122. The guide according to claim 121, wherein the nucleobase sequence of the region has 100% identity with the isochronous region of any one of SEQ ID NO: 554, 981 to 984, 1024 to 1026, 1127 to 1152 or 1343 to 1361.
123. The guide according to any one of claims 115 to 122, wherein D consists of 20 to 24 nucleosides having a nucleobase sequence that is 100% complementary to the nucleobase sequence of the target DNA.
124. The guide according to any one of claims 41 to 43, wherein the guide comprises a region having a nucleobase sequence having at least 90%, 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any one of SEQ ID NO: 121 to 141, 510 to 602, 961 to 1430, 1436 to 2833, wherein the region comprises at least 40, at least 50, at least 60, at least 70, at least 80, or at least 90 nucleosides.
125. The guide according to any one of claims 41 to 43, wherein the guide comprises a protein recognition region having a nucleobase sequence having at least 90%, 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any one of SEQ ID NO: 121 to 141, 510 to 602, 961 to 1430, 1436 to 2833.
126. The guide according to any one of claims 41 to 125, wherein the guide comprises a modified oligonucleotide.
127. The guide according to any one of claims 41 to 126, wherein the guide comprises at least one modified sugar portion.
128. The guide according to claim 127, wherein the modified sugar moiety is selected from 2'-OMe and 2'-F.
129. The guide according to claim 127, wherein the modified oligonucleotide comprises a 2'-OMe or 2'-F sugar moiety within the first five nucleotides at the 5' end of the guide or within the last five nucleotides at the 3' end of the guide.
130. The guide according to claims 127 to 129, wherein the guide comprises 2 to 3 2'-OMe sugar portions at the 5' end of the guide and 2 to 3 2'-OMe sugar portions at the 3' end of the guide.
131. The guide of claim 130, wherein each non-2'-OMe sugar moiety of the guide comprises an unmodified RNA sugar moiety.
132. The guide according to any one of claims 127 to 131, wherein the guide comprises a modified nucleoside inter-bond.
133. The guide according to claim 132, wherein the modified nucleoside inter-bond is a thiophosphate nucleoside inter-bond.
134. The guide according to claim 132 or 133, wherein the guide contains at least one modified internucleotide bond within the first five nucleotides at the 5' end of the guide or within the last five nucleotides at the 3' end of the guide.
135. The guide according to claim 134, wherein the guide comprises 2 to 3 thiophosphate nucleoside inter-bonds at the 5' end of the guide and 2 to 3 thiophosphate nucleoside inter-bonds at the 3' end of the guide.
136. The guide according to claim 135, wherein each remaining nucleoside inter-bond is an unmodified phosphodiester nucleoside inter-bond.
137. The guide according to any one of claims 132 to 136, wherein each nucleoside inter-bond of the guide is a standard-length nucleoside inter-bond.
138. The guide according to any one of claims 41 to 125, wherein the guide comprises an unmodified oligonucleotide.
139. The guide according to any one of claims 41 to 138, wherein the guide comprises a target recognition region that is at least 90%, at least 95%, or 100% complementary to the target DNA sequence.
140. The guide of claim 139, wherein the target recognition region comprises 17 to 25, 19 to 24, 20 to 24, 20 to 23, 20 to 22, 21 to 22, 21 to 23, 21 to 24, 20, 21, 22, 23 or 24 linked nucleosides.
141. An editing system comprising a polypeptide according to any one of claims 1 to 25 and a wizard according to any one of claims 41 to 139.
142. An editing system comprising: Nucleic acid according to any one of claims 26 to 40; and The guide according to any one of claims 41 to 140.
143. An editing system comprising: Nucleic acid according to any one of claims 26 to 40; and The nucleic acid encoding the guide according to any one of claims 41 to 125 or 138.
144. The editing system according to claim 142 or 143, wherein the nucleic acid encoding the polypeptide is exogenous mRNA.
145. The editing system according to claim 142 or 143, wherein the nucleic acid encoding the polypeptide is DNA.
146. The editing system according to any one of claims 143 to 145, wherein the nucleic acid encoding the wizard is DNA.
147. A composition comprising the editing system and LNP according to any one of claims 141 to 146.
148. A viral vector comprising the editing system according to claim 143.
149. A method for editing a target nucleic acid, the method comprising administering to a subject the composition of claim 146 or the viral vector of claim 148.
150. A method for editing target nucleic acid, the method comprising contacting a cell with the composition of claim 147 or the viral vector of claim 148.
151. A method for creating double-strand breaks or nicks in a target nucleic acid, the method comprising contacting a cell with the composition of claim 147 or the viral vector of claim 148.
152. The composition according to claim 147 or the viral vector according to claim 148 for use in therapy.
Citation Information
Patent Citations
Compositions and methods related to mRNA translational enhancer elements
EP2610340A1
Compositions and methods related to mRNA translational enhancer elements
EP2610341A1
Liposomal apparatus and manufacturing methods
US20040142025A1
Cationic lipids and methods of use
US20060083780A1
Lipid nanoparticle based compositions and methods for the delivery of biologically active molecules
US20060240554A1